Getting it into your agent
It runs from inside its repository, so the clone comes first — what it calls does not travel with the file alone.
git clone --depth 1 https://github.com/vasuag09/harness-claudenpx agentmods add skills/vasuag09/harness-claude/operateWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/vasuag09/harness-claude/operate)<a href="https://agentmods.dev/skills/vasuag09/harness-claude/operate"><img src="https://agentmods.dev/badge/skills/vasuag09/harness-claude/operate/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/vasuag09/harness-claude/operate"><img src="https://agentmods.dev/badge/skills/vasuag09/harness-claude/operate.svg" alt="Reviewed on agentmods" width="80" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00098 | $0.01355 |
| Opus 5 | $0.00049 | $0.00678 |
| Sonnet 5 | $0.00020 | $0.00271 |
| Haiku 4.5 | $0.00010 | $0.00136 |
Grade A, and why
operate scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 8d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 81 lines — stays where its author put it; the contents beside it link to each section on GitHub.
/operate — run long, without silent drift
Goal: let an agent work unattended across many iterations (or wake on a schedule) and
guarantee it halts rather than drifting — spinning on a dead end, burning budget, or
quietly breaking the repo. This skill is the discipline layer: the platform /loop and
/schedule are the engine; /harness-claude:operate adds guardrails, durable state, and a drift check
wired to the harness's own eval skills.
Opt-in. No default-pipeline skill or hook starts a run. Running long is always explicit. Git boundary: a run never commits or pushes unless you explicitly arm it for that in the objective; branch creation for the run's non-trivial work follows
rules/git.md(branch-at-first-write).
How it works
Each iteration the operator does one increment of work, then runs a checkpoint:
node scripts/operate/step.js --id <run-id> [--spec <path>] [--cmd '<check>' ...] \
[--max-iterations <n>] [--max-fails <n>] [--budget-ms <n>]
step.js loads the durable run state (.claude/runs/<id>.json), runs the drift check —
harness-claude:health (test/lint pulse) and, when --spec is set, harness-claude:eval
(acceptance-criteria gate) — updates and persists state, evaluates the guardrails, and exits:
- 0 = continue — schedule the next iteration.
- 1 = halt — a guardrail tripped (
drift|budget|iteration-cap); stop and report. - 2 = usage/error — fix the invocation (e.g. missing
--id).
The state file is the sole source of truth across firings — each /loop firing may be a
fresh context, so iteration count, budget spent, and the consecutive-fail counter persist there
and resume automatically. Config flags are honored only when the run is first created; later
firings ignore them so counts accumulate rather than reset.
Do this
- Frame the run. Name a one-line objective, a
--max-iterationsceiling, a drift threshold--max-fails(default 2 — one failure can be transient, two is a trend), and a--budget-mswall-clock cap. Point--specat the spec the run must keep satisfying when one exists (that arms theharness-claude:evalhalf of the drift check). - Pick a trigger.
- Interval / autonomous: drive iterations with the platform
/loop(fixed interval) or let the model self-pace; each firing advances the work one increment, then callsstep.js. - Scheduled / cron: use the platform
/scheduleto wake on a cadence; each wake runs one checkpointed iteration.
- Interval / autonomous: drive iterations with the platform
- Supervise by exit code. Continue on
0; on1, stop the loop, read the printed summary, and surface the halt reason and the failing criterion — never restart blindly past adrifthalt. Delegate the per-iteration work to theharness-claude:loop-operatoragent. - Report on halt. Always end with the run summary: iterations run, budget spent, last
verdict, halt reason. (
step.jsprints it; relay it.)
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 8d ago First seen · 81 lines · 98 tokens per session scan A a9ca47362386
operate is a skill published in the GitHub repository vasuag09/harness-claude (2 stars, last pushed 2mo ago), licensed MIT. It adds 98 tokens to every session and 1,355 once invoked, about $0.0005 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other skills, from other repositories
pr-triage
4-phase PR backlog management with audit, deep code review, validated comments, and optional worktree setup. Use when triaging pull requests, catching up on pending code reviews, or managing a backlog of open PRs. Args: 'all' to review all, PR numbers to focus (e.g. '42 57'), 'en'/'fr' for language, no arg = audit…
audit-agents-skills
Audit Claude Code agents, skills, and commands for quality and production readiness. Use when evaluating skill quality, checking production readiness scores, or comparing agents against best-practice templates.
eval-agents
Audit Claude Code agents defined in .claude/agents/ for description specificity, model tier appropriateness, tools scoping, and system prompt quality. Detects dispatch ambiguity between agents, flags over-permissive tool grants, and checks for human-in-the-loop patterns that break programmatic orchestration. Use when…
check-cache-bugs
Audit Claude Code setup for cache bugs (CC#40524): sentinel, --resume/--continue, attribution header + ArkNill B3/B4/B5.
issue-triage
3-phase issue backlog management with audit, deep analysis, and validated triage actions. Use when triaging GitHub issues, sorting bug reports, cleaning up stale tickets, or detecting duplicate issues. Args: 'all' to analyze all, issue numbers to focus (e.g. '42 57'), 'en'/'fr' for language, no arg = audit only.
git-ai-archaeology
Analyze AI config evolution in a git repo. Use when mapping AI adoption history, finding when configs were first introduced, charting commit velocity by month, or identifying maturity phases in a project's AI tooling.