repository-harness is a repository protocol that makes a software codebase easier for coding agents to understand and work on safely. It is for teams using Claude Code, Codex, Cursor, and similar agents who need authoritative documents, durable plans, explicit decision boundaries, and evidence for completed work. Its catalogue entries install the protocol's skills and instructions.
Borrowing it
Nothing to install: this file belongs to hoangnb24/repository-harness. Take a copy, put it at the same path in your own repository, and replace the rules that are about this project with yours.
curl -O https://raw.githubusercontent.com/hoangnb24/repository-harness/main/.agents/skills/improve-harness/SKILL.mdgit clone --depth 1 https://github.com/hoangnb24/repository-harnessWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/hoangnb24/repository-harness/improve-harness)<a href="https://agentmods.dev/skills/hoangnb24/repository-harness/improve-harness"><img src="https://agentmods.dev/badge/skills/hoangnb24/repository-harness/improve-harness/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/hoangnb24/repository-harness/improve-harness"><img src="https://agentmods.dev/badge/skills/hoangnb24/repository-harness/improve-harness.svg" alt="Reviewed on agentmods" width="80" height="20"></a>- NVIDIA SkillSpector pass
What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00076 | $0.00957 |
| Opus 5 | $0.00038 | $0.00478 |
| Sonnet 5 | $0.00015 | $0.00191 |
| Haiku 4.5 | $0.00008 | $0.00096 |
Grade A, and why
improve-harness scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 12d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 113 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Improve Harness
Improve one bounded future-agent behavior without turning every difficult task into permanent process. Keep consumer truth with its owner and require a fresh rerun before claiming improvement.
Establish Authority
- Read
AGENTS.md,docs/WORKFLOW.md, and applicable local instructions. - Confirm the request authorizes changing Harness behavior. Inspection or a request to report friction does not authorize edits.
- Record the initial repository root, revision, branch, status, and unrelated changes. Preserve all existing work.
- Treat invocation as authority for this bounded experiment, not for changing product policy, weakening proof, adding credentials, or mutating external systems.
1. Preserve The Baseline
Use an observed task trajectory when available. Record:
- the representative job and accepted outcome;
- the concrete failure and evidence;
- human steering, relay, or recovery required;
- the worker, repository revision, relevant external state, tools, and authority; and
- existing proof and known limitations.
Do not diagnose a worker limitation from one run. If no observed baseline exists, stop with an experiment proposal; do not manufacture one.
Copy docs/templates/harness-improvement.md to
docs/plans/active/harness-improvement-<slug>.md. Reuse an existing active
record for the same experiment.
2. Locate The Earliest Gap
Trace the failure upstream to the first owner that could have prevented or exposed it:
- Context: knowledge was absent, stale, overloaded, or not retrieved.
- Capability: discovery, invocation, interpretation, repair, or real-system verification failed.
- Domain ownership: no canonical type, API, state machine, or source owned the invariant.
- Authority: permission, approval, audit, or recovery was unclear.
- Proof: checks established a proxy rather than the accepted outcome.
- Environment: an external prerequisite was unavailable.
Assign the correction to repository-harness, the consumer repository, the
external environment, or a human decision. Do not copy consumer commands or
policy into a generic upstream template.
What ships with it
1 file beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 12d ago First seen · 113 lines · 76 tokens per session scan A 1061235593bc
improve-harness is a skill published in the GitHub repository hoangnb24/repository-harness (1,215 stars, last pushed 1mo ago), licensed MIT. It adds 76 tokens to every session and 957 once invoked, about $0.0004 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other skills, from other repositories
implementation-loop
Disciplined plan-build-verify loop for implementing a feature or change. Use when building something new or modifying existing behavior, especially across multiple files.
debugging-loop
Systematic reproduce-isolate-fix-verify loop for bugs, test failures, and unexpected behavior. Use when something is broken and the cause is not yet known.
lite-tools
Use compact wrappers when running routine Maven builds and tests, supported npm/Node test workflows, or Go tests. Select mvn-lite for supported Maven build and test workflows, npm-lite for npm run verify, npm run test:unit, or node --test, and go-lite for go test.
repo-map
Use when locating a registered or related local repository, or when discovering commands registered through repo-map. Do not use for source-tree, symbol, or architecture mapping.
steadyagent-workflow
Local-first Codex workflow for planning, debugging, reviewing, refactoring, improving AGENTS.md, building skills, publishing agent harness repositories, or running complex multi-step coding tasks that need staged diagnosis, context control, verification loops, review strategy, release evidence, and…
ackit-workflow
Enforce the ACKit docs-first, task-first workflow with one active checklist item and evidence-based completion.