repository-harness: Skill for Codex

.agents/skills/improve-harness/SKILL.md

improve-harness is a skill for Codex from hoangnb24/repository-harness. It costs 76 tokens per session (957 once invoked), scanned A, original, MIT.

A controlled workflow for making one evidence-based improvement to a coding agent’s instructions, tools, procedures, or checks.

In plain words
What is it for?
Use it to investigate repeated agent problems, preserve a baseline, propose or apply one bounded improvement, and verify whether that improvement works.
Why use it?
It prevents teams from turning a single difficult task into unnecessary permanent process. Changes require observed evidence and a fresh rerun before they are accepted.

Skill for Codex

Written for Codex: agents/openai.yaml present. Also seen: installed under .agents/ (shared by several agents); mentions AGENTS.md; $skill-name invocation.

This is hoangnb24/repository-harness's own configuration. It tells Codex how to work on repository-harness itself, so it is not a mod to install elsewhere. Copy it as a starting point and replace the rules that are about this project. Everything repository-harness configures →

About the project

repository-harness is a repository protocol that makes a software codebase easier for coding agents to understand and work on safely. It is for teams using Claude Code, Codex, Cursor, and similar agents who need authoritative documents, durable plans, explicit decision boundaries, and evidence for completed work. Its catalogue entries install the protocol's skills and instructions.

hoangnb24/repository-harness · 1,215 stars · on GitHub

Reuse

Borrowing it

Nothing to install: this file belongs to hoangnb24/repository-harness. Take a copy, put it at the same path in your own repository, and replace the rules that are about this project with yours.

Copy the file
curl -O https://raw.githubusercontent.com/hoangnb24/repository-harness/main/.agents/skills/improve-harness/SKILL.md
Clone the repo
git clone --depth 1 https://github.com/hoangnb24/repository-harness

Made for: Codex.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for improve-harness

README.md
[![agentmods](https://agentmods.dev/badge/skills/hoangnb24/repository-harness/improve-harness/github.svg)](https://agentmods.dev/skills/hoangnb24/repository-harness/improve-harness)
Your own site
<a href="https://agentmods.dev/skills/hoangnb24/repository-harness/improve-harness"><img src="https://agentmods.dev/badge/skills/hoangnb24/repository-harness/improve-harness/github.svg" alt="Measured on agentmods" height="20"></a>

Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.

agentmods 80×15 button for improve-harness

Your own site · 80×15
<a href="https://agentmods.dev/skills/hoangnb24/repository-harness/improve-harness"><img src="https://agentmods.dev/badge/skills/hoangnb24/repository-harness/improve-harness.svg" alt="Reviewed on agentmods" width="80" height="20"></a>
Per session 76 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 957 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. A grade says what 26 rules found in the file — not that it is safe. Third-party audits
  • NVIDIA SkillSpector pass 7 Sept 2026
How audits are shown
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5.1 $0.00076 $0.00957
Opus 5 $0.00038 $0.00478
Sonnet 5 $0.00015 $0.00191
Haiku 4.5 $0.00008 $0.00096

Measured 12d ago against content hash 1061235593bc, method: parsed. Prices are Anthropic first-party input rates as of 2026-09-12, from the pricing page.

Security

Grade A, and why

improve-harness scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 12d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

.agents/skills/improve-harness/SKILL.md · 113 lines

How it starts

The opening of the file, as written. The whole thing — 113 lines — stays where its author put it; the contents beside it link to each section on GitHub.

Improve Harness

Improve one bounded future-agent behavior without turning every difficult task into permanent process. Keep consumer truth with its owner and require a fresh rerun before claiming improvement.

Establish Authority

  • Read AGENTS.md, docs/WORKFLOW.md, and applicable local instructions.
  • Confirm the request authorizes changing Harness behavior. Inspection or a request to report friction does not authorize edits.
  • Record the initial repository root, revision, branch, status, and unrelated changes. Preserve all existing work.
  • Treat invocation as authority for this bounded experiment, not for changing product policy, weakening proof, adding credentials, or mutating external systems.

1. Preserve The Baseline

Use an observed task trajectory when available. Record:

  • the representative job and accepted outcome;
  • the concrete failure and evidence;
  • human steering, relay, or recovery required;
  • the worker, repository revision, relevant external state, tools, and authority; and
  • existing proof and known limitations.

Do not diagnose a worker limitation from one run. If no observed baseline exists, stop with an experiment proposal; do not manufacture one.

Copy docs/templates/harness-improvement.md to docs/plans/active/harness-improvement-<slug>.md. Reuse an existing active record for the same experiment.

2. Locate The Earliest Gap

Trace the failure upstream to the first owner that could have prevented or exposed it:

  • Context: knowledge was absent, stale, overloaded, or not retrieved.
  • Capability: discovery, invocation, interpretation, repair, or real-system verification failed.
  • Domain ownership: no canonical type, API, state machine, or source owned the invariant.
  • Authority: permission, approval, audit, or recovery was unclear.
  • Proof: checks established a proxy rather than the accepted outcome.
  • Environment: an external prerequisite was unavailable.

Assign the correction to repository-harness, the consumer repository, the external environment, or a human decision. Do not copy consumer commands or policy into a generic upstream template.

Read the full file on GitHub · 113 lines

Files

What ships with it

1 file beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 12d ago First seen · 113 lines · 76 tokens per session scan A 1061235593bc

Subscribe to this mod's changes

improve-harness is a skill published in the GitHub repository hoangnb24/repository-harness (1,215 stars, last pushed 1mo ago), licensed MIT. It adds 76 tokens to every session and 957 once invoked, about $0.0004 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.

Related

Other skills, from other repositories

implementation-loop

Disciplined plan-build-verify loop for implementing a feature or change. Use when building something new or modifying existing behavior, especially across multiple files.

KingEmma7/cursor-os · 33 tokens

debugging-loop

Systematic reproduce-isolate-fix-verify loop for bugs, test failures, and unexpected behavior. Use when something is broken and the cause is not yet known.

KingEmma7/cursor-os · 37 tokens

lite-tools

Use compact wrappers when running routine Maven builds and tests, supported npm/Node test workflows, or Go tests. Select mvn-lite for supported Maven build and test workflows, npm-lite for npm run verify, npm run test:unit, or node --test, and go-lite for go test.

ejboy/agent-scripts · 62 tokens

repo-map

Use when locating a registered or related local repository, or when discovering commands registered through repo-map. Do not use for source-tree, symbol, or architecture mapping.

ejboy/agent-scripts · 35 tokens

steadyagent-workflow

Local-first Codex workflow for planning, debugging, reviewing, refactoring, improving AGENTS.md, building skills, publishing agent harness repositories, or running complex multi-step coding tasks that need staged diagnosis, context control, verification loops, review strategy, release evidence, and…

Khalilzhang0825/boring-is-all-you-need · 64 tokens

ackit-workflow

Enforce the ACKit docs-first, task-first workflow with one active checklist item and evidence-based completion.

Cynrath/agent-context-kit · 26 tokens