propose-harness-change

propose-harness-change is a skill for Claude Code from Rockielab/rockie-claude. It costs 134 tokens per session (1,102 once invoked), scanned A, original, Apache-2.0.

A controlled process for proposing improvements to a coding-agent harness, such as a hook, script, or skill. It separates the person creating a patch from the fresh-context reviewer who checks it.

In plain words
What is it for?
Use it to prepare a reviewed patch, explain the reason for the change, provide a smoke test, and optionally prepare the change for a pull request to the upstream repository.
Why use it?
It reduces the risk that an automatic self-improvement changes the harness without independent review or proof that the change works.

Skill for Claude Code

Written for Claude Code: shipped in a Claude Code plugin. Also seen: mentions CLAUDE.md.

Part of the rockie-claude plugin — 29 skills, 1 MCP server shipped together

Good fit Use it to prepare a reviewed patch, explain the reason for the change, provide a smoke test, and optionally prepare the change for a pull request to the upstream repository.

Compare 6 skills from other repositories ↓
Install with agentmods
npx agentmods add skills/rockielab/rockie-claude/propose-harness-change
Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

Any agent
npx skills add Rockielab/rockie-claude --skill propose-harness-change
Clone the repo
git clone --depth 1 https://github.com/Rockielab/rockie-claude

Made for: Claude Code.

Or install rockie-claude, the plugin that ships this one along with the rest of its 29 skills, 1 MCP server.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for propose-harness-change

README.md
[![agentmods](https://agentmods.dev/badge/skills/rockielab/rockie-claude/propose-harness-change/github.svg)](https://agentmods.dev/skills/rockielab/rockie-claude/propose-harness-change)
Your own site
<a href="https://agentmods.dev/skills/rockielab/rockie-claude/propose-harness-change"><img src="https://agentmods.dev/badge/skills/rockielab/rockie-claude/propose-harness-change/github.svg" alt="Measured on agentmods" height="20"></a>

Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.

agentmods 80×15 button for propose-harness-change

Your own site · 80×15
<a href="https://agentmods.dev/skills/rockielab/rockie-claude/propose-harness-change"><img src="https://agentmods.dev/badge/skills/rockielab/rockie-claude/propose-harness-change.svg" alt="Reviewed on agentmods" width="80" height="20"></a>
Per session 134 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 1,102 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. A grade says what 26 rules found in the file — not that it is safe.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5.1 $0.00134 $0.01102
Opus 5 $0.00067 $0.00551
Sonnet 5 $0.00027 $0.00220
Haiku 4.5 $0.00013 $0.00110

Measured 12d ago against content hash 9f3f10a40cb8, method: parsed. Prices are Anthropic first-party input rates as of 2026-09-12, from the pricing page.

Security

Grade A, and why

propose-harness-change scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 12d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

Origin

Copies of this mod

1 near-identical copy found in the catalogue:

project-harness/skills/propose-harness-change/SKILL.md · 99 lines

How it starts

The opening of the file, as written. The whole thing — 99 lines — stays where its author put it; the contents beside it link to each section on GitHub.

/propose-harness-change — safe self-improvement

Autonomous research harnesses that let the agent edit themselves tend to drift (MINJA / eTAMP memory-poisoning, arXiv 2603.29231's finding that "memory scaffolds universally decrease long-horizon reliability", the Ouroboros "CLAUDE.md rewrite" footgun). rockie's discipline is Generator / Verifier / Updater separation — nobody is allowed to propose, verify, and commit in the same role.

The three roles

Generator (the proposing agent)

  • Writes the diff against a LOCAL CLONE of the rockie source repo.
  • Writes a short rationale: what pattern broke, why the fix composes with existing differentiators, what smoke-test assertion(s) prove it.
  • Never commits directly. Produces a patch file ~/rockie-proposals/<YYYY-MM-DD-slug>/patch.diff plus rationale.md, test.sh (the specific smoke-test snippet).

Verifier (fresh-context audit agent)

  • Dispatched via the Agent tool with NO prior context.
  • Reads the patch + rationale, the files being touched, and the CONTRIBUTING.md composition rules.
  • Must answer four questions with evidence:
    1. Does this compose with the existing differentiators, or duplicate one of them?
    2. Does the smoke test actually test the claimed improvement?
    3. Is the change local (one file) or does it ripple across the schema?
    4. Is there a path-traversal, SQL-injection, or shell-injection regression?
  • Returns APPROVE | CHANGES_REQUESTED | REJECT with a short report.

Updater (the human)

  • Reviews Verifier report + diff.
  • Runs bash tests/smoke-test.sh in the rockie clone — must be green.
  • If everything looks right, runs scripts/apply_upstream_patch.sh <proposal-dir> which commits the diff locally with the rationale as the commit message, then offers gh pr create.
  • The human — not the agent — chooses when to push.

Invocation

Normal flow (agent finds a harness-level improvement during work):

[LEARN harness-upstream] apply_patch.py should normalize Windows line endings before SEARCH match

Read the full file on GitHub · 99 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 12d ago First seen · 99 lines · 134 tokens per session scan A 9f3f10a40cb8

Subscribe to this mod's changes

propose-harness-change is a skill published in the GitHub repository Rockielab/rockie-claude (21 stars, last pushed 1mo ago), licensed Apache-2.0. It adds 134 tokens to every session and 1,102 once invoked, about $0.0007 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.

Related

Other skills, from other repositories

auto-run

Autonomous personalized research loop. Use when the user wants to research a topic autonomously, run a research loop, start adaptive research, or use presets like technique-scout or cross-domain. Triggers on: 'auto run', 'research loop', 'autonomous research', 'run research', 'start research', 'adaptive research'.

primeline-ai/claude-adaptive-research · 70 tokens

autoresearch

Orchestrates end-to-end autonomous AI research projects using a two-loop architecture. The inner loop runs rapid experiment iterations with clear optimization targets. The outer loop synthesizes results, identifies patterns, and steers research direction. Routes to domain-specific skills for execution, supports…

Orchestra-Research/AI-Research-SKILLs · 98 tokens

git-commit

A guided Git commit workflow that examines changes and creates a commit message using the Conventional Commits format, a shared style for labeling changes such as features, fixes, tests, or documentation.

adongwanai/learn-workbuddy · 0 tokens

pi-sync

Daily upstream-sync job for the pi Go port — fetch upstream pi, triage every change since the recorded pin, port what's in scope, verify idiomatic + parity via independent reviews, update the ledger, and push. Use for "sync with upstream", "porting job", or as the scheduled daily run.

sky-valley/pi · 66 tokens

pi-triage

Decide whether an upstream pi change needs porting to the Go port. Use when assessing upstream commits/PRs ("should we port X?"), or as the triage stage of /pi-sync. Outputs a WHY/WHAT/SCOPE verdict per change.

sky-valley/pi · 57 tokens

pi-parity-review

Adversarially verify that a ported change is faithful to the original pi implementation (TS source + published npm build). Use after porting upstream pi changes, or standalone on any area of this repo ("is X faithful to pi?").

sky-valley/pi · 54 tokens