Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add chipi/agentic-ai-homelab --skill replay-proven-fixgit clone --depth 1 https://github.com/chipi/agentic-ai-homelabWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/chipi/agentic-ai-homelab/replay-proven-fix)<a href="https://agentmods.dev/skills/chipi/agentic-ai-homelab/replay-proven-fix"><img src="https://agentmods.dev/badge/skills/chipi/agentic-ai-homelab/replay-proven-fix/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/chipi/agentic-ai-homelab/replay-proven-fix"><img src="https://agentmods.dev/badge/skills/chipi/agentic-ai-homelab/replay-proven-fix.svg" alt="Reviewed on agentmods" width="80" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00123 | $0.01307 |
| Opus 5.5 | $0.00049 | $0.00523 |
| Sonnet 5.5 | $0.00025 | $0.00261 |
| Haiku 4.5 | $0.00012 | $0.00131 |
Grade A, and why
replay-proven-fix scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 6d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 61 lines — stays where its author put it; the contents beside it link to each section on GitHub.
replay-proven-fix
The operator's standing requirement for a behaviour change: a replayed test case with the input, today's output, the problem, and the output after the fix. This skill is that loop, plus the traps found while running it (homelab, 2026-09-30 → 10-02).
repro-first covers "one bug, one failing test". This skill is for changes to
decisions over many inputs, where the risk is what else moves.
The loop
- Find it in real data, with counts. "10,764 of 11,700 rows repeat", not "lots of noise". Query the real system; never estimate.
- Freeze the evidence into a committed corpus: snapshots, payloads, and the external state the decision reads (e.g. GitHub issue states). Scrub secrets and user ids. History moves on; the corpus keeps the replay reproducible.
- Write the fix. Prefer a deterministic rule over a prompt or a threshold tweak.
- Write the replay: old code vs new code over the corpus.
- The old code comes from git at the commit before the fix. Never re-implement "how it used to work" by hand.
- Python:
scripts/new_replay.pyscaffolds this. - Any language:
eval "$(scripts/baseline_worktree.sh <BASE_REV>)"checks the old code out into$BASELINE_DIR; run the replay there and in the working tree.
- Python:
- First check the baseline reproduces the recorded history. If it can't, the replay isn't faithful, and nothing it says about the fix counts.
- Write explicit pass criteria and exit non-zero on failure. Include the invariant ("nothing outside the bug changes"), not only the headline ("the 6 alerts became 1").
- Stub external writes; serve reads from the corpus. A replay never sends anything.
- The old code comes from git at the commit before the fix. Never re-implement "how it used to work" by hand.
- Mutation-check the replay. Break the fix in the plausible ways it could be wrong, and confirm the replay fails each time:
python3 ~/.claude/skills/replay-proven-fix/scripts/mutate.py \ --file path/to/module.py --find '<exact fix text>' --replace '<broken version>' \ --precheck 'python3 -m py_compile path/to/module.py' -- <replay command>- Exit codes: 0 = killed (good), 1 = survived (the replay proves nothing; add the missing check), 2 = setup or invalid mutant. The file is always restored, verified by sha256.
- Read WHY it failed. A mutant can fail for the wrong reason, e.g. a syntax error your edit introduced;
--precheckcatches that.
- Read the changed cases, not just the totals, and judge them against explicit links, not titles or names. For example:
- a duplicate link (GitHub GraphQL
ClosedEvent.duplicateOf); - the stored signal or row behind the case;
- an id that may have been reused.
- a duplicate link (GitHub GraphQL
- Gate it: add the replay to the project's deploy or CI gate list, plus focused unit tests. Run every gate.
- Commit with the numbers: input, before, after, the mutants and their results. If a later finding contradicts what an earlier commit message claimed, say so in the next commit message. Push and deploy only with the operator's go.
- Verify live. "The replay passed" is not "it works in production". After deploying, exercise the real path: one real input pushed through by hand, or the real decision run read-only against live state with writes intercepted. If a deploy step loaded its own script before the pull (bash reads it first), run any newly added gate by hand once.
- Document: update the project's design guide (why) and runbook (how), and list what is still not covered.
What ships with it
3 files beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 6d ago First seen · 61 lines · 123 tokens per session scan A 2768b921ccba
replay-proven-fix is a skill published in the GitHub repository chipi/agentic-ai-homelab (1 stars, last pushed yesterday), licensed MIT. It adds 123 tokens to every session and 1,307 once invoked, about $0.0005 per session on Opus 5.5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-10-03.
Other skills, from other repositories
ios-simulator
Verify and debug native, React Native, Expo, or Flutter apps on an iOS Simulator with agent-device. Use when an agent needs to launch an app, inspect its live UI, tap, type, scroll, validate a code change, collect failure evidence, or reproduce a workflow on an iPhone or iPad Simulator.
ai-discover
Read-only discovery step of the auto-improvement loop. One bounded agent turn per cycle reads the focus files (Read/Grep/Glob only, no shell, no subagents) and replies with a JSON array of testable surfaces {file, line, symbol, rule, message, hypothesis}; the spine authors fixes and verifies them. Discovery only — no…
behavior-contract
Bug condition/postcondition formalization as testable Behavior Contracts. Defines invariants that must be preserved across fixes.
quality-hooks
Language-specific auto-lint/format/typecheck pipeline. Supports Python (ruff+pyright), TypeScript (prettier+eslint+tsc), Go (gofmt+golangci-lint). Auto-fix and convergence loops.
moai-workflow-loop
Ralph Engine - Automated feedback loop with LSP diagnostics and AST-grep integration for continuous code quality improvement. Use when implementing error-driven development, automated fixing, or continuous quality validation workflows.
moai-workflow-testing
Use when writing tests, measuring coverage, or running characterization, performance, or PR-review QA. Comprehensive specialist combining DDD testing, characterization tests, performance profiling, and TRUST 5 quality-assurance validation.