Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add skills/megaprompting/torque-loop/verifynpx skills add Megaprompting/torque-loop --skill verifygit clone --depth 1 https://github.com/Megaprompting/torque-loopWhat it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00086 | $0.00891 |
| Opus 5 | $0.00043 | $0.00445 |
| Sonnet 5 | $0.00017 | $0.00178 |
| Haiku 4.5 | $0.00009 | $0.00089 |
Grade A, and why
verify scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 89 lines — stays where its author put it; the contents beside it link to each section on GitHub.
/ratchet:verify — the embarrassment harness
Do not grade the artifact. Asking a model "how good is this?" gets you self-praise. This command instead builds a harness designed to embarrass the artifact, then runs it through. Validation, not vibes.
Step 0 — Load state and target
ratchet status
ratchet snapshot repo
Identify the change or artifact under test (usually the last artifact, or the current diff).
Build the harness
Construct all seven, then run the artifact against them:
- Acceptance criteria — the conditions that define "correct".
- Happy-path test — the intended use, working.
- Edge-case tests — boundaries: empty, huge, zero, negative, unicode, concurrent.
- Abuse / misuse tests — hostile or wrong input a real user will eventually send.
- Ambiguity tests — under-specified inputs where behavior is undefined.
- Regression tests — what previously worked and must still work.
- Fake-progress red flags — the tells that this is theater: passes only on the author's example, swallows errors, asserts nothing, tests the mock not the code.
Run the artifact through the harness for real. Where a runtime exists, execute it via Bash — do not simulate a pass in your head. Report the actual result.
Output contract
HARNESS: <the tests, briefly>
RESULTS:
- PASS: <checks that held>
- FAIL: <check> — <what happened> — severity: critical/high/medium/low
RED FLAGS: <any fake-progress tells found, or "none">
USABLE DESPITE FAILURES? <yes/no + one-line reason>
REQUIRED PATCHES: <smallest delta per failure>
Serialize
Run the harness BOUND to the artifact it is about, so the evidence names the exact bytes it was gathered against:
ratchet-evolve verify <target> --artifact <id> --test "<the command>" --json
That prints verifiedHash and verifiedRev. Carry both into the log append — it
recomputes them and refuses if the file moved (file changed after verification) or if
the artifact was revised after the harness ran (artifact revised after verification).
The hash alone cannot catch a metadata-only revision: retitle an artifact and the file
is untouched while the revision moves, so rev-1 evidence would be stamped onto rev 2.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 2d ago First seen · 89 lines · 86 tokens per session scan A 2d2b7d214568
verify is a skill published in the GitHub repository Megaprompting/torque-loop (5 stars, last pushed 1mo ago), licensed MIT. It adds 86 tokens to every session and 891 once invoked, about $0.0004 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other skills, from other repositories
new-plugin
Factory line for adding a new HAR verification plugin (like playwright or rocketsim) for any framework — research the framework docs, build the template under src/templates/plugins/, register it everywhere, validate on a real repository, and open a PR. Use when asked to add/create a plugin, plugin template, or…
factory-line
Factory line for executing one station of a declared multi-station program — read the installed line bundle (har line status), plan parallel work into isolated HAR slots, run the cumulative gate with har line gate, and hand off for human review. Use when asked to "run a factory line", "run the next station", "execute…
v1-milestone
Factory line for executing one milestone of the HAR v1.0.0 refactor (epic os-factory/har#225) — plan the wave of parallel subagents, implement each issue in its own HAR slot, ship stacked PRs, run the fixture-e2e milestone gate, and hand off for review. Use when asked to "run the next v1 milestone", "work on v1.0.0"…
ctx
Codebase intelligence and evidence-driven governance with the indexed ctx CLI. Use when exploring an unfamiliar repository, locating symbols or callers, checking for existing implementations, estimating change impact, enforcing architecture rules, scoring a branch, finding hotspots or duplication, or analyzing…
ctx
Codebase intelligence and evidence-driven governance with the indexed ctx CLI. Use when exploring an unfamiliar repository, locating symbols or callers, checking for existing implementations, estimating change impact, enforcing architecture rules, scoring a branch, finding hotspots or duplication, or analyzing…
golden-rss
Use when testing the rss golden build.