verify

A testing skill that tries to find defects in a code change or other work by running it through deliberately difficult tests. TDD, or test-driven development, is a way of designing code around tests; this tool instead focuses on checking completed work.

In plain words
What is it for?
Use it to define acceptance conditions, run happy-path and edge-case tests, test hostile or ambiguous inputs, check previous behavior, and record failures as defects.
Why use it?
It replaces vague self-assessment with tests for normal use, unusual inputs, misuse, unclear requirements, and regressions. It also looks for tests that only appear to pass without checking the real code.

Skill for Claude CodeCodex

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add skills/megaprompting/torque-loop/verify
Any agent
npx skills add Megaprompting/torque-loop --skill verify
Clone the repo
git clone --depth 1 https://github.com/Megaprompting/torque-loop

Made for: Claude Code, Codex.

Per session 86 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 891 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00086 $0.00891
Opus 5 $0.00043 $0.00445
Sonnet 5 $0.00017 $0.00178
Haiku 4.5 $0.00009 $0.00089

Measured 2d ago against content hash 2d2b7d214568, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

verify scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

skills/verify/SKILL.md · 89 lines

How it starts

The opening of the file, as written. The whole thing — 89 lines — stays where its author put it; the contents beside it link to each section on GitHub.

/ratchet:verify — the embarrassment harness

Do not grade the artifact. Asking a model "how good is this?" gets you self-praise. This command instead builds a harness designed to embarrass the artifact, then runs it through. Validation, not vibes.

Step 0 — Load state and target

ratchet status
ratchet snapshot repo

Identify the change or artifact under test (usually the last artifact, or the current diff).

Build the harness

Construct all seven, then run the artifact against them:

  1. Acceptance criteria — the conditions that define "correct".
  2. Happy-path test — the intended use, working.
  3. Edge-case tests — boundaries: empty, huge, zero, negative, unicode, concurrent.
  4. Abuse / misuse tests — hostile or wrong input a real user will eventually send.
  5. Ambiguity tests — under-specified inputs where behavior is undefined.
  6. Regression tests — what previously worked and must still work.
  7. Fake-progress red flags — the tells that this is theater: passes only on the author's example, swallows errors, asserts nothing, tests the mock not the code.

Run the artifact through the harness for real. Where a runtime exists, execute it via Bash — do not simulate a pass in your head. Report the actual result.

Output contract

HARNESS: <the tests, briefly>
RESULTS:
- PASS: <checks that held>
- FAIL: <check> — <what happened> — severity: critical/high/medium/low
RED FLAGS: <any fake-progress tells found, or "none">
USABLE DESPITE FAILURES? <yes/no + one-line reason>
REQUIRED PATCHES: <smallest delta per failure>

Serialize

Run the harness BOUND to the artifact it is about, so the evidence names the exact bytes it was gathered against:

ratchet-evolve verify <target> --artifact <id> --test "<the command>" --json

That prints verifiedHash and verifiedRev. Carry both into the log append — it recomputes them and refuses if the file moved (file changed after verification) or if the artifact was revised after the harness ran (artifact revised after verification). The hash alone cannot catch a metadata-only revision: retitle an artifact and the file is untouched while the revision moves, so rev-1 evidence would be stamped onto rev 2.

Read the full file on GitHub · 89 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 2d ago First seen · 89 lines · 86 tokens per session scan A 2d2b7d214568

Subscribe to this mod's changes

verify is a skill published in the GitHub repository Megaprompting/torque-loop (5 stars, last pushed 1mo ago), licensed MIT. It adds 86 tokens to every session and 891 once invoked, about $0.0004 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.

Related

Other skills, from other repositories

new-plugin

Factory line for adding a new HAR verification plugin (like playwright or rocketsim) for any framework — research the framework docs, build the template under src/templates/plugins/, register it everywhere, validate on a real repository, and open a PR. Use when asked to add/create a plugin, plugin template, or…

os-factory/har · 89 tokens

factory-line

Factory line for executing one station of a declared multi-station program — read the installed line bundle (har line status), plan parallel work into isolated HAR slots, run the cumulative gate with har line gate, and hand off for human review. Use when asked to "run a factory line", "run the next station", "execute…

os-factory/har · 101 tokens

v1-milestone

Factory line for executing one milestone of the HAR v1.0.0 refactor (epic os-factory/har#225) — plan the wave of parallel subagents, implement each issue in its own HAR slot, ship stacked PRs, run the fixture-e2e milestone gate, and hand off for review. Use when asked to "run the next v1 milestone", "work on v1.0.0"…

os-factory/har · 110 tokens

ctx

Codebase intelligence and evidence-driven governance with the indexed ctx CLI. Use when exploring an unfamiliar repository, locating symbols or callers, checking for existing implementations, estimating change impact, enforcing architecture rules, scoring a branch, finding hotspots or duplication, or analyzing…

agentis-tools/ctx · 60 tokens

ctx

Codebase intelligence and evidence-driven governance with the indexed ctx CLI. Use when exploring an unfamiliar repository, locating symbols or callers, checking for existing implementations, estimating change impact, enforcing architecture rules, scoring a branch, finding hotspots or duplication, or analyzing…

agentis-tools/ctx · 60 tokens

golden-rss

Use when testing the rss golden build.

yusufkaraaslan/Skill_Seekers · 12 tokens