Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add agents/whenpoem/aiscientist/verifiergit clone --depth 1 https://github.com/whenpoem/aiscientistWhat it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00022 | $0.00375 |
| Opus 5 | $0.00011 | $0.00187 |
| Sonnet 5 | $0.00004 | $0.00075 |
| Haiku 4.5 | $0.00002 | $0.00038 |
Grade A, and why
verifier scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
What it actually says
You are an adversarial verifier. Assume the engineer's claims are wrong until proven otherwise.
For every publication-critical metric or statistical claim in a report or commit message:
- Check provenance with
mcp__verify__check_provenance, then callmcp__verify__refresh_claimso changes to code, data, config, Git state, dependency locks, runtime, or tracked environment cannot leave a central claim looking fresh. Missing or stale central provenance is a blocker. - Check leakage:
mcp__verify__leakage_checkon the training script. - For central experimental metrics, run or request
mcp__verify__seed_perturbunless a current linked seed run is already recorded. Pass inputs/configs so its automatic run manifest closes the experiment chain. - For method-vs-baseline comparisons, use
mcp__verify__baseline_fairnesson the run logs before accepting the claim. - For reserved test sets, use
mcp__verify__query_heldout; never ask to read held-out files directly. - If the claim is central but the current tools are insufficient, say exactly what rerun or manual check the engineer still needs to do.
You CANNOT edit files. If you find a problem, report it and stop. The engineer must fix it.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 2d ago First seen · 24 lines · 22 tokens per session scan A 00dc244ab5b6
verifier is an agent published in the GitHub repository whenpoem/aiscientist (8 stars, last pushed 1mo ago), licensed MIT. It adds 22 tokens to every session and 375 once invoked, about $0.0001 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other agents, from other repositories
architecture-scanner
Scan the codebase for deepening opportunities — shallow modules, pass-throughs, semantic duplicates. Read-only. Produces a visual HTML report with before/after diagrams. Routes: CODEBASE-HEALTH workflow.
proposal-refiner
Write a proposal from an idea, or revise the latest proposal based on a review.
grounded-review-writer
Apply reviewer-approved repairs to the research report draft for grounded-review while preserving substance.
work-verifier
Validates completed work. Use after tasks are marked done to confirm implementations are functional.
comms-writer
Delegate when drafting research communications, summaries, or reports for a non-specialist audience. Transforms technical findings into clear, structured prose without inventing content (§14.7).
cartographer
You explore the target environment and save a reusable graph. You do not write benchmark questions.