Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add skills/gioviat/research-toolkit/math-verificationnpx skills add gioviat/research-toolkit --skill math-verificationgit clone --depth 1 https://github.com/gioviat/research-toolkitWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/gioviat/research-toolkit/math-verification)<a href="https://agentmods.dev/skills/gioviat/research-toolkit/math-verification"><img src="https://agentmods.dev/badge/skills/gioviat/research-toolkit/math-verification.svg" alt="Measured on agentmods" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00081 | $0.00541 |
| Opus 5 | $0.00041 | $0.00270 |
| Sonnet 5 | $0.00016 | $0.00108 |
| Haiku 4.5 | $0.00008 | $0.00054 |
Grade A, and why
math-verification scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 3d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 31 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Math verification
Before checking anything
- Restate every assumption the claim depends on, explicitly, in your own words. If an assumption is implicit in the source material, name it.
- State the claim itself in a form precise enough that "true" or "false" is unambiguous. If the claim as given is ambiguous, say so and verify the most natural reading, flagging the ambiguity.
- Check that all symbols match the project's canonical notation (papers/notation.md). Flag any symbol used inconsistently with that file.
Verification order (cheapest checks first)
- Dimensional / type consistency: do units, shapes, or types match on both sides?
- Boundary and limiting cases: does the claim reduce to a known result at an edge case (n=1, t=0, d→∞, etc.)?
- A small numeric or symbolic example, computed independently (SymPy, a script, or a proof assistant where applicable) — not by re-reading the existing derivation and agreeing with it.
- Full independent re-derivation of the result, without looking at the given derivation step-by-step. Only after finishing, compare against the original to see where they diverge.
Rules
- Never state that a derivation is correct because it "looks standard" or "follows the usual pattern." Redo the steps.
- If a step cannot be verified with available tools (e.g., it requires a proof assistant you don't have, or relies on an unavailable reference), say exactly that: "unverified — requires X." Never round this up to "likely correct."
- If you find an error, state the exact line/step where it occurs and what the correct expression should be, rather than a general "there may be an issue here."
- Distinguish clearly between "I verified this is correct," "I verified this is incorrect," and "I could not verify this."
Output format
For each claim checked, report:
- Claim (restated precisely)
- Assumptions used
- Method of verification (dimensional check / limiting case / independent derivation / symbolic computation)
- Verdict: correct / incorrect / unverified, with the specific reason
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 3d ago First seen · 31 lines · 81 tokens per session scan A 1e98a77adf45
math-verification is a skill published in the GitHub repository gioviat/research-toolkit (2 stars, last pushed 1mo ago), licensed MIT. It adds 81 tokens to every session and 541 once invoked, about $0.0004 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other skills, from other repositories
biopython
Comprehensive molecular biology toolkit. Use for sequence manipulation, file parsing (FASTA/GenBank/PDB), phylogenetics, and programmatic NCBI/PubMed access (Bio.Entrez). Best for batch processing, custom bioinformatics pipelines, BLAST automation. For quick lookups use gget; for multi-service integration use…
exploratory-data-analysis
Perform bounded, local exploratory analysis of explicitly supported scientific files. Use for redacted CSV/TSV/JSON profiles; optional NumPy, HDF5, FASTA/FASTQ, and basic image metadata inspection; missingness/leakage audits; outlier and transformation sensitivity; and rigorous EDA report scaffolds. Other domain…
nature-statistics
Audit, revise, or draft manuscript statistical reporting for Nature / high-impact journal submissions. Use when the user asks to check statistical analysis sections, p values, confidence intervals, sample size, biological versus technical replicates, randomization, blinding, multiple-comparison correction, model…
evaluating-with-leakage-gates
Evaluate an OpenMed de-identification or clinical NER model against the leakage-first release gates G1a through G8, which gate releases on residual PHI leakage rather than on F1. Use when the user wants to run the OpenMed eval harness on a synthetic golden set, decide whether a de-id model is RELEASABLE or…
mapping-to-snomed
Maps clinical concept spans extracted by OpenMed to SNOMED CT concepts through a USER-SUPPLIED terminology server (the user's own Ontoserver, Snowstorm, or UMLS/UTS), never a bundled vocabulary. Use when the user wants to code findings, disorders, procedures, body structures, or substances to SNOMED CT, run an ECL…
mixed-precision
Use FP16/BF16 mixed precision to accelerate training and reduce memory. Use when optimizing GPU performance.