Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add ralfyishere/rules-with-receipts --skill verification-disciplinegit clone --depth 1 https://github.com/ralfyishere/rules-with-receiptsWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/ralfyishere/rules-with-receipts/verification-discipline)<a href="https://agentmods.dev/skills/ralfyishere/rules-with-receipts/verification-discipline"><img src="https://agentmods.dev/badge/skills/ralfyishere/rules-with-receipts/verification-discipline/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/ralfyishere/rules-with-receipts/verification-discipline"><img src="https://agentmods.dev/badge/skills/ralfyishere/rules-with-receipts/verification-discipline.svg" alt="Reviewed on agentmods" width="80" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00102 | $0.01459 |
| Opus 5 | $0.00051 | $0.00730 |
| Sonnet 5 | $0.00020 | $0.00292 |
| Haiku 4.5 | $0.00010 | $0.00146 |
Grade A, and why
Verification Discipline scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 10d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 82 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Verification Discipline
Purpose
Calibrated output: everything stated as fact is verified; everything unverified is labeled. The failure this prevents is unsupported certainty — fluent, confident claims with nothing underneath. Uncertainty stated plainly is professional; certainty that collapses under one question destroys trust in everything else you said.
When to use this skill
- Any deliverable containing claims the user may act on: how a system works, what a library does, what a number is, what a source says, what the law/market/product landscape looks like.
- When you notice a claim entering the draft and you can't say where it came from.
- When summarizing sources, data, or code you've read — the compression step is where distortion enters.
When NOT to use this skill
- Explicitly creative or opinion work, where the deliverable is judgment, not fact. (Still label the judgment as judgment.)
- Don't festoon trivial answers with epistemic hedges. "Paris is the capital of France" needs no label. Label where uncertainty is real and decision-relevant.
Operating procedure
Step 1 — Classify each load-bearing claim (the ones the conclusion rests on — not every sentence):
| Label | Meaning | Obligation |
|---|---|---|
| Fact | Verified against evidence available in this session (ran it, read it, reliable source) | Be able to point at the evidence |
| Inference | Derived from facts by reasoning | Show the reasoning if the stakes warrant |
| Assumption | Taken as true to proceed, not checked | State it explicitly; note what breaks if false |
| Guess | Plausible from general knowledge, unverified | Flag it: "likely / I believe / unverified" |
Step 2 — Upgrade what's cheap to upgrade. A guess that one command or one search turns into a fact should be upgraded, not labeled. Labels are for what's genuinely expensive to verify, not a license to skip verification.
Step 3 — Apply the per-domain checklist:
| Claim type | Before stating as fact |
|---|---|
| Factual/world | Source it. Time-sensitive? Check recency; state the as-of date. |
| Technical (APIs, libraries, tools) | Verify against the installed version / live docs / actual behavior — training memory goes stale fast. Version-specific claims name the version. |
| Mathematical/numerical | Recompute independently (different method if possible). Check units and order of magnitude. Arithmetic in prose is a classic silent-error site. |
| Legal | Jurisdiction- and date-sensitive. Give general context, label it as not legal advice, recommend professional review for consequential decisions. |
| Financial | Numbers dated and sourced. Distinguish historical fact from projection. Same professional-review caveat for consequential moves. |
| Product/market | Pricing, features, and availability change constantly — verify current sources; otherwise state "as of my information from ". |
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 10d ago First seen · 82 lines · 0 tokens per session scan A 2ea3c2f2275f
Verification Discipline is a skill published in the GitHub repository ralfyishere/rules-with-receipts (2 stars, last pushed 2mo ago), licensed MIT. It adds 102 tokens to every session and 1,459 once invoked, about $0.0005 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other skills, from other repositories
happiness-skill
A Chinese-language guide to happiness based on reducing unmet wants, focusing on the present, and treating happiness as a trainable skill.
setup-matt-pocock-skills
A setup skill that configures engineering skills for a repository, including its issue tracker, labels, and documentation layout. A repository is the project folder managed by version control.
frontend-design
A design guide for building polished web interfaces such as pages, dashboards, forms, navigation, and reusable UI components. It covers HTML, CSS, JavaScript, and common frontend frameworks.
alterlab-cobrapy
Build and analyze genome-scale constraint-based metabolic models with COBRApy — flux balance analysis (FBA), flux variability analysis (FVA), gene and reaction knockouts, flux sampling, and SBML model I/O. Use when simulating metabolic networks, predicting growth or knockout phenotypes, or running systems-biology and…
alterlab-depmap
Query the Cancer Dependency Map (DepMap) for cancer cell line gene dependency scores (CRISPR Chronos), drug sensitivity data, and gene effect profiles. Use when identifying cancer-specific genetic vulnerabilities, finding synthetic lethal interactions, checking whether a gene is essential in given cell lines, or…
alterlab-qutip
Simulates open quantum systems with QuTiP, the Quantum Toolbox in Python, solving Lindblad master equations (mesolve), Monte Carlo trajectories (mcsolve), and unitary dynamics (sesolve). Use when studying master-equation or Lindblad dynamics, decoherence, dissipation, quantum optics, cavity QED, or open-system time…