Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add ralfyishere/rules-with-receipts --skill effort-calibrationgit clone --depth 1 https://github.com/ralfyishere/rules-with-receiptsWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/ralfyishere/rules-with-receipts/effort-calibration)<a href="https://agentmods.dev/skills/ralfyishere/rules-with-receipts/effort-calibration"><img src="https://agentmods.dev/badge/skills/ralfyishere/rules-with-receipts/effort-calibration/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/ralfyishere/rules-with-receipts/effort-calibration"><img src="https://agentmods.dev/badge/skills/ralfyishere/rules-with-receipts/effort-calibration.svg" alt="Reviewed on agentmods" width="80" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00116 | $0.01264 |
| Opus 5 | $0.00058 | $0.00632 |
| Sonnet 5 | $0.00023 | $0.00253 |
| Haiku 4.5 | $0.00012 | $0.00126 |
Grade A, and why
Effort Calibration scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 10d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 77 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Effort Calibration
Purpose
Spend effort where it buys outcome quality, and nowhere else. Under-effort produces confident wrong answers on hard problems; over-effort produces ceremony, latency, and bloated output on easy ones. Both are calibration failures. The tier decision takes seconds and governs everything downstream.
When to use this skill
- At the start of every task (the decision is cheap; skipping it is not).
- Mid-task, when stakes shift: an action turns out to be irreversible, an assumption breaks, scope grows, or the user signals urgency or importance.
- When torn between "just answer" and "go verify" — that tension is the trigger.
When NOT to use this skill
- Don't loop on it. Pick a tier, state nothing (Low/Medium) or one line (High/Critical), and move. Re-calibrate only on new information.
Operating procedure
Step 1 — Score two axes:
- Complexity: How many steps? How familiar? How much unknown?
- Stakes: What happens if this is wrong? Reversible or not? Who sees it?
Step 2 — Pick the tier (stakes win ties):
| Tier | Typical signals | Behavior |
|---|---|---|
| Low | Factual question, one-step edit, reversible, user wants speed | Answer directly. No plan. Verify only if a claim is load-bearing and cheap to check. Short output. |
| Medium | Multi-step but familiar; moderate blast radius; standard requests | Micro-plan (one paragraph). Verify key claims against live state. One quick self-review pass before finalizing. |
| High | Complex or unfamiliar; multi-file/multi-part; wrong answer costs real rework; user says "important", "production", "customer-facing" | Full plan-gate. live-state-truth for all state claims. adversarial-verify before presenting. Label remaining uncertainty. |
| Critical | Irreversible actions (deletes, sends, deploys, payments); legal/financial/medical territory; public-facing artifacts | Everything in High, plus: explicit assumption list, confirm with the user before the irreversible step, state confidence and what wasn't verified. Slow is correct here. |
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 10d ago First seen · 77 lines · 0 tokens per session scan A 17db72cd9aba
Effort Calibration is a skill published in the GitHub repository ralfyishere/rules-with-receipts (2 stars, last pushed 2mo ago), licensed MIT. It adds 116 tokens to every session and 1,264 once invoked, about $0.0006 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other skills, from other repositories
happiness-skill
A Chinese-language guide to happiness based on reducing unmet wants, focusing on the present, and treating happiness as a trainable skill.
setup-matt-pocock-skills
A setup skill that configures engineering skills for a repository, including its issue tracker, labels, and documentation layout. A repository is the project folder managed by version control.
frontend-design
A design guide for building polished web interfaces such as pages, dashboards, forms, navigation, and reusable UI components. It covers HTML, CSS, JavaScript, and common frontend frameworks.
alterlab-cobrapy
Build and analyze genome-scale constraint-based metabolic models with COBRApy — flux balance analysis (FBA), flux variability analysis (FVA), gene and reaction knockouts, flux sampling, and SBML model I/O. Use when simulating metabolic networks, predicting growth or knockout phenotypes, or running systems-biology and…
alterlab-depmap
Query the Cancer Dependency Map (DepMap) for cancer cell line gene dependency scores (CRISPR Chronos), drug sensitivity data, and gene effect profiles. Use when identifying cancer-specific genetic vulnerabilities, finding synthetic lethal interactions, checking whether a gene is essential in given cell lines, or…
alterlab-qutip
Simulates open quantum systems with QuTiP, the Quantum Toolbox in Python, solving Lindblad master equations (mesolve), Monte Carlo trajectories (mcsolve), and unitary dynamics (sesolve). Use when studying master-equation or Lindblad dynamics, decoherence, dissipation, quantum optics, cavity QED, or open-system time…