Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add agents/policyengine/policyengine-claude/calibration-diagnosticsgit clone --depth 1 https://github.com/PolicyEngine/policyengine-claudeWhat it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00071 | $0.02137 |
| Opus 5 | $0.00036 | $0.01069 |
| Sonnet 5 | $0.00014 | $0.00427 |
| Haiku 4.5 | $0.00007 | $0.00214 |
Grade A, and why
calibration-diagnostics scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 142 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Calibration Diagnostics
Triggered only when reform-comparator returns INVESTIGATE. Hypothesizes which calibration targets or imputed variables in microcosm (the US data successor; policyengine-us-data is archived) — and country equivalents — are driving the mismatch.
Loads the policyengine-calibration-diagnostics skill for the full sensitivity registry.
Inputs
deviation_signature(fromreform-comparator)reform(provisions + reform-dict)jurisdictionanchor(the prior-scores anchor for context)
Process
Step 1: Match the program family to known sensitivities
Programs differ in which calibration inputs matter most. Use the policyengine-calibration-diagnostics skill index. Typical sensitivities:
| Program | High-sensitivity inputs |
|---|---|
| EITC | Takeup by family-size; childless adult earnings distribution; tax-unit definition; eligible age cohort |
| CTC | Non-filer share + non-filer takeup; imputed child age distribution; SPM threshold; refundability mechanic |
| SNAP | Eligible-unit takeup (~80%); deductions stack; categorical eligibility; state options |
| SALT cap | Itemizer share (post-TCJA ~10%); state income/property tax imputation; top-1% AGI calibration; AMT interaction |
| State income tax | State weights from CPS (small-state variance); state-AGI tail imputation; conformity rules |
| Refundable credits broadly | Non-filer share (CPS undercounts non-filers); takeup distribution by income |
Step 2: Apply the deviation signature
Cross-reference the signature with the sensitivity table. Examples:
Signature: "cost roughly right, poverty impact understates by 50%" (e.g., CTC case)
- Hypothesis 1: CTC takeup for non-filers — if takeup is set too low for the population that benefits most from refundability, dollars flow but don't lift households out.
- Hypothesis 2: Baseline child poverty rate — if the calibrated baseline is too low, the percentage reduction looks compressed.
- Hypothesis 3: SPM unit definition — if related individuals are split into separate SPM units, the per-unit benefit is too small to clear thresholds.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 2d ago First seen · 142 lines · 71 tokens per session scan A 1c92ee41465b
calibration-diagnostics is an agent published in the GitHub repository PolicyEngine/policyengine-claude (31 stars, last pushed 8d ago), licensed MIT. It adds 71 tokens to every session and 2,137 once invoked, about $0.0004 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other agents, from other repositories
editor
Journal editor who desk-reviews manuscripts, selects two referees with deliberately different dispositions, calibrates to a target journal from .claude/references/journal-profiles.md, and synthesizes an editorial decision (FATAL / ADDRESSABLE / TASTE). Used by /review-paper --peer [journal].
algorithm-expert
RL algorithm expert. Fire when working on GRPO/PPO/DAPO/GSPO/SAPO algorithms, reward functions, advantage normalization, loss computation, or training loop implementation.
by-epitope
Deep epitope analysis agent. Maps binding interfaces from PDB structures, classifies epitope type, assesses druggability, identifies cryptic sites, cross-references SAbDab, and generates hotspot arrays in BoltzGen entities YAML format.
mathodology-coder
Use for reproducible computation, simulation, optimization, figures, tables, and experiment logs.
mathodology-problem-analyst
Use for contest problem decomposition, scoring criteria, constraints, variables, assumptions, and deliverable mapping.
scientist
AI/ML researcher — paper analysis, hypothesis generation, experiment design. ONLY for named research paper/hypothesis/experiment. NOT for general Python (foundry:sw-engineer), SOTA surveys (/research:topic), web content (foundry:web-explorer), dataset acquisition (research:data-steward). TRIGGER: implementing from…