policyengine-calibration-diagnostics

policyengine-calibration-diagnostics is a skill for Claude Code, Codex from PolicyEngine/policyengine-claude. It costs 188 tokens per session (2,653 once invoked), scanned A, original, MIT.

A diagnostic guide for explaining why PolicyEngine microsimulation results differ from expected values. Microsimulation estimates policy effects across many representative households; calibration adjusts the underlying data to match known totals.

In plain words
What is it for?
Use it when investigating an unexpected policy result, reviewing microsimulation changes, checking calibration targets, or understanding volatile state-level estimates.
Why use it?
It ranks likely causes of a mismatch using calibration diagnostics instead of relying only on intuition. It connects differences to policy programs, data assumptions, and measured calibration errors.

Skill for Claude CodeCodex

Part of the analysis-tools plugin — 11 skills shipped together , and of complete

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add skills/policyengine/policyengine-claude/policyengine-calibration-diagnostics
Any agent
npx skills add PolicyEngine/policyengine-claude --skill policyengine-calibration-diagnostics
Clone the repo
git clone --depth 1 https://github.com/PolicyEngine/policyengine-claude

Made for: Claude Code, Codex.

Or install analysis-tools, the plugin that ships this one along with the rest of its 11 skills.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for policyengine-calibration-diagnostics

README.md
[![agentmods](https://agentmods.dev/badge/skills/policyengine/policyengine-claude/policyengine-calibration-diagnostics.svg)](https://agentmods.dev/skills/policyengine/policyengine-claude/policyengine-calibration-diagnostics)
Your own site
<a href="https://agentmods.dev/skills/policyengine/policyengine-claude/policyengine-calibration-diagnostics"><img src="https://agentmods.dev/badge/skills/policyengine/policyengine-claude/policyengine-calibration-diagnostics.svg" alt="Measured on agentmods" height="20"></a>
Per session 188 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 2,653 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5.1 $0.00188 $0.02653
Opus 5 $0.00094 $0.01326
Sonnet 5 $0.00038 $0.00531
Haiku 4.5 $0.00019 $0.00265

Measured 5d ago against content hash a55216268497, method: parsed. Prices are Anthropic first-party input rates as of 2026-09-05, from the pricing page.

Security

Grade A, and why

policyengine-calibration-diagnostics scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 5d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

skills/policyengine-calibration-diagnostics/SKILL.md · 177 lines

How it starts

The opening of the file, as written. The whole thing — 177 lines — stays where its author put it; the contents beside it link to each section on GitHub.

PolicyEngine calibration diagnostics

Converts tribal "I'd check the takeup rate first" knowledge into a structured sensitivity registry, and pairs it with the live per-target calibration API so hypotheses are ranked against real relative_error numbers rather than assumptions. When an /analyze-policy comparison returns INVESTIGATE, this skill supplies the ranked candidate causes.

The calibrated microdata is now a Microcosm build (see the policyengine-data skill for how targets, weights, and L0 sparsity work). Calibration targets live in the Microcosm build's target set — not in a hand-maintained loss file — and their fit is queryable per release from the dashboard API below.

When to use

  • Stage 5.6 / Stage 6 of /analyze-policy — invoked by the calibration-diagnostics agent.
  • Code review of microsim PRs where the headline number differs from priors.
  • Designing or auditing a Microcosm calibration target.
  • Debugging why a state-level run looks volatile.

Top-level architecture

PolicyEngine microsim results depend on three layers:

  1. Country model logic (policyengine-us, policyengine-uk, policyengine-canada) — formulas, parameters.
  2. Calibrated microdata (Microcosm) — survey weights + imputations matched to administrative targets.
  3. Behavioral assumptions (takeup rates, labor-supply elasticities) — usually parameters but easy to overlook.

A magnitude mismatch is almost always rooted in layer 2 or 3, not layer 1 (layer-1 mismatches show up as outright simulation errors, not magnitude drift).

Live calibration API (check this first)

Per-target fit for the current Microcosm release, no auth, reads the release from Hugging Face:

BASE = https://calibration-diagnostics.vercel.app/calibration/dashboard/api/populace
GET {BASE}/target-diagnostics?source=<source>     # every target for a source, with relative_error
GET {BASE}/target-investigation?target=<id>        # full investigation packet for one target
GET {BASE}/releases                                # release ids for pinning

Read the full file on GitHub · 177 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 5d ago First seen · 177 lines · 188 tokens per session scan A a55216268497

Subscribe to this mod's changes

policyengine-calibration-diagnostics is a skill published in the GitHub repository PolicyEngine/policyengine-claude (32 stars, last pushed 3d ago), licensed MIT. It adds 188 tokens to every session and 2,653 once invoked, about $0.0009 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.

Related

Other skills, from other repositories

instrument-data-to-allotrope

Convert laboratory instrument output files (PDF, CSV, Excel, TXT) to Allotrope Simple Model (ASM) JSON format or flattened 2D CSV. Use this skill when scientists need to standardize instrument data for LIMS systems, data lakes, or downstream analysis. Supports auto-detection of instrument types. Outputs include full…

anthropics/knowledge-work-plugins · 123 tokens

exploratory-data-analysis

Perform bounded, local exploratory analysis of explicitly supported scientific files. Use for redacted CSV/TSV/JSON profiles; optional NumPy, HDF5, FASTA/FASTQ, and basic image metadata inspection; missingness/leakage audits; outlier and transformation sensitivity; and rigorous EDA report scaffolds. Other domain…

K-Dense-AI/scientific-agent-skills · 83 tokens

matlab

Build, review, migrate, and safely plan MATLAB or GNU Octave numerical workflows, including arrays, tabular/time data, tests, projects, graphics, MAT files, and explicit Python interoperability.

K-Dense-AI/scientific-agent-skills · 42 tokens

phylogenetics

Build and analyze phylogenetic trees using MAFFT (multiple alignment), IQ-TREE 2 (maximum likelihood), and FastTree (fast NJ/ML). Visualize with ETE3 or FigTree. For evolutionary analysis, microbial genomics, viral phylodynamics, protein family analysis, and molecular clock studies.

K-Dense-AI/scientific-agent-skills · 68 tokens

research-engineer

An uncompromising Academic Research Engineer. Operates with absolute scientific rigor, objective criticism, and zero flair. Focuses on theoretical correctness, formal verification, and optimal implementation across any required technology.

davila7/claude-code-templates · 43 tokens

mapping-to-snomed

Maps clinical concept spans extracted by OpenMed to SNOMED CT concepts through a USER-SUPPLIED terminology server (the user's own Ontoserver, Snowstorm, or UMLS/UTS), never a bundled vocabulary. Use when the user wants to code findings, disorders, procedures, body structures, or substances to SNOMED CT, run an ECL…

maziyarpanahi/openmed · 205 tokens