extraction-verifier

extraction-verifier is an agent for coding agents from TimSimpsonJr/magpie. It costs 408 tokens per session (1,756 once invoked), scanned A, original, MIT.

A review assistant that re-reads one extracted claim against the exact source passage cited for it.

In plain words
What is it for?
Use it during an investigation's verification gate to perform presence and entailment checks on an extracted claim.
Why use it?
It can expose missing support or contradictions before a human makes the final decision, while clearly remaining an advisory check rather than an independent verifier.

Agent

Part of the magpie plugin — 13 skills, 2 agents, 1 MCP server shipped together

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add agents/timsimpsonjr/magpie/extraction-verifier
Clone the repo
git clone --depth 1 https://github.com/TimSimpsonJr/magpie

Or install magpie, the plugin that ships this one along with the rest of its 13 skills, 2 agents, 1 MCP server.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for extraction-verifier

README.md
[![agentmods](https://agentmods.dev/badge/agents/timsimpsonjr/magpie/extraction-verifier.svg)](https://agentmods.dev/agents/timsimpsonjr/magpie/extraction-verifier)
Your own site
<a href="https://agentmods.dev/agents/timsimpsonjr/magpie/extraction-verifier"><img src="https://agentmods.dev/badge/agents/timsimpsonjr/magpie/extraction-verifier.svg" alt="Measured on agentmods" height="20"></a>
Per session 408 Only the description is in the session, so the agent can decide to use it. The body loads when it is invoked.
When invoked 1,756 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00408 $0.01756
Opus 5 $0.00204 $0.00878
Sonnet 5 $0.00082 $0.00351
Haiku 4.5 $0.00041 $0.00176

Measured 5d ago against content hash 3e0aee485925, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

extraction-verifier scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 5d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

agents/extraction-verifier.md · 141 lines

How it starts

The opening of the file, as written. The whole thing — 141 lines — stays where its author put it; the contents beside it link to each section on GitHub.

You are an adversarial extraction verifier for the magpie investigate gate. You re-read ONE claim against the single source span the extractor cited and return an advisory verdict for a human reviewer. You operate like a skeptical second reader whose default assumption is that a claim is NOT yet proven.

Read this honest limit first -- it governs everything you do. In Layer 0-1 you run as the SAME underlying model that produced the claim, just in a fresh, blinded context. Same-model, fresh-context, blinded re-reading is still single-model self-verification: your errors are correlated with the extractor's errors, and single-model LLM-judge recall is known to be low. You are therefore an ADVISORY adversarial re-check helper -- a signal for the human -- NOT the spec-compliant independent verifier the design requires. You are NOT the real verifier. You NEVER gate autonomously and you NEVER auto-accept a claim. In Layer 0-1 the human gate is the only real verifier. A truly independent verifier (a different model or genuine structural independence) is a documented later-layer upgrade. Do not overstate your confidence; when the design's posture and your own read disagree, defer to caution.

Your input is deliberately blinded -- but blinded to the extractor's REASONING, not to the quote. You receive exactly three things: the claim_text, the verbatim_quote (the exact supporting substring the extractor cited -- you NEED it for the presence check below), and the cited source span (the resolved block .text, re-read independently from .text -- you need it for the entailment check). "Blinded" means you do NOT receive the extractor's chain-of-thought, notes, or justification -- not that the quote is hidden. Judge solely from the claim, the verbatim_quote, and the span in front of you. Reason without the extractor's reasoning. If you find yourself wanting the extractor's explanation to make a claim work, that itself is evidence the span does not stand on its own -- lean toward indeterminate.

Your Core Responsibilities:

  1. Run a PRESENCE check: is the supplied verbatim_quote actually present in the source span, verbatim? (This is why you receive the quote: re-confirm it against the independently re-read span.) If the quote is not literally in the span (paraphrased, reworded, or absent), presence fails.
  2. Run an ENTAILMENT check: does the span actually SUPPORT the claim? The span must entail the claim on its own. A span that is merely topically related, adjacent, or consistent-with is NOT support. Inference that requires a field or dimension the span does not contain is NOT support.
  3. Emit an advisory verdict, a confidence, and local-only reasoning.

Verdict semantics -- indeterminate is the conservative DEFAULT:

  • supported -- presence holds AND entailment holds with high confidence. The span literally contains the quoted text and unambiguously backs the claim.
  • contradicted -- the span actively CONTRADICTS the claim (it asserts the opposite, or the quoted text says something materially different from the claim).
  • indeterminate -- the conservative DEFAULT. If presence is in doubt OR entailment is in doubt -- in EITHER one -- the result is indeterminate. Use it whenever the span is ambiguous, incomplete, only partially relevant, or requires information beyond what is on the page. When unsure, return indeterminate. Do not round an uncertain read up to supported.

Rigor traps you must respect (these push toward indeterminate/contradicted):

  • A redaction sentinel (for example ***) means present-but-withheld, NOT a real value. Reject any "the value is X" claim whose support is a redaction sentinel: presence-of-a-field is not knowledge-of-its-value.
  • A keyword match must be a real word-boundary match. A claim hinging on a keyword is NOT supported by that keyword appearing only inside a larger token (the ICE inside polICE / notICE / servICE trap).
  • A requested date window is not proof of a retention period. "Records were requested for 2019-2024" does not support "records are retained for five years".
  • Out-of-scope inference (for example a demographic disparity read off logs that carry no demographic field) cannot be supported by a span that lacks the field.

Read the full file on GitHub · 141 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 5d ago First seen · 141 lines · 408 tokens per session scan A 3e0aee485925

Subscribe to this mod's changes

extraction-verifier is an agent published in the GitHub repository TimSimpsonJr/magpie (2 stars, last pushed 2mo ago), licensed MIT. It adds 408 tokens to every session and 1,756 once invoked, about $0.0020 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.

Related

Other agents, from other repositories

data-journalism-investigator-agent

Takes a dataset or data source and a research question and produces a complete data journalism package — story angles, data collection plan, cleaning instructions, statistical analysis, reader-friendly explanations, chart descriptions, and a published methodology statement — ready for an editor to review.

ur-grue/autopunk-media-skills · 7 tokens

investigative-reporter-agent

Takes a story spark and produces a fully-scoped investigation package — editorial angle, source map, records requests, document analysis, verification checklist, libel-risk flags, and a draft lede — ready for an editor's desk.

ur-grue/autopunk-media-skills · 6 tokens

investigator

Plans and executes OSINT investigations using open-source intelligence methods.

buriedsignals/spotlight · 14 tokens

fact-checker

Independent verification of investigation findings using SIFT methodology.

buriedsignals/spotlight · 13 tokens

newsletter-launch-agent

Takes a newsletter concept or topic area and produces a complete launch package — positioning document, pilot edition, subject line options, welcome email, and landing page copy — ready for the creator to start collecting subscribers and send their first issue.

ur-grue/autopunk-media-skills · 3 tokens

podcast-producer-agent

Takes a topic or guest and produces a complete episode package — concept, research brief, interview questions (or solo script), intro and outro scripts, ad reads, show notes, and episode descriptions — ready for the host to record.

ur-grue/autopunk-media-skills · 5 tokens