shadow-verify

A verification workflow that independently challenges findings from investigations, reviews, audits, refactors, and similar work. It checks claims against concrete references such as files, lines, commits, or API routes.

In plain words
What is it for?
Use it after high-stakes analysis or when an agent makes a strong claim, to test whether each finding is supported by a specific piece of evidence.
Why use it?
It helps catch confident but unsupported conclusions before they influence decisions or code changes. Claims without evidence are separated instead of being treated as verified.

Skill for Claude CodeCodex

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add skills/griffinwork40/agent-framework/shadow-verify
Any agent
npx skills add griffinwork40/agent-framework --skill shadow-verify
Clone the repo
git clone --depth 1 https://github.com/griffinwork40/agent-framework

Made for: Claude Code, Codex.

Per session 136 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 2,693 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 1 finding. Scan, not verified.
Origin 100% copy Near-identical to another mod in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00136 $0.02693
Opus 5 $0.00068 $0.01347
Sonnet 5 $0.00027 $0.00539
Haiku 4.5 $0.00014 $0.00269

Measured 2d ago against content hash 9f7db5a0612c, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

shadow-verify scanned grade A with 1 finding against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Makes network callslowCapability

Not a fault in itself. Listed so you know the mod talks to something, and to what.

- **Default to `subagent_type: "research-agent"` (mechanically locked to Read/Grep/Glob/WebFetch/WebSearch — cannot Edit/commit/push).** If the claim requires Bash to verify (running a failing test, `gh pr view`, `git lo
Origin

This is a copy

100% identical to shadow-verify — 0 lines differ, which has more behind it and is treated as the original. This page carries a canonical link to it rather than competing with it.

skills/shadow-verify/SKILL.md · 96 lines

How it starts

The opening of the file, as written. The whole thing — 96 lines — stays where its author put it; the contents beside it link to each section on GitHub.

Sub-agent contract

/contract

When a sub-agent (or wave) returns investigation findings, code-review conclusions, audit claims, refactor plans, gap-analysis results, or counts that will drive user decisions or file changes, do NOT surface the report. Instead, run a shadow verification wave before merging.

Pre-flight: claim normalization and scale guard

Before dispatching verifiers, the coordinating agent normalizes each claim to a canonical form:

[CLAIM_ID] <subject> :: <predicate> :: <evidence_ref>

evidence_ref must be a concrete pointer — a file path, line number, git ref, config key, or API endpoint. Claims that cannot be anchored to a concrete evidence_ref are classified UNVERIFIABLE immediately and removed from the verification queue. They are preserved in the final output under a dedicated section with the reason they could not be anchored (no re-derivation sub-agent is dispatched for them).

Scale cap: if the normalized claim list exceeds 50 items, the coordinating agent halts and asks the operator to scope the investigation before proceeding. Verification at scale degrades into noise; a 50-item cap prevents a finding flood from becoming an echo-chamber rubber-stamp.

Select 2–3 of the highest-stakes normalized claims to send to the verifier wave. Prefer claims that are (a) decision-driving, (b) expressed with high-confidence language, or (c) hard to re-derive from a single artifact.


Wave 2 — Adversarial verifiers (parallel, independent):

  1. Extract 2–3 concrete, re-checkable claims from the returned report (e.g., "X function is unused", "file Y exceeds 300 lines", "PR targets main", "no tests cover Z") and normalize them to [CLAIM_ID] subject :: predicate :: evidence_ref form as described above.
  2. Dispatch one shadow sub-agent per claim, in parallel. Each receives the normalized claim text + the user's original goal + the search surface — the inventory of files, directories, or URLs the original investigation touched. It must NOT receive the original agent's reasoning, verdict, confidence language, or the specific line/region it concluded from. Withhold the conclusion, not the map. Withholding the map too does not buy extra independence — the verifier still has to reach the same evidence, it just spends its budget guessing paths to get there. Measured: one verifier denied the inventory spent 70 grep + 18 read_file calls re-locating files the parent already had paths for, guessed 5 nonexistent paths on the way, and hit its tool-loop ceiling before finishing. The independence that matters is epistemic (re-deriving the verdict), not navigational.
    • The inventory is a starting surface, not a boundary: it does not satisfy the composition-axis guard below, and a verifier that reads only inside it still returns evidence_base: artifact-internal. At least one primary source outside that surface is still required for independent-rederivation.
    • Default to subagent_type: "research-agent" (mechanically locked to Read/Grep/Glob/WebFetch/WebSearch — cannot Edit/commit/push). If the claim requires Bash to verify (running a failing test, gh pr view, git log origin/...), fall back to a Bash-capable subagent type with isolation: "worktree" and prepend this prefix to the prompt: "Verifier sub-agent — do not Edit, Write, commit, push, gh pr create, or curl. Return findings only."
    • Every verifier dispatch carries an explicit budgetmax_tool_use_iterations (a wave of 2–3 claim checks needs ~15–25 rounds each, not 50) plus the cheapest sufficient model. An unbudgeted verifier does not fail loudly: it exhausts the default tool-round ceiling, terminates stopReason: "tool_use_loop_capped", and emits its verdict from a tools-stripped wind-down round built on partial evidence. A CONFIRMED produced that way is indistinguishable from a real one and silently defeats the entire point of the wave. Check each returned verifier's stop reason before merging its verdict; treat a capped or wind-down verifier as UNVERIFIABLE, not as a verdict.
  3. Each verifier re-derives the verdict independently using tool calls only — never re-reading the original report's reasoning. Returns {claim_id, verifier_verdict, evidence_pointer, evidence_base}, where verifier_verdict is one of CONFIRMED, REFUTED, STALE, or UNVERIFIABLE, and evidence_base is independent-rederivation (read primary sources outside the cited artifact's boundary) or artifact-internal (re-read only the cited file/region). On REFUTED or STALE, the verifier also emits a corrected or updated finding.

Read the full file on GitHub · 96 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 2d ago First seen · 96 lines · 136 tokens per session scan A 9f7db5a0612c

Subscribe to this mod's changes

shadow-verify is a skill published in the GitHub repository griffinwork40/agent-framework (23 stars, last pushed 7d ago), licensed Apache-2.0. It adds 136 tokens to every session and 2,693 once invoked, about $0.0007 per session on Opus 5. A static security scan graded it A with 1 finding (makes network calls). It is 100% identical to shadow-verify, differing in 0 lines, and is treated as a copy.

Related

Other skills, from other repositories

systematic-debugging

Use when encountering any bug, test failure, or unexpected behavior, before proposing fixes.

obra/superpowers · 21 tokens

brainstorming

You MUST use this before any creative work - creating features, building components, adding functionality, or modifying behavior. Explores user intent, requirements and design before implementation.

obra/superpowers · 37 tokens

chat-pet-sprite-creation

Use when creating or changing VS Code chat pet sprite art, sprite sheets, state animations, eye treatments, Stable/Insiders variants, or pet transitions under src/vs/workbench/contrib/chat/browser/widget/media/chatPet.

microsoft/vscode · 53 tokens

cpu-profile-analysis

Analyze V8/Chrome CPU profiles (.cpuprofile) and DevTools trace files (Trace-.json). Use when: profiling performance, investigating slow functions, comparing code paths, finding bottlenecks, analyzing timeToRequest, understanding call trees from sampling profiler data, analyzing layout/paint/rendering, investigating…

microsoft/vscode · 71 tokens

agent-host-chat-contributions

Build and review cross-cutting agent-host chat behavior through lifecycle contributions. Use when adding turn lifecycle side effects, prompt or context injection, restored-history transformation, protocol-action observation, or when reviewing changes that add code to AgentSideEffects or AgentService.

microsoft/vscode · 56 tokens

auto-perf-optimize

Run agent-driven VS Code performance or memory investigations. Use when asked to launch Code OSS, automate a VS Code scenario, run the Chat memory smoke runner, capture renderer heap snapshots, take workflow screenshots, compare run summaries, or drive a repeatable scenario before heap-snapshot analysis.

microsoft/vscode · 62 tokens