Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add skills/griffinwork40/agent-afk/shadow-verifynpx skills add griffinwork40/agent-afk --skill shadow-verifygit clone --depth 1 https://github.com/griffinwork40/agent-afkWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/griffinwork40/agent-afk/shadow-verify)<a href="https://agentmods.dev/skills/griffinwork40/agent-afk/shadow-verify"><img src="https://agentmods.dev/badge/skills/griffinwork40/agent-afk/shadow-verify.svg" alt="Measured on agentmods" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00136 | $0.02693 |
| Opus 5 | $0.00068 | $0.01347 |
| Sonnet 5 | $0.00027 | $0.00539 |
| Haiku 4.5 | $0.00014 | $0.00269 |
Grade A, and why
shadow-verify scanned grade A with 1 finding against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 5d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Makes network callslowCapability
Not a fault in itself. Listed so you know the mod talks to something, and to what.
- **Default to `subagent_type: "research-agent"` (mechanically locked to Read/Grep/Glob/WebFetch/WebSearch — cannot Edit/commit/push).** If the claim requires Bash to verify (running a failing test, `gh pr view`, `git lo Copies of this mod
1 near-identical copy found in the catalogue:
- shadow-verify — 100% identical, 0 lines differ
How it starts
The opening of the file, as written. The whole thing — 96 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Sub-agent contract
/contract
When a sub-agent (or wave) returns investigation findings, code-review conclusions, audit claims, refactor plans, gap-analysis results, or counts that will drive user decisions or file changes, do NOT surface the report. Instead, run a shadow verification wave before merging.
Pre-flight: claim normalization and scale guard
Before dispatching verifiers, the coordinating agent normalizes each claim to a canonical form:
[CLAIM_ID] <subject> :: <predicate> :: <evidence_ref>
evidence_ref must be a concrete pointer — a file path, line number, git ref, config key, or API endpoint. Claims that cannot be anchored to a concrete evidence_ref are classified UNVERIFIABLE immediately and removed from the verification queue. They are preserved in the final output under a dedicated section with the reason they could not be anchored (no re-derivation sub-agent is dispatched for them).
Scale cap: if the normalized claim list exceeds 50 items, the coordinating agent halts and asks the operator to scope the investigation before proceeding. Verification at scale degrades into noise; a 50-item cap prevents a finding flood from becoming an echo-chamber rubber-stamp.
Select 2–3 of the highest-stakes normalized claims to send to the verifier wave. Prefer claims that are (a) decision-driving, (b) expressed with high-confidence language, or (c) hard to re-derive from a single artifact.
Wave 2 — Adversarial verifiers (parallel, independent):
- Extract 2–3 concrete, re-checkable claims from the returned report (e.g., "X function is unused", "file Y exceeds 300 lines", "PR targets main", "no tests cover Z") and normalize them to
[CLAIM_ID] subject :: predicate :: evidence_refform as described above. - Dispatch one shadow sub-agent per claim, in parallel. Each receives the normalized claim text + the user's original goal + the search surface — the inventory of files, directories, or URLs the original investigation touched. It must NOT receive the original agent's reasoning, verdict, confidence language, or the specific line/region it concluded from. Withhold the conclusion, not the map. Withholding the map too does not buy extra independence — the verifier still has to reach the same evidence, it just spends its budget guessing paths to get there. Measured: one verifier denied the inventory spent 70
grep+ 18read_filecalls re-locating files the parent already had paths for, guessed 5 nonexistent paths on the way, and hit its tool-loop ceiling before finishing. The independence that matters is epistemic (re-deriving the verdict), not navigational.- The inventory is a starting surface, not a boundary: it does not satisfy the composition-axis guard below, and a verifier that reads only inside it still returns
evidence_base: artifact-internal. At least one primary source outside that surface is still required forindependent-rederivation. - Default to
subagent_type: "research-agent"(mechanically locked to Read/Grep/Glob/WebFetch/WebSearch — cannot Edit/commit/push). If the claim requires Bash to verify (running a failing test,gh pr view,git log origin/...), fall back to a Bash-capable subagent type withisolation: "worktree"and prepend this prefix to the prompt: "Verifier sub-agent — do not Edit, Write, commit, push,gh pr create, orcurl. Return findings only." - Every verifier dispatch carries an explicit budget —
max_tool_use_iterations(a wave of 2–3 claim checks needs ~15–25 rounds each, not 50) plus the cheapest sufficient model. An unbudgeted verifier does not fail loudly: it exhausts the default tool-round ceiling, terminatesstopReason: "tool_use_loop_capped", and emits its verdict from a tools-stripped wind-down round built on partial evidence. ACONFIRMEDproduced that way is indistinguishable from a real one and silently defeats the entire point of the wave. Check each returned verifier's stop reason before merging its verdict; treat a capped or wind-down verifier asUNVERIFIABLE, not as a verdict.
- The inventory is a starting surface, not a boundary: it does not satisfy the composition-axis guard below, and a verifier that reads only inside it still returns
- Each verifier re-derives the verdict independently using tool calls only — never re-reading the original report's reasoning. Returns
{claim_id, verifier_verdict, evidence_pointer, evidence_base}, whereverifier_verdictis one ofCONFIRMED,REFUTED,STALE, orUNVERIFIABLE, andevidence_baseisindependent-rederivation(read primary sources outside the cited artifact's boundary) orartifact-internal(re-read only the cited file/region). OnREFUTEDorSTALE, the verifier also emits a corrected or updated finding.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 5d ago First seen · 96 lines · 136 tokens per session scan A 9f7db5a0612c
shadow-verify is a skill published in the GitHub repository griffinwork40/agent-afk (54 stars, last pushed today), licensed Apache-2.0. It adds 136 tokens to every session and 2,693 once invoked, about $0.0007 per session on Opus 5. A static security scan graded it A with 1 finding (makes network calls). No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other skills, from other repositories
ian-xiaohei-illustrations
生成 Ian 风格的中文正文配图。用于用户要求为中文文章、帖子、博客、Notion 文档、工作流文档、方法论、流程、结构、状态、隐喻或观点生成“怪诞”“小黑”“手绘”“正文配图”“文章插图”“配图建议”“shot list”“去标题/改图”等任务;默认使用小黑 IP、纯白手绘、少量红橙蓝批注、简洁清爽但天马行空的视觉风格。.
capabilities
Your capability catalog — read this at boot. Lists the temporal date-range skills and the external integrations (reached via the loopback broker) available to you as a spawned worker, and exactly how to call each. Read-only. Consult it whenever you're unsure what tools/integrations you have or how to invoke them.
temporal
Resolve ANY named time window — today, yesterday, thisWeek, lastWeek, last7Days, last30Days, last90Days, thisMonth, lastMonth, thisQuarter, lastQuarter, thisYear, lastYear, last12Months — or an arbitrary range (lastNdays / lastNweeks / lastNmonths) to a concrete ISO date range relative to your run time. Read-only: no…
last30Days
Resolve "last30Days" to a concrete ISO date range relative to your run time — a rolling 30-day window ending today. Returns inclusive civil dates plus exact UTC instants so you have temporal context without computing dates by hand. Read-only: no writes, no network. Use before a "last 30 days" / trailing-month task…
md-audit
Read-only code quality audit — scan the current working directory for common issues (bugs, dead code, security hotspots, missing error handling) and return a prioritised findings report. No files are edited. Use when asked to "audit the code", "quick audit", "find issues", "code scan", or "what's wrong with this…
thisQuarter
Resolve "thisQuarter" to a concrete ISO date range relative to your run time — this quarter so far (quarter start → today). Returns inclusive civil dates plus exact UTC instants so you have temporal context without computing dates by hand. Read-only: no writes, no network. Use before a quarter-to-date task (QTD…