Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add jmagly/aiwg --skill claims-validatorgit clone --depth 1 https://github.com/jmagly/aiwgWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/jmagly/aiwg/claims-validator)<a href="https://agentmods.dev/skills/jmagly/aiwg/claims-validator"><img src="https://agentmods.dev/badge/skills/jmagly/aiwg/claims-validator/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/jmagly/aiwg/claims-validator"><img src="https://agentmods.dev/badge/skills/jmagly/aiwg/claims-validator.svg" alt="Reviewed on agentmods" width="80" height="20"></a>- NVIDIA SkillSpector pass
What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00017 | $0.01701 |
| Opus 5 | $0.00009 | $0.00851 |
| Sonnet 5 | $0.00003 | $0.00340 |
| Haiku 4.5 | $0.00002 | $0.00170 |
Grade A, and why
claims-validator scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 6d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 264 lines — stays where its author put it; the contents beside it link to each section on GitHub.
claims-validator
Validate documentation for unsupported claims, made-up metrics, and unverifiable statements.
Triggers
Alternate expressions and non-obvious activations (primary phrases are matched automatically from the skill description):
- "fact-check this" → claim validation
- "verify [claim]" → specific claim check
Purpose
This skill identifies statements that make claims without evidence, including:
- Performance metrics without benchmarks or data
- Time/cost estimates without basis
- Percentage claims without citation
- Comparative statements without baselines
- Features described as implemented that don't exist
- Marketing superlatives presented as facts
Behavior
When triggered, this skill:
-
Scans for metric claims:
- Percentage improvements ("40% faster", "reduces by 60%")
- Time estimates ("saves 2-3 hours", "in minutes not hours")
- Cost projections ("$50-150/month", "ROI of 3x")
- Performance numbers ("99x faster", "sub-millisecond")
-
Identifies unsupported comparatives:
- "faster than", "better than", "more efficient"
- "best", "leading", "revolutionary", "game-changing"
- "comprehensive", "complete", "full-featured"
-
Checks for feature claims:
- Commands or flags mentioned that don't exist in codebase
- Features described in present tense that aren't implemented
- Integration claims without actual integration code
-
Validates citations:
- Claims that reference data should have sources
- Benchmarks should link to methodology
- Statistics should be reproducible
-
Generates report:
- List each claim found
- Classification (metric, comparative, feature, cost)
- Recommendation (remove, add citation, verify, rephrase)
Claim Categories
Metrics Without Data
# Flagged
"Time Saved: 92-96% (9-15 hours → 45-60 minutes)"
"99x faster routing"
"45x cache speedup"
# Problem
No benchmark data, methodology, or reproducible test
# Fix
Remove claim, or add: "Based on [benchmark/test], measured [how]"
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 6d ago First seen · 264 lines · 17 tokens per session scan A 674b8c61a18c
claims-validator is a skill published in the GitHub repository jmagly/aiwg (210 stars, last pushed today), licensed MIT. It adds 17 tokens to every session and 1,701 once invoked, about $0.0001 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-09-03.
Other skills, from other repositories
surge
Use when a user provides a PRD, spec, or detailed requirements document and needs a full project delivered through iterative expert orchestration — multi-round analyze/research/design/implement/QA cycles with convergence detection. NOT for: single-file edits, quick prototypes, simple Q&A, or tasks without a written…
goal-writer
Drafts a goal+rider document pair that briefs an autonomous coding agent on one round of work — a goal file under 4,000 characters (sized to fit the /goal command in both Claude Code and Codex) plus an unbounded rider with phased plans and named depth tests. Use when the user says "draft a goal", "write a goal+rider"…
horizon
Run a durable Horizon workflow for a multi-feature goal with bounded autonomous retries and an audit trail.
check-in
Record a Parallax protocol checkpoint with concrete evidence before gated implementation work.
debug
Perform an evidence-based Parallax diagnosis or post-build audit and verify the repair.
hyperplan
Harden a non-trivial plan through a three-round adversarial critique and evidence-based synthesis.