verification-bench

verification-bench is a skill for Claude Code, Codex from Dailyaiagents/daily-ai-agent-toolkit. It costs 31 tokens per session (159 once invoked), scanned A, original, Apache-2.0.

A test tool that runs a fixed set of artificial examples through a checker and reports which expected failures and passes it handled.

In plain words
What is it for?
Use it to measure caught failures, accepted passing cases, catch rate, and false-refusal rate for the bundled subject.
Why use it?
It provides repeatable results for a known test set without treating those results as proof that the checker works accurately in real production use.

Skill for Claude CodeCodex

Written for no agent in particular: nothing here depends on one.

Good fit Use it to measure caught failures, accepted passing cases, catch rate, and false-refusal rate for the bundled subject.

Compare 6 skills from other repositories ↓
Install with agentmods
npx agentmods add skills/dailyaiagents/daily-ai-agent-toolkit/verification-bench
Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

Any agent
npx skills add Dailyaiagents/daily-ai-agent-toolkit --skill verification-bench
Clone the repo
git clone --depth 1 https://github.com/Dailyaiagents/daily-ai-agent-toolkit

Made for: Claude Code, Codex.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for verification-bench

README.md
[![agentmods](https://agentmods.dev/badge/skills/dailyaiagents/daily-ai-agent-toolkit/verification-bench/github.svg)](https://agentmods.dev/skills/dailyaiagents/daily-ai-agent-toolkit/verification-bench)
Your own site
<a href="https://agentmods.dev/skills/dailyaiagents/daily-ai-agent-toolkit/verification-bench"><img src="https://agentmods.dev/badge/skills/dailyaiagents/daily-ai-agent-toolkit/verification-bench/github.svg" alt="Measured on agentmods" height="20"></a>

Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.

agentmods 80×15 button for verification-bench

Your own site · 80×15
<a href="https://agentmods.dev/skills/dailyaiagents/daily-ai-agent-toolkit/verification-bench"><img src="https://agentmods.dev/badge/skills/dailyaiagents/daily-ai-agent-toolkit/verification-bench.svg" alt="Reviewed on agentmods" width="80" height="20"></a>
Per session 31 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 159 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. A grade says what 26 rules found in the file — not that it is safe.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5.1 $0.00031 $0.00159
Opus 5 $0.00015 $0.00079
Sonnet 5 $0.00006 $0.00032
Haiku 4.5 $0.00003 $0.00016

Measured 10d ago against content hash 9010d9db7aa6, method: parsed. Prices are Anthropic first-party input rates as of 2026-09-10, from the pricing page.

Security

Grade A, and why

verification-bench scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 10d ago.

The scan reads SKILL.md. This mod also ships 2 executable files (scripts/run.sh, scripts/selftest.sh), listed below but not scanned — reading those needs a real analyzer, not pattern matching.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

skills/verification-bench/SKILL.md · 20 lines

What it actually says

Verification Bench

Run the bundled synthetic fixtures:

bash scripts/run.sh

The receipt reports every fixture's expected and observed outcome, the denominator, caught failures, accepted pass cases, catch rate, and false-refusal rate. The bundled subject rejects missing, empty, TODO, and PLACEHOLDER artifacts. Synthetic results do not establish production accuracy or the performance of a different checker.

Files

What ships with it

3 files beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 10d ago First seen · 20 lines · 31 tokens per session scan A 9010d9db7aa6

Subscribe to this mod's changes

verification-bench is a skill published in the GitHub repository Dailyaiagents/daily-ai-agent-toolkit (0 stars, last pushed 6d ago), licensed Apache-2.0. It adds 31 tokens to every session and 159 once invoked, about $0.0002 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.

Related

Other skills, from other repositories

browser-use

Direct browser control via CDP for web interaction: automation, scraping, testing, screenshots, and site/app work.

browser-use/browser-use · 26 tokens

temporal-python-testing

Test Temporal workflows with pytest, time-skipping, and mocking strategies. Covers unit testing, integration testing, replay testing, and local development setup. Use when implementing Temporal workflow tests or debugging test failures.

wshobson/agents · 45 tokens

mem0-test-integration

Verify a Mem0 integration produced by /mem0-integrate. Runs in the same workspace on the same branch (loose coupling) — installs dependencies, runs the repo's native test suite, then exercises a real end-to-end smoke flow against the user's API key. Produces a scorecard. TRIGGER when: user has just run /mem0-integrate…

mem0ai/mem0 · 207 tokens

code-review

Perform a structured code review of changes, checking for correctness, style, tests, and potential issues.

langchain-ai/deepagents · 23 tokens

test-corpus

The testdocuments submodule is a bucket-fetched fixture corpus that is not committed. This skill covers readtestfixture, missing fixtures, valid A/B controls, and submodule push order. Load before running Rust tests on a fresh clone, setting up an A/B control, adding a fixture-backed test, or diagnosing…

xberg-io/xberg · 72 tokens

typescript-providers

Implement, modify, test, or document TypeScript provider packages under ts/packages/providers, including framework adapters for OpenAI, Anthropic, Google, LangChain, Mastra, Vercel, LlamaIndex, Cloudflare, and Claude Agent SDK. Use for provider-specific TS work; do not use for core-only changes.

ComposioHQ/composio · 70 tokens