code-verification

code-verification is a skill for Claude Code from bhaumikmaan/claude-code-master-skills. It costs 63 tokens per session (1,695 once invoked), scanned A, original, MIT.

An adversarial checker for code changes that actively tries to break an implementation instead of only looking for signs that it works. It produces a PASS, FAIL, or PARTIAL result with supporting evidence.

In plain words
What is it for?
Use it after implementing a feature, when checking whether a fix really works, or before reporting that a coding task is complete.
Why use it?
It exposes failures that happy-path tests, mocked tests, or a quick code read may miss, such as broken buttons, lost state, or invalid input handling.

Skill for Claude Code

Written for Claude Code: user-invocable in frontmatter. Also seen: mentions subagents.

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add skills/bhaumikmaan/claude-code-master-skills/code-verification
Any agent
npx skills add bhaumikmaan/claude-code-master-skills --skill code-verification
Clone the repo
git clone --depth 1 https://github.com/bhaumikmaan/claude-code-master-skills

Made for: Claude Code.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for code-verification

README.md
[![agentmods](https://agentmods.dev/badge/skills/bhaumikmaan/claude-code-master-skills/code-verification.svg)](https://agentmods.dev/skills/bhaumikmaan/claude-code-master-skills/code-verification)
Your own site
<a href="https://agentmods.dev/skills/bhaumikmaan/claude-code-master-skills/code-verification"><img src="https://agentmods.dev/badge/skills/bhaumikmaan/claude-code-master-skills/code-verification.svg" alt="Measured on agentmods" height="20"></a>
Per session 63 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 1,695 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 1 finding. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5.1 $0.00063 $0.01695
Opus 5 $0.00032 $0.00847
Sonnet 5 $0.00013 $0.00339
Haiku 4.5 $0.00006 $0.00169

Measured 5d ago against content hash eb079d3a13c5, method: parsed. Prices are Anthropic first-party input rates as of 2026-09-05, from the pricing page.

Security

Grade A, and why

code-verification scanned grade A with 1 finding against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 5d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Makes network callslowCapability

Not a fault in itself. Listed so you know the mod talks to something, and to what.

- **Frontend**: Start dev server, use browser automation to navigate/click/screenshot, curl subresources (images, API routes, static assets -- HTML can serve 200 while everything it references fails), run frontend tests
skills/code-verification/SKILL.md · 132 lines

How it starts

The opening of the file, as written. The whole thing — 132 lines — stays where its author put it; the contents beside it link to each section on GitHub.

Code Verification

Your job is not to confirm the implementation works -- it's to try to break it.

CRITICAL: Reading code is not verification. Run it.

Known Failure Patterns

You have two documented failure modes. Recognize them and do the opposite:

  1. Verification avoidance: You read code, narrate what you would test, write "PASS," and move on without running anything.
  2. Seduced by the first 80%: You see a polished UI or passing test suite and feel inclined to pass, not noticing half the buttons do nothing, state vanishes on refresh, or the backend crashes on bad input. Your entire value is in finding the last 20%.

Rationalization Inventory

You will feel the urge to skip checks. These are the exact excuses you reach for:

  • "The code looks correct based on my reading" -- reading is not verification. Run it.
  • "The tests already pass" -- tests may be heavy on mocks, circular assertions, or happy-path coverage that proves nothing end-to-end. Verify independently.
  • "This is probably fine" -- probably is not verified. Run it.
  • "Let me check the code" -- no. Start the server and hit the endpoint.
  • "I don't have a browser" -- check for browser automation tools first. If present, use them.
  • "This would take too long" -- not your call.

If you catch yourself writing an explanation instead of a command, stop. Run the command.

Verification Strategy by Change Type

Adapt strategy based on what was changed:

  • Frontend: Start dev server, use browser automation to navigate/click/screenshot, curl subresources (images, API routes, static assets -- HTML can serve 200 while everything it references fails), run frontend tests
  • Backend/API: Start server, curl/fetch endpoints, verify response shapes against expected values (not just status codes), test error handling, check edge cases
  • CLI/script: Run with representative inputs, verify stdout/stderr/exit codes, test edge inputs (empty, malformed, boundary), verify --help/usage accuracy
  • Infrastructure/config: Validate syntax, dry-run where possible (terraform plan, kubectl --dry-run, docker build, nginx -t), check env vars are actually referenced not just defined
  • Library/package: Build, full test suite, import from fresh context and exercise public API as a consumer, verify exported types match docs
  • Bug fixes: Reproduce the original bug first, verify fix, run regression tests, check related functionality for side effects
  • Database migrations: Run migration up, verify schema matches intent, run migration down (reversibility), test against existing data not just empty DB
  • Refactoring (no behavior change): Existing test suite MUST pass unchanged, diff the public API surface (no new/removed exports), spot-check same inputs produce same outputs
  • Other: The pattern is always (a) exercise the change directly, (b) check outputs against expectations, (c) try to break it with inputs/conditions the implementer didn't test

Read the full file on GitHub · 132 lines

Files

What ships with it

1 file beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 5d ago First seen · 132 lines · 63 tokens per session scan A eb079d3a13c5

Subscribe to this mod's changes

code-verification is a skill published in the GitHub repository bhaumikmaan/claude-code-master-skills (3 stars, last pushed 5mo ago), licensed MIT. It adds 63 tokens to every session and 1,695 once invoked, about $0.0003 per session on Opus 5. A static security scan graded it A with 1 finding (makes network calls). No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.

Related

Other skills, from other repositories

metrics-instrumentation

Specification for instrumenting an opik-backend workflow with operational OpenTelemetry metrics — per-stage throughput/latency/error counters and native histograms, dimensioned per-customer (workspace). Use when a pipeline (scoring, ingestion, experiments, jobs) needs per-stage visibility. Covers metric emission only…

comet-ml/opik · 93 tokens

happiness-skill

当用户问「怎么才能更幸福/为什么得到了还不满足/怎么减少焦虑」时调用。 核心理念: 幸福是缺憾感清空的默认状态, 是可训练的技能; 欲望是与自己的契约(得到前不快乐), 同时只留一个重大欲望; 活在当下。 不适用于: 临床抑郁等需要专业治疗的场景(本书方法不能替代医疗)。 Triggers: 幸福/不快乐/欲望/焦虑/知足/活在当下/happiness/desire/anxiety.

kangarooking/cangjie-skill · 136 tokens

short-drama-storyboard

把剧本和视觉事实转成有镜头职责、空间连续性和可冻结起点的 剧集/ /分镜.md。 每镜使用二级标题 ## SHOT-...,同镜下用 ### 冻结关键帧提示词 写起始帧正文。.

zenstory-ai/drama-skills · 102 tokens

setup-matt-pocock-skills

为本仓库配置工程技能——设置其 issue tracker、分诊标签词汇表和领域文档布局。首次使用其他工程技能前运行一次。.

devcxl/mattpocock-skills-zh · 43 tokens

pixel-art-studio

Create production-quality pixel art and animations programmatically. Use when the user asks to "create pixel art", "draw a sprite", "make pixel animation", "generate sprite sheet", "convert image to pixel art", "pixelate this image", "make a pixel character", "пиксель арт", "пиксельная графика", "спрайт", "像素画"…

AnastasiyaW/codex-claude-code-config · 349 tokens

frontend-design

Создание высококачественных, визуально выдающихся фронтенд-интерфейсов. Используй ВСЕГДА когда пользователь просит создать веб-страницу, компонент, лендинг, дашборд, UI-кит, форму, карточки, навигацию, анимации, или любой другой веб-интерфейс. Скилл покрывает: HTML/CSS/JS компоненты, React/Vue/Svelte, Tailwind CSS…

AnastasiyaW/codex-claude-code-config · 264 tokens