verify

A verification workflow that requires fresh command output before claiming software is complete, fixed, or passing. It uses tests, linters, builds, requests, or other checks as evidence.

In plain words
What is it for?
Use it to run the relevant checks, read their output, and report whether the evidence supports claims such as tests passing or a build succeeding.
Why use it?
It prevents success claims based only on memory, assumptions, or stale results.

Skill for Claude CodeCodex

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add skills/sandsower/beislid/verify
Any agent
npx skills add sandsower/beislid --skill verify
Clone the repo
git clone --depth 1 https://github.com/sandsower/beislid

Made for: Claude Code, Codex.

Per session 31 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 589 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 1 finding. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00031 $0.00589
Opus 5 $0.00015 $0.00295
Sonnet 5 $0.00006 $0.00118
Haiku 4.5 $0.00003 $0.00059

Measured 2d ago against content hash 9762c473a6d0, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

verify scanned grade A with 1 finding against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Makes network callslowCapability

Not a fault in itself. Listed so you know the mod talks to something, and to what.

1. **Identify** the verification command (test suite, linter, build, curl, etc.)
skills/verify/SKILL.md · 61 lines

How it starts

The opening of the file, as written. The whole thing — 61 lines — stays where its author put it; the contents beside it link to each section on GitHub.

Check Done

No completion claims without fresh verification evidence. Period.

If the repo declares custom lifecycle hooks, read ../lifecycle-hooks.md and honor any phase-boundary hooks before and after verify.

The Gate

When reporting durable evidence, use the terse Verification report shape from artifact-templates.md.

For every positive claim about work state:

  1. Identify the verification command (test suite, linter, build, curl, etc.)
  2. Run it fresh — not from memory, not from a previous run
  3. Read the actual output
  4. Confirm the output supports the claim
  5. Then make the claim

What Requires Evidence

Claim Required evidence
"Tests pass" Test command output showing 0 failures
"Linter clean" Linter output showing 0 errors
"Build succeeds" Build command with exit code 0
"Bug is fixed" Failing test now passes (red to green)
"Feature works" Test or manual verification output

Proof Requirements

When a Work Contract includes proof_requirements, match verification evidence to that vocabulary before claiming done. A command log can satisfy command_gate, review output can satisfy review or fresh_eyes, CI status can satisfy ci_check, and a deck/screenshot can satisfy screenshot_show_me. Missing required proof means stop and report the missing proof or human interrupt; do not downgrade it to advisory in chat.

Red Flags

These words in your response without preceding evidence mean you're guessing:

  • "should work", "probably works", "seems to be working"
  • "I believe this fixes", "this should resolve"
  • "tests should pass now"

Replace with: run the command, paste the output, state the fact.

After Subagent Work

Agent success reports are not evidence. After a subagent completes:

  1. Check the VCS diff
  2. Run verification independently
  3. Then confirm

Read the full file on GitHub · 61 lines

Files

What ships with it

1 file beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 2d ago First seen · 61 lines · 31 tokens per session scan A 9762c473a6d0

Subscribe to this mod's changes

verify is a skill published in the GitHub repository sandsower/beislid (10 stars, last pushed 3d ago), licensed MIT. It adds 31 tokens to every session and 589 once invoked, about $0.0002 per session on Opus 5. A static security scan graded it A with 1 finding (makes network calls). No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.

Related

Other skills, from other repositories

systematic-debugging

Use when encountering any bug, test failure, or unexpected behavior, before proposing fixes.

obra/superpowers · 21 tokens

brainstorming

You MUST use this before any creative work - creating features, building components, adding functionality, or modifying behavior. Explores user intent, requirements and design before implementation.

obra/superpowers · 37 tokens

chat-pet-sprite-creation

Use when creating or changing VS Code chat pet sprite art, sprite sheets, state animations, eye treatments, Stable/Insiders variants, or pet transitions under src/vs/workbench/contrib/chat/browser/widget/media/chatPet.

microsoft/vscode · 53 tokens

cpu-profile-analysis

Analyze V8/Chrome CPU profiles (.cpuprofile) and DevTools trace files (Trace-.json). Use when: profiling performance, investigating slow functions, comparing code paths, finding bottlenecks, analyzing timeToRequest, understanding call trees from sampling profiler data, analyzing layout/paint/rendering, investigating…

microsoft/vscode · 71 tokens

agent-host-chat-contributions

Build and review cross-cutting agent-host chat behavior through lifecycle contributions. Use when adding turn lifecycle side effects, prompt or context injection, restored-history transformation, protocol-action observation, or when reviewing changes that add code to AgentSideEffects or AgentService.

microsoft/vscode · 56 tokens

auto-perf-optimize

Run agent-driven VS Code performance or memory investigations. Use when asked to launch Code OSS, automate a VS Code scenario, run the Chat memory smoke runner, capture renderer heap snapshots, take workflow screenshots, compare run summaries, or drive a repeatable scenario before heap-snapshot analysis.

microsoft/vscode · 62 tokens