Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add skills/arcblock/agent-skills/verificationnpx skills add ArcBlock/agent-skills --skill verificationgit clone --depth 1 https://github.com/ArcBlock/agent-skillsWhat it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00059 | $0.01963 |
| Opus 5 | $0.00030 | $0.00981 |
| Sonnet 5 | $0.00012 | $0.00393 |
| Haiku 4.5 | $0.00006 | $0.00196 |
Grade A, and why
verification scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 3d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 121 lines — stays where its author put it; the contents beside it link to each section on GitHub.
verification (agentloop engine)
Repo-agnostic. The check list and gate commands come from the consuming repo:
.claude/repo-profile.md(gate_mode,verification_entry,pre_merge_entry) and.claude/verify/config.ts. Paths shown as.claude/verify/...are arc's defaults.
Deterministic gate whose numbers the scripts measure — the agent chooses which scenario to run and reads the result, but never hand-fills a stat. This is the guardrail: a check's exit code decides pass/fail, not a narrative.
Two layers
- Engine (this plugin, repo-agnostic):
lib/report.ts(CheckResult + render),lib/comment.ts(sticky PR-comment upsert),lib/scenario.ts(runScenario+cmd()). Knows nothing about pnpm/turbo/paths. - Repo config (in the consuming repo):
.claude/verify/config.tsdeclares the check list. Command-checks are pure config (cmd({ command: "pnpm build" })); logic-checks import a repo-local module. A thin.claude/verify/pre-pr.tscallsrunScenario(config, process.argv).
How to run
The repo exposes a scenario entry (<verification_entry>). Common flags:
--comment [<pr#>] upsert the report onto the PR (run + post = one step)
--json machine-readable
--na "<reason>" write an N/A exemption (docs-only / native-only PRs)
--only a,b / --skip x,y scope the check set (unknown id → hard error, exit 2)
→ a scoped run is a DIAGNOSTIC, never a gate (see below)
--deliver-cached post the cached PASS report without re-running
Run the gate with --comment <pr#> so "run" and "post" are one step. Exit codes:
0 = PASS (and, when --comment/--post was requested, the report WAS delivered);
1 = verify FAIL; 2 = empty check set / unknown --only/--skip id (fails
loud, never silent-green); 4 = verified PASS but the requested report was NOT
delivered — the remedy is to retry / fall back the comment post (e.g. paste the
stdout sticky body via MCP), NOT to touch the diff. Do not hand-write the report or
substitute a single tsc/build command for the scenario script.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 3d ago First seen · 121 lines · 59 tokens per session scan A 4ec2409675ff
verification is a skill published in the GitHub repository ArcBlock/agent-skills (5 stars, last pushed 3d ago), licensed MIT. It adds 59 tokens to every session and 1,963 once invoked, about $0.0003 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other skills, from other repositories
systematic-debugging
Use when encountering any bug, test failure, or unexpected behavior, before proposing fixes.
brainstorming
You MUST use this before any creative work - creating features, building components, adding functionality, or modifying behavior. Explores user intent, requirements and design before implementation.
auto-perf-optimize
Run agent-driven VS Code performance or memory investigations. Use when asked to launch Code OSS, automate a VS Code scenario, run the Chat memory smoke runner, capture renderer heap snapshots, take workflow screenshots, compare run summaries, or drive a repeatable scenario before heap-snapshot analysis.
chat-perf
Run chat perf benchmarks and memory leak checks against the local dev build or any published VS Code version. Use when investigating chat rendering regressions, validating perf-sensitive changes to chat UI, or checking for memory leaks in the chat response pipeline.
chat-pet-sprite-creation
Use when creating or changing VS Code chat pet sprite art, sprite sheets, state animations, eye treatments, Stable/Insiders variants, or pet transitions under src/vs/workbench/contrib/chat/browser/widget/media/chatPet.
cpu-profile-analysis
Analyze V8/Chrome CPU profiles (.cpuprofile) and DevTools trace files (Trace-.json). Use when: profiling performance, investigating slow functions, comparing code paths, finding bottlenecks, analyzing timeToRequest, understanding call trees from sampling profiler data, analyzing layout/paint/rendering, investigating…