Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add skills/classmethod/tsumiki/dev-verifynpx skills add classmethod/tsumiki --skill dev-verifygit clone --depth 1 https://github.com/classmethod/tsumikiWhat it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00090 | $0.01614 |
| Opus 5 | $0.00045 | $0.00807 |
| Sonnet 5 | $0.00018 | $0.00323 |
| Haiku 4.5 | $0.00009 | $0.00161 |
Grade A, and why
dev-verify scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 163 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Dev Verify
Plan単位ですべてのタスクの完了状態を確認し、全テスト実行・ビルド・Lint・ファイルサイズチェックを行い、検証レポートを docs/dev/plans/<plan-name>/reports/ に出力する。
前提知識
dev-*スキルフロー内の位置
dev-context → dev-plan → dev-impl → [dev-verify]
↘ dev-debug
引数フォーマット
/dev-verify <plan-name>
ワークフロー
Step 1: タスク完了チェック
docs/dev/plans/<plan-name>/tasks/ 内の全タスクファイルを読み込み、フロントマターの status を確認する。
- すべて
done→ Step 2 へ進む pending/in_progressのタスクがある → ユーザーに未完了タスクの一覧を報告し、続行するか確認する(AskUserQuestion)
Step 2: 全テスト実行
docs/dev/context.md からテスト実行コマンドを取得し、全テストを実行する。
# context.md の Test Framework セクションからコマンドを取得
<test-command>
テスト結果を記録する: passed / failed / skipped の件数。
Step 3: カバレッジチェック
docs/dev/context.md の Test Framework セクションから Coverage Command と Coverage Threshold を取得し、プロジェクト全体のカバレッジを計測する。
- Coverage Command でプロジェクト全体のカバレッジを計測する
- パッケージごとにカバレッジ率を抽出し、閾値と比較する
[no test files]は 0% として扱う- 閾値未達のパッケージは Issues Found に記録する
Step 4: ビルド確認
コンパイル言語やビルドステップがある場合、ビルドコマンドを実行する(context.md から取得)。
インタープリタ言語でビルドステップがない場合はスキップする。
Step 5: Lint実行
context.md に Lint コマンドが記載されている場合、実行する。記載がない場合はスキップする。
Step 6: ファイルサイズチェック(500行ルール)
Planで変更・作成されたファイル(各タスクファイルの Files セクションから収集)の行数を確認する(絶対パスを使用):
wc -l "$(git rev-parse --show-toplevel)/<ファイルの相対パス>"
500行を超えるファイルがあれば警告として記録する。
Step 7: 信号機サマリー
Plan内の全タスクの確信度レポートを集約する:
- 🟡 妥当な推測の項目一覧(ユーザー確認推奨)
- 🔴 AI推論補完の項目一覧(人間の確認必須)
🔴 が残っている場合は明確に警告する。
Step 8: レポート出力
検証結果を docs/dev/plans/<plan-name>/reports/verify-YYYY-MM-DD.md に出力する。
# Verification Report - YYYY-MM-DD
## Plan: <plan-name>
## Summary
| Item | Result |
|------|--------|
| Tasks | X/Y completed |
| Tests | Z passed, W failed, V skipped |
| Coverage | X/Y packages above threshold (threshold: 80%) |
| Build | OK / NG / N/A |
| Lint | OK / NG / N/A |
| 500-line rule | OK / X files over limit |
## Task Status
| ID | Title | Status |
|----|-------|--------|
| 001 | ... | done |
| 002 | ... | done |
## Test Results
[テスト実行の出力サマリー]
[失敗テストがある場合は詳細]
## Coverage Results
| パッケージ | カバレッジ | 閾値 | 結果 |
|-----------|----------|------|------|
| internal/handler | XX.X% | 80% | OK / NG |
[閾値未達パッケージの詳細(ある場合)]
## File Size Check
[500行超過ファイルの一覧(ある場合)]
## Confidence Summary
### 🟡 妥当な推測(要確認)
- [ファイル](パス) — [確認すべき点]
### 🔴 AI推論補完(人間の確認必須)
- [ファイル](パス) — [判断が必要な理由]
## Issues Found
[検出された問題の詳細]
## Recommendations
[問題がある場合の次アクション提案]
- テスト失敗 → `/dev-debug` の使用を推奨
- 500行超過 → `/dev-impl` のリファクタリングを推奨
- 🔴 残存 → 該当ファイルの人間レビューを推奨
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 2d ago First seen · 163 lines · 90 tokens per session scan A 5ea3a6b07f70
dev-verify is a skill published in the GitHub repository classmethod/tsumiki (974 stars, last pushed 25d ago), licensed MIT. It adds 90 tokens to every session and 1,614 once invoked, about $0.0005 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other skills, from other repositories
systematic-debugging
Use when encountering any bug, test failure, or unexpected behavior, before proposing fixes.
next-cache-components-adoption
Turn on Cache Components in a Next.js app and resolve the blocking routes it surfaces. Use when the user wants to enable, adopt, or migrate to Cache Components, flip the cacheComponents flag, work through a flood of blocking-prerender / instant validation errors, run the cache-components-instant-false codemod, or…
babysit-pr
Babysit a GitHub pull request after creation by continuously polling review comments, CI checks/workflow runs, and mergeability state until the PR is merged/closed or user help is required. Diagnose failures, retry likely flaky failures up to 3 times, auto-fix/push branch-related issues when appropriate, and keep…
imagegen
Generate or edit raster images when the task benefits from AI-created bitmap visuals such as photos, illustrations, textures, sprites, mockups, or transparent-background cutouts. Use when Codex should create a brand-new image, transform an existing image, or derive visual variants from references, and the output…
cpu-profile-analysis
Analyze V8/Chrome CPU profiles (.cpuprofile) and DevTools trace files (Trace-.json). Use when: profiling performance, investigating slow functions, comparing code paths, finding bottlenecks, analyzing timeToRequest, understanding call trees from sampling profiler data, analyzing layout/paint/rendering, investigating…
next-cache-components-optimizer
Drive a Next.js route to instant navigation by setting up an agentic loop, under Cache Components / PPR, on initial load (hard navigation) and client-side navigation (soft navigation). Encode the goal as a failing @next/playwright instant() e2e and work it to green, one verified route at a time; the shipped test then…