Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add skills/andrewcigan/vibe-dev-plugin/validation-samplenpx skills add andrewcigan/vibe-dev-plugin --skill validation-samplegit clone --depth 1 https://github.com/andrewcigan/vibe-dev-pluginWhat it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00085 | $0.00559 |
| Opus 5 | $0.00043 | $0.00280 |
| Sonnet 5 | $0.00017 | $0.00112 |
| Haiku 4.5 | $0.00009 | $0.00056 |
Grade A, and why
validation-sample scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
What it actually says
/validation-sample
Построение эталонной выборки через validation-sample-builder agent.
Что происходит
- Subagent
validation-sample-builder(Sonnet) читает PRODUCT, ARCHITECTURE, domain-rules - Identify источники реалистичных сценариев:
- User-provided data (приоритет)
- Voice/chat logs old projects
- Competitor reviews (WebSearch)
- User-perspective-critic generated
- Synthetic (≤20%)
- Строит 50-100 scenarios в категориях:
- basic_intent 60-70%
- edge 15-20%
- error 10-15%
- Leak prevention:
docs/validation-scenarios/inputs/— только inputdocs/validation-scenarios/ground-truth/— expected (только judge видит)
- Запускает baseline run на текущей сборке (если есть код)
Output
docs/validation-sample.md— сводкаdocs/validation-scenarios/inputs/*.md— 50-100 сценариевdocs/validation-scenarios/ground-truth/*.md— expected./validation-runs/run-<ts>.jsonl— результаты
Critical gotchas
- Leak prevention: expected НЕ в одном файле с input
- Judge contains rule: YES если expected appears anywhere в Got, не exact match
- Truncate Got >500 chars запрещён
Pass thresholds
- ≥90% → ✓ /ship
- 80-89% → 🟡 /ship с warnings, failed → backlog
- <80% → ❌ stop, 5 Why на failed scenarios
Дальше
- Если pre-/feature loop — запоминаем baseline, /feature feat-001
- Если в /ship — финальный gate
Cost cap
$3. Может занять до часа compute на judge.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 2d ago First seen · 56 lines · 85 tokens per session scan A 6084473b1b99
validation-sample is a skill published in the GitHub repository andrewcigan/vibe-dev-plugin (5 stars, last pushed 1mo ago), licensed MIT. It adds 85 tokens to every session and 559 once invoked, about $0.0004 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other skills, from other repositories
harness-engineering-guide
Audit, design, and implement AI agent harnesses for any codebase. A harness is the constraints, feedback loops, and verification systems surrounding AI coding agents — improving it is the highest-leverage way to improve AI code quality. Three modes: Audit (scorecard), Implement (set up components), Design (full…
google-drive-sheets
Find, read, export, edit, and manage the user's Google Drive, Docs, Sheets, and Slides through per-user OAuth.
github-gitlab
Work with GitHub and GitLab repositories through resident gh/glab/git auth on the agent computer.
harness-creator
Build, audit, and improve harnesses that make AI coding agents reliable: AGENTS.md/CLAUDE.md instruction files, feature/state tracking, verification gates, scope boundaries, session handoff, memory persistence, context budgets, tool-permission safety, and multi-agent coordination. Use this whenever a coding agent is…
miniapp
Build a tiny interactive HTML playground only when someone asks to see, play with, or step through a mechanism.
commit-push-pr
Commit selected local changes, push the branch, and create or update a GitHub pull request with BitFun attribution. Use when the user asks to 提交 PR、提代码、commit and push、开 PR、create a pull request, or wants a Claude Code-like one-command PR publishing flow from BitFun.