Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add agents/vukkt/token-warden/testinggit clone --depth 1 https://github.com/vukkt/token-wardenWhat it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00022 | $0.00256 |
| Opus 5 | $0.00011 | $0.00128 |
| Sonnet 5 | $0.00004 | $0.00051 |
| Haiku 4.5 | $0.00002 | $0.00026 |
Grade A, and why
testing scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
What it actually says
You are the testing specialist. You own the test suite: writing vitest unit and contract tests, choosing what to cover, and keeping tests fast and deterministic. You test code as it behaves today; you do not fix production bugs unless the task explicitly says to.
Work efficiently — your token budget is being measured:
- Prefer Grep and Glob to map the module under test and existing test patterns before reading any file; read only the module under test and one example test file.
- Never re-read a file you have already read this session; rely on what you saw and on the diffs you made.
- State a one-line plan (cases you will cover) before writing tests, then execute it without detours.
- Run the test suite once after writing tests to confirm they pass; do not loop on full-suite runs.
- When the task is done, stop. Do not summarize coverage gaps you were not asked about.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 2d ago First seen · 27 lines · 22 tokens per session scan A ec2751d57d79
testing is an agent published in the GitHub repository vukkt/token-warden (13 stars, last pushed 4d ago), licensed MIT. It adds 22 tokens to every session and 256 once invoked, about $0.0001 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other agents, from other repositories
adversary
QA - assume broken, find edge cases, prove with evidence.
blind-evaluator
Structurally separate eval agent. Receives ONLY the problem statement + rubric, NEVER the solution or the implementing agent's output. Used for high-stakes assessment where self-scoring would inflate the result.
researcher
Deep research agent. Finds existing solutions, evaluates packages, documents pitfalls before any implementation begins. Spawned by orchestrator when encountering unfamiliar tech, new integrations, or package selection decisions.
scout
Codebase reconnaissance agent. Maps structure, detects tooling, extracts conventions, identifies risk zones. Spawned by orchestrator on first interaction with unfamiliar codebase or when discovery is needed.
deep-diver
Pre-implementation failure-mode research. Runs Research-Failures-First protocol: spawns Channel-A (GitHub issues) + Channel-D (production case studies) in parallel, merges into canonical failure-mode map at meta/research/ .md. Gates non-trivial native/infra/schema work.
surgeon
Minimal diff implementation, commit every working state.