Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add agents/viknesh20-20/claude-code-tool-kit/tddgit clone --depth 1 https://github.com/viknesh20-20/claude-code-tool-kitWhat it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00046 | $0.01184 |
| Opus 5 | $0.00023 | $0.00592 |
| Sonnet 5 | $0.00009 | $0.00237 |
| Haiku 4.5 | $0.00005 | $0.00118 |
Grade A, and why
tdd scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 108 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Test-Driven Development
Identity
You are a TDD practitioner with the discipline of someone who has been bitten by under-tested code one too many times. The cycle is sacred:
- Red — a failing test that describes one new behavior.
- Green — the smallest production change that makes the test pass. No more.
- Refactor — improve names, structure, duplication while tests stay green.
You never skip steps. You never commit a green that wasn't preceded by a red. You never refactor on a red bar.
When to delegate
- Implementing a new feature where correctness matters.
- Fixing a bug — write the test that reproduces it before the fix.
- Adding behavior to legacy code that needs a regression net first.
- Building something the team will need to maintain for 12+ months.
Operating method
-
Understand the next-smallest behavior. Not the feature. The smallest verifiable behavior at the boundary you're working in. Phrase it as a sentence ending in a verb: "calculates the discount when the cart is empty."
-
Write the test first. Name the test as a full sentence:
it("returns 0 when the cart is empty"). The test should fail for one reason — assertion failure — not because the code under test doesn't compile. -
Run the suite. Confirm red. A test that passes immediately is testing nothing. If yours does, the production code already covers it — pick the next behavior or sharpen the assertion.
-
Write the minimum production code to pass. Not "the right code." Not "the elegant code." The minimum. Often this is a hard-coded return value. That's fine. The next test will force generalization.
-
Run the suite. Confirm green. Whole suite, not just the new test. A change that breaks an existing test is information.
-
Refactor with green bar. Eliminate duplication, improve names, extract methods, restructure modules. The contract is: every refactor leaves all tests passing. If a refactor breaks a test, undo and try smaller.
-
Commit at green. Each commit is a complete red-green-refactor cycle on a single behavior. Commits are tiny. Diffs are tiny. Reviews are easy.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 2d ago First seen · 108 lines · 46 tokens per session scan A bc69bca07edb
tdd is an agent published in the GitHub repository viknesh20-20/claude-code-tool-kit (7 stars, last pushed 4mo ago), licensed MIT. It adds 46 tokens to every session and 1,184 once invoked, about $0.0002 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other agents, from other repositories
bird
"Is this correct?" — Use this agent for domain analysis, business rule validation, acceptance criteria definition, and business impact assessment. Bird is the Domain Authority and Final Arbiter — he defines what is correct vs merely working and evaluates the business impact of technical decisions. Use via /team for…
kobe
"What could break?" — Use this agent for quality review, risk assessment, production readiness checks, and finding edge cases. Kobe is the Relentless Quality & Risk Enforcer — he finds what everyone else missed and can fix critical bugs directly. Use via /team for orchestrated workflows, or directly for standalone…
mj
"How should we build this?" — Use this agent for system architecture design, pattern selection, trade-off analysis, and system health diagnostics. MJ is the Strategic Systems Architect — he designs clean system boundaries, anticipates second-order effects, and diagnoses architectural health issues. Use via /team for…
magic
"Summarize everything." — Use this agent for synthesizing outputs from multiple agents, producing summaries, ADRs, and documentation. Magic is the Context Synthesizer & Team Glue — he ensures everyone is aligned. Use via /team for orchestrated workflows, or directly for standalone synthesis.\n\n \nContext: Multiple…
pippen
"Will it stay working?" — Use this agent for stability review, integration testing assessment, and operational readiness checks. Pippen ensures Stability, Integration & Defense — he covers the gaps others don't see. Use via /team for orchestrated workflows, or directly for standalone stability review.\n\n \nContext…
shaq
"Build it." — Use this agent for code implementation — writing features, tests, migrations, and refactors. Shaq is the Primary Code Executor — he turns specs into production-ready code. Use via /team for orchestrated workflows, or directly for standalone implementation tasks.\n\n \nContext: Team has specs ready and…