Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add commands/ching-kuo/claude-codex/tdd-claude-codexgit clone --depth 1 https://github.com/ching-kuo/claude-codexWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/commands/ching-kuo/claude-codex/tdd-claude-codex)<a href="https://agentmods.dev/commands/ching-kuo/claude-codex/tdd-claude-codex"><img src="https://agentmods.dev/badge/commands/ching-kuo/claude-codex/tdd-claude-codex.svg" alt="Measured on agentmods" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00014 | $0.02295 |
| Opus 5 | $0.00007 | $0.01148 |
| Sonnet 5 | $0.00003 | $0.00459 |
| Haiku 4.5 | $0.00001 | $0.00230 |
Grade A, and why
tdd-claude-codex scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 5d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 219 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Deprecated: This command has been converted to a skill (
/tdd-claude-codex). The skill version is recommended for new usage and supports eval-based testing. This command is retained for backward compatibility (model pinning, tool restrictions).
TDD-Claude-Codex — Tests First, Claude Implements, Codex Reviews
$ARGUMENTS
Core Protocols
- TDD Mandate: Tests must be written and verified RED before any implementation begins
- Sovereignty: Claude implements; Codex is the external reviewer — do not skip review
- Stop-Loss: Do not proceed to the next phase until the current phase output is validated
- Language: Use English when calling tools/models; communicate with user in their language
- MCP for review: Always call Codex review via
mcp__codex__codex. Do NOT use Bashcodex review - Context Sanitization: Never pass
.env, secrets, tokens, API keys, or credentials to any external agent or MCP. Exclude files matching.env*,*secret*,*credential*,*.pem,*.key. Redact inline secrets before sending.
Execution Workflow
Phase 0: Read Plan
- If argument is a file path, read it and extract: task description, implementation steps, key files
- If no plan file, use the argument as the task description directly
- Confirm with user before proceeding if key context is missing
Phase 1: Context Retrieval
- Read key files using Read, Glob, Grep (sanitize before passing to agents)
- Identify existing test framework, patterns, test file conventions, and configuration
- Establish test baseline:
- Fast suite (<2 min): run the full suite, record pass/fail counts
- Slow suite (>2 min) or partially failing: run only tests related to the task scope, record known failures
- No test suite: skip — note "no baseline available", regression detection relies on new tests only
- Record
$START_SHAviagit rev-parse HEADfor diff scoping - If worktree is dirty (uncommitted changes), stop and ask the user to commit or stash their changes before running this command. Do not proceed with a dirty worktree.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 5d ago First seen · 219 lines · 14 tokens per session scan A a8478da24ccf
tdd-claude-codex is a command published in the GitHub repository ching-kuo/claude-codex (23 stars, last pushed 5mo ago), licensed MIT. It adds 14 tokens to every session and 2,295 once invoked, about $0.0001 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other commands, from other repositories
tdd-requirements
TDD開発の要件整理を行います。機能要件を明確化し、テスト駆動開発のための準備を行います。.
pm-auto-design2dev
設計レビューから実装完了まで完全自動化(設計レビュー→作業計画→TDD実装).
dashboard
Generar dashboard HTML local con métricas de eficiencia del proyecto SDD.
usage-add
PitWay: Accumulate measured planning or qa token usage onto a milestone.
hub-tdd
TDD workflow for MCP Hub implementation. Types → Tests (red) → Implementation (green) with git gates.
eval
Evaluate and improve one healthcare agent's system prompt. Run up to 5 iterations of: prepare fixed questions -> answer -> judge -> improve -> re-score -> commit if better.