Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add ching-kuo/claude-codex --skill tdd-claude-codexgit clone --depth 1 https://github.com/ching-kuo/claude-codexWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/ching-kuo/claude-codex/tdd-claude-codex)<a href="https://agentmods.dev/skills/ching-kuo/claude-codex/tdd-claude-codex"><img src="https://agentmods.dev/badge/skills/ching-kuo/claude-codex/tdd-claude-codex.svg" alt="Measured on agentmods" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00089 | $0.03123 |
| Opus 5 | $0.00044 | $0.01562 |
| Sonnet 5 | $0.00018 | $0.00625 |
| Haiku 4.5 | $0.00009 | $0.00312 |
Grade A, and why
tdd-claude-codex scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 7d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 269 lines — stays where its author put it; the contents beside it link to each section on GitHub.
TDD-Claude-Codex — Tests First, Claude Implements, Codex Reviews
Input
The user's message after the trigger is either:
- A file path to a plan (e.g.
.claude/plan/feature.md) — read it and extract task description, implementation steps, key files - A direct task description — use as-is
Confirm with user before proceeding if key context is missing.
Model Recommendation
Works with claude-sonnet-4-6 (default). Override with /model opus for tasks requiring deep architectural reasoning.
Tool Requirements
Advisory checklist — ensure these tools are available:
AskUserQuestion— confirm missing contextTask— dispatch tdd-guide and implementation subagentsRead,Glob,Grep— context retrievalWrite,Edit— write tests and implementBash— run tests/lint/git commandsmcp__codex__codexandmcp__codex__codex-reply— Codex test audit and implementation review
Core Protocols
- TDD Mandate: Tests must be written and verified RED before any implementation begins
- Sovereignty: Claude implements; Codex is the external reviewer — do not skip review
- Stop-Loss: Do not proceed to the next phase until the current phase output is validated
- Language: Use English when calling tools/models; communicate with user in their language
- MCP for review: Always call Codex review via
mcp__codex__codex. Do NOT use Bashcodex review - Context Sanitization: Never pass
.env, secrets, tokens, API keys, or credentials to any external agent or MCP. Exclude files matching.env*,*secret*,*credential*,*.pem,*.key. Redact inline secrets before sending.
Execution Workflow
Phase 0: Read Plan
- If argument is a file path, read it and extract: task description, implementation steps, key files
- If no plan file, use the argument as the task description directly
- Confirm with user before proceeding if key context is missing
Phase 1: Context Retrieval
- Read key files using Read, Glob, Grep (sanitize before passing to agents)
- Identify existing test framework, patterns, test file conventions, and configuration
- Establish test baseline:
- Fast suite (<2 min): run the full suite, record pass/fail counts
- Slow suite (>2 min) or partially failing: run only tests related to the task scope, record known failures
- No test suite: skip — note "no baseline available", regression detection relies on new tests only
- Record
$START_SHAviagit rev-parse HEADfor diff scoping - If worktree is dirty (uncommitted changes), stop and ask the user to commit or stash their changes before running this skill. Do not proceed with a dirty worktree.
What ships with it
3 files beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 7d ago First seen · 269 lines · 89 tokens per session scan A 4a593c8df004
tdd-claude-codex is a skill published in the GitHub repository ching-kuo/claude-codex (23 stars, last pushed 5mo ago), licensed MIT. It adds 89 tokens to every session and 3,123 once invoked, about $0.0004 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other skills, from other repositories
engram-testing-coverage
TDD and coverage standards for Engram. Trigger: When implementing behavior changes in any package.
nw-fp-clojure
Clojure language-specific patterns, data-first modeling, REPL-driven development, and spec.
strict-tdd
Strict RED->GREEN->REFACTOR test-driven development with enforcement. Never write production code before a failing test. Atomic commits per TDD cycle.
conductor-implement
Execute tasks from a track's implementation plan following TDD workflow.
tdd
This skill should be used when the user wants to implement features or fix bugs using test-driven development. Enforces the RED-GREEN-REFACTOR cycle with vertical slicing, context isolation between test writing and implementation, human checkpoints, and auto-test feedback loops. Uses multi-agent orchestration with the…
mobiai-mobile-tdd
You MUST use this before writing any implementation code for a mobile feature, bug fix, refactor, or behavior change. Tests come before implementation — no exceptions.