Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add skills/ching-kuo/claude-codex/tdd-execute-codexnpx skills add ching-kuo/claude-codex --skill tdd-execute-codexgit clone --depth 1 https://github.com/ching-kuo/claude-codexWhat it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00086 | $0.04527 |
| Opus 5 | $0.00043 | $0.02263 |
| Sonnet 5 | $0.00017 | $0.00905 |
| Haiku 4.5 | $0.00009 | $0.00453 |
Grade A, and why
tdd-execute-codex scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 3d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 331 lines — stays where its author put it; the contents beside it link to each section on GitHub.
TDD-Execute-Codex — Tests First, Smart Routing, Claude Reviews
Input
The user's message after the trigger is either:
- A file path to a plan (e.g.
.claude/plan/feature.md) — read it and extract task description, implementation steps, key files - A direct task description — use as-is
Confirm with user before proceeding if key context is missing.
Model Recommendation
Works with claude-sonnet-4-6 (default). No model override needed — routing handles complexity automatically.
Tool Requirements
Advisory checklist — ensure these tools are available:
AskUserQuestion— confirm missing contextmcp__codex__codexandmcp__codex__codex-reply— Codex test audit and Route B implementationTask— dispatch subagents (tdd-guide, code-reviewer, general-purpose)Read,Glob,Grep— context retrievalWrite,Edit— write tests and implement (Route A)Bash— run tests/lint/git commands
Core Protocols
- TDD Mandate: Tests must be written and verified RED before any implementation begins
- Code Sovereignty: Claude is the final authority — review and approve all changes before delivery
- Stop-Loss: Do not proceed to next phase until current phase output is validated
- Language: Use English when calling tools/models; communicate with user in their language
- Test Ownership: Claude owns all test files. Codex never modifies test files.
- Context Sanitization: Never pass
.env, secrets, tokens, API keys, or credentials to any external agent or MCP. Exclude files matching.env*,*secret*,*credential*,*.pem,*.key. Redact inline secrets before sending.
Execution Workflow
Phase 0: Read Plan
- If argument is a file path, read it and extract: task description, implementation steps, key files
- If no plan file, use the argument as the task description directly
- Confirm with user before proceeding if key context is missing
Phase 1: Context Retrieval
- Read key files using Read, Glob, Grep (sanitize before passing to agents)
- Identify existing test framework, patterns, test file conventions, and configuration
- Establish test baseline:
- Fast suite (<2 min): run the full suite, record pass/fail counts
- Slow suite (>2 min) or partially failing: run only tests related to the task scope, record known failures
- No test suite: skip — note "no baseline available", regression detection relies on new tests only
- Record
$START_SHAviagit rev-parse HEADfor diff scoping - If worktree is dirty (uncommitted changes), stop and ask the user to commit or stash their changes before running this skill. Do not proceed with a dirty worktree.
What ships with it
3 files beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 3d ago First seen · 331 lines · 86 tokens per session scan A e10ee4651897
tdd-execute-codex is a skill published in the GitHub repository ching-kuo/claude-codex (23 stars, last pushed 5mo ago), licensed MIT. It adds 86 tokens to every session and 4,527 once invoked, about $0.0004 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other skills, from other repositories
agent-integration
Run all three agent integration phases sequentially: research, write-tests, and implement using E2E-first TDD (unit tests written last). For individual phases, use /agent-integration:research, /agent-integration:write-tests, or /agent-integration:implement. Use when the user says "integrate agent", "add agent…
engram-testing-coverage
TDD and coverage standards for Engram. Trigger: When implementing behavior changes in any package.
code-assist
Guides implementation of code tasks using test-driven development in an Explore, Plan, Code, Commit workflow. Acts as a Technical Implementation Partner and TDD Coach — following existing patterns, avoiding over-engineering, and producing idiomatic, modern code.
tdd
Test-driven development. Use when the user wants to build features or fix bugs test-first, mentions "red-green-refactor", or wants integration tests.
css-design-tdd
Test-driven CSS design system modifications. Run checks before/after CSS changes to verify token usage, variable definitions, fallbacks, and consistency. Use when modifying CSS tokens, fixing design inconsistencies, or auditing CSS architecture.
mobiai-mobile-tdd
You MUST use this before writing any implementation code for a mobile feature, bug fix, refactor, or behavior change. Tests come before implementation — no exceptions.