tdd-cycle

A test-first coding workflow for one small behaviour change: write a test that fails, make it pass, then refactor. It first agrees on the public boundary being tested, such as a function, web route, or command-line interface.

In plain words
What is it for?
Adding a feature, fixing a bug, or writing code with TDD (test-driven development). It is intended for work outside a larger build process that already defines the change.
Why use it?
It gives a clear definition of done and helps ensure tests check real behaviour rather than implementation details. Working on one slice at a time keeps the change focused.

Skill for Claude CodeCodex

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add skills/raisedadead/claude-code-plugins/tdd-cycle
Any agent
npx skills add raisedadead/claude-code-plugins --skill tdd-cycle
Clone the repo
git clone --depth 1 https://github.com/raisedadead/claude-code-plugins

Made for: Claude Code, Codex.

Per session 91 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 683 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00091 $0.00683
Opus 5 $0.00046 $0.00342
Sonnet 5 $0.00018 $0.00137
Haiku 4.5 $0.00009 $0.00068

Measured 2d ago against content hash ce4a6d2f4abc, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

tdd-cycle scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.

The scan reads SKILL.md. This mod also ships 1 executable file (scripts/run_slice.sh), listed below but not scanned — reading those needs a real analyzer, not pattern matching.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

plugins/whetstone/skills/tdd-cycle/SKILL.md · 58 lines

How it starts

The opening of the file, as written. The whole thing — 58 lines — stays where its author put it; the contents beside it link to each section on GitHub.

tdd-cycle — one slice, red before green

Test-first for ad-hoc work that no ledger is driving. Agree the seam, watch the test fail, make it pass, keep the suite green. Refactoring is a separate pass.

When to use

  • Adding a feature or fixing a bug where the test should define "done".
  • Any behaviour-bearing edit outside a dossier:build / ck:build covenant. Under a covenant, dossier:build drives WHEN and WHAT and composes this skill's run_slice.sh as its RED/GREEN proof (dossier ADAPTERS §whetstone); the §T row already fixes the seam, so the interview is skipped. One ledger drives one edit.

The seam (agree it first)

Before the first assertion, name the seam: the public boundary you test through — a function signature, an HTTP route, a CLI. Get the user's agreement on it. Assert behaviour through that boundary, so the test survives a refactor and fails when the behaviour is wrong. A wrong seam buys tests that pass over broken behaviour, or break on every rename. Highest-value moment in the loop; spend it.

The loop (one slice at a time)

  1. RED — write one failing test. Then prove it fails for the right reason:

    "${CLAUDE_PLUGIN_ROOT}"/skills/tdd-cycle/scripts/run_slice.sh red <test-command>
    

    run_slice red exits non-zero when the test passed — a test that never failed characterises nothing. Check the test against reference/anti-patterns.md before running it.

  2. GREEN — minimum code to pass:

    run_slice.sh green <test-command>
    

    Just enough for this one test.

  3. Full-suite gate:

    run_slice.sh full <suite-command>
    

    The slice leaves everything else green.

  4. Next slice — one seam, one test, one implementation per cycle. Repeat from RED.

Refactor is a separate pass

Once GREEN and the suite passes, the slice is done. Cleanup and design improvement go to a review/simplify step, so a failure during red-green means exactly one thing.

Verification

Read the full file on GitHub · 58 lines

Files

What ships with it

2 files beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 2d ago First seen · 58 lines · 91 tokens per session scan A ce4a6d2f4abc

Subscribe to this mod's changes

tdd-cycle is a skill published in the GitHub repository raisedadead/claude-code-plugins (2 stars, last pushed 7d ago), licensed ISC. It adds 91 tokens to every session and 683 once invoked, about $0.0005 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.

Related

Other skills, from other repositories

agent-integration

Run all three agent integration phases sequentially: research, write-tests, and implement using E2E-first TDD (unit tests written last). For individual phases, use /agent-integration:research, /agent-integration:write-tests, or /agent-integration:implement. Use when the user says "integrate agent", "add agent…

entireio/cli · 89 tokens

engram-testing-coverage

TDD and coverage standards for Engram. Trigger: When implementing behavior changes in any package.

Gentleman-Programming/engram · 25 tokens

code-assist

Guides implementation of code tasks using test-driven development in an Explore, Plan, Code, Commit workflow. Acts as a Technical Implementation Partner and TDD Coach — following existing patterns, avoiding over-engineering, and producing idiomatic, modern code.

mikeyobrien/ralph-orchestrator · 53 tokens

tdd

Test-driven development. Use when the user wants to build features or fix bugs test-first, mentions "red-green-refactor", or wants integration tests.

GreyDGL/PentestGPT · 33 tokens

css-design-tdd

Test-driven CSS design system modifications. Run checks before/after CSS changes to verify token usage, variable definitions, fallbacks, and consistency. Use when modifying CSS tokens, fixing design inconsistencies, or auditing CSS architecture.

xiaolai/vmark · 49 tokens

mobiai-mobile-tdd

You MUST use this before writing any implementation code for a mobile feature, bug fix, refactor, or behavior change. Tests come before implementation — no exceptions.

ArisGuimera/MobiAI-Core · 38 tokens