test-desiderata

A test-review guide based on Kent Beck’s Test Desiderata, a set of 12 qualities that make automated tests more useful.

In plain words
What is it for?
Reviewing unit and integration tests, spotting quality problems, and suggesting specific changes to make tests clearer and more reliable.
Why use it?
It helps find tests that are fragile, unclear, dependent on order, or too tied to implementation details. It also helps prioritize practical improvements.

Skill for Claude CodeCodex

Part of the stepwise-core plugin — 13 skills, 5 agents shipped together

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add skills/nikeyes/stepwise-dev/test-desiderata
Any agent
npx skills add nikeyes/stepwise-dev --skill test-desiderata
Clone the repo
git clone --depth 1 https://github.com/nikeyes/stepwise-dev

Made for: Claude Code, Codex.

Or install stepwise-core, the plugin that ships this one along with the rest of its 13 skills, 5 agents.

Per session 64 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 1,849 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00064 $0.01849
Opus 5 $0.00032 $0.00924
Sonnet 5 $0.00013 $0.00370
Haiku 4.5 $0.00006 $0.00185

Measured 3d ago against content hash f1891bef9b6a, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

test-desiderata scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 3d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

core/skills/test-desiderata/SKILL.md · 217 lines

How it starts

The opening of the file, as written. The whole thing — 217 lines — stays where its author put it; the contents beside it link to each section on GitHub.

Test Desiderata

Analyze and improve tests using Kent Beck's Test Desiderata framework - 12 properties that make tests more valuable.

Attribution: All Test Desiderata concepts and principles are created by Kent Beck. Original content: https://testdesiderata.com/ and https://medium.com/@kentbeck_7670/test-desiderata-94150638a4b3

Analysis Workflow

When analyzing tests:

  1. Read the test code - Understand what's being tested and how
  2. Evaluate against principles - Assess each relevant Test Desiderata property
  3. Identify tradeoffs - Note where properties conflict or support each other
  4. Prioritize improvements - Focus on high-impact issues first
  5. Suggest specific changes - Provide concrete, actionable recommendations

The 12 Test Desiderata Properties

These properties make tests more valuable. Some support each other, some interfere, and sometimes properties only seem to interfere (that's where design improvements help).

1. Isolated

Tests return the same results regardless of execution order. Tests don't depend on shared state, previous test results, or external ordering.

Issues to detect:

  • Shared mutable state between tests
  • Tests that must run in specific order
  • Setup/teardown that affects other tests
  • Database state dependencies

2. Composable

Test different dimensions of variability separately and combine results. Break complex scenarios into independent, reusable test components.

Issues to detect:

  • Monolithic tests covering multiple concerns
  • Inability to test dimensions independently
  • Duplicated test setup across related tests
  • Tests that can't be combined or reused

3. Deterministic

If nothing changes, test results don't change. No randomness, timing dependencies, or environmental variations.

Issues to detect:

  • Random data generation
  • Time-dependent assertions
  • Flaky tests that pass/fail intermittently
  • Network or external service dependencies

4. Fast

Tests run quickly, enabling frequent execution during development.

Read the full file on GitHub · 217 lines

Files

What ships with it

1 file beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 3d ago First seen · 217 lines · 64 tokens per session scan A f1891bef9b6a

Subscribe to this mod's changes

test-desiderata is a skill published in the GitHub repository nikeyes/stepwise-dev (24 stars, last pushed 14d ago), licensed Apache-2.0. It adds 64 tokens to every session and 1,849 once invoked, about $0.0003 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.