testing

A set of general guidelines for writing and reviewing software tests. It covers test structure, useful cases, mocking, and how to think about coverage; TDD means writing tests as part of the development process.

In plain words
What is it for?
Use it when adding tests, practising TDD, reviewing coverage, or changing code that should be tested.
Why use it?
It helps avoid tests that only check implementation details or miss important inputs, errors, and edge cases.

Skill for Claude CodeCodex

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add skills/johanthoren/jeff/testing
Any agent
npx skills add johanthoren/jeff --skill testing
Clone the repo
git clone --depth 1 https://github.com/johanthoren/jeff

Made for: Claude Code, Codex.

Per session 79 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 1,258 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00079 $0.01258
Opus 5 $0.00039 $0.00629
Sonnet 5 $0.00016 $0.00252
Haiku 4.5 $0.00008 $0.00126

Measured 2d ago against content hash cdf8fcfb06a8, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

testing scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

skills/testing/SKILL.md · 62 lines

How it starts

The opening of the file, as written. The whole thing — 62 lines — stays where its author put it; the contents beside it link to each section on GitHub.

Testing Standards

Language-agnostic defaults. Language-specific skills override where they conflict.

Golden Rule: If you can't test it easily, refactor it.

Structure and Naming

  • Arrange, act, assert: set up the data, execute the code, verify the result. One behavior per test.
  • Names state the expectation: validateEmail returns false for invalid format, not it works or test user.

What to Test

  • DO: happy path; edge cases (boundaries, empty, null); error cases (invalid input, failures); business logic; public APIs.
  • DON'T: third-party libraries; framework internals; simple getters/setters; private implementation details.

Coverage

Coverage is an outcome of testing the right intersections, not a target to chase. Not every line, and not every one-liner, needs a test; a one-line pass-through does not. Cover each behavior and its real edges; a task's acceptance criteria are the floor.

Principles

  • Test behavior, not implementation: focus on what, not how. Assert results, never that a mock was called with particular arguments: expect(mock).toHaveBeenCalledWith(...) pins procedure and goes red on a behavior-preserving refactor.
  • Would it survive a refactor? If a behavior-preserving refactor would turn the test red, it tests procedure; rewrite it to assert the result. Procedure-coupled tests lock in design flaws.
  • Mock through injected boundaries: dependency injection makes a hand-rolled fake ({ findById: () => fixture }) sufficient; no framework magic required.
  • Independent tests: no shared state, any order, one assertion's worth of behavior each.
  • Fast and reliable: run tests frequently; fix failures immediately.

Change-Detector Tests Are a Banned Smell

A change-detector test asserts a value or call shape that no consumer observes, so it goes red on any edit to that value and catches no regression the edit would not. It is configuration duplicated into the test as a second place to edit. Do not write one.

Read the full file on GitHub · 62 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 2d ago First seen · 62 lines · 79 tokens per session scan A cdf8fcfb06a8

Subscribe to this mod's changes

testing is a skill published in the GitHub repository johanthoren/jeff (4 stars, last pushed 5d ago), licensed Apache-2.0. It adds 79 tokens to every session and 1,258 once invoked, about $0.0004 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.