mutation-testing

mutation-testing is a skill for Claude Code, Codex from vmobifystudio/app-dev-team. It costs 83 tokens per session (1,079 once invoked), scanned A, original, MIT.

A testing method that deliberately introduces defects into code or rules to see whether the test suite detects them.

In plain words
What is it for?
It helps evaluate gates, assertions, CI checks, and guard rules by reporting which deliberate changes were caught or survived.
Why use it?
A passing test suite is not useful if its checks cannot fail when the guarded behaviour breaks.

Skill for Claude CodeCodex

Part of the app-dev-team plugin — 32 skills, 27 commands, 30 agents, 2 hooks shipped together

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add skills/vmobifystudio/app-dev-team/mutation-testing
Any agent
npx skills add vmobifystudio/app-dev-team --skill mutation-testing
Clone the repo
git clone --depth 1 https://github.com/vmobifystudio/app-dev-team

Made for: Claude Code, Codex.

Or install app-dev-team, the plugin that ships this one along with the rest of its 32 skills, 27 commands, 30 agents, 2 hooks.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for mutation-testing

README.md
[![agentmods](https://agentmods.dev/badge/skills/vmobifystudio/app-dev-team/mutation-testing.svg)](https://agentmods.dev/skills/vmobifystudio/app-dev-team/mutation-testing)
Your own site
<a href="https://agentmods.dev/skills/vmobifystudio/app-dev-team/mutation-testing"><img src="https://agentmods.dev/badge/skills/vmobifystudio/app-dev-team/mutation-testing.svg" alt="Measured on agentmods" height="20"></a>
Per session 83 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 1,079 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00083 $0.01079
Opus 5 $0.00042 $0.00540
Sonnet 5 $0.00017 $0.00216
Haiku 4.5 $0.00008 $0.00108

Measured yesterday against content hash 77b9915b9654, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

mutation-testing scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured yesterday.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

skills/mutation-testing/SKILL.md · 81 lines

How it starts

The opening of the file, as written. The whole thing — 81 lines — stays where its author put it; the contents beside it link to each section on GitHub.

Mutation testing

defect-hunting §3 says a new rule is not done until you have watched it fail, in three steps. That instruction was followed by hand, when someone remembered. On 2026-07-29 the suite read 385 passed, 0 failed while containing a grep -E with a PCRE lookahead (a syntax error, stderr to /dev/null, || ok every time), two doctor assertions that a demoted gate could not turn red, and a hook that stood down in exactly the incident it was written for.

A green suite is evidence only to the extent its assertions can go red. scripts/mutate.sh is that, executable.

Running it

sh scripts/mutate.sh                # all 16 catalogued mutations (~1 min each — the whole suite runs)
sh scripts/mutate.sh --list         # the catalogue, plus what it cannot test and why
sh scripts/mutate.sh --only M04     # one mutation, while you iterate on an assertion
sh scripts/mutate.sh --sample 4     # what CI runs on every PR

Exit 0 all caught · 1 something SURVIVED · 2 could not run (baseline not green, anchor drifted).

Three verdicts matter:

  • CAUGHT (n assertions) — the gate bites.
  • SURVIVED — the gate can be deleted and the suite stays green. That is a hole, and it is a finding at the same severity as the bug the gate was supposed to catch.
  • CAUGHT, but NOT by the assertion written for it — some unrelated assertion noticed. Not a hole today; the named guard is decorative, and the next refactor that touches the unrelated one takes the coverage away silently. Fix the named assertion.

The rule

A new gate ships with a mutation proving its assertion bites. Adding a check to board-doctor, ship-gate, verify-done, spawn-gate, a hook, or scripts/test.sh is not done until scripts/mutate.sh --only <your-id> prints CAUGHT and names your assertion.

Adding one is four fields in the catalogue at the top of scripts/mutate.sh:

M17@@scripts/your-gate.sh@@<exact text to break>@@<replacement>@@<the test.sh label that must fail>

Read the full file on GitHub · 81 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. yesterday First seen · 81 lines · 83 tokens per session scan A 77b9915b9654

Subscribe to this mod's changes

mutation-testing is a skill published in the GitHub repository vmobifystudio/app-dev-team (4 stars, last pushed 25d ago), licensed MIT. It adds 83 tokens to every session and 1,079 once invoked, about $0.0004 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-09-03.

Related

Other skills, from other repositories

Test-first

Use before implementing a feature or bugfix — write the failing test before the code.

Kilbex/Vigla · 20 tokens

execute-task

Implement one task (or a cohesion bundle) from a signed-off spec (Ready or Active): recompute the execution freshness gate, write the verifying test first, implement to green, run the project's full CI with adaptive retry, converge via the configured reviewsequence (default /polish --nested), then open a draft PR…

inkatze/planwright · 111 tokens

builder

Detect a project's stack and recommend or apply the universal mechanical quality guards from planwright's core catalog (formatter, linters, type-checker, test runner, secret scan, commit hooks, CI gate), plus the growable breadth dimensions. Escalates stake-bearing decisions (auth, data modeling, security posture…

inkatze/planwright · 100 tokens

spec-walkthrough

Render a spec bundle (or a chosen slice) into a plain-language, didactic comprehension artifact a human reads and judges for themselves: an unaided cold read before kickoff, re-orientation mid-execution, or onboarding to a finished or abandoned spec. Standalone and strictly read-only: it renders any status, never…

inkatze/planwright · 100 tokens

drain

Run the on-demand drain pass over every spec bundle's Gate deferral entries: evaluate structured GATE(when:) conditions, surface date and free-text gates, report malformed ones, inventory each live bundle's [manual] test-spec entries, and surface the observations log's unmined state. Read-only; nothing is…

inkatze/planwright · 75 tokens

implementing-tasks

Use when executing a batch of TDD-sized tasks inside a running-an-iteration call — dispatches an implementer subagent per task following red-green-refactor discipline and returns per-task completion status.

prime-radiant-inc/iterative-development · 45 tokens