mutation-test-safety

mutation-test-safety is a skill for Claude Code, Codex from Grinv/mal-mcp. It costs 51 tokens per session (491 once invoked), scanned A, original, MIT.

A safety procedure for testing tools that change real account data. It requires checking the current state, making the smallest change, and verifying or undoing it.

In plain words
What is it for?
Use it before adding, updating, removing, deleting, posting, following, or favouriting data through a live account.
Why use it?
It reduces the risk of accidentally changing the wrong data or making a larger change than intended during live tests.

Skill for Claude CodeCodex

Written for no agent in particular: nothing here depends on one. Also seen: installed under .agents/ (shared by several agents).

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add skills/grinv/mal-mcp/mutation-test-safety
Any agent
npx skills add Grinv/mal-mcp --skill mutation-test-safety
Clone the repo
git clone --depth 1 https://github.com/Grinv/mal-mcp

Made for: Claude Code, Codex.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for mutation-test-safety

README.md
[![agentmods](https://agentmods.dev/badge/skills/grinv/mal-mcp/mutation-test-safety.svg)](https://agentmods.dev/skills/grinv/mal-mcp/mutation-test-safety)
Your own site
<a href="https://agentmods.dev/skills/grinv/mal-mcp/mutation-test-safety"><img src="https://agentmods.dev/badge/skills/grinv/mal-mcp/mutation-test-safety.svg" alt="Measured on agentmods" height="20"></a>
Per session 51 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 491 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5.1 $0.00051 $0.00491
Opus 5 $0.00026 $0.00246
Sonnet 5 $0.00010 $0.00098
Haiku 4.5 $0.00005 $0.00049

Measured 6d ago against content hash 83101f0a32ff, method: parsed. Prices are Anthropic first-party input rates as of 2026-09-06, from the pricing page.

Security

Grade A, and why

mutation-test-safety scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 6d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

.agents/skills/mutation-test-safety/SKILL.md · 37 lines

What it actually says

Mutation testing safety contract

Applies to any mutation tool (add/update/remove/delete/post/follow/favourite/ etc.) once the user has given explicit go-ahead to test it against a real account. For every mutation call:

  1. Capture the exact pre-state first (via the matching read tool) — not just an assumption of what it probably is. A target with no existing state has a clear pre-state too: "absent." This applies to every call to a mutation tool, including one you intend only as a validation-only probe — if the schema isn't .strict(), a bogus/typo'd field is silently dropped rather than rejected, and any other real field in the same call still applies for real.
  2. Make the smallest possible change that still exercises the behavior (e.g. one field, not a full rewrite). Never combine an unknown/invalid probe field with a real valid field in one call — if the probe field is ignored instead of rejected, the valid field mutates the account for real. Test unknown-param rejection with no other fields set (or on a read-only tool) instead.
  3. Verify the change landed by re-fetching via a read tool — a mutation's own echoed response is not always trustworthy (some tools have historically omitted fields they actually changed).
  4. Revert to the captured pre-state immediately, in the same turn, and verify the revert too. Don't batch several mutations and revert at the end — revert each one before moving to the next unrelated test.
  5. Never leave the target in a different state than you found it, even if a step errors partway through — check and clean up regardless.
  6. Never touch an uninvolved third party. Self-targeted mutations (a self-message, a self-created and immediately-deleted test post/ thread/comment) are fine; acting on a random other real user/account is not.
Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 6d ago First seen · 37 lines · 51 tokens per session scan A 83101f0a32ff

Subscribe to this mod's changes

mutation-test-safety is a skill published in the GitHub repository Grinv/mal-mcp (2 stars, last pushed 12d ago), licensed MIT. It adds 51 tokens to every session and 491 once invoked, about $0.0003 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.

Related

Other skills, from other repositories

live-audit

Audit anilist-mcp-server — build/test/lint gate, live MCP tool edge-case sweep (input validation, not-found paths, mutations with capture/revert), source-level code review, and docs/metadata consistency. Use when asked to test/audit the published or just-fixed anilist-mcp-server package, hunt for bugs/edge cases, or…

Grinv/anilist-mcp-server · 85 tokens

tool-description-check

Self-check a new or edited MCP tool description/field .describe() text before committing — verify every behavioral claim against live testing, check for contradictions with sibling tools, and score against Glama's Tool Definition Quality Score (TDQS) rubric. Use whenever a tool description or schema field description…

Grinv/anilist-mcp-server · 74 tokens

release

Cut a release of anilist-mcp-server — draft CHANGELOG entries, check docs/metadata consistency, then bump/tag/push. Use when asked to release, cut a version, or publish a new version of this package.

Grinv/anilist-mcp-server · 48 tokens

fixture-accuracy-check

Make sure a mocked-fetch test fixture mirrors AniList's real GraphQL response shape, not just whatever fields make the current code pass. Use before writing or changing a fixture in src/tests/.test.ts.

Grinv/anilist-mcp-server · 48 tokens

docs-consistency-check

Check README/manifest.json/server.json/CHANGELOG.md/AGENTS.md and docs/.md for drift against the actual registered tools and source. Use after adding, renaming, or removing a tool, or as part of a live-audit pass.

Grinv/anilist-mcp-server · 56 tokens

mutation-test-safety

The capture-state/smallest-change/verify/revert/verify-revert contract for live-testing a mutation tool against a real account. Use any time you're about to call a mutation tool live, not just during a full audit.

Grinv/anilist-mcp-server · 51 tokens