adversarial-testing

A plugin that creates adversarial tests: tests designed to challenge assumptions and expose bugs rather than only confirm expected behavior.

In plain words
What is it for?
Use it to generate adversarial test cases for software projects, including Python projects using pytest, and to investigate bugs, mutations, and quality gaps.
Why use it?
It helps reveal failures that ordinary, expected-case tests may miss by deliberately testing hostile or unusual inputs and situations.

Plugin for Claude Code

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

Claude Code
/plugin marketplace add jimmc414/claude-code-plugin-marketplace
agentmods
npx agentmods add plugins/jimmc414/claude-code-plugin-marketplace/adversarial-testing
Clone the repo
git clone --depth 1 https://github.com/jimmc414/claude-code-plugin-marketplace

Made for: Claude Code.

Per session not measured What this adds to a session before it is invoked.
When invoked not measured Not applicable: nothing here is loaded into a session.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Security

Grade A, and why

adversarial-testing scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

plugins/adversarial-testing/.claude-plugin/plugin.json · 21 lines

What it actually says

{
  "name": "adversarial-testing",
  "version": "1.0.0",
  "description": "Adversarial test generation that finds real bugs by inverting the reward structure",
  "author": {
    "name": "jimmc414",
    "email": "[email protected]"
  },
  "homepage": "https://github.com/jimmc414/claude-code-plugin-marketplace/tree/main/plugins/adversarial-testing",
  "repository": "https://github.com/jimmc414/claude-code-plugin-marketplace",
  "license": "MIT",
  "keywords": ["testing", "adversarial", "bugs", "mutation", "quality", "pytest"],
  "category": "testing",
  "dependencies": {},
  "commands": "./commands/",
  "agents": "./agents/",
  "skills": "./skills/",
  "hooks": "./hooks/hooks.json",
  "mcpServers": "./.mcp.json"
}
Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 2d ago First seen · 21 lines scan A 6354610c89f6

Subscribe to this mod's changes

adversarial-testing is a plugin published in the GitHub repository jimmc414/claude-code-plugin-marketplace (4 stars, last pushed 3d ago), licensed MIT. Its token cost is not measured: this kind of file is read by the harness, not the model. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.

Related

Other plugins, from other repositories

claude-ops

Claude Code operations toolkit. Twelve skills: audit-skill-visibility (audit whether each installed skill is actually VISIBLE to the model, and diagnose why most of a fleet never gets used — a skill is invisible when its description is dropped by Claude Code's skill-listing context budget, which sheds descriptions…

melodic-software/claude-code-plugins · not measured

autonomy

Governed autonomous agent operation: role-topology, binding-seam, wiring-vs-advisor, telemetry, return-accounting, trigger-dispatch, per-work-class guardrail-matrix, standing-routine-catalog, and design-only runner-charter contracts for climbing the AI-adoption ladder, plus a guided-setup skill that discovers an…

melodic-software/claude-code-plugins · not measured

claude-config

Nine configuration-health skills (plus setup) for a repo's Claude Code configuration: audit (settings.json / .mcp.json / hooks / plugins / permissions drift), audit-automation-gaps (evidence-gated verdicts on automation gaps), audit-permission-grants (allow-rule / allowed-tools grants for auto-mode durability and…

melodic-software/claude-code-plugins · not measured

context-guard

Per-session context-window observability plus the first shipped consumer: a statusline wrapper tees each session's contextwindow fields to a per-session snapshot file, a zone resolver classifies usage into smart/acceptable/dumb bands (percentage bands plus window-class token bands, conservative-min combination, zones.

melodic-software/claude-code-plugins · not measured

context-budget

Measure a Claude Code session's fixed startup context payload per item, on the consumer's machine at a pinned, version-stamped binary — including per-tool attribution of the built-in tool pools that /context reports only as lump sums, derived live by A/B bare-name-deny differencing with enforced comparability rules…

melodic-software/claude-code-plugins · not measured

adhd

Shape and restructure the assistant's output for a reader with ADHD — action-first, low-friction, and digestible. adhd:shape is a standing session posture: lead with the concrete next action, number multi-step work, restate state across turns, cap and rank lists, give concrete time estimates, make wins visible, and…

melodic-software/claude-code-plugins · not measured