Testing

18,225 mods in this category, of every kind an agent can take. Each one carries what it costs per session, what the scan found, and whether it is the original.

jamjet-labs/jamjet

Skill Claude CodeCodex

Production durability patterns for AI agents on the JVM — crash recovery, audit trails, human-in-the-loop, and replay testing using JamJet runtime.

not rated 21 7d ago A 36 tokens original Apache-2.0

eval

530

simple-agent-lab/AutoTrainess

Skill Claude CodeCodex

Use when evaluating a model on the current benchmark(s).

not rated 21 1mo ago A 13 tokens original MIT

edwardyap90/counterfactual-engineering-skill

Skill Codex

Use when a code change has multiple plausible implementation paths, high regression risk, uncertain root cause, or unclear tradeoffs. Creates isolated candidate implementations, runs comparable verification across each candidate, compares evidence, and applies the best proven approach while preserving user changes.

not rated 21 3mo ago A 56 tokens original MIT

yaleh/meta-cc

Agent Claude Code

Build automated validation tools to enforce API conventions at scale ensuring consistency without manual checks for Bootstrap-006.

not rated 21 17d ago A 24 tokens original MIT

tdd

533

kirillgreen/skills

Skill Claude CodeCodex

Spec-Driven TDD with a strict Red-Green-Refactor cycle using context-isolated subagents. Every feature starts with a structured spec; tests are generated from numbered acceptance criteria, with post-cycle three-dimension verification (Completeness + Traceability + Coherence), severity-tiered findings…

not rated 21 4d ago A SkillSpector: warn 170 tokens original CC0-1.0

tester

534

dkp-consult/dev-with-ai

Agent Claude Code

Conçoit, audite et écrit les tests d'une étape ou d'un module, à partir de la spec. Met aussi en place l'outillage de test d'un projet en mission "--setup". N'écrit jamais de code source — un test rouge est rapporté, jamais contourné.

not rated 21 3d ago A 63 tokens

androjack

535

VIKAS9793/AndroJack-mcp

MCP server Claude CodeCodexCursor +2

MCP server that validates AI-generated Android code against official SDK documentation. Runs locally from the androjack-mcp npm package.

not rated 21 +1 yesterday A tokens not measured

unity-coding-skills

536

nowsprinting/unity-coding-skills

Plugin Claude Code

Bundles 10 skills, 3 agents · 1,154 tokens together

Skills and subagents for developing Unity projects with Claude Code — maintainable test design and implementation, test-first workflow, coding guidelines, scene editing, and more.

not rated 21 8d ago A tokens not measured original Unlicense

test-case-design

537

91160/skills

Skill Claude CodeCodex

A functional test-case designer that turns approved requirements or design documents into structured QA test cases. QA, or quality assurance, is the process of checking that software works as intended.

not rated 21 +1 3mo ago A 366 tokens

playwright-test

538

s-hiraoku/vscode-sidebar-terminal

MCP server Claude CodeCodexCursor +2

A high-level API to automate web browsers. Runs locally from the playwright npm package.

not rated 21 2d ago A tokens not measured copy · 100% MIT

test-writer

539

oro-ad/nuxt-claude-devtools

Agent Claude Code

Creates Vitest unit and component tests. Use when writing tests for components or functions.

not rated 20 7mo ago A 21 tokens GPL-3.0

collect-evidence

542

currents-dev/currents-mcp

Skill Claude CodeCodex

Show that work you implemented actually works, or demo it, using artifacts from tests running in CI via Currents — before/after screenshots, text and JSON attachments, videos, traces, and GIFs. Use when asked to "collect evidence", "prove it works", "show me it works", "demo the feature", "capture a before/after", or…

not rated 20 4d ago A SkillSpector: warn 117 tokens original Apache-2.0

webdriver-management

543

tugkanboz/awesome-cursorrules

Cursor rule Cursor

Selenium WebDriver lifecycle management — driver factory, explicit waits, browser options, and teardown.

not rated 20 3d ago A 3,012 tokens original MIT

mnvsk97/agentbreak

Skill Codex

Orchestrates end-to-end resilience testing for LLM agents with AgentBreak, including LLM infrastructure failures, prompt injection, agent skill supply-chain risk, guardrail verification, and MCP server/tool failures. Use when the user asks to "test my agent for resilience", "chaos test this agent", "find failure modes…

not rated 20 3mo ago A 97 tokens original MIT

junjo-evaluation

546

mdrideout/junjo

Skill Codex

Turn a developer's product-quality objective into a complete Junjo Studio-backed evaluation workflow. Use when a coding agent needs to establish a baseline, design or generate typed cases, run or resume application-owned Node, Workflow, or Agent targets, inspect outcomes and exact trace evidence, compare a candidate…

not rated 20 7d ago A SkillSpector: pass 83 tokens original Apache-2.0

ctf-qa-validation

547

mr-pmillz/gogatoz

Skill Codex

QA testing and validation of GoGatoZ features against the local GoGatoZ CTF lab. Invoke for post-change testing, live flag validation, lab infrastructure checks, payload smoke tests, enumerate/attack/search/pivot/notify validation, or any request to confirm that GoGatoZ still works.

not rated 20 6d ago A 68 tokens

pxp

548

infews/pxp_skill

Plugin Claude Code

Bundles 3 skills · 180 tokens together

Pivotal Labs–style XP pairing workflow: epic refinement, story slicing, EARS specifications, and strict per-spec TDD with developer approval at every gate.

not rated 20 2mo ago A tokens not measured original MIT

selenium

549

abarrac/mcp-selenium

MCP server Claude CodeCodexCursor +2

MCP server "selenium" as configured in abarrac/mcp-selenium. Launched with java -jar ~/.mcp-selenium/mcp-selenium.jar. Needs 3 environment variables to run.

not rated 20 1y ago A tokens not measured original MIT

oci-skills

551

acedergren/oci-agent-skills

Plugin Claude Code

Bundles 13 skills · 687 tokens together

Complete OCI automation skills covering compute, networking, databases, monitoring, secrets, GenAI, IAM, IaC, FinOps, and best practices.

not rated 20 +1 4mo ago A tokens not measured original MIT

user-testing-agent

552

ncklrs/claude-chrome-user-testing

Plugin Claude Code

Bundles 14 skills, 10 commands · 392 tokens together

Persona-based user testing agent that simulates realistic user interactions with web applications. Embodies different user archetypes (Boomers, Millennials, Gen Z, Gen Alpha) with authentic behaviors, timing patterns, and frustration triggers to identify UX issues before real users do.

not rated 19 8mo ago A tokens not measured original MIT

At most 3 mods per repository are shown here, and a mod shipped inside a plugin is left to that plugin's page — the rest are on their repository pages: