Testing

18,481 mods in this category, of every kind an agent can take. Each one carries what it costs per session, what the scan found, and whether it is the original.

bruno-mcp

1081

jackmulligan-ire/bruno-mcp

MCP server Claude CodeCodexCursor +2

MCP server for Bruno API collections. Runs locally from the bruno-mcp Python package.

not rated 4 5mo ago A tokens not measured

ingpoc/ui-test-generation-mcp

MCP server Claude CodeCodexCursor +2

MCP server "playwright-test-generator-mcp" as configured in ingpoc/ui-test-generation-mcp. Runs locally from the playwright-test-generator-mcp npm package.

not rated 4 1y ago A tokens not measured

podium marketplace

1083

hoainho/podium-mcp

Plugin Claude Code

Lists 1 plugin

Podium — the mobile + canvas automation MCP for AI agents: iOS & Android control, native UI automation, Maestro E2E, React Native debugging, and a no-vision canvas/WebGL brain that taps game UIs like DOM elements — for Claude Code.

not rated 4 2mo ago A tokens not measured original MIT

coding-standards

1084

corbat-tech/coding-standards-mcp

MCP server Claude CodeCodexCursor +2

AI coding standards that enforce production-grade code with DDD, SOLID, TDD guardrails. Runs locally from the @corbat-tech/coding-standards-mcp npm package.

not rated 4 5d ago A tokens not measured original MIT

setup-skills-evals

1085

ahnafyy/skills-evals

Skill Claude CodeCodex

Set up the skills-evals library in a repository — discover agent artifacts, interview the user about what to test, scaffold eval cases, and wire CI and local runners. Use when the user wants to set up skills-evals, test their agent skills, add evals for skills or custom agents, check why a skill isn't triggering, or…

not rated 4 1mo ago A 87 tokens original MIT

ui-validation

1086

Dallionking/ui-validation-kit

Skill Claude Code

Validate UIs by clicking through them like a real user — iOS Simulator, Android emulator, tvOS, desktop, and web. Detects the platform and picks the right tool — agent-device (Callstack) for mobile/TV/desktop, agent-browser (Vercel Labs) for web, Maestro + Maestro Viewer for declarative cross-platform flows; raw xcrun…

not rated 4 2mo ago A 200 tokens original MIT

autoevolve-worker

1087

RightNow-AI/autoevolve

Skill Claude CodeCodex

Use this skill whenever asked to join an autoevolve run, evolve code toward a measured target, work an evolution population, mutate a candidate under EVOLVE-BLOCK rules, or report measured autoevolve progress and artifacts.

not rated 4 +1 1mo ago A 50 tokens original Apache-2.0

qa-test-agent

1088

Yaomeng1749/claude-code-digital-nomad-config-pack

Agent Claude Code

Use this agent when you need to validate software behavior, derive test cases from specifications, create reproducible bug reports, confirm fixes, or distinguish real failures from setup issues. This agent focuses on user-visible correctness, regressions, and edge cases.\n\nExamples:\n\n- user: "I just implemented the…

not rated 4 4mo ago A 408 tokens

mattpocock-skills

1089

baleen37/bstack

Plugin Claude Code

Bundles 26 skills · 896 tokens together

Matt Pocock's agent skills for real engineering: grilling, spec/ticket flows, TDD, code review, domain modelling and more. Plug-and-play, not vibe coding.

not rated 4 changed 2d ago A tokens not measured original MIT

tracegraph

1090

wayfarer-ai/tracegraph

MCP server Claude CodeCodexCursor +2

The graph your agent actually follows — induce a behavioral spec from agent traces, then check, diff, and gate against it. Runs locally from the tracegraph npm package.

not rated 3 1mo ago A tokens not measured copy · 78% Apache-2.0

a11ychk

1091

IsaacEryn/a11ychk

Plugin Claude Code

Bundles 2 skills, 1 MCP server · 201 tokens together

A plugin for checking website accessibility against WCAG 2.2 AA and KWCAG 2.2, two sets of guidelines for making websites usable by people with disabilities. It includes scanning, remediation guidance, and repeat checks.

not rated 3 today A tokens not measured AGPL-3.0

evals

1092

homemade-software-inc/completion-kit

MCP server Claude CodeCodexCursor +2

Prompt evals over MCP: run a prompt on your dataset, score each output 1-5 with an LLM judge. Remote server at completionkit.com.

not rated 3 7d ago A tokens not measured

forkmind

1093

Medhovarsh/forkmind

Plugin Claude Code

Bundles 2 skills, 1 command, 2 agents, 1 MCP server · 487 tokens together

Local-first LLM branching, debugging & context offloading. Treat AI context windows like a Git repo — capture, branch, regression-test LLM calls as a DAG, and archive context as encrypted, restorable capsules. Teaches Claude when and how to drive ForkMind.

not rated 3 24d ago A tokens not measured original MIT

docs

1094

SpecLeft/specleft

Skill Claude CodeCodex

Skill "docs" from SpecLeft/specleft, covering specleft cli reference, setup, workflow, quick checks and safety.

not rated 3 4mo ago A 0 tokens original Apache-2.0

TestAtlas.Mcp

1095

Karzone/TestAtlas

MCP server Claude CodeCodexCursor +2

MCP server "TestAtlas.Mcp" as configured in Karzone/TestAtlas. Launched with TestAtlas.Mcp.

not rated 3 6d ago A tokens not measured original MIT

dsh-verify

1096

263311487-ux/dsh-verify

MCP server Claude CodeCodexCursor +2

Agent-built web app quality gate. Real browser is the judge. PASS/FAIL with receipts. Runs locally from the dsh-verify npm package.

not rated 3 yesterday A tokens not measured original MIT

agent-gate

1097

Jott2121/agent-gate

MCP server Claude CodeCodexCursor +2

An MCP server that lets an AI agent gate its own work: deterministic checks, refute-first review, and tamper-evident honest receipts. Fleet Mode, as a tool. Runs locally from the mcp-agent-gate Python package.

not rated 3 24d ago A tokens not measured original MIT

linkinator-mcp

1098

JustinBeckwith/linkinator-mcp

MCP server Claude CodeCodexCursor +2

MCP server for link checking using linkinator. Runs locally from the linkinator-mcp npm package.

not rated 3 yesterday A tokens not measured original MIT

mushi-mushi

1099

kensaurus/mushi-mushi

Skill Claude CodeCodex

Set up, configure, and use Mushi Mushi — the AI-powered QA platform for automatic bug detection, user story mapping, TDD scenario generation, and PDCA auto-improvement. Use when setting up Mushi, configuring SDK/CLI/MCP, managing API keys, or asking how any Mushi feature works.

not rated 3 yesterday A 70 tokens original MIT

e2e-runner

1100

fastslack/mtw-e2e-runner

Plugin Claude Code

Bundles 1 skill, 4 commands, 1 MCP server · 76 tokens together

JSON-driven E2E browser test runner — no JavaScript test files needed. Parallel execution against a Chrome pool (browserless, CDP, Lightpanda, Obscura, Steel), 28+ built-in actions, visual verification, network debugging, flaky test detection, reusable modules, and a real-time dashboard. Includes 3 specialized agents.

not rated 3 3mo ago A tokens not measured original Apache-2.0

kilotest

1101

jrpool/kilotest

MCP server Claude CodeCodexCursor +2

Ensemble testing of web pages for accessibility, usability, and standards conformity. Remote server at kilotest.com.

not rated 3 today A tokens not measured original MIT

allure-testops-mcp

1102

mshegolev/allure-testops-mcp

MCP server Claude CodeCodexCursor +2

MCP server for Allure TestOps — projects, launches, test cases, test results, defect categories. Runs locally from the allure-testops-mcp Python package. Needs 3 environment variables to run.

not rated 3 1mo ago A tokens not measured original MIT

benchmark-adder

1103

tmuskal/arc-agi-benchmarker

Skill Claude Code

Given a benchmark repository URL (agentic envs like arc-agi, memory/eval benchmarks like longmemeval, QA/code/tool-use benchmarks, etc.), orchestrate the creation of a full Claude Code plugin that benchmarks the current harness setup against it. Wraps the babysitter:babysit skill with the benchmark-plugin-creator…

not rated 3 4mo ago A 74 tokens

bughunterpro

1104

sector-b79/web-hunter-pro

Skill Claude CodeCodex

Operate only in authorized scope: bug bounty targets explicitly in scope, owned systems, defensive reviews, or labs. Decline or pause on requests involving unauthorized access, stealth, persistence, service disruption, credential abuse, real data theft, or abuse of third-party systems.

not rated 3 4mo ago A 0 tokens

At most 3 mods per repository are shown here, and a mod shipped inside a plugin is left to that plugin's page — the rest are on their repository pages: