Testing

18,324 mods in this category, of every kind an agent can take. Each one carries what it costs per session, what the scan found, and whether it is the original.

build-judge

601

breethomas/bette-think

Skill Claude CodeCodex

Build an LLM-as-Judge evaluator for one specific failure mode. Binary pass/fail only. Use when a failure mode requires interpretation (tone, faithfulness, relevance, completeness) and cannot be checked with code. Do NOT use when the failure can be checked with regex, schema validation, or execution tests. Do NOT use…

not rated 16 6mo ago A 79 tokens

testing

603

aa-io/site

Cursor rule Cursor

This project uses Vitest for testing with React Testing Library for component testing. The testing stack includes.

not rated 16 2mo ago A 2,028 tokens

test

604

mdlmarkham/TailOpsMCP

Command Claude Code

Run tests, identify issues, and generate coverage reports.

not rated 16 8mo ago A 11 tokens

verify

605

meltforce/FreeReps

Skill Claude CodeCodex

Run the checks that have to pass before a FreeReps commit lands — Go build, vet, tests and golangci-lint, the frontend type check and build, and the document contract check. Triggers — "verify", "prüf das durch", "vor dem commit", "läuft das durch", "check before committing", "run the checks", "does CI pass". Not for…

not rated 16 28d ago A 97 tokens original MIT

mcp-tui-test

606

GeorgePearse/mcp-tui-test

MCP server Claude CodeCodexCursor +2

MCP server "mcp-tui-test" as configured in GeorgePearse/mcp-tui-test. Runs locally from the mcp-tui-test Python package.

not rated 16 +1 6mo ago A tokens not measured original MIT

vs-ui-explore

609

dhq-boiler/Unofficial-VS-MCP

Skill Claude CodeCodex

Autonomously explore the UI of a Visual Studio debuggee via vs-mcp UIA tools (uisnapshot, uifindelements, uiwait) and produce a structured bug/coverage report. Use when the user asks to "test all screens", "crawl the UI", "find UI bugs autonomously", or similar.

not rated 16 +2 13d ago A SkillSpector: warn 71 tokens original MIT

sldd

610

soujava/sldd-skills

Skill Claude CodeCodex

Start, resume, inspect, or continue SLDD spec-driven development workflows, including /sldd slash-style commands, gated intent/design/test/implementation steps, structured journals, and workflow kind routing.

not rated 16 +2 2mo ago A 44 tokens CC-BY-4.0

cursor-rules

611

willcoliveira/qualiow-playwright-skills

Cursor rule Cursor

Playwright E2E testing rules for {{PROJECTNAME}}; the full skill lives in .agents/skills/playwright-e2e.

not rated 16 +2 changed today A 423 tokens original MIT

perform-e2e-test

612

DauQuangThanh/hanoi-rainbow

Command Claude Code needs its repo

Execute end-to-end tests based on the E2E test plan and generate detailed test result reports.

not rated 16 +1 7mo ago A 20 tokens original MIT

sextant

614

hellotern/Sextant

Plugin Claude Code

Bundles 13 skills · 1,537 tokens together

Architecture-aware engineering principles framework for Claude Code. Provides systematic, tiered workflows for bug fixes, new features, refactoring, code review, test writing, requirements refinement, debugging, shipping, sprint planning, migrations, and security audits.

not rated 15 5mo ago A tokens not measured original MIT

e2e-playwright

615

burhankhatri/e2e-testing

Skill Claude CodeCodex

Battle-tested Playwright E2E testing patterns for Next.js/React apps. Use when writing, running, debugging, or fixing Playwright tests. Also triggers on 'e2e', 'end-to-end', 'playwright', 'browser test', 'UI test', 'integration test with browser', 'flaky test', 'test keeps failing'. Covers locators, assertions…

not rated 15 1mo ago A 103 tokens

agentv-eval-writer

616

EntityProcess/agentv

Skill Claude CodeCodex needs its repo

Write, edit, review, and validate AgentV EVAL.yaml / .eval.yaml evaluation files. Use when asked to create new eval files, update or fix existing ones, add or remove test cases, configure graders (llm-rubric, script), review whether an eval is correct or complete, convert between EVAL.yaml and evals.json using agentv…

not rated 15 2mo ago A 129 tokens original MIT

prove-it

617

Pablo-aps/prove-it

Skill Codex

Adversarially verify claims that code, fixes, tests, CI, deployments, logs, or systems are correct, complete, healthy, or safe to merge. Use when asked to prove, verify, validate, confirm, double-check, review the agent's own work, check whether a bug is actually fixed, or decide whether green signals justify a…

not rated 15 21d ago A SkillSpector: pass 120 tokens original Apache-2.0

ragscore

620

HZYAI/RagScore

MCP server Claude CodeCodexCursor +2

The Fastest Way to Audit Your RAG - Generate QA datasets & evaluate RAG systems in Colab, Jupyter, or CLI. Privacy-first, any LLM, visual reports. Runs locally from the ragscore Python package. Needs 2 environment variables to run.

not rated 15 3mo ago A tokens not measured original Apache-2.0

bmad-skills

621

bmad-labs/skills

Plugin Claude Code

Bundles 24 skills · 3,714 tokens together

Collection of skills including skill creation, MCP server development, E2E testing, and unit testing guides for TypeScript/NestJS projects.

not rated 15 12d ago A tokens not measured original MIT

barrhawk-premium-e2e

622

barrhawk/barrhawk_premium_e2e_mcp

MCP server Claude CodeCodexCursor +2

MCP server "barrhawk-premium-e2e" as configured in barrhawk/barrhawkpremiume2emcp. Runs locally from the barrhawk-premium-e2e npm package.

not rated 15 1mo ago A tokens not measured

otg-mcp

623

h4ndzdatm0ld/otg-mcp

MCP server Claude CodeCodexCursor +2

Open Traffic Generator - Model Context Protocol. Runs locally from the otg-mcp Python package.

not rated 15 6mo ago A tokens not measured original Apache-2.0

At most 3 mods per repository are shown here, and a mod shipped inside a plugin is left to that plugin's page — the rest are on their repository pages: