Testing

18,346 mods in this category, of every kind an agent can take. Each one carries what it costs per session, what the scan found, and whether it is the original.

antrieb

673

jade-pico/antrieb-mcp-server

MCP server Claude CodeCodexCursor +2

Validates AI infra code on real VMs. Self-corrects until it works. No containers, no sandboxes. Remote server at antrieb.sh.

not rated 12 4mo ago A tokens not measured original Apache-2.0

verify-cli

674

NimbleBrainInc/mpak

Skill Claude Code

Run smoke tests against the mpak CLI to verify all commands work correctly before publishing a new release.

not rated 12 1mo ago A 23 tokens

plan-validator

676

sonomirco/agents-and-commands

Agent Claude Code

Use this agent when you have created a plan (e.g., implementation plan, architecture design, refactoring strategy, feature specification) and need to validate and iteratively improve it before execution. This agent should be invoked:\n\n- After drafting any significant technical plan that will guide implementation…

not rated 12 6mo ago A 0 tokens original Apache-2.0

quality-check-command

677

Kazuki-tam/next-stage

Command Claude CodeCursor

This command performs comprehensive code quality checks. Use it before commits or when implementation is complete.

not rated 12 1mo ago A 0 tokens original MIT

ai_tdd_workflow

678

LuthienResearch/luthien_control

Cursor rule Cursor

This rule outlines the strict Test-Driven Development process to be followed by the AI assistant when implementing new features or fixing bugs. This complements the broader developmentworkflow by providing specific TDD execution steps for the AI.

not rated 12 9mo ago A 0 tokens

captain-obvious

679

shmulc8/captain-obvious

Plugin Claude Code

Bundles 1 skill, 1 hook · 203 tokens together

Deterministic scanner that finds and deletes tests that can never fail — assertions the type checker already guarantees, tautologies, mock-echo tests, dead/swallowed assertions, and duplicates. TypeScript (Jest/Vitest/bun:test) and Python (pytest + mypy).

not rated 12 23d ago A tokens not measured original MIT

harmonyos-dev-mcp

680

Deslord319/harmonyos-dev-mcp

MCP server Claude CodeCodexCursor +2

HarmonyOS MCP service for device automation, app deployment, UI interaction, E2E support, and log validation. Runs locally from the harmonyos-dev-mcp Python package.

not rated 12 1mo ago A tokens not measured

bagisto-theme-testing

681

bagisto/agent-skills

Skill Codex

Audit and prove Bagisto storefront themes with source inspection, ownership mapping, admin-to-storefront mutation tests, and Playwright commerce journeys. Use when checking that visible content is dynamic and merchant-controlled; validating theme customizations, channels, CMS, categories, products, search, filters…

not rated 12 today A 97 tokens

cypress-mcp

682

yashpreetbathla/cypress-mcp

MCP server Claude CodeCodexCursor +2

MCP server for AI-driven Cypress test execution. Run, debug, and iterate on E2E tests directly from your AI agent. Runs locally from the cypress-mcp npm package.

not rated 12 6mo ago A tokens not measured

debug-e2e-workflow

683

djscheuf/agentic-dev-ecosystem-template

Skill Claude Code needs its repo

Complete E2E test debugging workflow (composite orchestrator). Starts by reviewing the provided test failure evidence, then forms hypotheses, applies fixes, and verifies results for a presumed E2E playwright test suite.

not rated 12 today A SkillSpector: pass 49 tokens original MIT

swagger-mcp

684

amrsa1/swagger-mcp

MCP server Claude CodeCodexCursor +2

MCP Server for Swagger/OpenAPI documentation and API testing. Runs locally from the swagger-mcp npm package.

not rated 12 1y ago A tokens not measured copy · 100% MIT

thkt/dotclaude

Skill Claude Code

A test-first coding process based on TDD, or test-driven development: write a failing test, make it pass, then improve the code.

not rated 12 today A SkillSpector: warn 21 tokens original MIT

add-test

686

open-metadata/ai-sdk

Skill Claude Code

Use when adding unit or integration tests. Provides test patterns, naming conventions, and fixtures for Python (pytest), TypeScript (vitest), Java (JUnit/Mockito), and Rust.

not rated 12 +1 yesterday A 40 tokens

e2e-test-conventions

687

agentmantis/test-skills

Skill Claude Code

Core conventions and rules for Playwright E2E testing with TypeScript. Covers project structure, naming conventions, selector strategy, authentication, navigation, environment configuration, test independence, and parallelism. Automatically loaded when writing or modifying E2E tests. Use when: generating E2E tests…

not rated 12 5mo ago A 84 tokens original MIT

gedd-chat

688

aws-samples/sample-GEDD

Command Claude Code ✓ vendor

You are a GEDD coaching assistant. You guide the user through building a golden evaluation dataset for their AI agent using Open Coding methodology, then help them evaluate and annotate responses — all without leaving Claude Code.

not rated 12 1mo ago A 0 tokens MIT-0

full-development

689

Deepank308/hermes-swe-agent

Skill Claude CodeCodex

End-to-end workflow for feature requests, enhancements, refactors, and tasks. Covers planning, TDD implementation, verification, integration testing, and PR creation.

not rated 12 5mo ago A 35 tokens original MIT

relentless-tester

690

readysettech/rdst

Skill Claude Code needs its repo

Autonomous QA tester that systematically tests every rdst command via tmux harness, applies a quality rubric, and files bugs in beads.

not rated 12 changed 4d ago A SkillSpector: warn 32 tokens original MIT

pdlc

691

kanfu-panda/pdlc-skills

Plugin Claude Code

Bundles 38 skills · 617 tokens together

Turn AI software engineering into an auditable, on-disk state machine. A staged PDLC workflow (PRD, design, TDD, implement, review, ship, retro) enforces hard contracts — artifacts on disk, per-feature state machine, tests-before-code, objective checks from real command exit codes, single-shot auto-repair — so AI work.

not rated 12 +1 changed 4d ago A tokens not measured original MIT

e2e-skills

692

voidmatcha/e2e-skills

Plugin Claude Code

Bundles 6 skills, 2 agents · 1,019 tokens together

Four agent skills for Playwright and Cypress end-to-end tests: generate new Playwright coverage with live exploration only on local/disposable or externally isolated approved non-production targets, review existing specs or PR diffs, and debug failed runs.

not rated 12 changed 5d ago A tokens not measured original Apache-2.0

manta-e2e-smoke

693

antoinedc/MantaUI

Skill Claude CodeCodex

Drive MANTA's built Electron app in a real renderer context (Playwright's electron launcher) and assert that key UI surfaces render correctly — no crash, no blank screen, sidebar/chat/terminal present. Load BEFORE marking any frontend/UI task done, and when manta-pr-workflow or manta-handle-reviewer-return verifies a…

not rated 12 +3 2d ago A SkillSpector: pass 102 tokens original MIT

prove-it

694

Pablo-aps/prove-it

Skill Codex

Adversarially verify claims that code, fixes, tests, CI, deployments, logs, or systems are correct, complete, healthy, or safe to merge. Use when asked to prove, verify, validate, confirm, double-check, review the agent's own work, check whether a bug is actually fixed, or decide whether green signals justify a…

not rated 11 +4 21d ago A SkillSpector: pass 120 tokens original Apache-2.0

forge

696

zjio26/forge

Plugin Claude Code

Bundles 1 skill, 4 agents · 93 tokens together

Multi-agent collaborative development workflow: Planner → Dev → Test → Learn with automatic bug-fix loops, test layering, and experience accumulation.

not rated 11 4mo ago A tokens not measured original MIT

At most 3 mods per repository are shown here, and a mod shipped inside a plugin is left to that plugin's page — the rest are on their repository pages: