Testing

18,394 mods in this category, of every kind an agent can take. Each one carries what it costs per session, what the scan found, and whether it is the original.

bpe

817

MasonEgger/bpe-claude-code-plugin

Plugin Claude Code

Bundles 13 skills, 3 agents · 739 tokens together

BPE (Brainstorm-Plan-Execute) development workflow: spec-driven TDD, autonomous multi-step runs with a validation gate, and cross-session continuity.

not rated 7 8d ago A tokens not measured original MIT

beat

818

kirkchen/beat

Plugin Claude Code

Bundles 8 skills, 1 hook · 181 tokens together

Agent-driven BDD workflow using Gherkin feature files. Rhythm for your development: explore → design → plan → apply → verify → archive.

not rated 7 2mo ago A tokens not measured original MIT

ripplo

819

ripplo/claude-plugin

Plugin Claude Code

Bundles 2 skills · 143 tokens together

Ripplo reviews pull requests by driving your app end to end in a real browser and reporting what broke, with the evidence. This plugin connects an app to Ripplo and lets Claude Code fix what a review found.

not rated 7 13d ago A tokens not measured

julia-repl

820

seabbs/claude

Skill Claude CodeCodex

Evaluate Julia through the warm AgentREPL MCP session rather than julia -e, hot-reload edits with Revise, and filter TestItemRunner suites while iterating. Use for any Julia evaluation, package iteration, or test run, and to decide when a fresh process is needed instead.

not rated 7 13d ago A 65 tokens

chaosmesh-mcp-server

821

RadiumGu/Chaosmesh-MCP

MCP server Claude CodeCodexCursor +2

MCP server "chaosmesh-mcp-server" as configured in RadiumGu/Chaosmesh-MCP. Runs locally from the chaosmesh-mcp-server Python package.

not rated 7 5mo ago A tokens not measured

mumez/pharo-agentic-browser

Skill Claude CodeCodex

Generates an AgenticBrowser Scripting DSL orchestration that implements a feature end-to-end — plan (if needed), TDD implementation, tests, and a lint/style-guide review pass — previews it as docs/scripting-features/feature- .scripting.md, and on user approval runs it via st-eval. Use this whenever the user wants to…

not rated 7 +1 changed 6d ago A 224 tokens

implementer

823

LH8PPL/core-memory-kit

Agent Claude Code

Deep implementation work delegated by the lead — writing kit code, writing tests (five-exit-doors discipline), debugging failures, and the implementer self-review pass on a diff. Use for any coding work bigger than a trivial edit. Runs on Opus.

not rated 7 6d ago A 55 tokens original MIT

probe

824

nikzlabs/shipit

Skill Claude CodeCodex

How to run the test-plugin probe and read its report — which field verifies which part of the docs/262 plugin usage contract.

not rated 7 +1 changed 6d ago A 28 tokens original Apache-2.0

jest

825

anivar/jest-skill

Skill Claude Code

Jest best practices, patterns, and API guidance for JavaScript/TypeScript testing. Covers mock design, async testing, matchers, timer mocks, snapshots, module mocking, configuration, and CI optimization. Baseline: jest ^30.0.0. Triggers on: jest imports, describe, it, test, expect, jest.fn, jest.mock, jest.spyOn…

not rated 7 +1 1mo ago A 97 tokens original MIT

saleem-daqa/qa-test-case-generation

Skill Claude CodeCodex

Use when generating manual QA test cases from requirements, BRDs, user stories, acceptance criteria, spreadsheets, live mockup/prototype URLs, uploaded mockups, screenshots, wireframes, or existing test case templates; especially when coverage, deduplication, traceability, validations, permissions, workflows, or edge…

not rated 7 3mo ago A 70 tokens original MIT

jmh

828

umit/skills

Skill Claude CodeCodex

Write Java microbenchmarks with JMH (Java Microbenchmark Harness) that produce trustworthy numbers — not numbers distorted by JIT dead-code elimination, constant folding, insufficient warmup, or single-fork JIT contamination. Use this skill whenever the user writes @Benchmark, mentions JMH, microbenchmark, throughput…

not rated 7 +1 4mo ago A 273 tokens original MIT

mospira/walkforward-audit

Skill Claude CodeCodex

Audits time-based machine learning backtests, walk-forward validation, rolling retraining, forecasting evaluations, and temporal train/test pipelines for data leakage, faulty split logic, invalid feature timing, target leakage, calibration/tuning leakage, and misleading experiment comparisons. Use when agent needs to…

not rated 6 2mo ago A 99 tokens original MIT

agileteam

831

DYAI2025/Plumbline

Command Claude Code

Orchestrate an autonomous, defense-in-depth TDD multi-agent team (requirements → spec-sanity gate → planner → coder/reviewer loop → verification/security/validation/judgment gates → human acceptance → retrospective) to build a feature end-to-end against fully verified, independently validated requirements.

not rated 6 7d ago A 58 tokens

kodama-verification

832

amergrgic/kodama

Skill Claude CodeCodex

Define measurable success criteria and collect targeted test, build, lint, type-check, or smoke-test evidence before claiming work is complete.

not rated 6 today A 31 tokens original MIT

respect-the-oracle

833

chris-short/respect-the-oracle

Skill Claude CodeCodex

Use when working against a test suite, spec, or graded harness you do not own - especially inside an automated loop scored on how many tests pass - and tempted to change the tests, weaken assertions, hardcode expected outputs, special-case inputs, or overfit the visible examples to turn things green.

not rated 6 2mo ago A 65 tokens original MIT

playwright-skill

835

appautomaton/playwright-skill

Skill Claude CodeCodex

Complete browser automation with Playwright. Auto-detects dev servers, writes clean test scripts to $TMPDIR (or /tmp). Test pages, fill forms, take screenshots, check responsive design, validate UX, test login flows, check links, automate any browser task. Use when user wants to test websites, automate browser…

not rated 6 8mo ago A 83 tokens

bug-check

836

yeaight7/agent-powerups

Command Claude Code

Run automated tests and build checks first, then agent code review. For each bug found, propose or document a regression test.

not rated 6 1mo ago A 25 tokens original Apache-2.0

shipproof

837

WhorideChicken/shipproof

Plugin Claude Code

Bundles 1 skill · 125 tokens together

Evidence-driven acceptance gate for AI-generated software changes. Reviews PRs, migrations, and repos with adaptive scope and issues a PASS / CONDITIONALPASS / FAIL / INCONCLUSIVE verdict backed by verifiable evidence.

not rated 6 19d ago A tokens not measured original MIT

cloudml-eval-ops

840

MiaoDX/roboclaws

Skill Codex

Run frozen Roboclaws Eval Harness rows on CloudML with bounded parallelism, official cml lifecycle commands, executor-backed JuiceFS transfer, durable task receipts, verified collection, and explicit retry/preemption evidence. Use when a user asks to run, refresh, resume, monitor, collect, or debug a Roboclaws…

not rated 6 8d ago A SkillSpector: warn 99 tokens original MIT

At most 3 mods per repository are shown here, and a mod shipped inside a plugin is left to that plugin's page — the rest are on their repository pages: