Testing

18,481 mods in this category, of every kind an agent can take. Each one carries what it costs per session, what the scan found, and whether it is the original.

stm

1033

mdsohaib/screenshot-time-machine

Skill Claude Code

Screenshot every page of the running localhost dev server and report which pages changed since the last snapshot. Use after editing anything users can see (pages, components, CSS/Tailwind, layouts, templates) to visually verify before saying you're done, and when the user says "check the UI", "does it look right"…

not rated 4 7d ago A 112 tokens original MIT

exosuit

1034

joris887/exosuit

Plugin Claude Code

Bundles 44 skills, 1 command, 9 agents, 10 hooks · 1,703 tokens together

Suit up. Build anything. A full engineering organization for Claude Code — deep idea interrogation, TDD sprints, enforced quality gates. 43 skills, 8 agents, 13 hooks, path-scoped rules.

not rated 4 23d ago A tokens not measured original MIT

eval-coach

1035

BayramAnnakov/eval-coach

Plugin Claude Code

Bundles 1 skill · 20 tokens together

AI evaluation strategy design assistant using Evaluation-Driven Development (EDD).

not rated 4 8mo ago A tokens not measured original MIT

looptimal

1036

Renn-Labs/Looptimal

Plugin Claude Code

Bundles 1 skill · 505 tokens together

Turns an objective into a delivered, VERIFIED OUTCOME: frames a hash-pinned sealed acceptance suite, designs the right loop (the loop-design wizard, formerly LoopPrint), war-games it forward, executes with dynamic domain-expert sub-agents (maker != checker), and gates completion on a separate verifier re-running the…

not rated 4 4d ago A tokens not measured original MIT

agent-engineering

1037

endorphin-ai/hasbrains-agent-kit

Plugin Claude Code

Bundles 6 skills · 1,051 tokens together

Engineering discipline for AI agent systems — four always-on principles, test-driven development, verification-before-completion, system-architecture playbooks, a scope-disciplined product manager, and docs/-native project management.

not rated 4 yesterday A tokens not measured original MIT

persona-ux-test

1039

Ericwong5021/persona-ux-test

Skill Claude CodeCodex

Run an isolated persona-based UX test through the real desktop or browser UI. Creates a precise non-developer user persona, gives the tester only an approved product introduction, prevents source-code and design-document leakage, and produces an evidence-based Chinese evaluation report. Use when the user asks for…

not rated 4 1mo ago A 106 tokens original MIT

red-handed

1041

sjh9714/red-handed

Plugin Claude Code

Bundles 2 skills, 2 hooks · 164 tokens together

Your agent said the tests pass. This checks that a test actually ran. Installing turns on two hooks: one strips the tail or grep that eats test results before a test command runs, one audits the session against git the moment it ends and warns only when something was caught. It also adds /red-handed:audit and…

not rated 4 1mo ago A tokens not measured original MIT

faf

1042

Wolfe-Jam/faf-skills

Plugin Claude Code

Bundles 7 skills · 532 tokens together

Claude Code skills for AI-context, testing, and MCP development. IANA-registered format (application/vnd.faf+yaml). Create .faf project DNA, score AI-readiness, sync with CLAUDE.md, build MCP servers, generate test suites.

not rated 4 24d ago A tokens not measured original MIT

devtest

1043

AILiteracyLab/Claude-Build-Test-Loop

Skill Claude CodeCodex

Spec-Build-Test loop — the user defines a spec, then three agents iterate (Builder implements, Tester validates, Supervisor monitors for freezes) until the result matches. Works for any digital function — UI, APIs, CLI tools, conversational AI, data pipelines, and more.

not rated 4 6mo ago A 58 tokens original MIT

agent-evidence-capture

1044

xiexie-qiuligao/agent-evidence-mcp

Skill Claude CodeCodex

Capture screenshots, short recordings, and milestone evidence during long-running agent tasks. Use when an agent is asked to perform multi-step browser, desktop, QA, troubleshooting, deployment, or admin workflows where the user wants checkpoint artifacts, progress evidence, error snapshots, or a final timeline of…

not rated 4 4mo ago A 65 tokens original MIT

rag-eval

1045

LucasSantana-Dev/hitgate

Skill Claude CodeCodex needs its repo

Run the retrieval regression gate against the current repo state and report whether a recent change helped, hurt, or held steady.

not rated 4 1mo ago A 0 tokens original MIT

jacoco

1046

alexmond/jhelm

Skill Claude Code

Check JaCoCo code coverage for jhelm modules.

not rated 4 5d ago A 13 tokens original Apache-2.0

Driftya/code-meridian

Skill Codex

Plan focused tests with CodeMeridian by finding relevant test shields, coverage gaps, impacted behavior, and the smallest useful test set before implementation.

not rated 4 2d ago A 35 tokens original MIT

unreal-playtest-agent

1048

dcc-mcp/dcc-mcp-unreal

Skill Claude Code

Domain skill - run screenshot-light PIE playtest episodes with structured entity observations, bounded semantic actions, transition polling, and in-memory traces for QA and external policy or RL runners.

not rated 4 changed 7d ago A 41 tokens

mcplab-assistant

1049

inspectr-hq/mcplab

Skill Claude CodeCodex

Operator guide for MCPLab config authoring, Test Case Assistant workflows, execution, and result analysis. Use when users need to create or refine test cases from runs/traces, suggest deterministic checks or value capture, write or debug MCPLab eval YAML, run or queue evaluations, troubleshoot failures, or compare…

not rated 4 changed 7d ago A 71 tokens original

looper-qa

1050

quangdang46/looper_rust

Skill Claude CodeCodex

Use when a Looper-managed GitHub repo needs scheduled pre-merge QA — a PR carries the looper:qa label, the spec stage reaches looper:spec-ready, or the Looper reviewer loop requests an independent second pass. Runs the full QA cycle (Looper state probe → PR checkout → ffs code review → language-specific test suite →…

not rated 4 24d ago A 125 tokens original MIT

coco-delivery

1051

pcopu/coco

Skill Claude CodeCodex

Implement and verify CoCo features end-to-end (Telegram commands, callbacks, app-server transport, queueing, watchdogs, approvals, and tests). Use when changing this repository's bot behavior and needing repo-specific file targets, workflows, and validation commands. NOT for generic Python tasks outside CoCo.

not rated 4 9d ago A 65 tokens original MIT

run-qa-gate

1052

Espenandreass1/agentslice

Skill Claude CodeCodex

Independently verify a formal PR with risk-based QA and record the decision in the PR.

not rated 4 changed 5d ago A 25 tokens original MIT

bench-contribute

1053

elbalen/skopus

Command Claude Code

Generate benchmark scenarios from your real corrections for the Correction-Persistence dataset.

not rated 4 4mo ago A 13 tokens original MIT

opentester

1054

kznr02/OpenTester-Skills

Skill Claude Code

Automated testing execution using OpenTester DSL. Use when the user wants to create tests, run tests, validate test syntax, or manage test projects. Supports CLI testing with a YAML-based DSL.

not rated 4 6mo ago A 43 tokens original MIT

pux-ci-context

1055

Blackrose-blackhat/pux

Cursor rule Cursor

Before diagnosing, explaining, or fixing a CI, build, test, or deployment failure, read .ai-context/failure.md when it exists. It contains the latest captured CI incident and relevant repository context.

not rated 4 1mo ago A 107 tokens original MIT

push_rules

1056

leaveanest/slack-utils-channel

Cursor rule Cursor needs its repo

A project rule requiring checks before code is committed and pushed. For this Deno project, the required checks are formatting, linting, and the full test suite; new functions also need documentation and tests.

not rated 4 2mo ago A 2,210 tokens original MIT

At most 3 mods per repository are shown here, and a mod shipped inside a plugin is left to that plugin's page — the rest are on their repository pages: