Testing

18,481 mods in this category, of every kind an agent can take. Each one carries what it costs per session, what the scan found, and whether it is the original.

testforge

1201

whaojie797-design/Novera-AI-skills

Skill Claude CodeCodex

A test-generation tool for Python that creates runnable pytest tests. Pytest is a Python tool for writing and running automated tests.

not rated 2 1mo ago A 73 tokens

portable-skill

1202

daijx-ai/skillwitness

Skill Claude CodeCodex

Validate a synthetic Agent Skill when testing SkillWitness locally or in CI.

not rated 2 2mo ago A 18 tokens original MIT

rot-dtd-goal

1203

Nova-Violet-Role/RoT-DTD-GOAL

Plugin Claude Code

Bundles 9 commands, 8 agents, 31 hooks · 533 tokens together

A /goal engine whose completion is EARNED, not announced. Acceptance criteria are real shell commands; the Stop hook re-runs every one of them and only exits 0 lets a goal finish. Verification defends itself: a sealed integrity ledger re-hashes each criterion before the gate trusts it, a negative-control red team…

not rated 2 19d ago A tokens not measured

caliper

1204

zhengbowenai-cmd/caliper

Skill Claude Code

Statistically test whether a prompt or SKILL.md change is actually better than the old version. Use when the user asks to compare two prompts, A/B test a prompt change, check if a recent edit really improved things, find rule conflicts in a long prompt, or identify which sections of a prompt are pulling weight.…

not rated 2 2mo ago A 93 tokens original MIT

tester

1205

erikfiala/e2e-tester

Skill Cursor

Run end-to-end web quality audits with Playwright and Lighthouse using existing project scripts first. Use when the user says /tester, e2e test, smoke test, Playwright audit, Lighthouse audit, performance audit, accessibility audit, dark mode audit, mobile audit, or asks for 100/100 Lighthouse improvement guidance.

not rated 2 5mo ago A 67 tokens original MIT

mcp-server-evaluations

1206

mcp-com-ai/mcp-server-evaluations-skills

Skill Claude Code

Test MCP servers for quality and reliability. Verify tool functionality, test error handling, generate tests, and assess response quality with no dependencies other than curl. Use this when validating MCP server implementations, testing OpenAPI-to-MCP conversions, or assessing API tool quality.

not rated 2 7mo ago A 59 tokens original MIT

evolve-skill

1207

taneltaluri/evolve-skill

Skill Claude CodeCodex

Evolve Skill: measurement-first skill optimizer. Evaluates SKILL.md files against an anchored 9-dimension rubric, validates that the rubric itself is stable (test-retest), optimizes with a hill-climbing loop that only accepts improvements larger than measurement noise, protects against overfitting with train/holdout…

not rated 2 4mo ago A 128 tokens original MIT

skill-check

1208

JckJhns/skill-check

Skill Claude Code

Comprehensive testing and validation of Claude skills. Use this skill whenever the user wants to test, validate, audit, or quality-check a skill — whether they say "test my skill", "check this skill works", "validate my skill", "run skill-check", or anything similar. Also trigger when the user asks things like "does…

not rated 2 4mo ago A 192 tokens original MIT

kimtth/azure-ml-finetuning-eval-skills

Skill Claude CodeCodex

Generate synthetic and simulated datasets for evaluation and fine-tuning using Azure AI Foundry simulators. Create non-adversarial task data, adversarial safety data, and conversation datasets without manual data collection.

not rated 2 9mo ago A 48 tokens

rigor

1210

olimxonuz0-lab/rigor

Skill Claude CodeCodex

Use for any non-trivial coding, engineering, or deliverable-producing task — building a feature, fixing a bug, refactoring, designing an architecture, writing a script someone will run, or drafting a document someone will use. Enforces upfront planning before acting, rejects placeholder/stub/TODO code and unhandled…

not rated 2 25d ago A 181 tokens original MIT

agent-reliability

1212

ByteStack-Labs/claude-plugins

Plugin Claude Code

Bundles 4 skills · 933 tokens together

Claude skills for AI agent and ML reliability: reproduce the eval-to-production gap, catch confidently-wrong outputs, and prove root cause with verified numbers. Start with production-autopsy.

not rated 2 2mo ago A tokens not measured original MIT

cantrips

1213

toverux/cantrips

Plugin Claude Code

Bundles 24 skills · 803 tokens together

The core engineering loop for coding agents (Claude Code, Codex CLI): grill, spec, tickets, implement with TDD at agreed seams, review, commit, plus a user-gated compound step that turns session learnings into durable project memory. Basic spells a caster always has prepared.

not rated 2 changed 7d ago A tokens not measured original MIT

browser-test-executor

1214

dangnhit/qa-tester

Skill Claude CodeCodex

Execute approved bounded browser Test DSL cases with fresh isolated contexts and auditable attempts. Use when running browser tests, reruns, regression checks, or blocked execution diagnostics.

not rated 2 1mo ago A 38 tokens original Apache-2.0

Coverage Guard

1215

rolecraft-sh/skills

Skill Claude CodeCodex

Use when the user wants to check test coverage, enforce 100% coverage, find uncovered code, add missing tests, or increase code coverage. Works with vitest, jest, react-scripts, and other test runners. Also for 'coverage', 'test coverage', 'cover', 'untested', 'uncovered', 'add tests for', 'increase coverage'…

not rated 2 26d ago A 90 tokens original MIT archived

pg-skill-forge

1216

gaoguo/pg-skill-forge

Skill Claude Code

A toolkit for creating, testing, and improving skills for the opencode coding agent. A skill is a reusable set of instructions for a particular workflow.

not rated 2 2mo ago A 111 tokens original MIT

tcr

1217

xpepper/tcr-skill

Skill Claude CodeCodex

Guide users through TCR (Test && Commit || Revert), TCRDD, and git-gamble workflows. ALWAYS trigger when a user mentions TCR, TCRDD, "test commit revert", "git gamble", or "git-gamble". Trigger when a user wants to combine TDD with automatic commits/reverts, enforce baby steps via a commit-or-revert loop, or asks…

not rated 2 5mo ago A 164 tokens original MIT

qa-testing-kit

1218

SoftwareOneHN/qa-testing-kit

Skill Claude CodeCodex

Lifecycle-driven QA workflow — from requirements analysis to test reports, with state tracking and impact analysis.

not rated 2 3mo ago A 5 tokens original MIT

pw-playwright-fieldkit

1219

jpbaking/playwright-fieldkit

Skill Codex

Explore, debug, audit, compare, record, and test live websites with deterministic Playwright scripts and QE workflows. Use for requests to map a site, find broken pages or links, reproduce browser bugs, discover hidden or role-gated features, audit accessibility/performance, compare crawls, design or review test cases…

not rated 2 1mo ago A 143 tokens original MIT

mx-generate

1220

Claritune/mutantx

Skill Claude Code

MutantX Phase 2 — Generate realistic code mutants as unified diff patches.

not rated 2 1mo ago A 20 tokens original MIT

ios-dev

1221

AlphaSquadTech/ios-dev

Skill Claude Code

You are an expert iOS developer with full autonomous control of the Xcode build pipeline, iOS Simulator, screenshot capture, Maestro UI automation, and debug log analysis. Follow these procedures exactly.

not rated 2 6mo ago B 95 tokens original MIT

cloud

1222

ShiplightAI/agent-skills

Skill Claude CodeCodex

Sync local tests with Shiplight cloud — push and pull YAML test cases, templates, and functions between your repo and the cloud. Requires a Shiplight cloud subscription.

not rated 2 2mo ago A Socket: warnSnyk: pass 35 tokens original MIT archived

playwright-cli

1223

sonofmagic/skills

Skill Claude Code

Automate browser interactions, test web pages and work with Playwright tests.

not rated 2 changed 8d ago A 19 tokens copy · 100% MIT

proofrag

1224

unshDee/proofrag

Plugin Claude Code

Bundles 1 skill, 1 command · 117 tokens together

Evaluate a RAG/LLM app: generate a golden set from your docs, run LLM-as-judge + retrieval metrics, and produce a shareable HTML scorecard with a CI gate.

not rated 2 1mo ago A tokens not measured original MIT

At most 3 mods per repository are shown here, and a mod shipped inside a plugin is left to that plugin's page — the rest are on their repository pages: