Testing

26,146 mods in this category, of every kind an agent can take. Each one carries what it costs per session, what the scan found, and whether it is the original.

llm-testing

241

Eyadkelleh/awesome-skills-security

Skill Claude CodeCodex

Comprehensive LLM security testing prompts for bias detection, data leakage, alignment testing, and adversarial prompt resistance.

not rated 376 +3 2mo ago A 27 tokens

tae0y/real-estate-mcp

Instructions file Claude Code

Instructions for tae0y/real-estate-mcp, covering claude.md, project overview, commands, run a single test file and run a single test by name.

not rated 373 1mo ago A 1,480 tokens original MIT

currents-dev/playwright-best-practices-skill

Skill Claude CodeCodex

Use when writing Playwright tests, fixing flaky tests, debugging failures, implementing Page Object Model, configuring CI/CD, optimizing performance, mocking APIs, handling authentication or OAuth, testing accessibility (axe-core), file uploads/downloads, date/time mocking, WebSockets, geolocation, permissions…

not rated 373 +4 1mo ago A 214 tokens original MIT

self-iteration

244

anomalyco/browser-control

Skill OpenCode

Dogfood a project-owned tool in a realistic user flow and evaluate the agent experience. Use when the user asks to dogfood or self-iterate, or when testing or evaluating a tool built by this project.

not rated 372 +3 3d ago A 46 tokens original MIT

LambdaTest/agent-skills

Skill Claude CodeCodex

Automatically generate comprehensive test cases from API definitions, endpoint descriptions, OpenAPI/Swagger specs, Postman collections, or raw HTTP request/response examples. Use this skill whenever the user mentions generating tests from APIs, writing test cases for REST endpoints, API testing, creating test suites…

not rated 366 +2 1mo ago A 182 tokens original MIT

test-repl

246

AlexGladkov/claude-in-mobile

Command Claude Code

Run a REPL test scenario against the claude-in-mobile REPL plugin (python/node/bash/...).

not rated 364 +1 5d ago A 21 tokens

leeguooooo/claude-code-usage-bar

Instructions file CodexOpenCode

AGENTS.md instructions for leeguooooo/claude-code-usage-bar, covering repository guidelines, project structure & module organization, build, test, and development commands, coding style & naming conventions and testing guidelines.

not rated 364 +4 yesterday A 636 tokens original MIT

tile-ai/tilelang-ascend

Skill Claude CodeCodex

A test-design skill for TileLang-Ascend operators, which are custom AI computations for Huawei Ascend chips. It creates or reviews test cases from design documents, example code, user descriptions, or existing tests.

not rated 363 yesterday A 102 tokens original MIT

playwright-cli

249

testdino-hq/playwright-skill

Skill Claude CodeCodex

Automates browser interactions for testing and validating your own web applications using playwright-cli. Use when you need terminal-first browser control for navigation, form filling, screenshots, tracing, bound browser sessions, debugging, or generating Playwright test code. Only use against applications you own or…

not rated 359 +4 2mo ago A 64 tokens original MIT

codspeed AGENTS.md

250

CodSpeedHQ/codspeed

Instructions file CodexOpenCode

AGENTS.md instructions for CodSpeedHQ/codspeed, covering agents.md, common development commands, building and testing, build the project and build in release mode.

not rated 353 +72 yesterday A 663 tokens original Apache-2.0

add-tests-and-ci

251

THUDM/DeepDive

Skill Claude Code

Guide for adding or updating slime tests and CI wiring. Use when tasks require new test cases, CI registration, test matrix updates, or workflow template changes.

not rated 346 2mo ago A 36 tokens

zkbugs AGENTS.md

252

zksecurity/zkbugs

Instructions file CodexOpenCode

Instructions for zksecurity/zkbugs, covering repository guidelines, project structure & module organization, build, test, and development commands, coding style & naming conventions and testing guidelines.

not rated 345 17d ago A 997 tokens original MIT

make-e2e-live

253

receptron/mulmoclaude

Skill Claude Code

A development guide for adding or improving real-LLM end-to-end tests. End-to-end tests exercise a complete feature path, while a real-LLM test uses an actual language model rather than a fake response.

not rated 345 +1 yesterday A 103 tokens original MIT

aminueza/terraform-provider-minio

Instructions file CodexOpenCode

AGENTS.md instructions for aminueza/terraform-provider-minio, covering repository guidelines, quick reference (read first), project structure, build & development commands and run single test.

not rated 342 2d ago A 6,191 tokens AGPL-3.0

systematic-debugging

255

dirge-code/dirge

Skill Claude CodeCodex

Use when encountering any bug, test failure, or unexpected behavior, before proposing fixes — find the root cause first instead of guessing at symptom patches.

not rated 340 +1 3d ago A 34 tokens GPL-3.0

benchflow-ai/benchflow

Skill Codex

Review Benchflow or SkillsBench task-run trajectories and integration-test Benchflow code changes. Use this skill whenever the user asks to audit traj health, failed or timed-out runs, healthy pass/fail/timeout status, no-skill leakage, skill loading, reward hacking, verifier isolation, metadata completeness, token…

not rated 340 +1 yesterday A 99 tokens original Apache-2.0

blazediff

257

teimurjan/blazediff

Skill Claude CodeCodex

Run, author, or update BlazeDiff visual regression tests. Trigger on "visual test", "screenshot regression", "blazediff", "/blazediff".

not rated 338 +1 7d ago A 37 tokens original MIT

huxtable AGENTS.md

258

hughjonesd/huxtable

Instructions file CodexOpenCode

AGENTS.md instructions for hughjonesd/huxtable, covering instructions to llm agents, documentation, writing code, tests and docs, testing your work and installing extra components.

not rated 332 5d ago A 834 tokens

flounder AGENTS.md

259

adshao/flounder

Instructions file CodexOpenCode

AGENTS.md instructions for adshao/flounder, covering agents.md, project defaults, thin-layer agentic mode, map → dig deep coverage and operating the deep audit.

not rated 330 +1 2d ago A 6,358 tokens AGPL-3.0

statewave AGENTS.md

260

smaramwbc/statewave

Instructions file CodexOpenCode

AGENTS.md instructions for smaramwbc/statewave, covering agents.md — guide for contributors and coding agents, what this repo is, setup, build, test, conventions and pull requests.

not rated 322 +1 2d ago A 933 tokens original Apache-2.0

verify

261

Housetan218/claude-code-haha

Skill Claude CodeCodex

Verify a code change does what it should by running the app.

not rated 321 5mo ago A 15 tokens

testing-preview

262

boringcomputers/nehemiah

Skill Claude CodeCodex

Test the preview proxy feature end-to-end. Use when verifying preview URL changes, auth changes on the web proxy route, or networking-related fixes.

not rated 320 +9 25d ago B 32 tokens original Apache-2.0

tester

263

IncomeStreamSurfer/claude-code-agents-wizard-v2

Agent Claude Code

Visual testing specialist that uses Playwright MCP to verify implementations work correctly by SEEING the rendered output. Use immediately after the coder agent completes an implementation.

not rated 318 +1 9mo ago A 32 tokens

style

264

Ayanami1314/swe-pruner

Cursor rule Cursor

Cursor rule "style" from Ayanami1314/swe-pruner, covering style guide, bad, good, test style and bad.

not rated 316 +2 2mo ago A 409 tokens original MIT

At most 3 mods per repository are shown here, and a mod shipped inside a plugin is left to that plugin's page — the rest are on their repository pages: