Testing

18,225 mods in this category, of every kind an agent can take. Each one carries what it costs per session, what the scan found, and whether it is the original.

adversarial-verify

483

NUX-Design/claude-skills-fable-opus

Skill Claude CodeCodex

Use after completing any substantive piece of work (code change, analysis, document, configuration, answer to a hard question) and BEFORE presenting it as done. Switches you from author to attacker - you try to refute your own work and only present it if it survives. Do not skip because the work "looks clean"…

not rated 24 2mo ago A 82 tokens

eferro/augmentedcode-skills

Skill Claude CodeCodex

Mutation testing patterns for Python using mutmut. Use when analyzing Python code to find weak or missing tests, verifying pytest effectiveness, strengthening Python test suites, or validating TDD workflows in Python projects.

not rated 24 7mo ago A 43 tokens

@ioloro/ios-testing

486

ioloro/iOS-Testing

Plugin Claude Code

Bundles 1 skill · 129 tokens together

Swift Testing, XCTest, and XCUITest skill for writing correct, modern iOS/macOS tests.

not rated 24 1mo ago A tokens not measured original MIT

ci-workflow

487

VilnaCRM-Org/php-service-template

Skill Claude Code

Run the current php-service-template verification stack and fix failures without lowering quality thresholds.

not rated 24 2mo ago A 20 tokens original CC0-1.0

harness-gen

488

specula-org/SysMoBench

Skill Claude CodeCodex

Trace harness generation for SysMoBench. Use when bootstrapping a new task: clone the system into artifacts/ /, instrument it to emit NDJSON traces at the task-required granularity, write a run.sh, and produce INSTRUMENTATION.md. One-time work per task; the resulting harness is reused for every spec evaluation via the…

not rated 24 1mo ago A 79 tokens original Apache-2.0

precision-bio-tests

489

lynnlangit/precision-medicine-mcp

Skill Claude CodeCodex

Expert guide for testing bioinformatics MCP servers. Covers pytest, DRYRUN modes, and PatientOne simulation scenarios.

not rated 24 12d ago A 28 tokens original Apache-2.0

devtestops

490

hexagon-codes/hexclaw-desktop

Skill Claude CodeCodex

A full-process testing and release checklist, written in Chinese, for software work from clarifying requirements through post-release monitoring. DevTestOps means linking development, testing, and operational checks across the whole delivery process.

not rated 24 yesterday A 60 tokens original Apache-2.0

qa-test-skills

492

Kokxi/qa-test-skills

Plugin Claude Code

Bundles 49 skills · 6,196 tokens together

A collection of 49 skills for planning, designing, running, and assessing software tests, with guidance intended for both AI agents and people new to testing. It covers testing methods across areas such as web, mobile, desktop, and small-program applications.

not rated 24 +1 changed 2d ago A tokens not measured original MIT

AyobamiH/openclaw-operator

Skill Claude CodeCodex

Bootstrap a new realtime eval folder inside this cookbook repo by choosing the right harness from examples/evals/realtimeevals, scaffolding prompt/tools/data files, generating a useful README, and validating it with smoke, full eval, and test runs. Use when a user wants to start a new crawl, walk, or run realtime eval…

not rated 24 +1 4d ago A 76 tokens

sage-eval-gauge

494

MythicAgents/sage

Skill Claude CodeCodex

Repo-local Sage eval-gauge / hill-climbing Phase-0 toolkit. Use when an operator, Claude Code, or Codex needs to measure Sage capability with a GROUND-TRUTHED gauge (not substring eval scores), run the Gate Experiment (does the eval track reality?), compare bare-model vs harness, get the noise floor /…

not rated 24 +3 11d ago A 108 tokens GPL-3.0

presentations

495

brightwave-inc/tidebreak

Skill Claude CodeCodex

Build PowerPoint (PPTX) decks — pptxgenjs when Node is present, python-pptx when not; edit existing decks in place via OOXML — with visual QA before delivery.

not rated 24 +2 2d ago A 44 tokens original Apache-2.0

api-testing

496

fishzjp/qa-skills

Skill Claude CodeCodex

API testing checks a software service directly through its endpoints, using OpenAPI or Swagger documentation or automated test cases. It can produce and run scripts that test requests and responses.

not rated 24 +2 changed 3d ago A 117 tokens original MIT

tovimx/maestro-mobile-testing-skill

Skill Claude CodeCodex

Maestro mobile E2E testing patterns for React Native/Expo apps: YAML test flows, testID selectors, adaptive auth state, optimistic update verification, GraalJS scripting, cross-platform stability, CI/CD integration, Maestro Cloud, and MCP server integration.

not rated 23 7mo ago C 56 tokens

Veridise/audithub-skills

Skill Codex

Coordinate full or multi-stage AuditHub OrCa fuzzing campaigns for Solidity projects. Use when a user asks to run an autonomous OrCa campaign, plan and execute setup/tuning/spec/long-run loops, resume an OrCa campaign, or coordinate target selection, deployment setup, smoke runs, callmetrics.json analysis, tuning, [V]…

not rated 23 19d ago A 89 tokens

tdd-execute-codex

499

ching-kuo/claude-codex

Skill Claude CodeCodex

Full TDD with smart routing: Claude writes tests first, Codex audits tests, then routes implementation by size (Claude small / Codex large), code-reviewer reviews. Best for: TDD on larger tasks where Codex should handle heavy implementation. Triggers on: /tdd-execute-codex, TDD with routing, test-driven execute, TDD…

not rated 23 5mo ago A 86 tokens original MIT

adacore

500

AdaCore/skills

Plugin Claude Code

Bundles 5 skills · 419 tokens together

AdaCore agent skills for Ada and SPARK development: Alire (package management), GNATprove (SPARK formal verification), GNATfuzz (coverage-guided fuzzing), GNATtest (unit testing), and GNATdoc (documentation generation).

not rated 23 1mo ago A tokens not measured original Apache-2.0

rgr

501

kingbootoshi/rgr

Plugin Claude Code

Bundles 2 skills · 248 tokens together

Strict Red-Green-Refactor proof for coding agents.

not rated 23 3mo ago A tokens not measured

autoresearch

502

lunchpaillola/lola-opencode

Skill Claude CodeCodex

Autonomously optimize any Claude Code skill by running it repeatedly, scoring outputs against binary evals, mutating the prompt, and keeping improvements. Based on Karpathy's autoresearch methodology. Use when: optimize this skill, improve this skill, run autoresearch on, make this skill better, self-improve skill…

not rated 23 1mo ago A 100 tokens

danialhasan/ticket-swarm-workflow

Skill Codex

Use when Squad agent behavior, prompts, tool manifests, eval traces, trajectories, datasets, or applied-AI harness changes need to be engineered through source traces, golden/regression/challenge/canary cases, scorers or judges, and regression receipts before product claims or implementation closeout.

not rated 23 3mo ago A 65 tokens

rudder

504

RudderCode/Rudder

Plugin Claude Code

Bundles 6 skills, 2 hooks · 491 tokens together

Maintain a device-local spec from coding-session intent, then generate focused unit tests with your existing coding agent.

not rated 23 3d ago A tokens not measured original Apache-2.0

At most 3 mods per repository are shown here, and a mod shipped inside a plugin is left to that plugin's page — the rest are on their repository pages: