Testing

18,217 mods in this category, of every kind an agent can take. Each one carries what it costs per session, what the scan found, and whether it is the original.

altirra-bridge

361

ilmenit/AltirraSDL

Skill Claude CodeCodex

Use this skill when controlling the AltirraSDL Atari emulator programmatically — driving the simulator from a script, capturing screenshots from a running Atari, injecting joystick/keyboard input, frame-stepping for deterministic testing, reading or writing CPU/memory/chip state, setting breakpoints or watchpoints…

not rated 57 5d ago A 190 tokens GPL-2.0

processmission/oh-my-qemu

Skill Codex

Use for QEMU peripheral, accelerator, MMIO, qdev, or SysBusDevice modeling with register contracts, an explicit register framework decision, and qtest-backed verification.

not rated 56 1mo ago A 42 tokens original MIT

mobile-pro-max

364

vurakit/agentup

Skill Claude CodeCodex

Mobile development intelligence. 12 domains, 9 stacks, 250+ entries. Actions: plan, build, create, design, implement, review, fix, improve, optimize, enhance, refactor, check mobile code. Domains: pattern, package, error, perf, test, security, interface, arch, idiom, anti, tooling, dependency. Stacks: flutter…

not rated 56 5mo ago A 116 tokens original MIT

Advance-Technologies-Foundation/clio

Skill Codex

Create, update, and validate custom Creatio Configuration Web Services and their tests. Use when users need to expose custom backend endpoints in Creatio, wire service contracts/implementations, return structured success/error results, or verify endpoints with integration-style and unit tests.

not rated 56 today A 58 tokens

js

366

r33drichards/mcp-js

MCP server Cursor

MCP server "js" as configured in r33drichards/mcp-js. Launched with /Users/robertwendt/mcp-v8/server/target/debug/server --s3-bucket test-mcp-js-buc.

not rated 56 +2 today A tokens not measured AGPL-3.0

evalbench-review

367

GoogleCloudPlatform/evalbench

Skill Claude Code

Review a change in the EvalBench repo for (a) does it actually work — verified by running the tests and style checks, (b) does it follow EvalBench architecture — base-class contracts, config-key registration, PYTHONPATH-relative imports, sandbox isolation, concurrency safety, docs, (c) does it still build and deploy …

not rated 55 yesterday A 170 tokens original Apache-2.0

cyboflow-verify-setup

368

kesteva/cyboflow

Agent Claude Code

Verify Setup subagent. Surveys a project for how its UI actually stands up (scripts, framework, Electron vs web, isolation levers, existing runbook), then drafts a portable verification runbook per modality with a required attestation channel, machine-local bindings, and the lowest-rung repo changes that make it work.…

not rated 55 +1 yesterday A 104 tokens original MIT

cf7-sort-fuzzer

369

FlashNightModReborn/CrazyFlashNight

Skill Claude CodeCodex

A TypeScript fuzzing and benchmarking workflow for a sorting router. Fuzzing means sending varied or unexpected inputs to find errors; it runs without Flash.

not rated 55 today A 0 tokens GPL-3.0

workshop-testing

370

dotnet-presentations/ai-workshop

Skill Claude CodeCodex

Walk through the .NET AI Workshop as an attendee to validate that the READMEs, commands, and code snapshots still work. USE FOR: testing the workshop, testing a specific Part, dry-running the labs, verifying a README against its snapshot, reconciling or refreshing code snapshots, producing a workshop test report. DO…

not rated 55 +1 11d ago A 95 tokens original MIT

using-axis

371

netlify/axis

Skill Claude CodeCodex

Run AXIS, read its reports, navigate its project layout, and interpret scores. Use when the user asks to run AXIS, invoke the CLI, compare runs, explain a score, find a regression, manage baselines, or understand where AXIS writes its files.

not rated 55 +1 yesterday A 58 tokens original MIT

tricorder-cli

372

tweag/tricorder

Skill Claude Code

Full CLI reference for the tricorder daemon (build diagnostics, test results, source lookup, eval comments, logs). This is the fallback invoked by the tricorder skill when the tricorder-mcp MCP tools aren't available or aren't working — invoke tricorder first; it decides whether this is needed.

not rated 55 +1 yesterday A 67 tokens

write-test-plan

373

andresharpe/dotbot

Skill Claude CodeCodex

Generate a QA/UAT test plan from product specifications and task definitions, covering acceptance testing, integration flows, and exploratory testing. Unit tests are out of scope (handled by write-unit-tests skill).

not rated 54 10d ago A 43 tokens original MIT

uxaudit

374

gotalab/uxaudit

Plugin Claude Code

Bundles 1 skill, 7 agents, 1 hook · 838 tokens together

A UX regression testing skill for browser-based and webview-based apps. Runs after E2E to catch usability, journey, accessibility, and interface-quality issues before they ship.

not rated 54 +2 4mo ago A tokens not measured original Apache-2.0

harness-init

375

Gizele1/harness-init

Plugin Claude Code

Bundles 1 skill · 47 tokens together

Makes a repo agent-ready: AGENTS.md, boundary tests, CI pipeline, GC scripts — based on OpenAI's harness engineering methodology.

not rated 53 5mo ago A tokens not measured original MIT

burner-phone

376

SouthpawIN/burner-phone

Skill Claude CodeCodex

Universal Android device control with vision feedback. Supports Termux phones, ADB-only devices, and emulators. Use for phone automation, AI companionship, or mobile app testing.

not rated 53 1mo ago B 39 tokens original MIT

merlinhu1/codex-game-studio

Skill Claude Code

Use for architecture review tasks that review architecture for layer violations, scalability risks, engine misuse, testing seams, and production readiness; produce verification evidence, changed or proposed files, and handoff boundaries.

not rated 53 +3 27d ago A 45 tokens original MIT

ai-evals

378

liqiongyu/lenny_skills_plus

Skill Claude CodeCodex

Create an AI Evals Pack (eval PRD, test set, rubric, judge plan, results + iteration loop). See also: building-with-llms (build), ai-product-strategy (strategy).

not rated 52 3mo ago A 46 tokens original Apache-2.0

spring-jpa-testing

379

spring-ai-community/spring-testing-skills

Skill Claude CodeCodex

@DataJpaTest requires TestEntityManager + flush()/clear() before assertions — without this, tests read from Hibernate L1 cache and pass falsely. Read before writing any JPA test. Triggers: @DataJpaTest, JpaRepository, TestEntityManager, @ServiceConnection, Testcontainers, @Modifying, Hibernate 6/7…

not rated 52 4mo ago A 96 tokens

build-task-benchmarks

380

radiantlogicinc/fastworkflow

Skill Claude CodeCodex

Build a two-tier conversation benchmark for a fastWorkflow workflow: short single-errand conversations that act as unit tests and reusable blocks, then long real-world tasks of 20-100 steps composed from those blocks. Covers turn and conversation structure, seeding objects with enough structure to sustain a chain…

not rated 52 2d ago A 161 tokens original Apache-2.0

scout

381

tester-army/scout

Skill Claude CodeCodex

Safely explore and adversarially test an authorized HTTP API using the scout CLI, with or without an OpenAPI spec. Use when asked to test, probe, validate, or explore an API, whether or not an OpenAPI/Swagger spec is available. Scout is the harness; you are the operator.

not rated 52 +1 1mo ago A 65 tokens original MIT

codex-test-bridge

382

msrv-tech/skills

Skill Claude CodeCodex

An HTTP bridge for demo and test 1C databases. It can replace a COM connection when working with test data and external 1C reports or print forms.

not rated 52 +1 changed 3d ago A 75 tokens

rewardkit

383

pku-liang/hwe-bench

Skill Claude CodeCodex

Write Harbor task verifiers using Reward Kit. Use when creating or editing a task's tests/ directory, adding grading criteria, setting up LLM/agent judges, or designing verifiers that produce a reward score.

not rated 52 +2 1mo ago A 46 tokens copy · 91% Apache-2.0

test-quality-tools

384

rollinsio/beyond-test-coverage

Plugin Claude Code

Bundles 2 skills · 335 tokens together

Two skills: test-quality (audit/harden/generate mutation-resistant tests against a quality scorecard) and results-dashboard (render a scorecard JSON into an interactive HTML readout).

not rated 52 2mo ago A tokens not measured original MIT

At most 3 mods per repository are shown here, and a mod shipped inside a plugin is left to that plugin's page — the rest are on their repository pages: