Testing skills

16,505 tagged Testing, measured the same way as everything else here.

Browse within: agentic 49javascript 38android 37openai 37agent-orchestration 34agent-browser 32ai-testing 31hacktoberfest 31coding-agent 29dotnet 28nextjs 28agentic-workflow 27antigravity 27flutter 27

dual-replay

97

MystenLabs/sui

Skill Claude CodeCodex

Run Sui dual execution replay between base and tip commits, recover failed steps, build, and commit replay instrumentation.

7.7k 3d ago A 27 tokens original Apache-2.0

e2e-test

98

maximhq/bifrost

Skill Claude CodeCodex

Skill "e2e-test" from maximhq/bifrost, covering playwright e2e testing, usage, workflow overview, auto-update workflow (sync mode) and step 0: detect what changed.

7.7k +80 yesterday A 0 tokens original Apache-2.0

automation

99

vercel-labs/native

Skill Claude CodeCodex ✓ vendor

Automation and verification guide for running Native SDK apps. Use when the user asks to test a running app, inspect runtime state, list windows, wait for readiness, drive widgets, take deterministic screenshots, send bridge commands, debug why automation is not connected, create smoke tests, or verify a Native SDK…

7.6k +14 8d ago A 69 tokens original Apache-2.0

regression-tests

100

telepresenceio/telepresence

Skill Claude CodeCodex

Run, scope, or debug telepresence regression tests under regressiontest/ — the integration-level suite. Use when the user wants to run an area, suite, or single test, debug a failure, or says "/regression-tests". Runs go test ./regressiontest scoped with -run, in the background, writing to a log file so heavy output…

7.3k 8d ago A 82 tokens original Apache-2.0

CursorTouch/Windows-MCP

Skill Claude CodeCodex

Automated testing skill for Windows-MCP tools. Use this skill whenever the user wants to test, validate, benchmark, or evaluate any Windows-MCP tool (App, PowerShell, Screenshot, Snapshot, Click, Type, Scroll, Move, Shortcut, Wait, MultiSelect, MultiEdit, Clipboard, Process, Notification, FileSystem, Registry…

6.9k +14 yesterday A 139 tokens original MIT

GreptimeTeam/greptimedb

Skill Claude CodeCodex

Investigate a failed GreptimeDB fuzz CI target link by downloading GitHub Actions job logs plus fuzz artifacts such as kind logs, monitor dumps, and CSV dumps, then correlate the failure with local GreptimeDB source code. Use when the user provides a failed fuzz CI target/job URL or asks to diagnose GreptimeDB fuzz CI…

6.6k +5 yesterday A 80 tokens original Apache-2.0

cursor/plugins

Skill Claude CodeCodex ✓ vendor

Reproduce triaged Slack bugs through a configured app-control adapter, verify existing fixes, and open a bounded draft pull request only after before-and-after proof. Use only from the configured Benny repro automation.

6.5k +283 yesterday A 48 tokens

mobile-automation

104

mobile-next/mobile-mcp

Skill Claude CodeCodex

Control Android and iOS devices, emulators and simulators — launch apps, tap, swipe, type, take screenshots, read the accessibility tree. Use when a task involves a mobile device or app, mobile UI testing, or reproducing a bug on a phone.

6.3k +63 yesterday A 58 tokens original Apache-2.0

solopi-ai

105

alipay/SoloPi

Skill Claude CodeCodex

A command-line framework for testing Android apps and devices with SoloPi, including on-device or cloud AI decision models. It manages devices, test cases, recorded interactions, replays, performance history, and evidence.

6.2k +10 14d ago A 127 tokens original Apache-2.0

driver-test-runner

106

rivet-dev/actors

Skill Claude CodeCodex

Methodically run the RivetKit driver test suite file by file across the native (NAPI) and wasm runtimes, tracking progress in /.agents/notes/driver-test-progress.md. Use when you need to validate the driver test suite after changes, bring up a new driver, or debug test failures systematically.

6.1k +3 yesterday A 68 tokens original Apache-2.0

Azure/azure-sdk-for-net

Skill Claude CodeCodex ✓ vendor

Discovers and implements gaps in Spector test coverage for the Azure C# HTTP client emitter. Use when asked to find missing Spector scenarios, add Spector test coverage, or implement a specific Spector spec for the Azure C# emitter. Can also compare coverage between the Azure dashboard and the Standard (TypeSpec core)…

6.0k +2 yesterday A 78 tokens original MIT

swift-testing-pro

108

jacklandrin/OnlySwitch

Skill Claude CodeCodex

Writes, reviews, and improves Swift Testing code using modern APIs and best practices. Use when reading, writing, or reviewing projects that use Swift Testing.

5.9k +2 3d ago A 34 tokens original MIT

add-unit-tests

109

areal-project/AReaL

Skill Claude CodeCodex

Guide for adding unit tests to AReaL. Use when user wants to add tests for new functionality or increase test coverage.

5.7k +4 yesterday A 30 tokens original Apache-2.0

kenn-io/agentsview

Skill Claude CodeCodex

Use when creating, editing, fixing, or reviewing tests; when adding mocks, fakes, assertions, unit tests, PG integration tests, frontend component tests, or Playwright e2e tests; or when changing tests after failures.

5.7k +75 yesterday A 54 tokens original MIT

verification

111

ageerle/ruoyi-ai

Skill Claude CodeCodex

Prove that a coding task is actually complete. Use this after meaningful code changes, when tests/builds fail or are skipped, before marking a plan or goal complete, and whenever acceptance depends on runtime, security, recovery, performance, or cross-module evidence.

5.7k +3 17d ago A 54 tokens original MIT

apify/apify-mcp-server

Skill Claude CodeCodex

Use when adding Langfuse workflow evals for a tool family of the Apify MCP server ("create evals for the storage tools"), when eval cases fail and you must decide whether the case, the tool, or its description is at fault, or when eval runs show tool errors in Langfuse traces.

5.6k +221 yesterday A 69 tokens original MIT

feature-walkthrough

113

chrisleekr/binance-trading-bot

Skill Claude CodeCodex

Autonomously test the running binance-trading-bot app in a real browser. The agent drives the browser itself via the Playwright MCP - logs in, looks at each screen, and works through every feature end-to-end like a real operator, finding and fixing bugs. Use when asked to test the app, smoke-test or walk through the…

5.5k +2 yesterday A 94 tokens original Apache-2.0

clawteam-dev

114

HKUDS/ClawTeam

Skill Claude CodeCodex

Use this skill when working inside the ClawTeam repository itself: local development, debugging, reviewing, testing, validating multi-agent flows, or checking whether a code change actually works end-to-end. Use the repository bootstrap scripts to standardize the local clawteam command and to wire project-local…

5.5k +2 3mo ago A 94 tokens original MIT

test-plan

115

cloudflare/agents

Skill Claude CodeCodex ✓ vendor

Produce a focused test plan for a change. Use when the user asks how to test a feature, what cases to cover, or for a QA checklist before shipping.

5.5k +5 yesterday A 36 tokens original MIT

vllm-project/semantic-router

Skill Claude CodeCodex

Calibrates routing changes against a live router endpoint with executable probes, local DSL validation, versioned deploys, and structured failure review. Use when tuning signals, projections, decisions, or maintained route examples against a real apiserver.

5.5k +56 yesterday A 53 tokens original Apache-2.0

benchmarking

117

Blaizzy/mlx-vlm

Skill Claude CodeCodex

Part of mlx-vlm-skills

Use this skill when the user wants to benchmark an MLX-VLM change and present the numbers in a PR — fork-vs-main A/B comparisons, isolated-module micro-benchmarks, median-of-N timing with warmup, peak-memory reporting, correctness checks, parameter sweeps, and self-contained reproducible bench scripts to paste into a…

5.5k +18 yesterday A 74 tokens original MIT

loopx-benchmark

118

huangruiteng/loopx

Skill Claude CodeCodex

Use when a LoopX-managed goal runs, tracks, scores, or analyzes a benchmark experiment through benchmark-toolkit, including experiment-board rows, solver arms, integrity qualification, matched comparisons, or case insights. Do not use for casual benchmark discussion, ordinary software microbenchmarks, or eval mentions…

5.4k +81 changed yesterday A 69 tokens original Apache-2.0

web-contrib

119

remsky/Kokoro-FastAPI

Skill Claude CodeCodex

Contributing to the Kokoro-FastAPI web player: vanilla JS constraints, MSE/audio gotchas, unit and e2e test setup. Use when changing anything under web/.

5.4k +11 yesterday A 41 tokens original Apache-2.0

next-qa

120

breaking-brake/cc-wf-studio

Skill Claude CodeCodex

Run one unattended iteration of the QUALITY-ASSURANCE loop — steward any in-flight QA PR, then build ONE queued qa issue (test infrastructure, unit tests, regression tests for known bugs) on a branch off auto-qa and open a PR that squash-merges on green CI. Adds tests and tooling only; never edits product source. Use…

5.4k 3d ago A 105 tokens

At most 3 mods per repository are shown here, and a mod shipped inside a plugin is left to that plugin's page — the rest are on their repository pages: