Testing skills

17,307 tagged Testing, measured the same way as everything else here.

Browse within: agentic 60android 44javascript 43ai-skills 38openai 37agent-orchestration 35agentic-workflow 34agent-browser 32ai-testing 31antigravity 31hacktoberfest 31coding-agent 28nextjs 28copilot 27

apify/apify-mcp-server

Skill Claude CodeCodex

Use when adding Langfuse workflow evals for a tool family of the Apify MCP server ("create evals for the storage tools"), when eval cases fail and you must decide whether the case, the tool, or its description is at fault, or when eval runs show tool errors in Langfuse traces.

5.6k +221 yesterday A 69 tokens original MIT

feature-walkthrough

122

chrisleekr/binance-trading-bot

Skill Claude CodeCodex

Autonomously test the running binance-trading-bot app in a real browser. The agent drives the browser itself via the Playwright MCP - logs in, looks at each screen, and works through every feature end-to-end like a real operator, finding and fixing bugs. Use when asked to test the app, smoke-test or walk through the…

5.5k +2 yesterday A 94 tokens original Apache-2.0

clawteam-dev

123

HKUDS/ClawTeam

Skill Claude CodeCodex

Use this skill when working inside the ClawTeam repository itself: local development, debugging, reviewing, testing, validating multi-agent flows, or checking whether a code change actually works end-to-end. Use the repository bootstrap scripts to standardize the local clawteam command and to wire project-local…

5.5k +2 3mo ago A 94 tokens original MIT

test-plan

124

cloudflare/agents

Skill Claude CodeCodex ✓ vendor

Produce a focused test plan for a change. Use when the user asks how to test a feature, what cases to cover, or for a QA checklist before shipping.

5.5k +5 yesterday A 36 tokens original MIT

vllm-project/semantic-router

Skill Claude CodeCodex

Calibrates routing changes against a live router endpoint with executable probes, local DSL validation, versioned deploys, and structured failure review. Use when tuning signals, projections, decisions, or maintained route examples against a real apiserver.

5.5k +56 yesterday A 53 tokens original Apache-2.0

benchmarking

126

Blaizzy/mlx-vlm

Skill Claude CodeCodex

Part of mlx-vlm-skills

Use this skill when the user wants to benchmark an MLX-VLM change and present the numbers in a PR — fork-vs-main A/B comparisons, isolated-module micro-benchmarks, median-of-N timing with warmup, peak-memory reporting, correctness checks, parameter sweeps, and self-contained reproducible bench scripts to paste into a…

5.5k +18 yesterday A 74 tokens original MIT

loopx-benchmark

127

huangruiteng/loopx

Skill Claude CodeCodex

Use when a LoopX-managed goal runs, tracks, scores, or analyzes a benchmark experiment through benchmark-toolkit, including experiment-board rows, solver arms, integrity qualification, matched comparisons, or case insights. Do not use for casual benchmark discussion, ordinary software microbenchmarks, or eval mentions…

5.4k +81 changed yesterday A 69 tokens original Apache-2.0

web-contrib

128

remsky/Kokoro-FastAPI

Skill Claude CodeCodex

Contributing to the Kokoro-FastAPI web player: vanilla JS constraints, MSE/audio gotchas, unit and e2e test setup. Use when changing anything under web/.

5.4k +11 yesterday A 41 tokens original Apache-2.0

next-qa

129

breaking-brake/cc-wf-studio

Skill Claude CodeCodex

Run one unattended iteration of the QUALITY-ASSURANCE loop — steward any in-flight QA PR, then build ONE queued qa issue (test infrastructure, unit tests, regression tests for known bugs) on a branch off auto-qa and open a PR that squash-merges on green CI. Adds tests and tooling only; never edits product source. Use…

5.4k 3d ago A 105 tokens

create-skill-test

130

dotnet/skills

Skill Claude CodeCodex

Scaffolds eval.yaml evaluation specs for agent skills in the dotnet/skills repository. Use when creating skill tests, writing evaluation stimuli, defining graders and rubrics, sizing an eval for statistical power, or setting up test fixture files. Handles the Vally eval.yaml schema, fixture organization, and…

5.3k +17 yesterday A 98 tokens original MIT

write-contract

131

internet-court/internet-court-skill

Skill Claude CodeCodex

Part of internet-court

Write production-quality GenLayer intelligent contracts. Always pins concrete GenVM runner version hashes and never uses local-only test/latest runner aliases. Covers equivalence principles, storage rules, LLM resilience, and cross-contract interaction.

5.3k +193 14d ago A 46 tokens

tech-leads-club/agent-skills

Skill Claude CodeCodex

Creates comprehensive Technical Design Documents (TDD) with mandatory and optional sections through interactive discovery. Use when user asks to "write a design doc", "create a TDD", "technical spec", "architecture document", "RFC", "design proposal", or needs to document a technical decision before implementation. Do…

5.1k +11 2d ago A 86 tokens

crabbox

133

openclaw/Peekaboo

Skill Claude CodeCodex

Use the Crabbox wrapper for OpenClaw remote validation across Linux, macOS, Windows, and WSL2, including delegated Blacksmith Testbox proof. Report the actual provider and id.

5.1k +22 2d ago B 43 tokens original MIT

agent-integration

134

entireio/cli

Skill Claude CodeCodex

Run all three agent integration phases sequentially: research, write-tests, and implement using E2E-first TDD (unit tests written last). For individual phases, use /agent-integration:research, /agent-integration:write-tests, or /agent-integration:implement. Use when the user says "integrate agent", "add agent…

5.0k +7 yesterday A 89 tokens original MIT

test-repo

135

entireio/cli

Skill Claude CodeCodex

Use this skill to test strategy changes against a fresh test repository. Invoke when the user asks to "test against a test repo", "validate the changes", or wants to verify session hooks, commits, and checkpoint creation work correctly.

5.0k +7 changed yesterday A 50 tokens original MIT

kiln-prerelease-check

136

Kiln-AI/Kiln

Skill Claude CodeCodex

Run the Kiln pre-release smoke test suite plus the standard CI checks (checks.sh), diagnose every prerelease test that broke (and why), and write a clean readable report with recommended actions. Read-only — it never edits code. Use when the user wants to validate a release candidate, run prerelease tests, or asks for…

5.0k +2 yesterday A 84 tokens

agent-device-evidence

137

Expensify/App

Skill Claude CodeCodex

Records iOS/Android native MP4 evidence for test/repro flows extracted from an Expensify GitHub PR or issue. Use when the user asks to "record the flow for PR.

5.0k +2 yesterday A 43 tokens original MIT

agent-device

138

Expensify/App

Skill Claude CodeCodex

Drive iOS and Android devices for the Expensify App - testing, debugging, performance profiling, bug reproduction, and feature verification. Use when the developer needs to interact with the mobile app on a device.

5.0k +2 yesterday A 45 tokens original MIT

zep-eval-harness

139

getzep/zep

Skill Claude CodeCodex

Run and manage the Zep eval harness pipeline — document chunking, user ingestion, document ingestion, evaluation, graph inspection, and results analysis. Use when the user asks to run eval harness scripts, use the Zep eval harness, get terminal commands for eval harness operations, chunk documents, ingest users or…

4.9k +4 yesterday A 139 tokens original Apache-2.0

create-adapter

140

harbor-framework/harbor

Skill Claude CodeCodex

Scaffold a new Harbor benchmark adapter by running harbor adapter init and then guide implementation using the Adapters Agent Guide as the authoritative spec.

4.9k +72 yesterday A 34 tokens original Apache-2.0

rewardkit

141

harbor-framework/harbor

Skill Claude CodeCodex

Write Harbor task verifiers using Reward Kit. Use when creating or editing a task's tests/ directory, adding grading criteria, setting up LLM/agent judges, or designing verifiers that produce a reward score.

4.9k +72 changed yesterday A 46 tokens original Apache-2.0

agent-release-gate

142

Agenta-AI/agenta

Skill Claude CodeCodex

Run the agent release gate — a portable, wire-level QA harness for the agent runtime. Drives the same product endpoint the playground drives and asserts on the SSE frame stream and real side effects, never on model prose, so it works against any deployment (cloud or self-hosted) from three env vars. Use before an…

4.7k +28 changed yesterday A 126 tokens

prest/prest

Skill Claude CodeCodexCursor

Guides writing and reviewing pREST Docker/network integration tests under integration/ so each request is human-readable via step comments or table-driven descriptions. Use when adding or editing integration//test.go, HTTP controller E2E coverage, make test-integration, test-integration-postgres…

4.6k +2 5d ago A 78 tokens original MIT

bat-adhoc

144

homeassistant-ai/ha-mcp

Skill Claude CodeCodex

Run bot acceptance tests to validate MCP tools work correctly from a real AI agent's perspective. Use when testing PRs, detecting regressions, or verifying tool changes end-to-end with Claude/Gemini CLIs.

4.6k +29 yesterday A 48 tokens original MIT

At most 3 mods per repository are shown here, and a mod shipped inside a plugin is left to that plugin's page — the rest are on their repository pages: