Testing skills

16,411 tagged Testing, measured the same way as everything else here.

Browse within: agent-browser 89ai-coding 82ai-testing 80javascript 52agentic-coding 50openclaw 49agentic-workflow 47openai 47android 44claude-code-plugin 44cypress 41static-analysis 38skill-scanner 35agentic-framework 34

code_learning_doc

265

clojurewasm/ClojureWasm

Skill Claude CodeCodex

Write Japanese per-task notes under private/notes/ during the per-task TDD loop. The per-concept chapter half (docs/ja/learnclojurewasm/NNNN.md) is DORMANT per ADR-0025 — no new chapters land, pre-commit gate is a no-op, existing chapters live in docs/ja/archive/. Re-activates by a future ADR.

not rated 186 23d ago A 88 tokens EPL-2.0

noqa-testing

266

noqa-ai/noqa

Skill Claude CodeCodex

Use this skill when the user wants to boot and interact with iOS or Android devices/simulators — inspect the screen, execute actions, generate or edit test cases, or run UI tests via the noqa platform.

not rated 185 1mo ago A 47 tokens original MIT

mcp-testing

267

adeze/raindrop-mcp

Skill Claude CodeCodex

MCP Testing Strategies with Vitest, Inspector, and Integration Tests.

not rated 183 +3 1mo ago A 17 tokens original MIT

ahk-test

268

enmanuelmag/agent-harness-kit

Skill Claude CodeCodex

Design, write, and run behavior-focused tests for an objective or existing code. Writes test files only and reports evidence. No tasks created, no harness tracking.

not rated 180 +1 yesterday A 36 tokens original Apache-2.0

android

269

yang1ming/android-harness

Skill Claude CodeCodex

Direct Android device control through ADB. Use for authorized device automation, testing, screenshots, UI inspection, and app interaction.

not rated 174 1mo ago A 27 tokens original MIT

test-fixer

270

mitchdenny/hex1b

Skill Claude CodeCodex

Agent for diagnosing and fixing flaky terminal UI tests in the Hex1b test suite. Use when tests pass locally but fail in CI, or when tests exhibit timing-sensitive behavior.

not rated 173 6d ago A 39 tokens original MIT

device-test

271

Mahdi-mortazavi/relay

Skill Claude CodeCodex

Run Relay against the phone plugged into this laptop and the Windows app installed on it, find what breaks, and fix it. Use when the user asks to test on real hardware, test on their phone, check pairing on their own Wi-Fi, or verify a release on their own machine. Only meaningful in a session running locally on the…

not rated 171 +3 2d ago A 83 tokens GPL-3.0

tdd

272

wbern/agent-instructions

Skill Claude CodeCodex

Remind agent about TDD approach and continue conversation.

not rated 168 3mo ago A 13 tokens original MIT

dev-cycle

273

xberg-io/crawlberg

Skill Claude CodeCodex

Crawlberg iteration loops codified as Taskfile tasks — alef install/generate/format/bump, core and binding builds, e2e generate/build/test cycles, cleanup tiers, and the mock-server / stale-.so / precompiled-NIF / generated-e2e gotchas. Load when running or debugging crawlberg build, alef regeneration, or e2e…

not rated 168 +3 yesterday A 89 tokens original MIT

implement-feature

274

tddworks/SkillsManager

Skill Claude CodeCodex

Guide for implementing features following architecture-first design, TDD, rich domain models, and Swift 6.2 patterns. Use this skill when: (1) Adding new functionality to a Swift app (2) Creating domain models that follow user's mental model (3) Building SwiftUI views that consume domain models directly (4) User asks…

not rated 165 3mo ago A 103 tokens

dogfood

275

paiml/paiml-mcp-agent-toolkit

Skill Claude CodeCodex

Dogfood pmat — rebuild, install, exercise every CLI command against pmat's own repo, check output integrity + self-quality, find next work. Read-only audit; files issues for bugs.

not rated 164 +1 yesterday A 40 tokens original MIT

qa

276

eric-tramel/slop-guard

Skill Claude CodeCodex

Black-box QA audit of slop-guard across MCP, CLI, fit, docs, agent workflows, and writing-effectiveness. Files GitHub issues for real problems found.

not rated 163 1mo ago A 37 tokens original MIT

cleo-validator

277

kryptobaseddev/cleo

Skill Claude CodeCodex

Independent IVTR peer-reviewer role. The Validator is spawned by a Lead AFTER a Worker reports an implementation candidate to verify that every Acceptance Criterion is satisfied by programmatic evidence. The Validator MUST be a different agent instance from the Worker who built the change — same-agent self-attestation…

not rated 160 15d ago A 196 tokens original MIT

api-testing

278

idavidov13/agentic-playwright

Skill Claude CodeCodex

API testing patterns for Playwright -- apiRequest fixture usage, Zod response schema creation and validation, test.step wrapping for multi-call tests, per-field negative/validation testing, path parameter fuzzing, and helper fixtures for shared setup/teardown. Use when writing or updating API test specs, adding tests…

not rated 159 +25 yesterday A 132 tokens original MIT

vdjdb-proofread

279

antigenomics/vdjdb-db

Skill Claude CodeCodex

Run QC scripts on a VDJdb chunk, report every error with a suggested fix, verify the output of previous /extract and /format steps, estimate confidence scores, and flag gaps in current pysrc QC coverage.

not rated 155 9d ago A 50 tokens AGPL-3.0

nexus-eval-harness

281

ProfSynapse/nexus

Skill Claude CodeCodex

Work on the Nexus LLM eval harness in tests/eval/ — author or fix a scenario fixture, write an eval config, change the executors, assertions or reports, or explain a run that produced nothing, everything-fails, or numbers that disagree. Use when an eval scenario is wrong, a run behaves oddly, or the harness itself…

not rated 154 2d ago A 96 tokens original MIT

glance-test

282

DebugBase/glance

Skill Claude CodeCodex

Run E2E browser tests on any web application using Glance MCP. Use when the user says "test this page," "check this URL," "run E2E tests," "browser test," "test the login flow," "check if the site works," "visual regression," or "screenshot this page." Also use for post-deploy verification and smoke tests.

not rated 151 4mo ago A 80 tokens original MIT

test-assess

283

ambient-code/agentready

Skill Claude CodeCodex

Test agentready assess against real GitHub repositories to validate assessor changes. Selects repos relevant to the change being tested, clones them to a temp directory, runs the local checkout's assessor, reports results and output locations, then cleans up. Use when testing a new or modified assessor, verifying a…

not rated 151 2d ago A 75 tokens original MIT

code-guidelines-go

284

dimetron/pi-go

Skill Claude CodeCodex

Go 1.24–1.27 coding guidelines for the dimetron/pi-go AI agent runtime. Use this skill whenever writing, reviewing, or refactoring ANY Go code in pi-go. This covers idiomatic style, error handling, concurrency, project layout, testing (table-driven, fuzz, benchmarks, synctest), new stdlib usage, golangci-lint v2…

not rated 151 +3 yesterday A 124 tokens original MIT

fork-test

285

OriginProtocol/origin-dollar

Skill Claude CodeCodex

Generate Foundry fork tests for contracts requiring integration testing against real on-chain state.

not rated 151 yesterday A 16 tokens original MIT

suite-converter

286

Margin-Lab/evals

Skill Claude CodeCodex

Converts test suites from external eval frameworks into the Margin Eval suite format. Use this skill whenever the user wants to import, convert, translate, or migrate an eval dataset or test suite into Margin Eval format, or when they mention converting tasks from other benchmarking frameworks into Margin's structure.

not rated 149 +2 1mo ago A 61 tokens AGPL-3.0

spec-linked-docs

287

bobmatnyc/claude-mpm

Skill Claude CodeCodex

Spec-Linked Documentation (SLD): Language-agnostic discipline for maintaining bidirectional traceability between functional specifications and source-code docstrings via stable identifiers and CI validation. Optional/opt-in adoption. Builds on OpenFastTrace and DO-178C Requirements Traceability Matrix traditions.

not rated 149 4d ago A 61 tokens

add-library-test

288

osama-raddad/FireCrasher

Skill Claude CodeCodex

Add or update a Robolectric JVM unit test for the FireCrasher library. Use when changing recovery logic, crash handling, back-stack counting, exit-info reporting, or the recovery-state codec, and a test should cover it.

not rated 148 2mo ago A 51 tokens original Apache-2.0

At most 3 mods per repository are shown here, and a mod shipped inside a plugin is left to that plugin's page — the rest are on their repository pages: