Testing skills

28,298 tagged Testing, measured the same way as everything else here.

Browse within: ai-coding 73agentic 72android 54agent-orchestration 50claude-ai 50javascript 50ai-skills 37agent-browser 35agentic-coding 35agent-framework 34ai-testing 34openai 34agentic-workflow 33openclaw 31

issue-review

145

mock-server/mockserver-monorepo

Skill Claude CodeCodex

Reviews a GitHub issue to determine validity, classify as user error or bug, check if already fixed, and take appropriate action. For user errors, improves error messages and documentation. For real bugs, implements a fix following the full commit workflow including tests and adversarial review. Closes the issue with…

not rated 5.0k +4 today A 106 tokens original Apache-2.0

zep-eval-harness

146

getzep/zep

Skill Claude CodeCodex

Run and manage the Zep eval harness pipeline — document chunking, user ingestion, document ingestion, evaluation, graph inspection, and results analysis. Use when the user asks to run eval harness scripts, use the Zep eval harness, get terminal commands for eval harness operations, chunk documents, ingest users or…

not rated 4.9k +4 yesterday A 139 tokens original Apache-2.0

create-adapter

147

harbor-framework/harbor

Skill Claude CodeCodex

Scaffold a new Harbor benchmark adapter by running harbor adapter init and then guide implementation using the Adapters Agent Guide as the authoritative spec.

not rated 4.9k +72 yesterday A 34 tokens original Apache-2.0

rewardkit

148

harbor-framework/harbor

Skill Claude CodeCodex

Write Harbor task verifiers using Reward Kit. Use when creating or editing a task's tests/ directory, adding grading criteria, setting up LLM/agent judges, or designing verifiers that produce a reward score.

not rated 4.9k +72 changed yesterday A 46 tokens original Apache-2.0

agent-release-gate

149

Agenta-AI/agenta

Skill Claude CodeCodex

Run the agent release gate — a portable, wire-level QA harness for the agent runtime. Drives the same product endpoint the playground drives and asserts on the SSE frame stream and real side effects, never on model prose, so it works against any deployment (cloud or self-hosted) from three env vars. Use before an…

not rated 4.7k +28 changed yesterday A 126 tokens

mobile-pentest

150

Awarexone/Agentic-Bug-Hunter

Skill Claude CodeCodex

Mobile app pentest for bug bounty (Android APK + iOS IPA) — runtime-first workflow: install app, proxy through Burp/mitmproxy, drive the UI, capture packets, then test the API exactly like a web target; escalate to decompile (apktool/jadx) and Frida/objection only when traffic is SSL-pinned, encrypted, or absent.…

not rated 4.7k +30 2d ago A 205 tokens original MIT

prest/prest

Skill Claude CodeCodexCursor

Guides writing and reviewing pREST Docker/network integration tests under integration/ so each request is human-readable via step comments or table-driven descriptions. Use when adding or editing integration//test.go, HTTP controller E2E coverage, make test-integration, test-integration-postgres…

not rated 4.6k +2 6d ago A 78 tokens original MIT

bat-adhoc

152

homeassistant-ai/ha-mcp

Skill Claude CodeCodex

Run bot acceptance tests to validate MCP tools work correctly from a real AI agent's perspective. Use when testing PRs, detecting regressions, or verifying tool changes end-to-end with Claude/Gemini CLIs.

not rated 4.6k +29 yesterday A 48 tokens original MIT

bat-story-eval

153

homeassistant-ai/ha-mcp

Skill Claude CodeCodex

Compare MCP tool behavior between target and baseline versions using pre-built and custom stories with diff-based triage.

not rated 4.6k +29 yesterday A 26 tokens original MIT

evaluate-environments

154

PrimeIntellect-ai/verifiers

Skill Claude CodeCodex

Run and evaluate verifiers tasksets. Set up the necessary config files and observe the runs and their results.

not rated 4.6k +7 yesterday A 26 tokens original MIT

python-unit-tests

155

dimensionalOS/dimos

Skill Claude CodeCodex

Use when writing, fixing, or reviewing Python pytest unit tests, fixtures, mocks, or test PR feedback.

not rated 4.4k +8 yesterday A 26 tokens

sanity-check

156

rivet-dev/agentos

Skill Claude CodeCodex

Run the deferred AgentOS E2E smoke test from public npm packages. Use when the user asks to sanity check, smoke test, or verify a release works.

not rated 4.4k +9 2d ago A 37 tokens original Apache-2.0

browser-smoke-review

157

fallow-rs/fallow

Skill Claude CodeCodex

Use browser automation to review docs pages, preview URLs, rendered output, or web-facing fallow surfaces. Use when the user wants a screenshot-based review, browser smoke test, docs site check, or preview deployment inspection.

not rated 4.4k +9 yesterday A 49 tokens original MIT

contributing

158

crmne/ruby_llm

Skill Claude CodeCodex

Contribute to RubyLLM - set up the repo, run and record specs, add providers or chat options, work on the Rails integration, and edit docs. Use when fixing a bug, building a feature, writing specs, or changing documentation in the RubyLLM codebase.

not rated 4.3k +7 2d ago A 60 tokens original MIT

android-emulator

159

callstack/agent-device

Skill Claude CodeCodex

Verify and debug native, React Native, Expo, or Flutter apps on an Android Emulator with agent-device. Use when an agent needs to launch an app, inspect its live UI, tap, type, scroll, validate a code change, collect failure evidence, or reproduce a workflow on an Android virtual device.

not rated 4.3k +28 yesterday A 65 tokens original MIT

seed-ssim-references

160

hao-ai-lab/FastVideo

Skill Claude CodeCodex

Seed HF reference artefacts for a single newly-added SSIM test (pixel .mp4 for runtexttovideosimilaritytest-style tests, or latent .pt for runtexttolatentsimilaritytest-style tests). Runs the test on Modal L40S, downloads the generated artefacts via modal volume get, pauses for the user to verify (visual eyeball for…

not rated 4.3k +75 yesterday A 146 tokens original Apache-2.0

feishu-e2e-test

161

m1heng/clawdbot-feishu

Skill Claude CodeCodex

Local E2E debug and test framework for clawd-feishu plugin development. Use when debugging message flow, testing bot responses, verifying Feishu web UI interactions, or performing end-to-end validation of the OpenClaw-Feishu integration during development.

not rated 4.2k +1 5mo ago A 58 tokens original MIT

archestra-dev-testing

162

archestra-ai/archestra

Skill Claude CodeCodex

Use when deciding whether a change needs a test and at which level — unit, backend route-level integration, MSW-backed frontend integration, or e2e — or when reviewing tests for the "fluff test" anti-pattern. Start here before archestra-dev-backend-tests or archestra-dev-e2e.

not rated 4.2k +4 2d ago A 68 tokens

oss-fuzz

163

apache/tika

Skill Claude CodeCodex ✓ vendor

Run Tika's OSS-Fuzz Jazzer targets locally against a working-tree checkout — build the image, build fuzzers from local source, fuzz a target, run a corpus as a regression pass, reproduce a crash, and add seeds. Use for "fuzz the OneNote parser", "run OneNoteParserFuzzer against these files", "reproduce an OSS-Fuzz…

not rated 4.0k +10 yesterday A 90 tokens original Apache-2.0

skill-creator

164

OpenBMB/PilotDeck

Skill Claude CodeCodex

Create new skills, modify and improve existing skills, and measure skill performance. Use when users want to create a skill from scratch, edit, or optimize an existing skill, run evals to test a skill, benchmark skill performance with variance analysis, or optimize a skill's description for better triggering accuracy.

not rated 4.0k +4 2d ago A 64 tokens AGPL-3.0

tool-skill

165

dromara/liteflow

Skill Claude CodeCodex

Skill that binds a Java tool for LiteFlow ReAct agent tests.

not rated 3.8k +5 1mo ago A 17 tokens original Apache-2.0

bean-tool-skill

166

dromara/liteflow

Skill Claude CodeCodex

Skill that binds a Spring-managed Java tool for DI verification tests.

not rated 3.8k +5 1mo ago A 17 tokens original Apache-2.0

release-sample-sweep

167

Atmosphere/atmosphere

Skill Claude CodeCodex

Run the pre-release end-to-end sweep of every user-facing surface — the 33 samples under samples/ (booted from their packaged artifacts and driven in a real browser via chrome-devtools MCP), the Expo/React Native client, and the atmosphere CLI. Use before cutting a release, and after any change to the Console bundle…

not rated 3.8k +8 2d ago A 143 tokens original Apache-2.0

test

168

OffchainLabs/prysm

Skill Claude CodeCodex

Run Prysm unit tests with Bazel for affected packages, or a given target.

not rated 3.8k yesterday A 19 tokens GPL-3.0

At most 3 mods per repository are shown here, and a mod shipped inside a plugin is left to that plugin's page — the rest are on their repository pages: