Testing skills

18,641 tagged Testing, measured the same way as everything else here.

Browse within: ai-coding 82agent-browser 80ai-testing 71javascript 54agentic-coding 51openclaw 48agentic-workflow 47openai 47claude-code-plugin 42android 39static-analysis 38skill-scanner 35agentic-framework 34hacktoberfest 34

benchflow-ai/benchflow

Skill Claude CodeCodex

Review Benchflow or SkillsBench task-run trajectories and integration-test Benchflow code changes. Use this skill whenever the user asks to audit traj health, failed or timed-out runs, healthy pass/fail/timeout status, no-skill leakage, skill loading, reward hacking, verifier isolation, metadata completeness, token…

not rated 339 today A 99 tokens original Apache-2.0

test-macafm

218

scouzi1966/maclocal-api

Skill Claude CodeCodex

Run the maclocal-api (AFM/MLX) test suite — automated assertions and smart analysis. Use when asked to test, validate, regression-check, or benchmark AFM before release, after code changes, or for model onboarding.

not rated 337 today A 54 tokens original MIT

elastic/integrations

Skill Claude CodeCodex ✓ vendor

Migrate an Elastic integration package from a legacy inline agent template to integrations with required input dependencies (requires.input, streams[].package). Gathers developer decisions on dataset naming, variable overrides, stack constraints, and tests before applying changes. Use when the user asks to migrate an…

not rated 334 today A 91 tokens

examples-qa

220

cognesy/instructor-php

Skill Claude CodeCodex

Verify Instructor behavior through the repository's ./examples/ suite in pass, live record, or hermetic replay mode. Use when running selected examples or the corpus, capturing and reusing recorded LLM HTTP responses, diagnosing hub results, or distinguishing real errors, assertion failures, skipped examples, and…

not rated 326 3d ago A 67 tokens original MIT

Asymptote-Labs/agent-beacon

Skill Claude CodeCodex

Verify a Beacon change end to end by running a real Claude Code session inside a disposable Linux cloud sandbox and checking that Beacon captured what the agent actually did. Use when asked to verify, validate, test, or prove that a Beacon change works for real rather than just compiling; when asked whether telemetry…

not rated 325 +4 yesterday A 114 tokens original MIT

gleam-practice

222

mizchi/skills

Skill Claude CodeCodex

Use when writing or reviewing Gleam code on the Erlang target — especially Wisp + Mist HTTP services, OTP-based processes (genserver / supervisor analogues), justfile workflows, gleeunit testing, gleam format, GitHub Actions CI, and performance measurement. Trigger on gleam.toml, .gleam files, or Erlang/BEAM-related…

not rated 324 1mo ago A 98 tokens

verify

223

Housetan218/claude-code-haha

Skill Claude CodeCodex

Verify a code change does what it should by running the app.

not rated 321 5mo ago A 15 tokens

phoronix-test-suite

224

intel/intel-performance-skills

Skill Claude CodeCodex ✓ vendor

Install, run, parse, and optimize benchmarks from the Phoronix Test Suite (PTS). Use this skill whenever the user mentions "phoronix", "pts/", or "phoronix-test-suite", or asks to run, measure, improve, or optimize a PTS test — e.g., "run pts/mt-dgemm", "optimize pts/compress-zstd", "what score does pts/x265 get".…

not rated 318 +1 2mo ago A 137 tokens

testing-preview

225

boringcomputers/nehemiah

Skill Claude CodeCodex

Test the preview proxy feature end-to-end. Use when verifying preview URL changes, auth changes on the web proxy route, or networking-related fixes.

not rated 320 +2 23d ago B 32 tokens original Apache-2.0

stove

226

Trendyol/stove

Skill Claude CodeCodex

Use when configuring, writing, or debugging Stove end-to-end tests; choosing JVM, process, container, or provided-application runners; wiring Stove systems; enabling tracing, dashboard, or MCP; or extending Stove with custom systems.

not rated 310 changed today A 49 tokens original Apache-2.0

trailblaze-author

227

block/trailblaze

Skill Claude CodeCodex ✓ vendor

Use when turning a captured human demonstration (a Trail Runner demonstration bundle: demo.yaml + actions.ndjson + per-action screenshots and view hierarchies) into a durable, independently runnable Trailblaze trail. Trigger when a prompt hands you a demonstration bundle directory and asks you to author, generate, or…

not rated 310 +2 2d ago A 99 tokens original Apache-2.0

droid-ash/finalrun-agent

Skill Claude CodeCodex

Generate test and suite specifications in the strict FinalRun YAML format. Handles automated test planning, folder grouping by feature, repo app configuration, environment-specific overrides in .finalrun/env/.yaml, and validation via finalrun check.

not rated 305 1mo ago A 51 tokens original Apache-2.0

playwright-openwebui

229

Fu-Jie/openwebui-extensions

Skill Claude CodeCodex

Use when inspecting UI bugs in OpenWebUI plugins, taking screenshots of plugin output, capturing console errors, testing Action/Filter/Pipe plugins in the chat interface, or verifying plugin installation in the Admin panel. Triggered by: plugin UI bug, Action HTML output, screenshot, console error, plugin not working…

not rated 302 +1 1mo ago A 80 tokens original MIT

n9e-config-driven-e2e

230

n9e/fe

Skill Claude CodeCodex

Maintain config-driven Nightingale E2E tests that convert JSON config data into UI-readable normalized values, drive Playwright + Midscene interactions, and verify persistence through APIs.

not rated 303 today A 43 tokens original Apache-2.0

exploit-xss

231

crazyMarky/pentest-skills

Skill Claude CodeCodex

Cross-site scripting (XSS) vulnerability detection and exploitation. Supports reflected XSS, stored XSS, DOM-based XSS, and blind XSS testing. Use this skill when user mentions XSS, cross-site scripting, script injection, or needs to test JavaScript injection in parameters, forms, headers, or DOM sources.

not rated 301 +3 3mo ago A 71 tokens original Apache-2.0

vizra-ai/vizra-adk

Skill Claude CodeCodex

Test and evaluate AI agents with automated evaluations, assertions, and LLM-as-a-Judge patterns.

not rated 295 +1 11d ago A 26 tokens original MIT

frontend-testing

233

PageAI-Pro/ralph-loop

Skill Claude CodeCodex

Generate Vitest + React Testing Library tests for frontend components, hooks, and utilities. Triggers on testing, spec files, coverage, Vitest, RTL, unit tests, integration tests, or write/review test requests.

not rated 293 4d ago A 48 tokens original MIT

stress-test

234

sgl-project/rbg

Skill Claude CodeCodex

RBG controller stress test - deploy kwok in existing cluster, run create/update/delete load tests with real controller pod, collect pprof profiling and logs, generate analysis report with recommendations.

not rated 289 yesterday A 39 tokens original Apache-2.0

visual-acceptance

235

huiliyi37/Tianshu-harness

Skill Claude CodeCodex

A visual final-review method for user-interface changes. It checks the rendered result with repeatable screenshots, pixel values, computed styles, and CSS layering checks across themes.

not rated 315 yesterday A 89 tokens original Apache-2.0

test-skill

236

joshuadavidthomas/opencode-agent-skills

Skill Claude CodeCodex

A test skill to verify all plugin tools work correctly - useskill, readskillfile, runskillscript, findskills.

not rated 271 9d ago A 29 tokens original MIT

autoresearch

237

ruizrica/agent-pi

Skill Claude CodeCodex

Autonomous Goal-directed Iteration. Apply Karpathy's autoresearch principles to ANY task. Loops autonomously — modify, verify, keep/discard, repeat. Invoke with /skill:autoresearch or when user says "work autonomously", "iterate until done", "keep improving", or "run overnight".

not rated 267 +1 1mo ago A 66 tokens original MIT

xmpp-enumeration

238

blacklanternsecurity/red-run

Skill Claude CodeCodex

XMPP/Jabber service enumeration for Openfire, ejabberd, Prosody, and other XMPP servers. Trigger when ports 5222 (client), 5223 (legacy TLS), or 5269 (server-to-server) are found open. Covers authentication testing, user enumeration, MUC room discovery, and server fingerprinting. Do NOT use for AD enumeration or…

not rated 266 +3 5mo ago A 92 tokens GPL-3.0

edge-python

239

dylan-sutton-chavez/edge-python

Skill Claude CodeCodex

Write, run, test and package Edge Python programs with the edge CLI. Use when editing .py files in an Edge Python project or when the user asks for Edge Python code.

not rated 264 changed today A 39 tokens original Apache-2.0

MingYuePop/SpecForge

Skill Claude CodeCodex

A tool for implementing planned software features with TDD, a method where tests guide the code changes through failing, passing, and cleanup stages.

not rated 263 5mo ago A 48 tokens

At most 3 mods per repository are shown here, and a mod shipped inside a plugin is left to that plugin's page — the rest are on their repository pages: