Testing skills

16,505 tagged Testing, measured the same way as everything else here.

Browse within: agentic 49javascript 38android 37openai 37agent-orchestration 34agent-browser 32ai-testing 31hacktoberfest 31coding-agent 29dotnet 28nextjs 28agentic-workflow 27antigravity 27flutter 27

test-review

73

datahub-project/datahub

Skill Claude CodeCodex

You are an expert DataHub test reviewer. Your role is to evaluate smoke tests and integration tests against established testing standards, identify issues, and provide actionable feedback.

13k +11 yesterday A 0 tokens original Apache-2.0

qa-test-planner

74

meshery/meshery

Skill Claude CodeCodex

Generate comprehensive test plans, manual test cases, regression test suites, and bug reports for QA engineers. Includes Figma MCP integration for design validation.

12k +26 yesterday A 34 tokens original Apache-2.0

sdk-e2e-cases

75

TencentCloud/CubeSandbox

Skill Claude CodeCodex

Write, review, and execute CubeSandbox SDK compatibility E2E pytest cases. Use when the user asks to add, design, review, debug, or run SDK E2E cases under tests/e2e/sdkcompat, or mentions lifecycle, network policy, sandbox templates, backend compatibility, pytest markers, or live E2E validation.

11k 5d ago A 76 tokens

pxi-eval-dataset

76

Arize-ai/phoenix

Skill Claude CodeCodex

Generate synthetic evaluation datasets for the PXI eval harness (evals/pxi/). Use whenever the user asks to create, author, draft, expand, or audit an eval dataset for a PXI tool, skill, or behavior — including phrases like "write evals for ", "test PXI behavior", "synthetic dataset for PXI", "cover this tool with…

11k +28 yesterday A 139 tokens

playwright-cli

77

VoltAgent/voltagent

Skill Claude CodeCodex

Automates browser interactions for web testing, form filling, screenshots, and data extraction. Use when the user needs to navigate websites, interact with web pages, fill forms, take screenshots, test web applications, or extract information from web pages.

11k +53 6d ago A 52 tokens original MIT

github/copilot-sdk

Skill Claude CodeCodex ✓ vendor

Use this skill when creating a new Java E2E integration test (failsafe IT) that requires a new replay proxy YAML snapshot file in test/snapshots/.

10k 3d ago A 44 tokens original MIT

winui-runtime-tests

79

unoplatform/uno

Skill Claude CodeCodex

Build, install, and run runtime tests against the WinUI (WinAppSDK) SamplesApp on Windows. Use when testing against native WinUI to validate parity with Uno.

10k +6 yesterday A 35 tokens original Apache-2.0

launch

80

microsoft/vscode-copilot-chat

Skill Claude CodeCodex ✓ vendor

Launch and automate VS Code Insiders with the Copilot Chat extension using agent-browser via Chrome DevTools Protocol. Use when you need to interact with the VS Code UI, automate the chat panel, test the extension UI, or take screenshots. Triggers include 'automate VS Code', 'interact with chat', 'test the UI', 'take…

10.0k +4 3mo ago A 80 tokens original MIT archived

a11y-flow-audit

81

gitkraken/vscode-gitlens

Skill Claude CodeCodex ✓ vendor

Use to audit a page, view, or composed flow for WCAG 2.1 AA compliance at the composition level - landmarks, heading hierarchy, tab order across components, focus handoff on modal open/close, live-region conflicts. Scope is page/view, NOT component internals. Safety-first - refuses to emit fixes that would create a…

9.9k +3 yesterday A 107 tokens

live-exercise

82

gitkraken/vscode-gitlens

Skill Claude CodeCodex ✓ vendor

Use whenever any UI-bearing work touches a running instance — building or fixing a feature, ship-gating, auditing, OR debugging visible bugs (flaky behavior, intermittent rendering, "sometimes does X" reports, hover/focus/animation glitches, layout overflow). Adaptive depth from tactical fix-loop to ship-gate audit.…

9.9k +3 yesterday A 76 tokens

run-load-test

83

omnigent-ai/omnigent

Skill Claude CodeCodex

Run the Omnigent load test and produce a results file explaining the latencies. Load when the user wants to load-test / stress-test / benchmark Omnigent under concurrency ("load test omnigent", "stress test the server", "how many hosts/sessions/turns can it handle", "load test real agent turns / conversations", "run a…

9.6k +82 yesterday A 168 tokens original Apache-2.0

bitwarden/android

Skill Claude CodeCodex

This skill should be used when writing or reviewing tests for Android code in Bitwarden. Triggered by "BaseViewModelTest", "BitwardenComposeTest", "BaseServiceTest", "stateEventFlow", "bufferedMutableSharedFlow", "FakeDispatcherManager", "expectNoEvents", "assertCoroutineThrows", "createMockCipher", "createMockSend"…

9.3k 4d ago A 111 tokens GPL-3.0

qa-cli-mcp-api

85

fastrepl/anarlog

Skill Claude CodeCodex

Select and run explicitly requested, risk-based QA for Anarlog's CLI, webhooks, stdio MCP, hosted Cloud API, and remote MCP. Test only affected lanes unless comprehensive coverage is requested.

9.2k +24 yesterday A 47 tokens original MIT

debugging-failures

86

onsi/ginkgo

Skill Claude CodeCodex

Part of ginkgo

Diagnose a failing Ginkgo suite as an agent — always run with --json-report into a predictable temp/gitignored location, read the terminal verdict line, then use jq to extract structured failure details (name, message, file:line, panic value, captured logs). Covers the panicked-vs-failed trap, panic locations pointing…

9.0k +3 23d ago A 122 tokens original MIT

filtering

87

onsi/ginkgo

Skill Claude CodeCodex

Part of ginkgo

Run a subset of a Ginkgo suite — Pending/PIt/XIt, runtime Skip, programmatic Focus/FIt (and ginkgo unfocus), Label with the --label-filter query language and label sets, suite-level labels, SemVerConstraint/--sem-ver-filter, --focus/--skip and --focus-file/--skip-file, the filtering precedence rules, and…

9.0k +3 23d ago A 117 tokens original MIT

ordering-and-flakes

88

onsi/ginkgo

Skill Claude CodeCodex

Part of ginkgo

Control spec ordering and manage flaky specs — Serial, Ordered containers with BeforeAll/AfterAll/ContinueOnFailure, OncePerOrdered, SpecPriority, plus FlakeAttempts/--flake-attempts, MustPassRepeatedly, --repeat, and --until-it-fails. Use when specs must run in a fixed order, you need once-per-group setup, you're…

9.0k +3 23d ago A 111 tokens original MIT

verify

89

usekaneo/kaneo

Skill Claude CodeCodex

Build/launch/drive recipe for verifying Kaneo changes end-to-end on a local dev instance.

8.9k +118 yesterday A 21 tokens original MIT

dark-mode-check

90

anyproto/anytype-ts

Skill Claude CodeCodex

Audit SCSS and TSX files for dark mode issues — missing variable usage, hardcoded colors, icon gaps, selector misuse, and inline dark overrides outside the theme folder.

8.7k +8 yesterday A 38 tokens

qa-engineer

91

anyproto/anytype-ts

Skill Claude CodeCodex

Analyze code changes and generate Playwright E2E tests in anytype-desktop-suite. Run after implementing features or modifying editor/component behavior to ensure new functionality has test coverage.

8.7k +8 yesterday A 39 tokens

jesse-ai/jesse

Skill Claude CodeCodex

Use when writing or modifying tests for Jesse's backend — especially behavior tied to a strategy (entries/exits, take-profit/stop-loss, position lifecycle hooks, closed-trade metrics). Documents this repo's strategy-driven test pattern: a thin test in tests/testparentstrategy.py that runs singleroutebacktest('Name')…

8.4k +8 6d ago A 0 tokens original MIT

instatic-user-e2e

93

CoreBunch/Instatic

Skill Claude CodeCodex

Run user-facing Instatic E2E audits with a real browser and disposable local data. Use when asked to test the app as a user, run an agent-browser pass, perform a fresh-install smoke test, audit UX friction, verify setup/login/edit/publish/public-page flows, retest E2E issues, or update the Instatic E2E protocol and…

8.4k +58 yesterday A 83 tokens original MIT

prek

94

j178/prek

Skill Claude CodeCodex

Use when setting up or running hooks with prek in any repository. prek is a Rust drop-in alternative to pre-commit for running Git hooks that check, format, lint, and validate code and repository files. Prefer it when speed, workspace mode, built-in hooks, shared toolchains, or native TOML configuration matter.

8.4k 2d ago A 67 tokens original MIT

no-mistakes

95

kunchenguid/no-mistakes

Skill Claude CodeCodex

Validate your code changes through the no-mistakes pipeline - automated code review, tests, lint, docs, push, PR, and CI - before they reach the configured push target. Use when the user asks to run no-mistakes, gate or ship or validate their changes, push safely, asks you to do a task and then validate it, or invokes…

8.2k +275 2d ago A 84 tokens original MIT

verify

96

anthropics/claude-agent-sdk-python

Skill Claude CodeCodex ✓ vendor

Drive this repo's build/release scripts end-to-end without touching the network. Use when verifying changes to scripts/downloadcli.py, scripts/buildwheel.py, scripts/updatecliversion.py, or scripts/cliversionvalidation.py.

8.0k +14 yesterday C 48 tokens original MIT

At most 3 mods per repository are shown here, and a mod shipped inside a plugin is left to that plugin's page — the rest are on their repository pages: