Testing skills

11,747 tagged Testing, measured the same way as everything else here.

Browse within: LLM 180agents 133agentic-ai 129cli 101ai-coding 99agent 80skills 72javascript 57agent-browser 54openai 51ai-testing 46agentic-workflow 41agent-orchestration 40claude-code-plugin 40

MadAppGang/magus

Skill Claude CodeCodex

Verify, screenshot, AND drive a native macOS app on the same machine the user is actively working on, WITHOUT stealing their mouse, keyboard, or window focus. Use this whenever you need to screenshot, inspect, click through, or confirm the behavior of a running Mac app (Swift/SwiftUI, Electron, AppKit, a dev build, a…

not rated 9 2d ago A 200 tokens original MIT

playwright-core

482

hzijad/playwright-agent-skills

Skill Claude CodeCodex

Conducts rigorous authoring and review of Playwright E2E tests. Enforces accessibility-first locators, web-first assertions, strict isolation, and DAMP architecture. Use when generating, refactoring, or reviewing any Playwright test code.

not rated 8 1mo ago A 53 tokens original MIT

aiopshwang/verify-regression-tests

Skill Codex

Verify that a regression test actually detects the defect it claims to guard against. Use after adding or reviewing a bug-fix guard, reproducing the original failure after a fix, or investigating a suspiciously green regression test. Do not use for general TDD, broad test-suite audits, mutation-score optimization…

not rated 8 10d ago A 75 tokens original MIT

qa-bach

484

DouyuShinyruo/One-Person-Company-Skill

Skill Claude CodeCodex

A quality-assurance guide based on James Bach's approach to software testing. It treats testing as learning about risks and unexpected behavior, not just checking a fixed list of expected results.

not rated 8 3mo ago A 34 tokens original MIT

validate-artifact

486

forthends/clockwork

Skill Claude CodeCodex

A checking procedure for work products such as project documents. It checks that required sections exist, contain real content, and point to files or references that actually exist.

not rated 8 2mo ago A 0 tokens original MIT

taskboard-feature-dev

487

oikon48/claudecode-book

Skill Claude Code

A development procedure for adding features to a TaskBoard application using TDD, or test-driven development. TDD means writing a failing test first, then implementing the feature and checking it in the browser.

not rated 9 +3 24d ago A 67 tokens original MIT

vetto-sandbox

488

shleder/vetto

Skill Codex

Enforce zero-daemon Landlock/Seatbelt security boundaries, network isolation, and subagent capability controls when executing untrusted commands or running subagents. Use when running terminal commands, testing untrusted scripts, isolating AI subagent workflows, or performing read-only session recovery for Codex and…

not rated 8 +5 yesterday A 67 tokens original Apache-2.0

hamza-ali-shahjahan/hamzaish

Skill Claude CodeCodex

Conducts multi-axis code review. Use before merging any change. Use when reviewing code written by yourself, another agent, or a human. Use when you need to assess code quality across multiple dimensions before it enters the main branch.

not rated 8 today A 51 tokens AGPL-3.0

flanliulf/aiforge

Skill Claude CodeCodex

Create system-level or epic-level test plans. Use when the user says "lets design test plan" or "I want to create test strategy".

not rated 8 3mo ago A 36 tokens copy · 78% MIT

PMDevSolutions/Aurelius

Skill Claude Code needs its repo

Orchestrates end-to-end Figma-to-React conversion pipeline with enforced TDD, automated pixel-diff visual QA, E2E testing, and app-type awareness (web apps, Chrome extensions, PWAs). Keywords: Figma to React, design tokens, autonomous component generation, Figma conversion, Tailwind config, component library, TDD…

not rated 8 21d ago A 85 tokens original MIT

casely

493

JohnWayneeee/casely-qa-skill

Skill Claude CodeCodex

Virtual QA Lead that turns requirement documents into review-ready, TestRail-importable test cases in one conversation — no commands to memorize. Use this skill whenever the user has requirements, a spec, a user story, or acceptance criteria (PDF, DOCX, XLSX, TXT, MD, or pasted text) and wants test cases, a test plan…

not rated 7 changed today A 233 tokens original MIT

specmint-tdd

494

ngvoicu/kluris

Skill Claude CodeCodex

TDD-first spec management for AI coding workflows. Use this skill when the user explicitly mentions specs, forging, or structured planning: says "forge", "forge a spec", "write a spec for X", "create a spec", "plan X as a spec", "resume", "what was I working on", "spec list/status/pause/switch/activate", "implement…

not rated 7 2mo ago A 146 tokens original MIT

backtest-data-prep

495

rgourley/quant-garage

Skill Claude CodeCodex

Build a clean, point-in-time, ready-to-backtest OHLCV dataset for a US equity universe across an arbitrary date window. Emits parquet plus a manifest plus an edge-case log, with corporate actions reconciled, survivorship treatment documented, holidays and half-days preserved correctly, and any IPO partial coverage or…

not rated 7 1mo ago A 108 tokens

eval-driven-dev

497

yiouli/pixie-qa

Skill Claude CodeCodex

Improve AI application with evaluation-driven development. Define eval criteria, instrument the application, build golden datasets, observe and evaluate application runs, analyze results, and produce a concrete action plan for improvements. ALWAYS USE THIS SKILL when the user asks to set up QA, add tests, add evals…

not rated 7 4mo ago A 89 tokens copy · 100% MIT

testing-with-pitlane

498

pitlane-ai/pitlane

Skill Claude CodeCodex

Design and create pitlane eval benchmarks that measure whether an AI coding skill or MCP server actually improves assistant performance. Use when the user wants to test a skill, evaluate an MCP server, create a pitlane eval YAML, benchmark an AI assistant, or compare baseline vs challenger configurations. Covers eval…

not rated 7 2mo ago A 76 tokens

cuj-guardian

499

aiatelie/ai-atelie

Skill Claude Code

Run and triage AI Atelie's Critical User Journey (CUJ) for every PR — the single end-to-end test that proves a user can open the app, create a project, drive the Claude Code agent, and see the canvas render. Before running, gate by inspecting the PR diff for changes that plausibly affect the journey (routes…

not rated 7 3mo ago A 137 tokens original MIT

test-personas

500

archugunov/pm-job-search

Skill Claude Code

Internal test harness for plugin maintainers. End users should not invoke this. Runs synthetic personas through critical journeys against the plugin and produces LLM-judge findings reports. Trigger words "/test-personas", "run the test harness", "test the plugin end-to-end", "run the test personas".

not rated 7 11d ago A 64 tokens original MIT

PowerAppsControl

501

ilyafainberg/PowerAppsControl

Skill Claude CodeCodex

Drive the PowerAppsControl MCP server to UX-test a Power App end to end: open and verify an app URL, let the user choose a mode (smoke test = in-depth read-only exploration that produces a repeatable natural-language test plan; or run my test plan), then run it in a recorded session and produce a video + HTML report…

not rated 7 2mo ago A 177 tokens GPL-3.0

katalon-labs/true-skills

Skill Claude CodeCodex

Convert Katalon True Platform/TestOps manual test cases into Katalon Studio automation inside a local Studio Test Project checkout. Use when you need to author or extend a .tc test case file and its paired Groovy script under Scripts/, keep test case variable GUIDs consistent with the .ts test suite bindings that read…

not rated 7 12d ago A 204 tokens original MIT

browser-automation

503

EmilLindfors/rust-browser-mcp

Skill Claude Code

Enterprise-grade browser automation using WebDriver protocol. Use when the user needs to automate web browsers, perform web scraping, test web applications, fill forms, take screenshots, monitor performance, or execute multi-step browser workflows. Supports Chrome, Firefox, and Edge with connection pooling and health…

not rated 7 8mo ago A 61 tokens

buildwright

504

raunakkathuria/buildwright

Skill Claude CodeCodex

Lightweight engineering workflow for agent-led development. Provides plan, work, verify, ship, and analyse commands with TDD, documentation discipline, security review, code review, and quality gates.

not rated 7 1mo ago A 41 tokens

At most 3 mods per repository are shown here, and a mod shipped inside a plugin is left to that plugin's page — the rest are on their repository pages: