Testing skills

11,755 tagged Testing, measured the same way as everything else here.

Browse within: LLM 188agentic-ai 140agents 134cli 107ai-coding 93agent 81skills 66javascript 56openai 52agentic-workflow 41agent-orchestration 40claude-code-plugin 37agent-browser 36software-architecture 35

mrd-to-code-v2

721

system777/dev-workflow

Skill Claude CodeCodex

An end-to-end workflow that turns a product requirements document, or MRD, into planned and tested code. It can clarify requirements, create a more detailed plan, design the technical approach, generate code and tests, and archive the work.

not rated 2 4mo ago A 135 tokens

vibe-agent-toolkit

722

jdutton/vibe-agent-toolkit

Skill Claude CodeCodex

Use when starting VAT work or deciding which VAT sub-skill applies. Router that points at sub-skills for adoption, skill/agent authoring, audit, distribution, RAG, knowledge resources, skill review, and enterprise org admin.

not rated 2 changed yesterday A 54 tokens original MIT

behave

724

projectbluefin/testsuite

Skill Claude CodeCodex

How to write behave scenarios and step definitions for the testsuite repo. Load when editing .feature files or steps.py.

not rated 2 today A 27 tokens

browser-qa

725

liatrio-labs/ai-prompts

Skill Codex

Use when you need lightweight browser QA for a web page, local HTML file, or app: inspect console errors, broken assets, keyboard/focus behavior, viewport readability, and publish evidence-backed findings JSON through a local HTML report viewer.

not rated 2 12d ago A 51 tokens original Apache-2.0

browser-test-progress

726

Datuoba/browser-test-progress

Skill Claude CodeCodex

Show and maintain a discrete-node progress panel while ChatGPT uses its built-in browser for an authorized website acceptance test. Use when the user starts an @Browser acceptance task, asks to test a website with the in-app browser, or asks to show browser test progress. Do not use for ordinary browsing, research, or…

not rated 2 1mo ago A 72 tokens original MIT

lab2

727

Bruce1986/vibe-to-prod-lab

Skill Claude Code

A Traditional Chinese teaching guide for a software lab on golden datasets and prompt regression. A golden dataset is a set of expected examples used to check whether changes still produce acceptable results.

not rated 2 17d ago A 45 tokens original MIT

linkskill-bench

728

orbitlinktracer/linkskill-bench

Skill Claude CodeCodex

Benchmark and evaluate the quality of any AgentSkill. Use when asked to test, evaluate, benchmark, or assess a skill's effectiveness. Triggers on phrases like 'benchmark this skill', 'evaluate skill quality', 'test this skill', 'how good is this skill', 'skill audit', 'skill assessment'. Works by generating test…

not rated 2 5mo ago A 115 tokens

hetzner-ansible-lab

729

Bitbull-Ideas/hermes.skills

Skill Claude CodeCodex

Provision temporary Hetzner Cloud VMs for Ansible role QA, bootstrap OS prerequisites such as Rocky 8 Python 3.9, run verified work, clean up cloud resources, and report sanitized evidence.

not rated 2 2mo ago A 51 tokens AGPL-3.0

eval-skills

730

JarvixGaby/eval-skill

Skill Claude CodeCodex

Evaluate, benchmark, or test-drive an unfamiliar AI agent skill, tool bundle, prompt workflow, or capability package. Use when the user asks whether a downloaded/shared .skill, SKILL.md, agent workflow, or reusable AI capability actually helps, is worth installing, works as advertised, performs better than asking an…

not rated 2 1mo ago A 114 tokens original MIT

alltest

731

Epsilondelta-ai/alltest

Skill Claude CodeCodex

Full project test coverage. Analyzes the entire codebase, identifies coverage gaps, and writes tests to achieve 100% coverage (minimum 80%). Ensures every source file with exportable logic has at least one test. Use when the user mentions 'alltest', wants comprehensive testing, or asks for full test coverage.

not rated 2 6mo ago A 69 tokens original MIT

mvp-test-strategist

732

mkashyap00/business-idea-claude-skills

Skill Claude Code

Selects the optimal testing mechanism, designs the minimum viable test, and actively executes the creation of the test assets (e.g., landing pages, scripts) using tool calling. Reads the GTM plan and outputs a live test environment and execution timeline.

not rated 2 3mo ago A 57 tokens

dstest

734

bxrne/dstest

Skill Claude CodeCodex

Deterministic simulation testing for containerized services. Write Lua scripts to inject chaos (pause, kill, resource deprivation) into Docker containers with reproducible, seeded fault injection. Use when writing chaos experiments, testing service resilience, or debugging distributed systems.

not rated 2 changed 4d ago A 53 tokens original MIT

aoa-evals-skills

735

8Dionysus/aoa-evals

Skill Codex

Route the aoa-evals skill family for central proof selection, review, evolution, named results or verdicts, source-linked reports, Eval Forge owner review, and proof lifecycle. Hand repository-local eval selection, application, intake/design, or session-hit classification to aoa-eval. Candidates, readiness checks…

not rated 2 2d ago A 84 tokens original Apache-2.0

tdd

736

JustinThomas2/agentrc

Skill Claude CodeCodex

How to write good tests - behavior-focused assertions, the AAA structure, red/green/refactor, small vertical slices, and what to mock versus leave real. Use whenever writing a new test, modifying or fixing an existing test, reviewing someone else's tests, deciding whether something needs a test, choosing what to mock…

not rated 2 1mo ago A 76 tokens original MIT

refactor-py

737

meshulga/agents-doc

Skill Claude Code

Apply standard Python refactoring patterns and re-run pytest.

not rated 2 3mo ago A 16 tokens original MIT

andrejkaxz/Vanessa_for_AI

Skill Codex

A skill for creating, editing, checking, and diagnosing Russian-language Turbo Gherkin `.feature` files for Vanessa Automation, a tool for testing 1C applications. It covers UI and export scenarios without MCP.

not rated 2 12d ago A 105 tokens

careerchain-ys/stdd

Skill Claude Code

A guide for documenting and testing an existing software feature or page by studying how its code already works. It creates requirement, design, and test documents that describe the current behavior rather than an imagined future version.

not rated 2 1mo ago A 137 tokens original Apache-2.0

e2e-tests

740

GlamgarOnDiscord/claude-saas-blueprint

Skill Claude Code needs its repo

Tests E2E Playwright pour SaaS Next.js : setup, Page Object Model, auth state, flows critiques (login, billing, onboarding), CI GitHub Actions.

not rated 2 3mo ago A 40 tokens original MIT

flomeile/leo-starter

Skill Claude CodeCodex

Trigger: "system-optimierung", "system optimieren", "optimierungslauf", "pruefset fahren", "prüfset fahren", "messlauf", "regeltreue messen". Misst mit kalten Prüfset-Läufen (dein Prüfset, beim ersten Mal aus der Vorlage 10System\Pruefset-Vorlage.md angelegt), ob das System seine eigenen Regeln einhält, leitet aus…

not rated 2 8d ago A 139 tokens

ping-pong-tdd

742

danethurber/.dotfiles

Skill Claude CodeCodex

Pair on an implementation via ping-pong TDD — alternating red/green rounds between the agent and the user. Use when the user says "ping-pong" or asks to alternate writing failing tests and making them pass.

not rated 2 3d ago A 51 tokens

atlas-benchmark

743

madaeroblade/atlas

Skill Claude CodeCodex

Benchmark whether Atlas measurably improves coding-agent behavior, by running every scenario twice — with and without Atlas — in clean contexts and scoring the pair blind.

not rated 2 27d ago A 35 tokens original MIT

qa-automation

744

azam-sdet/10xquality

Skill Claude CodeCodex

AI-powered QA test automation — record browser flows, generate test cases from PRDs or Figma designs, execute natural language test scripts, and convert to Playwright + Cucumber BDD tests. Use when asked to "record a test", "create test cases", "execute test script", "automate UI test", "convert to BDD", "generate…

not rated 2 6mo ago A 89 tokens original MIT

At most 3 mods per repository are shown here, and a mod shipped inside a plugin is left to that plugin's page — the rest are on their repository pages: