Testing skills

11,709 tagged Testing, measured the same way as everything else here.

Browse within: LLM 179agentic-ai 140agents 130cli 103ai-coding 92agent 82skills 71javascript 57openai 52agent-browser 50ai-testing 42agentic-workflow 41agent-orchestration 40claude-code-plugin 37

e2e-alertmanager-test

601

conallob/o11y-analysis-tools

Skill Claude CodeCodex needs its repo

Render end-to-end previews of what an alert notification will actually look like (plain-text email, HTML email, Slack attachment JSON, raw webhook JSON) by replaying a Prometheus unit-test file's expected alerts through a live Alertmanager. Effectively a "print preview" for alerts. Use to review notification…

not rated 4 1mo ago A 153 tokens original BSD-3-Clause

stm

602

mdsohaib/screenshot-time-machine

Skill Claude Code

Screenshot every page of the running localhost dev server and report which pages changed since the last snapshot. Use after editing anything users can see (pages, components, CSS/Tailwind, layouts, templates) to visually verify before saying you're done, and when the user says "check the UI", "does it look right"…

not rated 4 2d ago A 112 tokens original MIT

thinking-kernel

603

tienenwu/fables

Skill Claude CodeCodex

Use when working in a technical domain with no existing playbook, writing new coding guidelines / review criteria / checklists from scratch, judging whether a verification method actually proves a claim (test passed but does it count?), or when repeated fixes keep failing and the direction feels wrong. Domain-agnostic…

not rated 4 1mo ago A 82 tokens original MIT

persona-ux-test

604

Ericwong5021/persona-ux-test

Skill Claude CodeCodex

Run an isolated persona-based UX test through the real desktop or browser UI. Creates a precise non-developer user persona, gives the tester only an approved product introduction, prevents source-code and design-document leakage, and produces an evidence-based Chinese evaluation report. Use when the user asks for…

not rated 4 1mo ago A 106 tokens original MIT

devtest

605

AILiteracyLab/Claude-Build-Test-Loop

Skill Claude CodeCodex

Spec-Build-Test loop — the user defines a spec, then three agents iterate (Builder implements, Tester validates, Supervisor monitors for freezes) until the result matches. Works for any digital function — UI, APIs, CLI tools, conversational AI, data pipelines, and more.

not rated 4 6mo ago A 58 tokens original MIT

xiexie-qiuligao/agent-evidence-mcp

Skill Claude CodeCodex

Capture screenshots, short recordings, and milestone evidence during long-running agent tasks. Use when an agent is asked to perform multi-step browser, desktop, QA, troubleshooting, deployment, or admin workflows where the user wants checkpoint artifacts, progress evidence, error snapshots, or a final timeline of…

not rated 4 4mo ago A 65 tokens original MIT

rag-eval

607

LucasSantana-Dev/hitgate

Skill Claude CodeCodex needs its repo

Run the retrieval regression gate against the current repo state and report whether a recent change helped, hurt, or held steady.

not rated 4 25d ago A 0 tokens original MIT

jacoco

608

alexmond/jhelm

Skill Claude Code

Check JaCoCo code coverage for jhelm modules.

not rated 4 5d ago A 13 tokens original Apache-2.0

Driftya/code-meridian

Skill Codex

Plan focused tests with CodeMeridian by finding relevant test shields, coverage gaps, impacted behavior, and the smallest useful test set before implementation.

not rated 4 2d ago A 35 tokens original MIT

unreal-playtest-agent

610

dcc-mcp/dcc-mcp-unreal

Skill Claude Code

Domain skill - run screenshot-light PIE playtest episodes with structured entity observations, bounded semantic actions, transition polling, and in-memory traces for QA and external policy or RL runners.

not rated 4 changed 2d ago A 41 tokens

mcplab-assistant

611

inspectr-hq/mcplab

Skill Claude CodeCodex

Operator guide for MCPLab config authoring, Test Case Assistant workflows, execution, and result analysis. Use when users need to create or refine test cases from runs/traces, suggest deterministic checks or value capture, write or debug MCPLab eval YAML, run or queue evaluations, troubleshoot failures, or compare…

not rated 4 changed 2d ago A 71 tokens original

looper-qa

612

quangdang46/looper_rust

Skill Claude CodeCodex

Use when a Looper-managed GitHub repo needs scheduled pre-merge QA — a PR carries the looper:qa label, the spec stage reaches looper:spec-ready, or the Looper reviewer loop requests an independent second pass. Runs the full QA cycle (Looper state probe → PR checkout → ffs code review → language-specific test suite →…

not rated 4 19d ago A 125 tokens original MIT

coco-delivery

613

pcopu/coco

Skill Claude CodeCodex

Implement and verify CoCo features end-to-end (Telegram commands, callbacks, app-server transport, queueing, watchdogs, approvals, and tests). Use when changing this repository's bot behavior and needing repo-specific file targets, workflows, and validation commands. NOT for generic Python tasks outside CoCo.

not rated 4 4d ago A 65 tokens original MIT

opentester

614

kznr02/OpenTester-Skills

Skill Claude Code

Automated testing execution using OpenTester DSL. Use when the user wants to create tests, run tests, validate test syntax, or manage test projects. Supports CLI testing with a YAML-based DSL.

not rated 4 6mo ago A 43 tokens original MIT

code-review

615

SebaBoler/vanguard

Skill Claude CodeCodex

Use when reviewing a code change or diff for correctness, security, missing tests, and convention violations before opening or approving a PR. Review independently and adversarially, then fix high-confidence issues.

not rated 4 2d ago A 42 tokens original MIT

quality-check

616

Crearize/ai-dev-helm

Skill Codex needs its repo

A required pre-merge quality gate that runs local static checks, tests, and reviews. A merge is allowed only after the checks pass.

not rated 4 changed 2d ago A 84 tokens original MIT

check-work

617

codingmydna/grokers

Skill Claude CodeCodex

Check your work with a verification subagent that reviews diffs, runs builds and tests, and evaluates correctness. Read this file for instructions. Use when asked to "check work", "verify changes", "self-verify", "/check-work", "/check", "/verify", or "/self-verify".

not rated 4 1mo ago A 64 tokens copy · 100% Apache-2.0

validate

618

aiocean/claude-plugins

Skill Claude Code needs its repo

Run the marketplace validation script to check plugin integrity. Use when finishing plugin work, after adding or modifying plugins, or before committing plugin changes.

not rated 4 4d ago A 30 tokens original MIT

flutter-testing

619

jyotiraditya-chauhan/test-kit

Skill Claude Code

Writes unit, widget, golden, and integration tests for Flutter projects, detecting whether BLoC, Riverpod, Provider, or GetX is in use from pubspec.yaml before writing any business-logic test. Use when the user asks to write Flutter tests, add test coverage to a widget/bloc/provider/repository, test a .dart file, or…

not rated 4 17d ago A 96 tokens original MIT

test-first-bugfix

620

The-Artificer-of-Ciphers-LLC/skills-from-the-artificer

Skill Claude CodeCodex

Test-driven bug fixing — reproduce before you fix. Use this skill whenever the user reports a bug, describes unexpected behavior, says something is broken, mentions a regression, or asks you to fix an error. This includes phrases like "this is broken", "X doesn't work", "there's a bug in", "getting an error when", "it…

not rated 4 +1 5d ago A 138 tokens original MIT

setup-skills-evals

621

ahnafyy/skills-evals

Skill Claude CodeCodex

Set up the skills-evals library in a repository — discover agent artifacts, interview the user about what to test, scaffold eval cases, and wire CI and local runners. Use when the user wants to set up skills-evals, test their agent skills, add evals for skills or custom agents, check why a skill isn't triggering, or…

not rated 4 +1 1mo ago A 87 tokens original MIT

ui-validation

622

Dallionking/ui-validation-kit

Skill Claude Code

Validate UIs by clicking through them like a real user — iOS Simulator, Android emulator, tvOS, desktop, and web. Detects the platform and picks the right tool — agent-device (Callstack) for mobile/TV/desktop, agent-browser (Vercel Labs) for web, Maestro + Maestro Viewer for declarative cross-platform flows; raw xcrun…

not rated 4 +1 1mo ago A 200 tokens original MIT

test-in-serum

623

Celian-mrc/serum-mcp

Skill Claude Code needs its repo

Run the pre-real-Serum-test checklist on one or more .SerumPreset files -- automated CBOR wire-type scan plus a human-readable summary -- then hand off to the user for the real load-it-in-Serum test. Use this after any generatepreset/editpreset call on serum-mcp, or whenever a preset file needs to be verified before…

not rated 3 5d ago A 108 tokens original MIT

docs

624

SpecLeft/specleft

Skill Claude CodeCodex

Skill "docs" from SpecLeft/specleft, covering specleft cli reference, setup, workflow, quick checks and safety.

not rated 3 4mo ago A 0 tokens original Apache-2.0

At most 3 mods per repository are shown here, and a mod shipped inside a plugin is left to that plugin's page — the rest are on their repository pages: