bmad-eval-runner
577Skill Claude Code
Run a skill's evals and report results. Use when the user wants to evaluate a skill, run evals, benchmark a skill, validate triggers, optimize a description, or grade skill outputs.
11,717 tagged Testing, measured the same way as everything else here.
Browse within: LLM 179agentic-ai 140agents 130cli 103ai-coding 92agent 82skills 71javascript 57openai 52agent-browser 50ai-testing 42agentic-workflow 41agent-orchestration 40claude-code-plugin 37
Skill Claude Code
Run a skill's evals and report results. Use when the user wants to evaluate a skill, run evals, benchmark a skill, validate triggers, optimize a description, or grade skill outputs.
Skill Claude Code
Generate test suites from implementation files. Supports Jest, Vitest, and pytest. Produces AAA-structured tests covering happy path, error paths, and edge cases.
Skill Claude Code
Use when a user wants to evaluate or test a Rasa assistant using the eval scenario framework — at any stage of the process: setting up the evaluation infrastructure for the first time, deciding which scenario types to cover for a flow, writing or generating scenario YAML files, learning about supported assertion types…
Skill Claude CodeCodex
Use when implementing features, fixing bugs (including P0 hotfixes and production incidents), refactoring, removing dead code, or deleting legacy paths. Use when user mentions TDD, red-green-refactor, test-first, or when tempted to write code before tests. Use when fixing any bug — even trivial ones. Use when…
Skill Claude CodeCodex
Analyze Sui Move test coverage, identify untested code, write missing tests, and perform security audits. Includes Python tools for parsing coverage output and generating reports.
Skill Claude CodeCodex
Produces practical, risk-based testing guidance and minimal test plans for features or changes. Use when user asks what to test, how to pick test cases (boundaries, permissions, state machines), how to improve weak tests, or to review existing tests. Covers equivalence partitions, boundary values, decision tables, and…
Skill Claude CodeCodex
Drives development with tests. Use when implementing any logic, fixing any bug, or changing any behavior. Use when you need to prove that code works, when a bug report arrives, or when you're about to modify existing functionality.
iamBrzDev/enterprise-agent-skills
Skill Claude CodeCodex
Implements enterprise-grade testing strategies for .NET applications using xUnit, Moq, FluentAssertions, and WebApplicationFactory. Use when writing unit tests, integration tests, or API tests in .NET or C#, configuring test projects, deciding what to mock vs what to test with real dependencies, setting up shared test…
StuMinch/sauce-labs-agent-skills
Skill Claude CodeCodex
Sauce Labs VDC capability authoring skill for desktop and virtual devices. Use when asked to generate or validate W3C/Appium capabilities for Sauce Labs.
Skill Claude CodeCodex
Tests and evaluates any Claude Code skill for structural validity, quality, and trigger accuracy. Implements the cc-plugin-eval 4-stage pipeline (Analysis → Generation → Execution → Evaluation) and the 4D scoring rubric (Documentation/Code/Completeness/Usability 25% each). Use before packaging or deploying any skill.
Skill Codex
Create reproducible before-and-after evidence for software changes using annotated PNG, JPEG, MP4, backend HTTP scenarios, verified web flows, or Android/iOS app checks. Use for Playwright or Chrome DevTools MCP verification, UI/design fixes, visual regressions, interaction recordings, Android adb or iOS Simulator…
Skill Claude CodeCodex
Use when you need to generate a complete test suite for a project — dispatches test-strategist, unit-test-writer, integration-test-writer, e2e-test-writer, and test-reviewer in sequence to produce strategy, implementation, and review artifacts. Covers test pyramid design, unit/integration/E2E implementation, coverage…
Skill Claude CodeCodex needs its repo
Use this harness to validate Agentweaver's MCP protocol surface, capture the complete tools/call request/response evidence, and emit a normalized agentweaver.persona-judge-verdict/v1 JSON verdict. It is for MCP end-to-end validation, MCP tool-contract regression checks, and investigation of an MCP-reported issue; use…
Skill Claude CodeCodex
Analyze PR or local code changes and find Zebrunner test cases to run for regressions and new coverage gaps. Use when the user mentions test impact, PR test planning, which tests to run, regression coverage, sprint PR rollups, or Zebrunner + pull request / code changes.
shawnclybor/clybor-claude-tooling
Skill Claude CodeCodex
Run the Yellow Sheet chain end to end against ANY matter, from the original PDFs, and read the result honestly. Use whenever someone asks what has been tested end to end, whether a matter produces a usable sheet, how to onboard a NEW matter, or asks to re-run after a code change. Carries the intake mechanism that…
Skill Claude CodeCodex needs its repo
Writing or modifying tests for a feature, Hono route, repository, Vue page, or CLI command.
Skill Claude CodeCodex
Audit steam-games-mcp — build/test/lint gate, live MCP tool edge-case sweep (input validation, SteamID64/vanity/appid edge cases, key-gating), and source-level code review. Use when asked to test/audit the published or just-fixed steam-games-mcp package, hunt for bugs/edge cases, or repeat "the same kind of testing as…
Skill Claude CodeCodex
Acts as the QAI Consultant marketing/PR specialist. Use whenever Gabi asks to write, draft, or plan a social media post (LinkedIn, Facebook, Instagram) promoting the QAI Consultant app or MCP server - release announcements, educational QA tips, milestones, case studies, or behind-the-scenes. Covers Romanian and…
Skill Claude Code
Give an agent skill's decision logic a test suite and a deterministic implementation — a labeled validation dataset plus versionable Python — via an adversarial multi-persona loop that runs natively in Claude Code: proposer plus persona subagents through the Task tool, entirely on your Claude Code subscription, no…
Skill Claude Code
Generate integration tests for a Java class using Testcontainers. Supports Spring Boot (3.1+), Quarkus (3.0+), and Micronaut (4.0+). Detects class type (repository, controller, service) and applies framework-appropriate test patterns. Requires Java 17+.
Skill Claude CodeCodex
Quality harness for design-dna Phase 3 output with browser-based visual verification. After an agent generates a design from a Design DNA JSON + user content, this skill acts as a verification and scoring layer — collecting all page resources via console/network inspection, performing section-by-section screenshot…
Skill Codex
Test Agent Skills from tests/skillcases.yaml files. Use when validating whether a Skill triggers correctly, produces a dry-run plan, satisfies output contracts, handles edge/failure cases, or when asked to run Skill Test cases, judge Agent responses against YAML expectations, or produce Skill test reports.
Skill Claude CodeCodex needs its repo
Write OpenAPI examples that work with Microcks dispatchers for API mocking. Covers example pairing rules, JSON body dispatching, Groovy script dispatching, and dispatcher configuration via the Microcks API.
Skill Claude Code
Configure Bun's built-in test runner with Jest-compatible APIs. Use when setting up testing infrastructure, writing unit/integration/snapshot tests, migrating from Jest, or configuring test coverage. 3-10x faster than Jest.
At most 3 mods per repository are shown here, and a mod shipped inside a plugin is left to that plugin's page — the rest are on their repository pages: