Testing agents

5,616 tagged Testing, measured the same way as everything else here.

Browse within: ai-coding 105code-quality 55agent-orchestration 54agentic-coding 47agentic-workflow 46spec-driven-development 40playwright 36Multi-Agent 34ai-security 32cybersecurity 32github-copilot 30agentic 29context-engineering 28ai-development 27

vbw-qa

457

swt-labs/vibe-better-with-claude-code-vbw

Agent

Part of vbw

Verification agent using goal-backward methodology to validate completed work. Read-only (permissionMode plan). Persists verification results via write-verification.sh through Bash.

80 2mo ago A 36 tokens original MIT

soldier

458

kucherenko/gangsta

Agent

Part of gangsta

Use this agent for stateless code execution — implementing a single work package with TDD enforcement. Receives a work package brief, writes failing test, implements minimal code, verifies test passes, returns tribute.

80 20d ago A 44 tokens original MIT

dom-extraction-tester

459

WebMCP-org/npm-packages

Agent Claude Code

Use this agent when you need to test progressive DOM reading implementations by navigating to websites and extracting specific information. This agent works as a driver that receives instructions from a navigator AI about what website to visit and what data to extract, then attempts the extraction and reports back on…

80 3d ago A 275 tokens original MIT

js-unit-test-writer

460

woocommerce/woocommerce-paypal-payments

Agent Claude Code

Jest test writer for the plugin's JavaScript/TypeScript (React components, data-store selectors, handlers, helpers). MUST BE USED whenever JS/TS unit tests need to be written or updated - the colocated .test.js files next to module resources/js sources.

80 2d ago A 65 tokens GPL-2.0

harness-implementer

461

panayiotism/claude-harness

Agent

Part of claude-harness

Implements a single claude-harness feature end-to-end in an isolated context - acceptance tests first (ATDD), implementation, verification, checkpoint (commit/push/PR via gh), optional merge. Spawned by the /claude-harness:flow skill with a structured feature prompt; not intended for ad-hoc use.

79 1mo ago A 73 tokens original MIT

browser-tester-v2

462

lipas-liikuntapaikat/lipas

Agent Claude Code

Use this agent to perform manual browser testing of implemented features using Claude in Chrome (MCP). Delegate to this agent when you need to verify that a feature works correctly in the browser, test UI interactions, check for console errors, or validate user flows. Provide context about what was implemented and…

78 3d ago A 69 tokens original MIT

wshaddix/dotnet-skills

Agent

Part of dotnet-skills

WHEN designing .NET benchmarks, reviewing benchmark methodology, or validating measurement correctness. Avoids dead code elimination, measurement bias, and common BenchmarkDotNet pitfalls. Triggers on: design a benchmark, review benchmark, benchmark pitfalls, how to measure, memory diagnoser setup.

78 6mo ago A 61 tokens

ahtavarasmus/lightfriend

Agent Claude Code

Use this agent when the user needs to write or improve unit tests for the SMS assistant functionality, particularly for validating tool calls, agent behavior, and proper handling of external service responses. This agent should be used when:\n\n \nContext: User has just implemented a new tool call for the SMS…

77 3d ago A 0 tokens AGPL-3.0

delphi-tester

466

adrianosantostreina/delphi-dev

Agent

Part of delphi-dev

Subagente especializado em implementacao de testes unitarios DUnitX para projetos Delphi. Opera em dois modos: MODO EXPLICITO: Use quando o usuario solicitar /tdd, "crie testes", "implemente testes", "quero cobertura de testes", "teste unitario", "DUnitX". Nesse modo, analisa o projeto completo e gera a suite de…

79 2d ago A 291 tokens original MIT

corpus-prover

468

assaio/assaio

Agent Claude Code

Proves a change on the maintainer's real local corpus — what number moved, by how much, and that nothing else did. Use before calling any measurement change done. Never writes to the real store.

76 5d ago A 47 tokens original Apache-2.0

ai-evaluator

473

Productfculty-aipm/PM-Copilot-by-Product-Faculty

Agent

Part of PM-Copilot-by-Product-Faculty

Designs and runs AI product evaluation frameworks: error analysis, eval suite design, LLM-as-judge pipelines, human eval protocols, regression testing plans, and improvement flywheels. Use this agent when the user is building an AI-powered feature and needs to define how to measure quality, catch regressions, or…

75 4mo ago A 282 tokens

audio-quality-check

474

drakulavich/kesha-voice-kit

Agent Claude Code

Use after any commit touching rust/src/tts/ to objectively check that the TTS engine still produces sane audio for a fixed Russian + English corpus. Computes RMS, silence ratio, sample rate, channel count, length-vs-text ratio. Flags suspicious WAVs (all-silence, clipping, wrong rate, monosamples, length 10x off vs…

73 3d ago A 105 tokens original MIT

integ-test-runner

476

Apra-Labs/apra-fleet

Agent

Runs integ-test-playbook.md per cycle to close or assess this cycle's implemented features and verify-set beads (any issuetype, all children closed) against real evidence; closes passing ones, files [integ] bugs for failures.

72 6d ago A 53 tokens

harness-coder

477

Junhanliu-dev/espalier-engineering

Agent

Part of espalier-engineering

Implementation agent for {projectname} — writes code that follows the project's Espalier rules, layer specs, and Solution Selection Ladder (conventions first, correctness within them, clarity then brevity break ties). Spawned by the pipeline at Stage 3 (implementation — under folded test-mode this includes writing the…

72 changed today A 129 tokens original MIT

harness-reviewer

478

Junhanliu-dev/espalier-engineering

Agent

Part of espalier-engineering

Review agent for {projectname} — checks a diff against the project's Espalier conventions, layer boundaries, runtime surfaces, production-readiness seeds, test meaningfulness, and (advisory) minimalism + readability. Spawned fresh by the pipeline each Stage 4 review round (code AND its tests, one verdict) and for the…

72 changed today A 114 tokens original MIT

done-gate

479

edutrul/drupal-ai

Agent Claude Code

Runtime validator — runs builds, tests, and drush cr. Checks deliverable completeness and documentation. Run after quality-gate passes.

71 2mo ago A 31 tokens original MIT

review-runner

480

sunnykgupta/AI-Helpers

Agent

Full code review agent that runs linting, typechecks, and tests. Provides comprehensive feedback on code quality, correctness, and test coverage.

70 1mo ago A 28 tokens original MIT

At most 3 mods per repository are shown here, and a mod shipped inside a plugin is left to that plugin's page — the rest are on their repository pages: