Testing agents

5,574 tagged Testing, measured the same way as everything else here.

Browse within: code-quality 55agent-orchestration 54agentic-workflow 46agentic-coding 42spec-driven-development 40playwright 36Multi-Agent 34ai-security 32cybersecurity 32github-copilot 30agentic 29context-engineering 28ai-development 27ai-skills 27

hk-test-writer

433

deepklarity/harness-kit

Agent Claude Code

You are a test-writing specialist for the harness-kit monorepo. You receive a function, module, or feature to test and produce high-quality tests that follow this project's conventions exactly.

95 1mo ago A 0 tokens original MIT

eval-writer

434

GoogleCloudPlatform/cxas-scrapi

Agent Codex

Generate eval YAMLs for one entire eval type (all goldens, all sims, all tooltests, or all callbacktests) in a single dispatch. Reads the TDD's Coverage Map, the agent's actual tools and variables, then writes the appropriate file(s) — see "File layout per type" for what each type requires (sims are one file by runner…

94 6d ago A 113 tokens original Apache-2.0

review

436

tony/claude-code-riper-5

Agent Claude Code

Validation and quality assurance - ruthlessly verify implementation against plan.

93 17d ago A 13 tokens original MIT

code-reviewer

437

UnpaidAttention/fable5-methodology

Agent

Part of fable5-methodology

Adversarially reviews a diff cold — without the reasoning that produced it — for correctness, safety, design, and scope, hunting specifically for fake progress, silently dropped requirements, weakened tests, and scope creep. Delegate to code-reviewer for any non-trivial diff before it is accepted or committed…

92 1mo ago A 0 tokens

qa-verifier

438

UnpaidAttention/fable5-methodology

Agent

Part of fable5-methodology

Independently verifies a completed change against its acceptance criteria by running the tests/build/lint itself and probing edge cases — never trusting the implementer's claims. Delegate to qa-verifier after any builder (or your own) implementation, before accepting it as done. Requires the change and its acceptance…

92 1mo ago A 92 tokens

e2e-runner

439

krishnakanthb13/everything-antigravity

Agent

End-to-end testing specialist using Vercel Agent Browser (preferred) with Playwright fallback. Use PROACTIVELY for generating, maintaining, and running E2E tests. Manages test journeys, quarantines flaky tests, uploads artifacts (screenshots, videos, traces), and ensures critical user flows work.

90 6mo ago A 69 tokens

false-green-hunter

440

guardana/guardana

Agent Claude Code

Read-only adversarial reviewer for Guardana. Hunts the failure this project exists to prevent — code that compiles, types, tests green, and quietly reports "all clear" about something it never examined. Use before a release, after a subsystem lands, or when a green gate needs to be distrusted on purpose.

89 3d ago A 69 tokens original Apache-2.0

gate-runner

441

guardana/guardana

Agent Claude Code

Runs Guardana's full gate — lint, format, strict types, import contract, tests with coverage floors, dogfood, generated docs and the three isolated example suites — and reports what actually passed. Use when the answer to "is this green" has to be trustworthy, and to keep a long, noisy run out of the main conversation.

89 3d ago C 73 tokens original Apache-2.0

strategy-fidelity-voc

442

rwliebs/Dossier

Agent Cursor

Evaluates app fidelity and completion against docs/SYSTEMARCHITECTURE.md and domain references. Serves as voice of customer: defines user workflows and outcomes, then validates implementation against them. Use proactively before releases, after major changes, or when validating feature completeness.

88 1mo ago A 0 tokens

qa

443

developer2013/bricks-mcp-open

Agent

Du bist der QA Agent — testet Bricks-Pages auf Qualität, Accessibility und Performance.

85 12d ago A 0 tokens original MIT

grader

445

zby/commonplace

Agent

Evaluate expectations against an execution transcript and outputs.

85 3d ago A 0 tokens CC-BY-4.0

cook

446

Rune-kit/rune

Agent

Part of rune

Feature implementation orchestrator — handles 70% of requests. Full TDD cycle: understand → plan → test → implement → verify → commit. Use for ANY code modification (features, bugs, refactors, security).

84 18d ago A 46 tokens original MIT

andrewstellman/quality-playbook

Agent

Part of quality-playbook

Prompt template for the AI session driving an end-to-end QPB calibration cycle. The orchestrator AI executes Steps 1-12 from aicontext/CALIBRATIONPROTOCOL.md, spawns playbook subprocesses per benchmark, and writes the cycle audit + Lever Calibration Log entry. Designed for Claude Code sessions but will work in any…

83 1mo ago B 0 tokens original Apache-2.0

quality-playbook

448

andrewstellman/quality-playbook

Agent

Part of quality-playbook

AUTOMATION ONLY — DO NOT INVOKE FROM AN INTERACTIVE CODING SESSION. Run a complete quality engineering audit on any codebase. Orchestrates six phases — explore, generate, review, audit, reconcile, verify — each in its own context window via sub-agents. Then runs iteration strategies to find even more bugs. Finds the…

83 1mo ago A 87 tokens original Apache-2.0

quality-playbook

449

andrewstellman/quality-playbook

Agent

Part of quality-playbook

AUTOMATION ONLY — DO NOT INVOKE FROM AN INTERACTIVE CODING SESSION. Run a complete quality engineering audit on any codebase. Orchestrates six phases — explore, generate, review, audit, reconcile, verify — each in its own context window for maximum depth. Then runs iteration strategies to find even more bugs. Finds…

83 1mo ago A 86 tokens original Apache-2.0

coder

450

snyk/snyk-ls

Agent Cursor

Implementation specialist that writes production code using TDD and commits changes. Use proactively when implementing features, fixing bugs, or writing code for a confirmed plan. Delegates to planner when requirements are unclear or need updating. Hands over to qa when implementation is complete.

83 4d ago A 53 tokens original Apache-2.0

qa

451

snyk/snyk-ls

Agent Cursor

QA specialist that deeply analyzes code produced by the coder agent. Runs the verification skill, traces code paths, checks logic for gaps, unintended changes, edge cases, and omissions. Use proactively after implementation is complete, when coder says "done", or when asked to review/verify code quality.

83 4d ago A 60 tokens original Apache-2.0

feature-tester-e2e

452

AIBiz-Automatyzacje/claude-code-starter

Agent Claude Code

Weryfikuje scenariusze E2E w przeglądarce przez agent-browser. Uruchamia scenariusze checkboxów [E2E] (oba prefiksy: Test: i Weryfikacja:) z checklist zadań — responsywność, interakcje, nawigację klawiaturą, visual regression — i zwraca przebieg PASS/FAIL/SKIP per checkbox z dowodem. Nie pisze seedów, nie modyfikuje…

82 9d ago A 136 tokens

gsd-verifier

453

itsjwill/gsd-pro

Agent

Verifies phase goal achievement through goal-backward analysis. Checks codebase delivers what phase promised, not just that tasks completed. Creates VERIFICATION.md report.

81 5mo ago A 36 tokens original MIT

green-agent

454

mloda-ai/mloda

Agent Claude Code

TDD Green Phase specialist - writes minimal code to make failing tests pass.

81 4d ago A 17 tokens original Apache-2.0

red-agent

455

mloda-ai/mloda

Agent Claude Code

TDD Red Phase specialist - writes failing tests that define requirements.

81 4d ago A 15 tokens original Apache-2.0

qa

456

flplima/tmuxy

Agent Claude Code

QA agent that runs rotating test styles and creates/updates GitHub Issues for findings.

81 2d ago A 18 tokens original MIT

At most 3 mods per repository are shown here, and a mod shipped inside a plugin is left to that plugin's page — the rest are on their repository pages: