Testing agents

5,338 tagged Testing, measured the same way as everything else here.

Browse within: code-quality 57agent-orchestration 47harness 40spec-driven-development 40agentic-workflow 39Multi-Agent 38playwright 36agentic-coding 32github-copilot 31rtl 31verification 31agentic 29copilot 29context-engineering 27

stress-batch-tests

217

modelstudioai/cli

Agent

A manually run test system for sending many concurrent requests to different AI capabilities and producing Markdown or HTML reports. It is separate from the normal automated test suite, which checks smaller individual cases.

320 5d ago A 0 tokens original Apache-2.0

tester

218

IncomeStreamSurfer/claude-code-agents-wizard-v2

Agent Claude Code

Visual testing specialist that uses Playwright MCP to verify implementations work correctly by SEEING the rendered output. Use immediately after the coder agent completes an implementation.

317 +1 8mo ago A 32 tokens

panel-validator

219

revfactory/webtoon-harness

Agent Claude Code

A webtoon panel checker that reviews each rendered panel for visual consistency, readable and accurate Korean text, dialogue flow, and technical problems.

310 +2 2mo ago A 161 tokens original MIT

Plugin Tester

220

Fu-Jie/openwebui-extensions

Agent

End-to-end plugin testing agent for OpenWebUI. Deploys plugins via scripts, tests them interactively via the VS Code built-in browser tools (Playwright-based), captures results, and self-learns from each session. Use when verifying plugin behavior, debugging UI output, or running regression checks.

302 +1 1mo ago A 63 tokens original MIT

grader

222

neithhogg/lair

Agent Codex

Evaluate expectations against an execution transcript and outputs.

297 5mo ago A 0 tokens

horang-labs/tessera

Agent

The checked-in Windows launcher is the fail-closed boundary between an agent shell and every isolated packaged Electron child. Before each Start-Process, scripts/launch-electron-test-instances.ps1 snapshots and removes inherited agent-runtime state. Its finally block restores every saved process value after both…

297 6d ago A 0 tokens AGPL-3.0

dashclaw-gate-runner

224

ucsandman/DashClaw

Agent Claude Code

Runs the DashClaw verification gates (lint, full vitest suite, build, contract checks) and returns ONLY the failures plus a pass/fail verdict. Use to verify a change without dragging multi-hundred-line build/test logs into the main thread. Delegate gate-running here instead of running it inline.

296 4d ago A 69 tokens original MIT

platform-validator

226

mylee04/code-notify

Agent Claude Code

Validates notification functionality across platforms (macOS, Windows, Linux, WSL).

288 5d ago A 20 tokens original MIT

product-verifier

227

Undertone0809/rudder

Agent

Use this prompt for a spawned verifier child after writer implementation and writer checks. The verifier answers whether the product path works from the actor's side. It does not review architecture and it does not fix failures.

287 3d ago A 0 tokens original Apache-2.0

grader

229

cluesmith/codev

Agent Claude Code

Evaluate expectations against an execution transcript and outputs.

286 6d ago A 0 tokens copy · 100% Apache-2.0

tester

230

sigcli/sigcli

Agent Claude Code

Use this agent to write and run tests for the sigcli project. This agent creates unit tests using vitest and MemoryStorage, following existing test patterns, and runs the full test suite. Examples.

286 9d ago A 42 tokens original MIT

contract-review

231

Intelligent-Internet/zenith

Agent

Read-only adversarial contract reviewer. Reviews the full contract set against user scope, inventory, playbook rules, evidence feasibility, shortcut risk, and old-harness-style atomic assertion coverage before tasks are trusted.

283 +1 26d ago A 44 tokens original Apache-2.0

flow-validator

232

Intelligent-Internet/zenith

Agent

Leaf real-surface validation lane for a bounded subset of engineering assertions. Exercises assigned behavior through a parent-specified browser, API, CLI, background, artifact, data, library, parity, or caller-provided tool surface; writes evidence only to assigned paths.

283 +1 26d ago A 55 tokens original Apache-2.0

pavel-molyanov/molyanov-ai-dev

Agent

Reviews user-spec document quality: structure, interview coverage, acceptance-criteria testability, edge-case presence, contradictions, and template compliance. Use when: the user-spec is ready for pre-approval document review; solution adequacy and factual codebase claims are out of scope.

283 +2 10d ago A 62 tokens original MIT

finalize

234

vibe-motion/auto-motion

Agent Claude Code

Snapshot visual QA + one in-place fix pass + render. Dispatched only when Step 6 lint/inspect reports issues, or to do the final render.

281 +1 1mo ago A 0 tokens

AGENTS

235

jxnl/dots

Agent

Agent "AGENTS" from jxnl/dots, covering python, testing, git workflow, writing and content and autonomy.

277 1mo ago A 0 tokens

mutation-kill

236

bdfinst/agentic-dev-team

Agent

Autonomous mutation survivor-reduction loop — runs a scoped mutation tool, generates targeted tests for survivors in priority order, verifies they compile and pass, commits, and repeats until survivors stop decreasing. Gates on hard kills only (timeouts excluded). Complements the advisory /mutation-testing skill.

277 yesterday A 60 tokens original MIT

qa-lead

238

codemie-ai/codemie-code

Agent Claude Code

Use this agent when implementation is complete and code needs to be verified before committing or creating a PR. Triggers on phrases like "run quality gates", "check code quality", "run qa", "verify my changes", "pre-commit checks", "qa check", "act as qa lead", or when tech-lead or another agent suggests quality…

276 3d ago A 289 tokens original Apache-2.0

qa-reviewer

239

tmdgusya/roach-pi

Agent

Goal-completion QA + fraud gate. Independently re-verifies delivered behavior against the goal's successCriteria, hunts edge cases and regressions, and hunts fake implementations (stubs, hardcoding, assertion theater); returns PASS or FAIL via a fixed C1–C12 checklist.

275 1mo ago A 61 tokens

test-architect

240

CloudAI-X/opencode-workflow

Agent

Test strategy designer focusing on coverage, test design, and testing best practices. Use PROACTIVELY when adding tests, reviewing test coverage, or designing test approaches.

275 +1 7mo ago A 33 tokens original MIT