Testing agents

5,338 tagged Testing, measured the same way as everything else here.

Browse within: code-quality 57agent-orchestration 47harness 40spec-driven-development 40agentic-workflow 39Multi-Agent 38playwright 36agentic-coding 32github-copilot 31rtl 31verification 31agentic 29copilot 29context-engineering 27

agent-sdk-verifier-py

265

fazxes/Claude-code

Agent

Part of agent-sdk-dev

Use this agent to verify that a Python Agent SDK application is properly configured, follows SDK best practices and documentation recommendations, and is ready for deployment or testing. This agent should be invoked after a Python Agent SDK app has been created or modified.

233 5mo ago A 55 tokens

checker

266

cordwainersmith/Claudoscope

Agent

Fresh-context adversarial verification of completed work. Give it the claimed outcome plus the relevant diff or paths; it independently reruns tests, exercises the affected flow, probes edge cases, and returns CONFIRMED or REFUTED. Read-and-run only; it never plans, edits, or fixes anything.

233 +1 7d ago A 63 tokens original MIT

grader

267

lexler/skill-factory

Agent

Evaluate expectations against an execution transcript and outputs.

231 +1 7d ago A 0 tokens copy · 100% Apache-2.0

evaluator

268

OdradekAI/bundles-forge

Agent

Part of bundles-forge

Use when running one side of an A/B skill evaluation or chain verification. Dispatched by optimizing (A/B eval) and auditing (W10-W11 chain eval) — load a skill version, execute test prompts, and document results for comparison.

230 4mo ago A 53 tokens original Apache-2.0

grader

272

JimLiu/science-skills

Agent

Evaluate expectations against an execution transcript and outputs.

224 +1 2mo ago A 0 tokens

testing-pr-security

273

icoretech/airbroke

Agent

Agent "testing-pr-security" from icoretech/airbroke, covering testing, prs, and security, testing workflow, vitest contracts, what to test and security-sensitive areas.

222 3d ago A 0 tokens original MIT

pgEdge/pgedge-postgres-mcp

Agent Claude Code

Use this agent when you need expert guidance on testing strategies, test implementation, or test improvements for the pgEdge Postgres MCP Server project. Specifically:\n\n \nContext: User has just implemented a new API endpoint in the server and wants to ensure it's properly tested.\nUser: "I've added a new endpoint…

221 +1 7d ago A 0 tokens PostgreSQL

embedded-qa

275

DunCanYounG-1/auto-embedded

Agent

Use when verifying embedded competition project at CP-3: static analysis, MIL/SIL/PIL three-tier validation, 5-tuple scoring checklist verification, and one-click firmware pipeline. Independent verifier — never implements code, only audits.

218 +2 1mo ago A 51 tokens original MIT

output-ux-reviewer

276

ItamarZand88/CLI-Anything-WEB

Agent

Part of cli-anything-web-plugin

Review a cli-web- CLI from the end-user perspective by RUNNING it. Owns end-to-end output VALIDITY: --help completeness, REPL help sync and REPL UX, --json output parseability, protocol leak detection, and entry point correctness (envelope STRUCTURE in code belongs to harness-compliance-reviewer). Returns scored…

216 9d ago A 85 tokens original MIT

gem-implementer

277

mubaidr/gem-team

Agent

TDD code implementation: features, bugs, refactoring. Never reviews own work.

215 5d ago A 22 tokens original Apache-2.0

grader

278

CamusGIT/EvoQuant

Agent

Evaluate expectations against an execution transcript and outputs.

215 +11 16d ago A 0 tokens copy · 100% Apache-2.0

grader

279

spytensor/openmozi

Agent

Evaluate expectations against an execution transcript and outputs.

210 +18 26d ago A 0 tokens copy · 100% MIT

JayCRL/MobileVC

Agent Claude Code

Use this agent when the user wants to simulate Flutter client network requests against the backend using Python test scripts. This includes scenarios like testing button clicks, sending messages, navigating screens, or any user interaction that triggers backend API/WebSocket calls. The agent first reads Flutter code…

209 2mo ago B 355 tokens original MIT

comparator

281

DandanLLab/legadoSkill

Agent

Compare two outputs WITHOUT knowing which skill produced them.

203 +2 5mo ago A 0 tokens copy · 100% MIT

grader

282

DandanLLab/legadoSkill

Agent

Evaluate expectations against an execution transcript and outputs.

203 +2 5mo ago A 0 tokens copy · 100% MIT

tdd-guide

283

ab604/claude-code-r-skills

Agent

Part of claude-code-r-skills

Test-driven development specialist for R. Enforces test-first development with testthat. Use when writing new features, fixing bugs, or refactoring code.

197 5mo ago A 34 tokens original MIT

quality-assurance

284

agents-universe/agents-universe

Agent

Business-oriented QA agent – verify user and business outcomes, design and run automated tests, preserve execution evidence, report concise Jira results.

195 3d ago A 26 tokens original Apache-2.0

test-coverage-agent

286

kid-sid/claude-spellbook

Agent Claude Code

Use this agent to analyse an entire module or directory for missing test coverage, then generate the missing tests. Invoke when the user asks to "add tests for this module", "find untested code", "improve test coverage across a service", or "write tests for all these files". Prefer this over the inline /test-gen…

187 28d ago A 84 tokens original MIT

product-architect

287

iusztinpaul/squid

Agent

Part of squid

Grooms raw tasks into agent-ready specs (acceptance criteria + BDD scenarios) AND does final user-perspective acceptance review after the Tester passes. Use whenever a task needs to be turned into something the SWE can build, or whenever a task needs the final "is this actually right for users?" review before commit.

184 +1 1mo ago A 68 tokens original Apache-2.0

software-engineer

288

iusztinpaul/squid

Agent

Part of squid

Implements a single groomed task assigned by the orchestrator. Writes code and tests locally. Does NOT commit until the Tester has reviewed and approved. Use when a task is groomed and ready for implementation, or when the Tester has returned feedback that needs to be addressed.

184 +1 1mo ago A 59 tokens original Apache-2.0

At most 3 mods per repository are shown here, and a mod shipped inside a plugin is left to that plugin's page — the rest are on their repository pages: