Testing agents

7,605 tagged Testing, measured the same way as everything else here.

Browse within: code-quality 54agent-orchestration 45spec-driven-development 43agentic-workflow 39harness 36playwright 36Multi-Agent 35context-engineering 34agentic-coding 32ai-coding-assistant 32github-copilot 31rtl 31verification 30copilot 27

nspady/google-calendar-mcp

Agent Claude Code

PROACTIVELY use when adding new calendar features or modifying event handlers. Specializes in writing comprehensive test suites for Google Calendar MCP tools, including edge cases like timezone conversions, recurring events, multi-calendar scenarios, and error conditions. Ensures >90% code coverage.

1.2k 3mo ago A 59 tokens original MIT

qa

122

sheshbabu/zen

Agent Claude Code

Use this agent when you need to test recent code changes using Playwright automation. Examples: Context: The user has just implemented a new login feature and wants to test it. user: "I just added a new login validation feature, can you test it?" assistant: "I'll use the qa agent to test your recent changes with…

1.2k +3 4d ago A 0 tokens AGPL-3.0

browser-screenshots

123

PostHog/posthog.com

Agent

Every visual change needs before/after screenshots in four states: light and dark, narrow window and wide window. Anything that moves needs a before/after GIF as well. This guide explains how to capture them, and how to put them in the PR.

1.1k 2d ago A 0 tokens

vc-validate-agent

124

withkynam/vibecode-pro-max-kit

Agent Claude Code

VALIDATE MODE - Convert a written plan into an executable contract. Runs two-layer parallel fan-out (infra, test coverage, breaking changes, security + per-section feasibility agents), synthesizes findings, presents validate-menu to user, then writes validate-contract section into the plan file. Mandatory phase…

1.1k +7 2mo ago A 74 tokens original MIT

e2e-verifier

125

K9i-0/ccpocket

Agent Claude Code

An end-to-end testing agent for Flutter apps, meaning tests that operate the app as a user would on a simulator.

1.0k 3d ago A 65 tokens original MIT

Code Review (Gemini)

126

microsoft/agentrc

Agent ✓ vendor

Code review following VS Code contribution standards — correctness, lifecycle, naming, layering, accessibility, and security.

1.0k +3 7d ago A 26 tokens copy · 97% MIT

Code Review (Opus)

127

microsoft/agentrc

Agent ✓ vendor

Code review following VS Code contribution standards — correctness, lifecycle, naming, layering, accessibility, and security.

1.0k +3 7d ago A 26 tokens copy · 94% MIT

example-generator

128

galkahana/PDF-Writer

Agent Claude Code

Creates example test files for PDF-Writer community requests from GitHub issues or email content.

1.0k 2mo ago B 20 tokens original Apache-2.0

qa

129

piratuks/invoice-builder

Agent

Use this agent when reviewing a change for correctness, regressions, and verification in Invoice Builder.

970 5d ago A 21 tokens original MIT

ap-implementer

130

Spielewoy/autoprompt-skill

Agent

L3 executor - G4 IMPLEMENT. Builds one feature from its approved executable roadmap item or conditional frozen plan using strict TDD and real test runs; coverage >=95% on changed lines. Reports PLAN-CONFLICT rather than improvising.

959 +17 3d ago A 52 tokens original MIT

api-tester

131

ccplugins/awesome-claude-code-plugins

Agent

Use this agent for comprehensive API testing including performance testing, load testing, and contract testing. This agent specializes in ensuring APIs are robust, performant, and meet specifications before deployment. Examples:\n\n \nContext: Testing API performance under load.

929 +3 20d ago A 54 tokens original Apache-2.0

psi-oss/get-physics-done

Agent

Designs numerical experiments, parameter sweeps, convergence studies, and statistical analysis pipelines for physics computations.

920 +5 1mo ago A 26 tokens

validate-module

133

Auties00/Cobalt

Agent Claude Code

Validates one WA module against its Cobalt counterpart(s) across static parity AND observable live-runtime parity by writing and executing a Java scratch file, then applies fixes.

915 +1 1mo ago A 36 tokens original MIT

pbip-validator

134

data-goblin/power-bi-agentic-development

Agent

Validate Power BI Project (PBIP) file structure, TMDL syntax, and PBIR JSON schemas. Dispatch when the user asks to "validate my PBIP project", "check if the rename cascade is complete", "is this visual.json valid", or "my PBIP won't open".

886 +2 24d ago A 63 tokens GPL-3.0

CTO Agent

135

Salomondiei08/oh-my-hermes

Agent

You move a product through Understand, Design, Build, Check, Ship, and Learn. You coordinate specialists, maintain focus, and communicate decisions to the founder. You own the whole product lifecycle, not only engineering. A pull request is evidence inside the build stage, not the goal.

857 +2 1mo ago A 3 tokens

Pipelex/pipelex

Agent

Audience. Future-me (or any agent) the next time a make agent-test / pytest -n auto run in this repo hangs without finishing. The common causes here are xdist worker crash-and-replace cycles and fixture-teardown hangs; the iteration loop below generalizes to any hanging suite.

852 2d ago A 0 tokens original MIT

report-writer

137

H-mmer/pentest-agents

Agent Claude Code

Security report generation agent. Use for compiling findings into formal penetration test reports, executive summaries, technical write-ups, and bug bounty submissions. Provide the findings directory or list of vulnerabilities to document.

815 +2 2mo ago A 42 tokens

ssti-hunter

138

H-mmer/pentest-agents

Agent Claude Code

Server-Side Template Injection specialist. Covers Jinja2 (H1 #74), Twig, Velocity, FreeMarker, ERB, Handlebars, Thymeleaf. Use for any rule-engine, comment/message rendering, PR automation, admin template, or user-customizable template surface. Systematic blocklist mapper + CVE bypass runner + runtime-vs-parse…

815 +2 2mo ago A 83 tokens

lighteval-porter

139

groq/openbench

Agent Claude Code

Use this agent when you need to port an evaluation benchmark from the LightEval framework to openbench. This includes converting LightEval task definitions, dataset loaders, metrics, and scoring functions to the Inspect AI framework used by openbench. The agent should be invoked when the user mentions porting…

813 7d ago A 305 tokens original MIT

agent_b

140

NVIDIA/TileGym

Agent ✓ vendor

You are Agent B (Convert & Compile). You write the device kernel only: kernel.rs, the standalone Cargo project for the in-Rust pipeline test, generated canonical IR, and concise reports.

806 2d ago A 0 tokens

agent_e

141

NVIDIA/TileGym

Agent ✓ vendor

You are Agent E (Benchmark). Run tilegym pytest --print-record and report. Do NOT edit kernel or host code. One narrow exception (STEP 0.5): if testperf is missing cutile-rs in its backend parametrize, add it yourself — do NOT route to another agent.

806 2d ago A 0 tokens

f1-test-drive

142

cyrusagents/cyrus

Agent Claude Code

Orchestrate F1 test drives to validate the Cyrus agent system end-to-end. Use this agent to run comprehensive test drives that verify issue-tracker, EdgeWorker, and renderer components.

792 +1 3d ago A 43 tokens original Apache-2.0

tool-ui-reviewer

143

assistant-ui/tool-ui

Agent Claude Code

Quality gate for Tool UI components. Use after examples and documenter complete to verify pattern compliance and run checks.

773 3mo ago A 27 tokens original MIT

trellis-implement

144

fy-agent/fyagent

Agent Cursor

Trellis implementation agent. Use this exact agent for Trellis task implementation, implement.jsonl context injection, and hook-injection tests. Do not use generic/default/generalPurpose agents for Trellis implementation. No git commit allowed.

755 2d ago A 51 tokens