Testing agents

5,338 tagged Testing, measured the same way as everything else here.

Browse within: code-quality 57agent-orchestration 47harness 40spec-driven-development 40agentic-workflow 39Multi-Agent 38playwright 36agentic-coding 32github-copilot 31rtl 31verification 31agentic 29copilot 29context-engineering 27

eval-author

337

dynamics365ninja/d365fo-mcp-server

Agent Claude Code

Authoring role for the D365FO agent eval loop catalog. Drafts a new eval/cases/ .json spec (valid against eval/cases/schema.json), scaffolds its golden folder, and sets goldenpending until the golden is captured on the VM. Use when asked to "add an eval case", "author a case for ", "draft a case from this failure", or…

136 3d ago A 92 tokens original MIT

eval-implementer

338

dynamics365ninja/d365fo-mcp-server

Agent Claude Code

Implementer role of the D365FO agent eval loop. Runs ON THE VM (mcp-server in full mode + C# bridge) against the Contoso sandbox model. Takes an eval case id, drives the grounded MCP tool path to implement it, builds, scores against the golden/SysTest oracle, writes a corpus record, and rolls back. Use when asked to…

136 3d ago A 104 tokens original MIT

eval-improver

339

dynamics365ninja/d365fo-mcp-server

Agent Claude Code

Improver role of the D365FO agent eval loop. Reads the corpus of run records, ranks failure clusters, reproduces a TOOLDEFECT/KNOWLEDGEGAP/VALIDATORGAP as a minimal VM-free repo test, fixes it, validates against the held-out split, and opens a PR citing corpus evidence. Runs in the repo (never touches the VM). Use…

136 3d ago A 111 tokens original MIT

txtify-qa

340

lkmeta/txtify

Agent Claude Code

QA gatekeeper for Txtify. Use to run the full verification ladder on the current tree and report evidence — before merging, releasing, or when asked "does everything still work?".

135 18d ago A 41 tokens original Apache-2.0

grader

341

zjunlp/DataMind

Agent Claude Code

Evaluate expectations against an execution transcript and outputs.

135 3d ago A 0 tokens

harness-enhancer

342

ZhangShenao/harness9

Agent Claude Code

A full-repository quality review and improvement agent for a Go project. Go is a programming language, and a repository is the project folder containing its source code and supporting files.

135 4d ago A 97 tokens original MIT

build-verifier

343

vincenthopf/My-Claude-Code

Agent

Verifies implementation work by running code, executing tests, and generating adversarial test cases. Checks against the problem definition and principles, not against the proposal document. Approval is gated on execution results, not code reading.

135 5mo ago A 47 tokens fork archived

verifier

344

maxwell2732/paper-replicate-agent-demo

Agent Claude Code

End-to-end verification agent. Checks that slides compile, render, deploy, and display correctly. Use proactively before committing or creating PRs.

134 3mo ago A 31 tokens

testing

345

windviki/vBookmarks

Agent

Agent "testing" from windviki/vBookmarks, covering testing & real-browser harness (detail), unit tests (detail), manual testing checklist and headless smoke test (docker).

134 +1 3d ago A 0 tokens original MIT

grader

346

raphaelmansuy/edgeparse

Agent Codex

Evaluate expectations against an execution transcript and outputs.

133 4mo ago A 0 tokens

developer

347

bonigarcia/context-engineering

Agent

Implement Idea and scoreidea in src/backlog.py, and verify them with tests/testbacklog.py. Use the formula impact 5 + strategicfit 3 - effort 2.

132 4d ago A 0 tokens original Apache-2.0

jira-nyquist

348

koolamusic/claudefiles

Agent

Part of jira

Validates that the executed sprint actually meets PLAN.md's Nyquist criteria. Fills coverage gaps by writing tests for any criterion not already covered. Implementation files are READ-ONLY. Spawned by /jira:execute after jira-executor finishes.

132 4d ago A 54 tokens original MIT

grader

349

chaohong-ai/ai-auto-work

Agent Claude Code

Evaluate expectations against an execution transcript and outputs.

132 +1 1mo ago A 0 tokens

n8n-webhook-tester

350

czlonkowski/n8n-mcp-cc-buildier

Agent Claude Code

Use this agent when you need to test n8n webhook endpoints with specific input data and verify their execution results. This agent creates bash test scripts (handling JWT tokens when needed), executes them, and cleans up afterward. Keeps test execution isolated from the main conversation context.

132 10mo ago A 62 tokens

MCP-Tester

351

jamesmontemagno/ohmyposh-configurator

Agent

Tests the Oh My Posh Configurator MCP server. Use this agent to validate MCP tools like listing segments, creating configurations, validating configs, and exporting in different formats.

132 1mo ago A 39 tokens

browser-check

352

bitsocialnet/5chan

Agent Claude Code

Verifies UI changes in the browser using playwright-cli across Blink, Gecko, and WebKit. Use after making visual or interaction changes to React components, CSS, layouts, or routing to confirm they render and behave correctly.

132 2d ago A 47 tokens GPL-3.0

unit-test-writer

353

haacked/dotfiles

Agent

Use this agent when you need to write comprehensive unit tests for existing code, when implementing test-driven development, when code coverage needs improvement, or when refactoring requires test safety nets. Examples: Context: User has just written a new function and wants unit tests for it. user: 'I just wrote this…

131 5d ago A 228 tokens

fuzzer

355

ccashwell/evm-cortex

Agent

Foundry fuzz testing, Foundry invariant testing, Medusa and Echidna stateful fuzzing campaigns.

127 23d ago A 24 tokens original MIT

protocol-builder

356

zhensherlock/protocol-launcher

Agent

A coding workflow for adding new Protocol Launcher tools, which open applications through special links. It creates the implementation, exports, configuration entries, and unit tests needed for a new protocol.

127 6d ago A 18 tokens original MIT

php-test-validator

359

aaddrick/claude-pipeline

Agent Claude Code

Validates PHPUnit test comprehensiveness and integrity. Use after code review to audit PHP/Laravel tests for cheating, TODO placeholders, insufficient coverage, or hollow assertions. Reports failures requiring developer subagent correction.

126 6mo ago A 46 tokens original MIT

grader

360

panaversity/ksor

Agent Codex

Evaluate expectations against an execution transcript and outputs.

126 3d ago A 0 tokens copy · 100% Apache-2.0

At most 3 mods per repository are shown here, and a mod shipped inside a plugin is left to that plugin's page — the rest are on their repository pages: