Testing agents

5,338 tagged Testing, measured the same way as everything else here.

Browse within: code-quality 57agent-orchestration 47harness 40spec-driven-development 40agentic-workflow 39Multi-Agent 38playwright 36agentic-coding 32github-copilot 31rtl 31verification 31agentic 29copilot 29context-engineering 27

verifier

361

madebyaris/advance-minimax-m3-cursor-rules

Agent Cursor

Validates completed work. Use after tasks are marked done to confirm implementations are functional. Invoke with /verifier when you need to verify code actually works.

125 2mo ago A 34 tokens original MIT

stax-verifier

362

cesarferreira/stax

Agent Claude Code

Verifies stax code changes by running cargo check, clippy, and targeted nextest runs. Reports exact errors with file:line references and actionable fix suggestions.

123 3d ago A 38 tokens original MIT

pr-test-reviewer

363

go-to-k/cdkd

Agent Claude Code

Review test adequacy for a PR — find coverage gaps, mock anti-patterns, fixture realism issues. Read-only — never writes or edits.

122 yesterday A 34 tokens original Apache-2.0

typst-verify

365

lucifer1004/claude-skill-typst

Agent

Verify Typst document output against requirements. Use after compilation when you need to confirm content, structure, or visual layout correctness.

121 9d ago A 30 tokens original MIT

uc-coverage

366

AI-Unified-Process/marketplace

Agent

Part of aiup-vaadin-jooq

Read-only auditor that checks whether a use case (UC-XXX) or test case (TC-XXX) is completely implemented and completely tested against its specification. Use it during or after implementation and during or after writing tests: it maps every main success scenario step, alternative flow, business rule, precondition…

120 +2 4d ago B 108 tokens original Apache-2.0

debugger

367

vinilana/invokta

Agent

Use for reproducing failures, isolating root causes, and implementing evidence-backed regression fixes.

120 +3 9d ago A 21 tokens original MIT

CODEX

368

UiPath/coder_eval

Agent

Run OpenAI Codex as the agent under evaluation in Coder Eval — installation, authentication, task configuration, and how Codex telemetry maps to sandboxed, weighted scoring.

119 5d ago A 35 tokens original Apache-2.0

a11y-auditor

369

Houseofmvps/ultraship

Agent

Part of ultraship

Runs the static accessibility (WCAG 2.2) audit using the a11y-scanner tool. Dispatched by /ship for scorecard generation.

119 1mo ago A 39 tokens original MIT

judge

370

avibebuilder/claude-prime

Agent Claude Code

Decide whether the skill improvement loop should continue or stop. You are independent from the agent that wrote the improvements — your only job is to look at the evidence and make an honest call.

119 3mo ago A 0 tokens original MIT

python-test-engineer

371

anam-org/metaxy

Agent Claude Code

Part of metaxy

Use this agent when you need to create new tests, fix failing tests, refactor test code, or improve test organization and maintainability. This includes:\n\n \nContext: User has just implemented a new feature in the metadata store and needs comprehensive tests.\nuser: "I've added a new fallback store chain feature.…

119 12d ago A 358 tokens original Apache-2.0

qa

372

anam-org/metaxy

Agent Claude Code

Part of metaxy

Use this agent when:\n\n1. A logical unit of work has been completed (feature implementation, bug fix, refactoring)\n2. Code changes are ready for review before committing or creating a pull request\n3. You need to verify that acceptance criteria and definition of done are met\n4. After making changes to test files to…

119 12d ago A 494 tokens original Apache-2.0

test-validator

373

revfactory/claude-code-harness

Agent Claude Code

A test and usage layer for a small programming language, with four suites covering the lexer, parser, interpreter, and integration. It also includes example programs and a REPL, an interactive prompt for entering code and seeing results.

118 6mo ago A 0 tokens

integration-tester

374

revfactory/claude-code-harness

Agent Claude Code

An end-to-end integration test suite for five order scenarios: success, insufficient stock, payment failure, simultaneous orders, and timeout. End-to-end tests check the complete flow across connected parts of a system.

118 6mo ago A 0 tokens

test-generator

375

jellydn/my-ai-tools

Agent

Generates comprehensive, meaningful tests for code changes. Focuses on testing behavior and edge cases rather than implementation details.

118 4d ago A 26 tokens original MIT

qa

376

espennilsen/pi

Agent

QA specialist that tests running web applications against acceptance criteria.

117 10d ago A 12 tokens original MIT

integration-tester

377

ammarion/waf-detector

Agent Claude Code

Use this agent when you need to perform comprehensive end-to-end testing of a system with multiple components (CLI, web interface, APIs). This includes testing functionality, integration points, error handling, and accuracy validation. The agent should be invoked after significant code changes, before releases, or…

116 14d ago A 0 tokens

grader

378

Changan-Su/Forsion

Agent

Evaluate expectations against an execution transcript and outputs.

114 4d ago A 0 tokens

test-porter

379

openshift-eng/ai-helpers

Agent

Part of ci

Automated Ginkgo e2e test porting agent. Ports tests from openshift-tests-private to openshift/origin, creates PRs, monitors CI, responds to review feedback, pushes fixes, and escalates to humans when needed. Use this agent for any task related to porting tests between these repos.

114 6d ago A 67 tokens original Apache-2.0

ChrisRoyse/610ClaudeSubagents

Agent

Expert in API and microservice integration testing, contract validation, and service mesh testing. Orchestrates comprehensive API testing strategies including REST, GraphQL, gRPC, and event-driven architectures with advanced contract testing and service virtualization.

114 1y ago A 52 tokens

app-operator

382

bladeofgod/flutter-ai-harness

Agent Claude Code

A runtime UI verification agent for apps using Marionette, a tool for connecting to and controlling an app during development. It follows an already approved behavior specification and writes a run report.

113 22d ago A 49 tokens original MIT

grader

383

zhangdszq/teamclaw

Agent

Evaluate expectations against an execution transcript and outputs.

113 5mo ago A 0 tokens copy · 100% Apache-2.0

At most 3 mods per repository are shown here, and a mod shipped inside a plugin is left to that plugin's page — the rest are on their repository pages: