Testing agents

5,348 tagged Testing, measured the same way as everything else here.

Browse within: claude-code-plugin 142ai-coding 100code-quality 51agentic-coding 50agent-orchestration 45ai-coding-assistant 45ai-skills 42Multi-Agent 36ai-development 36copilot 35agent-tools 34agentic 34context-engineering 33agentic-workflow 32

os-dev

529

objectstack-ai/objectstack

Agent Claude Code

Developer agent for exactly ONE GitHub issue, dispatched by the /pm-dispatch PM loop. Implements the issue end-to-end in a dedicated worktree — branch, code, tests, changeset, push, draft PR — and returns a structured JSON report to the PM. Use only with a single fully-specified issue as input; never for open-ended or…

not rated 48 changed today C 80 tokens original Apache-2.0

screenshot-review

530

YuDefine/nuxt-supabase-starter

Agent Claude Code

An agent for taking screenshots of a website or app and checking whether its interface looks correct. It returns a report based on the requested visual checks.

not rated 45 today A 73 tokens original MIT

implementer

531

rjmurillo/ai-agents

Agent Codex

Ship production code from approved plans. Tests alongside code. Atomic commits.

not rated 45 changed today A 17 tokens original MIT

e2e-coverage-checker

532

incubateur-ademe/benefriches

Agent Claude Code

Analyzes branch changes vs main and determines whether e2e tests need to be created or updated. Produces a structured report with coverage gaps and recommendations.

not rated 45 2d ago A 39 tokens original MIT

api-security-tester

533

agigante80/actual-mcp-server

Agent Claude Code

Generates and runs comprehensive, non-destructive security tests for the actual-mcp-server MCP transport covering OWASP API Top 10, JSON-RPC injection, auth bypass (OIDC + static bearer), per-user budget ACL (IDOR), ActualQL/SQL injection, malformed input, and error-leakage. Use when writing security tests, expanding…

not rated 48 today A 94 tokens original MIT

code-reviewer

534

agigante80/actual-mcp-server

Agent Claude Code

Elite code review expert for security vulnerabilities, correctness bugs, performance, and maintainability. Runs the project's static analysis, security scanning, and tests as part of the review. Use PROACTIVELY for code quality assurance.

not rated 48 today A 48 tokens original MIT

test-specialist

535

martinkup/symfony-profiler-optimization-advisor-bundle

Agent Claude Code

PHPUnit testing specialist who creates and maintains tests through comprehensive test strategies, fixture management, and quality validation. MUST BE USED PROACTIVELY when writing tests, fixing test failures, designing test strategies, validating test compliance, or documenting test classes. Can run concurrently with…

not rated 45 3mo ago A 63 tokens

TheAstrelo/Claude-Pipeline

Agent Claude Code

Multi-perspective critique to stress-test designs. Runs 3 critic passes with different viewpoints to surface blind spots.

not rated 44 today A 28 tokens original MIT

refactor-engineer

537

gracefullight/stock-checker

Agent Codex

Behavior-preserving refactoring specialist. Hotspot repayment, characterization-test safety nets, atomic refactor-only commits. Never changes observable behavior.

not rated 44 today A 32 tokens

tool-developer

538

undergroundrap/UEFN-TOOLBELT

Agent Claude Code

Builds new UEFN Toolbelt tools autonomously. Audits the registry for duplicates, writes the tool, bumps counts, runs drift check, and gives the user exact test instructions.

not rated 44 2d ago A 38 tokens

e2e-tester

541

loonghao/auroraview

Agent

Visual E2E testing agent using ProofShot and agent-browser for automated UI verification, regression testing, and self-iterating bug detection.

not rated 44 1mo ago A 34 tokens original MIT

grader

542

1024XEngineer/bytemind

Agent

Evaluate expectations against an execution transcript and outputs.

not rated 43 1mo ago A 0 tokens copy · 100% MIT

openapi-specialist

545

infobip/infobip-openapi-mcp

Agent Claude Code

Use this agent when tasks involve reading, analyzing, debugging, or transforming OpenAPI specifications — especially in the context of how this framework converts OpenAPI operations into MCP tools. Examples: Context: Developer is adding a new test OpenAPI spec and wants to understand how it will be parsed. user: "Will…

not rated 43 yesterday A 436 tokens original MIT

phpt-author

546

lisachenko/native-php-matrix

Agent Claude Code

Use to write new .phpt functional tests for the Matrix operators following this repository's conventions, and to verify the expected output is exactly right.

not rated 43 15d ago A 33 tokens original MIT

test-runner

547

lisachenko/native-php-matrix

Agent Claude Code

Use to run the .phpt functional suite (or a single test) and report exactly which behaviours broke and why, including engine-level crashes.

not rated 43 15d ago A 33 tokens original MIT

tdd-guide

548

VenTheZone/pi-dots

Agent

Test-driven development specialist enforcing write-tests-first workflows with comprehensive coverage. Use when developing new features, fixing bugs with regression tests, guiding developers through the red-green-refactor cycle, or ensuring 80%+ unit, integration, and E2E coverage.

not rated 43 27d ago A 55 tokens

feature-implementer

549

DGouron/review-flow

Agent Claude Code

Implement a ReviewFlow feature via TDD inside-out. Receives a validated plan (docs/plans/ .plan.md) and spec (docs/specs/ .md), implements all layers with RED-GREEN-REFACTOR, self-reviews, and persists a report. Triggers when the user says "implement feature", "start implementation", "implement spec", or "run…

not rated 42 today A 89 tokens original MIT

tdd

550

DGouron/review-flow

Agent Claude Code

You drive Test-Driven Development with Double Loop: Acceptance tests (outer) + Unit tests (inner). You operate autonomously.

not rated 42 today A 0 tokens original MIT

comparator

551

devdavv/unity-ai-workflow

Agent Claude Code

Compare two outputs WITHOUT knowing which skill produced them.

not rated 41 5mo ago A 0 tokens copy · 100% MIT

At most 3 mods per repository are shown here, and a mod shipped inside a plugin is left to that plugin's page — the rest are on their repository pages: