Testing agents

8,775 tagged Testing, measured the same way as everything else here.

Browse within: code-quality 54agent-orchestration 44spec-driven-development 43agentic-workflow 39harness 36playwright 36Multi-Agent 35context-engineering 34agentic-coding 32ai-coding-assistant 32github-copilot 31rtl 31verification 30copilot 27

composio-community/awesome-claude-plugins

Agent

Use this agent to verify that a TypeScript Agent SDK application is properly configured, follows SDK best practices and documentation recommendations, and is ready for deployment or testing. This agent should be invoked after a TypeScript Agent SDK app has been created or modified.

1.9k +5 1mo ago A 56 tokens

backend-architect

98

Dicklesworthstone/pi_agent_rust

Agent

Expert backend architect specializing in scalable API design, microservices architecture, and distributed systems. Masters REST/GraphQL/gRPC APIs, event-driven architectures, service mesh patterns, and modern backend frameworks. Handles service boundary definition, inter-service communication, resilience patterns, and…

1.7k +10 today A 72 tokens

Dicklesworthstone/pi_agent_rust

Agent

Build production-ready monitoring, logging, and tracing systems. Implements comprehensive observability strategies, SLI/SLO management, and incident response workflows. Use PROACTIVELY for monitoring infrastructure, performance optimization, or production reliability.

1.7k +10 today A 49 tokens

arm-cortex-expert

100

Dicklesworthstone/pi_agent_rust

Agent

Senior embedded software engineer specializing in firmware and driver development for ARM Cortex-M microcontrollers (Teensy, STM32, nRF52, SAMD). Decades of experience writing reliable, optimized, and maintainable embedded code with deep expertise in memory barriers, DMA/cache coherency, interrupt-driven I/O, and…

1.7k +10 today A 72 tokens

e2e-tests-engineer

101

happier-dev/happier

Agent Claude Code

Deprecated placeholder. Do not use; follow root AGENTS.md and package instructions instead.

1.6k +15 today A 24 tokens original MIT

test

102

EtienneLescot/n8n-as-code

Agent

Give your AI agent n8n superpowers. 537 nodes with full schemas, 7,700+ templates, Git-like sync, and TypeScript workflows.

1.6k +1 4d ago not scanned tokens not measured original MIT

tdd-guide

103

rohitg00/skillkit

Agent

Test-Driven Development specialist. Write tests first, then implement minimal code to pass.

1.5k 3mo ago A 20 tokens original Apache-2.0

02-palette

104

inkline/inkline

Agent

Seat: the design language. Owns: ui/components/ — the single-source component catalog: headless parts, styled compositions, .styleframe.ts styling, colocated tests. (Story files under stories/ are @herald's craft.).

1.5k 2d ago A 0 tokens

consistency-qa

105

hyhmrright/brooks-lint

Agent Claude Code

The brooks-lint verification gate. Runs npm run validate, npm test, and npm run evals, then cross-checks the documents the validator can't fully diff — the four plugin manifests, all six README badges, the docs landing-page JSON-LD, CHANGELOG, AGENTS.md, GEMINI.md, and the derived book count — for drift. Reports…

1.4k +10 yesterday A 123 tokens original MIT

eval-curator

106

hyhmrright/brooks-lint

Agent Claude Code

Authors and maintains the brooks-lint eval suite in evals/evals.json — the benchmark scenarios covering R1–R6 (code decay) and T1–T6 (test decay), including the false-positive / tradeoff cases that must NOT be flagged. Ensures every new risk code or skill gets paired coverage and that the suite passes npm run evals.…

1.4k +10 yesterday A 97 tokens original MIT

benchmark-reviewer

107

evo-hq/evo

Agent

Reviews an evo benchmark in two modes. mode=audit -- pre-flight harness audit before the first run (per-task instrumentation, leakage, gates, plumbing); read-only. mode=review-experiment -- post-commit per-task failure analysis for a specific experiment; reads per-task traces and the eval-runner log, writes per-task…

1.4k +6 1mo ago A 112 tokens original Apache-2.0

verifier

108

evo-hq/evo

Agent

Read-only audit of one evo experiment for design-time cheating (pre-phase) or result-time validity (post-phase). Catches test-set leakage in training data, subsetted eval commands, missing gates for new artifacts, generic hypotheses, cache short-circuits, fake artifacts, and score-reproducibility failures. Returns…

1.4k +6 1mo ago A 146 tokens original Apache-2.0

PolyArch/humanize

Agent

Analyzes a plan and generates multiple-choice technical comprehension questions to verify user understanding before RLCR loop. Use when validating user readiness for start-rlcr-loop command.

1.4k +6 5d ago A 39 tokens

ci-watcher

110

nrwl/nx-console

Agent Cursor

Polls Nx Cloud CI pipeline and self-healing status. Returns structured state when actionable. Spawned by /monitor-ci command to monitor CI Attempt status.

1.4k 4d ago A 35 tokens copy · 98% MIT

grader

111

Prismer-AI/PrismerCloud

Agent

Evaluate expectations against an execution transcript and outputs.

1.4k 26d ago A 0 tokens copy · 100% MIT

e2e-runner

113

cfrs2005/claude-init

Agent Claude Code

An end-to-end testing assistant built around Playwright, a tool that controls real browsers to test complete user journeys.

1.4k +1 5mo ago A 70 tokens original MIT

qdhenry/Claude-Command-Suite

Agent Claude Code

Azure DevOps and cloud infrastructure specialist with comprehensive knowledge of all Azure services. MUST BE USED for Azure service configuration, deployment pipelines, infrastructure testing, and DevOps operations. Expert in using Azure CLI (az command) via Bash for all Azure operations, Azure Resource Manager, and…

1.3k 6mo ago A 67 tokens

test-team-leader

117

nrslib/takt

Agent

You are a team leader for E2E testing. Your job is to decompose a task into independent subtasks.

1.3k 3d ago A 0 tokens original MIT

verifier

118

higress-group/himarket

Agent

Verify that implemented features actually work by executing realistic functional scenarios against a running application.

1.3k +1 13d ago A 0 tokens

grader

119

CreminiAI/skillpack

Agent

Evaluate expectations against an execution transcript and outputs.

1.2k +1 20d ago A 0 tokens copy · 100% MIT

Ido-Levi/Hephaestus

Agent Claude Code

Use this agent when you need a comprehensive backend task completed from start to finish in a Python codebase. This includes implementation, testing, validation, and proper integration. Examples:\n\n \nContext: User needs a new API endpoint implemented with full test coverage.\nuser: "I need to add a POST /api/users…

1.2k +1 9mo ago A 0 tokens