Testing agents

5,338 tagged Testing, measured the same way as everything else here.

Browse within: code-quality 57agent-orchestration 47harness 40spec-driven-development 40agentic-workflow 39Multi-Agent 38playwright 36agentic-coding 32github-copilot 31rtl 31verification 31agentic 29copilot 29context-engineering 27

craftsmanship-scout

289

gotgenes/pi-packages

Agent

Fresh-context craftsmanship scout — reads the largest test files and sweeps method-level design, naming, and test-code quality into a scored debt inventory for phase planning.

184 3d ago A 31 tokens

verifier

291

techygarg/lattice

Agent

Part of lattice

Runs a project's configured verification stages (build/unit/integration/etc.) from .lattice/verification.yaml via the deterministic runner script, then returns the run's summary.json verbatim. Invoke before declaring work done, to confirm a change actually works, or whenever a faithful execution report is needed…

183 +2 4d ago A 101 tokens original MIT

verification

292

techygarg/lattice

Agent

Part of lattice

Why the verifier exists, the cost model behind it, and how to wire it into your own sessions.

183 +2 4d ago A 0 tokens original MIT

idor-agent

293

BugTraceAI/BugTraceAI-CLI

Agent

The IDOR Agent (Insecure Direct Object Reference) is a specialist agent in BugTraceAI that detects and exploits IDOR vulnerabilities. It uses a WET→DRY two-phase pipeline with LLM-powered deduplication and optional deep exploitation analysis.

183 +3 4d ago A 0 tokens AGPL-3.0

provider-tester

294

nesquikm/mcp-rubber-duck

Agent Claude Code

Use this agent when you need to test, debug, or validate LLM provider configurations. This includes verifying API keys work correctly, checking provider connectivity, testing model availability, debugging authentication issues, validating base URLs for custom providers, or troubleshooting why a specific provider isn't…

176 14d ago A 362 tokens original MIT

validator

295

TechDufus/oh-my-claude

Agent

Part of oh-my-claude

Mission-first validator that runs relevant checks, reports evidence, and returns a binary pass/fail verdict with next steps.

176 1mo ago A 23 tokens original MIT

act

297

al3rez/ooda-subagents

Agent Claude Code

OODA Act phase - Implements the decided solution with precision, tests thoroughly, and validates results.

174 1y ago A 20 tokens

player

298

tettethu/VibeGame

Agent

Runtime verification teammate for one code task. Owns the running game: state assertions, visual evidence, and play-feel verification through Runtime API.

172 +46 3d ago A 31 tokens original Apache-2.0

gitwand-verifier

300

devlint/GitWand

Agent Claude Code

Adversarially verifies a GitWand change before PR — runs the test suites, checks AGENTS.md compliance, and reviews the diff. Read-only (no Edit/Write) so the review stays honest. Use after the executor finishes, before opening a PR.

167 5d ago A 59 tokens original MIT

mcp-testing-engineer

301

softspark/ai-toolkit

Agent

Part of app

MCP protocol testing expert. Use for MCP server testing, protocol compliance, transport validation, integration testing. Triggers: mcp test, protocol compliance, mcp validation, transport testing.

167 4d ago A 44 tokens original Apache-2.0

AgriciDaniel/skill-forge

Agent

Part of skill-forge

Blind comparison agent for A/B testing skill versions. Evaluates outputs from two skill versions without knowing which is "new" vs "old" to eliminate bias. User says: "compare these two skill versions" User says: "run a blind A/B test on the skill".

166 +3 4mo ago A 70 tokens original MIT

e2e-tester

305

vericontext/vibeframe

Agent Claude Code

End-to-end tester for the current VibeFrame CLI. Use when asked to test everything, run full tests, or verify the repo works.

165 1mo ago A 35 tokens original MIT

bineval

306

smixs/skill-conductor

Agent

Part of skill-conductor

Evaluate a skill artifact with atomic binary yes/no questions, one answer (1/0) per question, each preceded by a written critique grounded in evidence from the skill's own files. Aggregate to per-dimension scores in [0,1]; the orchestrator turns your answers into the overall score and the pass/fail gate.

163 1mo ago A 0 tokens original MIT

outlmd/outl

Agent Claude Code

Validates that changes in outl-core preserve the tree CRDT invariants (convergence, idempotency, no-cycle, no-silent-loss). Use PROACTIVELY after any edit in crates/outl-core/src/tree.rs, log.rs, op.rs, or in tree CRDT tests. Rejects PRs that break any invariant.

161 yesterday A 77 tokens original MIT

outlmd/outl

Agent Claude Code

Ensures that the .md ↔ ops ↔ .md pipeline is stable and that blocks never disappear silently during external matching. Use PROACTIVELY after changes in outl-md (parse, render, sidecar, matching). Runs roundtrip + matching suite and reports divergences.

161 yesterday A 64 tokens original MIT

wio-candidate-scout

310

workersio/skills

Agent

Part of wio

Read-only WIO subagent for discovering high-value test or workload candidates before implementation. Use during $wio scan, $wio workload, or the discovery stage of $wio test.

159 1mo ago A 46 tokens original MIT

wio-strategy-critic

311

workersio/skills

Agent

Part of wio

Read-only WIO subagent for challenging the selected testing strategy before implementation. Use after candidate selection and before editing test files.

159 1mo ago A 32 tokens original MIT

wio-test-reviewer

312

workersio/skills

Agent

Part of wio

Read-only WIO subagent for reviewing a written test and deciding KEEP, REDO, or REMOVE. Use after $wio test edits a test, or when asked whether a test is valuable.

159 1mo ago A 47 tokens original MIT

At most 3 mods per repository are shown here, and a mod shipped inside a plugin is left to that plugin's page — the rest are on their repository pages: