grader
409Agent Claude Code
Evaluate expectations against an execution transcript and outputs.
5,554 tagged Testing, measured the same way as everything else here.
Browse within: agent-orchestration 54code-quality 52agentic-workflow 46agentic-coding 41spec-driven-development 40playwright 36Multi-Agent 34harness 34ai-security 32cybersecurity 32github-copilot 31rtl 31verification 31agentic 29
Agent Claude Code
Evaluate expectations against an execution transcript and outputs.
Agent
An integration reviewer that checks whether work across multiple software modules is complete and ready for a final go or no-go decision.
Agent Claude Code
Run a live smoke test of the /ask endpoint (SSE-streamed RAG). Boots fireseqsearchserver via tests/runlogseq.sh, runs tests/testask.py (protocol/invariant assertions) and tests/testendpoints.py --ask against a user-supplied question, and reports on answer grounding, citation validity, source quality, streaming…
Agent Claude Code
Run a live smoke test of fireseqsearchserver against an Obsidian vault. Drives tests/runsmoke.sh in one of two modes — lite (committed astro-wiki-lite fixture; fast, proves the plumbing) or full (the real 366-note AstroWiki2.0 vault; the only mode that can grade whether score priority and /ask answers are correct).…
Agent
Comprehensive code quality specialist ensuring production-ready code through formatting, linting, and testing with zero violations.
Agent
Investigates and fixes bugs. Takes a symptom, reproduces it, identifies root cause, writes the fix and a regression test.
Agent
Expert assistant for web accessibility (WCAG 2.1/2.2), inclusive UX, and a11y testing.
Agent Claude Code
Use proactively to validate functionality using project tools, perform systematic root cause analysis without changing test subjects, identify gaps in testing harness and observability, and write meaningful, fast, isolated unit tests that provide real value while avoiding fragile tests.
Agent
Integration test specialist -- Scenario builder, MockConnection, Tracer assertions, 8 test categories, Rust in-memory tests.
Agent
Compare two outputs WITHOUT knowing which skill produced them.
Agent
Implement tasks via TDD and commit small changes.
javiarmesto/ALDC-AL-Development-Collection
Agent Claude Code
Orchestrates Planning, Implementation, Review, and Commit cycle for AL Development. Enforces TDD and quality gates for Business Central extensions. Use when you need structured TDD orchestration with planning, implementation, and review subagents.
Agent
Use this agent when you need to perform any browser-based operations, including automated testing, debugging web applications, performance investigations, web scraping, or any other browser automation tasks. This agent should ALWAYS be used instead of the main agent when browser operations are required to prevent…
SalesforceAIResearch/agentforce-adlc
Agent
Part of agentforce-adlc
Tests Agentforce agents and optimizes based on session trace analysis.
debs-obrien/playwright-movies-app
Agent
Use this agent when you need to create automated browser tests using Playwright Examples: Context: User wants to generate a test for the test plan item.
debs-obrien/playwright-movies-app
Agent
Use this agent when you need to debug and fix failing Playwright tests.
debs-obrien/playwright-movies-app
Agent
Use this agent when you need to create comprehensive test plan for a web application or website.
Agent
Use this agent when you need to test routes after implementing or modifying them. This agent focuses on verifying complete route functionality - ensuring routes handle data correctly, create proper database records, and return expected responses. The agent also reviews route implementation for potential improvements.…
Agent Codex
Evaluate expectations against an execution transcript and outputs.
HermeticOrmus/LibreUIUX-Claude-Code
Agent Claude Code
Grand orchestrator of all LibreUIUX plugins. Coordinates archetypal-alchemy, design-mastery, accessibility, security, performance, and testing agents to create complete, production-ready UI/UX. The conductor of the plugin symphony. Use PROACTIVELY for any comprehensive UI/UX work.
Agent
An agent that reads source code, generates tests for it, and writes those tests to disk.
LuckyKuang/codex-tokens-compress
Agent Codex
Evaluate expectations against an execution transcript and outputs.
Agent
QC sub-agent. Executes tests, static analysis, and security tools. Asks user for permission before installing missing dependencies.
zoharbabin/due-diligence-agents
Agent Claude Code
Runs tests and reports results with diagnostics.
At most 3 mods per repository are shown here, and a mod shipped inside a plugin is left to that plugin's page — the rest are on their repository pages: