comparator
553Agent Claude Code
Compare two outputs WITHOUT knowing which skill produced them.
5,309 tagged Testing, measured the same way as everything else here.
Browse within: claude-code-plugin 133ai-coding 87code-quality 50agent-orchestration 46ai-coding-assistant 46ai-skills 45agentic-coding 44copilot 40Multi-Agent 39agentic 38agent-tools 35ai-development 35agentic-workflow 33ai-security 32
Agent Claude Code
Compare two outputs WITHOUT knowing which skill produced them.
Agent
Automated unit test generation for code with comprehensive coverage.
ncoevoet/claude-markdown-health-check
Agent Claude Code
Analyzes logs for anomalies. Use when investigating errors or unexpected behavior in application logs.
ncoevoet/claude-markdown-health-check
Agent Claude Code
Deploys the application to staging. Use when the user asks to push a build to the staging environment.
ncoevoet/claude-markdown-health-check
Agent Claude Code
Reviews diffs for correctness and security issues. Use after code edits to verify quality before a commit.
Agent Claude Code
Use this agent when the user wants to check out a branch or PR in a separate git worktree, set up alongside the current repository. This includes when the user mentions a branch name, PR number, or asks to 'check out', 'review', or 'test' a branch/PR in isolation. The agent creates the worktree as a sibling directory…
Agent Claude Code
Use this agent when code changes have been made and tests need to be executed to verify correctness. This agent should be invoked after completing a logical chunk of work such as implementing a feature, fixing a bug, or refactoring code. Examples:\n\n \nContext: User has just implemented a new WebSocket message…
Agent
Comprehensive UI/UX testing specialist using Playwright for enterprise-grade user experience validation.
Agent Claude Code
Use this agent when you need to design, implement, review, or improve test suites and testing strategies. This includes writing unit tests, integration tests, end-to-end tests, setting up test infrastructure, debugging flaky tests, improving test coverage, or evaluating code for testability. Also use when reviewing…
Agent
Alpha Squad - mechanisms and confounding lens. Orthogonalizes signals via Frisch-Waugh-Lovell, runs Double ML, designs placebo tests. Correlation is unobserved confounding until proven otherwise.
Agent
Evaluate expectations against an execution transcript and outputs.
Agent
Evaluate expectations against an execution transcript and outputs.
Agent
Your job is to perform blind A/B comparison of two skill outputs without knowing which is which.
Agent
Your job is to evaluate skill outputs against a set of assertions. For each assertion, determine if it passed or failed based on the actual output.
Agent Claude Code
Proactively use this agent to find existing similar endpoint groups that can be used as copy templates for new Umbraco Management API endpoint implementations for testing including builders and helpers. This agent identifies the best existing patterns to replicate. It can also be used for incomplete endpoint groups to…
Agent Claude Code
Proactively Use this agent when you need to create integration tests for MCP tools where all prerequisites are complete (builders, helpers, and tools exist with passing tests). This agent focuses solely on creating comprehensive integration test suites and should be used when:\n\n- \n Context: User has Document Type…
Agent Claude Code
QA agent that AUTOMATICALLY runs after /migrate-tools or /migrate-tests commands complete to validate the migration was done correctly. Use this agent proactively (without user asking) when you detect these migration commands have just completed.
fieldsphere/cursor-team-marketplace-template
Agent
Adversarial review plus acceptance verification. Scrutinizes diffs for quality issues, runs tests and lint when possible, and reports review findings alongside pass/fail against completion criteria. Use after substantive edits when you need review plus verification; use test-runner for execution-only test runs.
Agent
Implementation agent. Writing code, generating boilerplate, scaffolding components, implementing features from specs, writing tests, standard bug fixes.
Agent
Canonical reference for agents writing tests for device identification, doctor checks, and readiness pipelines. Read this before touching @podkit/device-testing, any file named .e2e.test.ts, or tasks in milestone m-19.
Agent
Vollständige Validierung des Fullstack Monorepos (Python/FastAPI Backend + React/Vite Frontend). Trigger: (1) vor Commits, (2) nach Feature-Completion, (3) vor Deployments, (4) bei Projektaudits. Context: User hat ein neues Feature im Backend implementiert. user: "Ich habe den neuen Analyse-Endpoint fertig. Validiere…
Agent Claude Code
Test runner agent that uses IDE bridge tools to run tests, analyze failures, fix broken tests, and ensure code quality. Use when you need tests run and failures fixed autonomously.
CodelyTV/agentic_programming-course
Agent Codex
Use when creating or modifying tests: unit tests, Object Mothers for test data, or hand-written Mock Objects for domain interfaces. Follows the project's testing conventions with jest, should pattern mocks, and faker-based mothers.
Agent
Validation agent for testing Copilot CLI agent delegation and structured output.
At most 3 mods per repository are shown here, and a mod shipped inside a plugin is left to that plugin's page — the rest are on their repository pages: