evaluator
73Agent Claude Code
Evaluates solution quality and completeness (MAP).
2,166 tagged ai coding, measured the same way as everything else here.
Browse within: agentic-coding 191claude-ai 125agent-orchestration 120code-quality 120copilot 117agentic-workflow 116claude-code-agents 93Code Generation 87ai-coding-assistant 80ai-coding-agent 64ai-skills 63claude-code-hooks 63agentic-engineering 60ai-agent-templates 60
Agent Claude Code
Evaluates solution quality and completeness (MAP).
Agent Claude Code
Reviews code for correctness, standards, security, and testability (MAP).
Agent Claude Code
Predicts consequences and dependency impact of changes (MAP).
Agent Codex
Delivery validation coordinator. Activated at Tier 2+ only (Tier 1 skips Red Team). Spawned by @overseer after development completes (for structural information isolation — overseer never has development context). Independently verifies the delivered product works correctly by dispatching validators…
Agent Codex
Independent quality gate authority. Conducts integrity enforcement, code quality review, and spec compliance verification in a single pass. Read-only — produces verdicts, never code. Its FAIL cannot be overridden by any agent; only the user can override.
Agent Codex
Scope card owner for multi-domain cards. Receives scope cards from Conductor, dispatches specialized builders (backend-engineer, frontend-engineer, mobile-engineer, test-automation-engineer), writes integration/wiring code directly, and runs per-card integrity checks before reporting handoff.
Agent
Reviews code for bugs, security, performance, and spec compliance. Run in fresh context without implementation bias.
Agent Claude Code
Use for one independent blind judge seat scoring an award-level submission against contest rubric criteria.
Agent Claude Code
Use for independent review of assumptions, model validity, evidence, reproducibility, writing, and final submission risk.
Agent Claude Code
Use for award-level paper structure, abstract, narrative, equations, figures, captions, and final polish.
Agent
An AI development-rules assistant with an automatic mode that can be enabled explicitly or through configured aliases. It keeps that mode active during a conversation and can proceed automatically in approved paths while enforcing safety rules.
Agent
An AI assistant for software development that identifies the kind of help needed and routes the request to a matching workflow.
Agent
Creates implementation plans from context and requirements.
Agent
Code review specialist for quality and security analysis.
Agent
Fast codebase recon that returns compressed context for handoff to other agents.
Agent
Agent bodies load on dispatch only. This index is the surface scan; never bulk-read agents/.md. An agent is a reviewer or author dispatched by a skill — never routed to, never "triggered.".
Agent
INTERNAL SMARTS analyst dispatched by the decision-variance skill. Produces a SMARTS analysis and recommendation for one (artifact-position, scaffold-evidence) pair. Never decides — the user decides. Never dispatch directly.
Agent
INTERNAL evidence-gatherer dispatched by the decision-variance and context-creation skills. Scans an assigned code scope and reports evidence of architectural decisions — file paths and line numbers only. Never dispatch directly.
Agent Claude Code
Analyze blind comparison results to understand WHY the winner won and generate improvement suggestions.
Agent Claude Code
Compare two outputs WITHOUT knowing which skill produced them.
Agent Claude Code
Evaluate expectations against an execution transcript and outputs.
CodeAlive-AI/ai-driven-development
Agent
Token-isolated deep research agent for academic papers. Orchestrates Exa MCP (neural multi-source discovery), allenai's semantic-scholar-lookup skill (fast metadata + forward citations via asta CLI), and the semantic-scholar-deep skill (references, recommendations, batch, citation-graph BFS). Use when the user asks…
Agent
Generate a custom checklist for the current feature based on user requirements.
Agent
Identify underspecified areas in the current feature spec by asking up to 5 highly targeted clarification questions and encoding answers back into the spec.