evals agents

15 tagged evals, measured the same way as everything else here.

analyzer

01

benchflow-ai/benchflow

Agent Codex

Analyze blind comparison results to understand WHY the winner won and generate improvement suggestions.

335 2d ago A 0 tokens copy · 100% Apache-2.0

comparator

02

benchflow-ai/benchflow

Agent Codex

Compare two outputs WITHOUT knowing which skill produced them.

335 2d ago A 0 tokens copy · 100% Apache-2.0

grader

03

benchflow-ai/benchflow

Agent Codex

Evaluate expectations against an execution transcript and outputs.

335 2d ago A 0 tokens copy · 100% Apache-2.0

AgentEval Dev

04

AgentEvalHQ/AgentEval

Agent

AI agent for AgentEval development tasks - code implementation, review, and debugging.

138 yesterday A 19 tokens original MIT

AgentEval DocWriter

05

AgentEvalHQ/AgentEval

Agent

AI agent for AgentEval documentation - writing, reviewing, and maintaining docs with brand consistency.

138 yesterday A 22 tokens original MIT

critic

07

rennf93/opus-fable-playbook

Agent

Adversarial verifier. Use PROACTIVELY before claiming a nontrivial change is done, fixed, or passing — give it the claim plus the relevant diff/paths and it attempts to refute the claim with evidence.

33 2mo ago A 48 tokens original MIT

vibe-checker

08

hev/vibecheck

Agent Claude Code

Use this agent when you need to create, run, or improve integration tests using the vibecheck evaluation platform. This includes writing YAML evaluation suites, running checks via the CLI, debugging failing evals, optimizing check patterns, or designing comprehensive test strategies. The agent is particularly useful…

22 7mo ago A 0 tokens

test-writer

09

ahnafyy/skills-evals

Agent Claude Code

Generates unit tests for existing code. Use when the user asks to add tests for a module or increase coverage.

3 1mo ago A 27 tokens original MIT

code-reviewer

10

ahnafyy/skills-evals

Agent

Reviews pull requests for correctness, style, and security issues. Use when the user asks for a code review or PR feedback.

3 1mo ago A 29 tokens original MIT

docs-writer

11

ahnafyy/skills-evals

Agent

Writes and updates documentation pages. Use when the user asks to document a feature or update the docs site.

3 1mo ago A 25 tokens original MIT

haiku-executor

12

Abhillashjadhav/AI-PM-essential-skills

Agent

Executes low-complexity tasks (cosmetic edits, formatting, simple lookups, mechanical transformations) delegated by the model-complexity-router. Fast and cheap.

2 yesterday A 39 tokens original MIT

opus-executor

13

Abhillashjadhav/AI-PM-essential-skills

Agent

Executes high-complexity tasks (architecture, multi-file refactors, novel tradeoff analysis, high error-cost decisions) delegated by the model-complexity-router.

2 yesterday A 38 tokens original MIT

sonnet-executor

14

Abhillashjadhav/AI-PM-essential-skills

Agent

Executes mid-complexity tasks (single features, standard analyses, drafts with known patterns) delegated by the model-complexity-router. Balanced cost and quality.

2 yesterday A 38 tokens original MIT