agent evaluation plugins

14 tagged agent evaluation, measured the same way as everything else here.

ifixai-community

01

ifixai-ai/iFixAi

Plugin Claude Code

iFixAi tools for Claude Code. Claude guides you through an independent, open-source audit of your own agent: is it doing the job it's supposed to do, given your business rules and org structure? Test any model (Anthropic, OpenAI, Gemini, Azure, Bedrock, …) graded by the judge(s) of your choice.

12k 3d ago A tokens not measured original Apache-2.0

ifixai

02

ifixai-ai/iFixAi

Plugin Claude Code

Independent auditing for AI agents, open source. Claude investigates your agent, authors an iFixAi fixture from its setup, and audits it against the one question other tools skip: is it doing the job it's supposed to do, given your business rules and org structure? Adversarial depth, assurance discipline. Test any…

12k 3d ago A tokens not measured original Apache-2.0

eval-view

03

hidai25/eval-view

Plugin Claude Code

Open-source testing and regression detection framework for AI agents. Golden baseline diffing, CI/CD integration, works with LangGraph, CrewAI, OpenAI, Anthropic Claude, HuggingFace, Ollama, and MCP.

132 9d ago A tokens not measured original Apache-2.0

coder-eval

04

UiPath/coder_eval

Plugin Claude Code

Test whether your Claude Code skills actually trigger, and benchmark AI coding agents against your own tasks.

119 3d ago A tokens not measured original Apache-2.0

coder-eval

05

UiPath/coder_eval

Plugin Claude Code

Test whether your Claude Code skills actually trigger, and benchmark any coding agent — author, run, and analyze the eval suite that proves it, locally or as a CI gate.

119 3d ago A tokens not measured original Apache-2.0

godmode

06

thiientv/godmode

Plugin Claude Code

Plugin marketplace listing 1 plugin: godmode.

93 6d ago A tokens not measured original MIT

godmode

07

thiientv/godmode

Plugin Claude Code

Composable engineering workflows and expert capabilities for AI coding agents.

93 6d ago A tokens not measured original MIT

self-care

08

Not-Diamond/self-care

Plugin Claude Code

Self-Care — Agent trace analysis and context remediation plugin.

28 4mo ago A tokens not measured original MIT

self-care

09

Not-Diamond/self-care

Plugin Claude Code

Trace analysis and context remediation for AI agents.

28 4mo ago A tokens not measured original MIT

iris-eval

10

iris-eval/mcp-server

Plugin Claude Code

Plugin marketplace listing 1 plugin: iris-eval.

7 8d ago A tokens not measured original MIT

iris

11

iris-eval/mcp-server

Plugin Claude Code

Stop shipping agents on vibes. Score every agent output for quality, safety, and cost.

7 8d ago A tokens not measured original MIT

iris-eval

12

iris-eval/mcp-server

Plugin Claude Code

Agent eval for any MCP workflow — score output quality, catch PII and prompt injection, verify citations, and enforce cost budgets. Adds the Iris MCP server (9 tools) plus an agent-eval skill that guides eval-driven development.

7 8d ago A tokens not measured original MIT

mutagent

14

mutagent-io/skills

Plugin Claude Code

Official MutagenT CLI skill for setup, package installation, usage discovery, and product feedback.

4 26d ago A tokens not measured original MIT