evaluation framework plugins

9 tagged evaluation framework, measured the same way as everything else here.

deepeval-plugins

01

confident-ai/deepeval

Plugin Claude Code

DeepEval plugins for LLM evaluation, tracing, and testing in Claude Code.

18k 2d ago A tokens not measured original Apache-2.0

deepeval

02

confident-ai/deepeval

Plugin Claude Code

Skills for adding DeepEval evaluations, tracing, datasets, Confident AI reports, and iterative improvement loops to AI applications.

18k 2d ago A tokens not measured original Apache-2.0

coder-eval

03

UiPath/coder_eval

Plugin Claude Code

Test whether your Claude Code skills actually trigger, and benchmark AI coding agents against your own tasks.

119 3d ago A tokens not measured original Apache-2.0

coder-eval

04

UiPath/coder_eval

Plugin Claude Code

Test whether your Claude Code skills actually trigger, and benchmark any coding agent — author, run, and analyze the eval suite that proves it, locally or as a CI gate.

119 3d ago A tokens not measured original Apache-2.0

agentv

05

EntityProcess/agentv

Plugin Claude Code

Evaluate and optimize AI agents.

15 1mo ago A tokens not measured original MIT

agentic-engineering

06

EntityProcess/agentv

Plugin Claude Code

Design and review AI agent systems: architecture patterns, workflow design, and plugin quality review.

15 1mo ago A tokens not measured original MIT

agentv-dev

07

EntityProcess/agentv

Plugin Claude Code

AgentV CLI skills for evaluating, optimizing, and governing AI agents.

15 1mo ago A tokens not measured original MIT

auto-itera

08

clfhaha1234/auto-itera

Plugin Claude Code

Plugin marketplace listing 1 plugin: auto-itera.

6 3mo ago A tokens not measured original MIT

auto-itera

09

clfhaha1234/auto-itera

Plugin Claude Code

Autonomous experimentation engine for AI engineering decisions. Define a goal + candidates + threshold, get a defensible ship-or-kill verdict in hours — sealed held-out test, pre-registered metric, sprint-and-generalize iteration.

6 3mo ago A tokens not measured original MIT