evaluation plugins

31 tagged evaluation, measured the same way as everything else here.

Browse within: browser-automation 11

langwatch

01

langwatch/langwatch

Plugin Claude Code

Plugin marketplace listing 1 plugin: langwatch.

3.5k yesterday A tokens not measured original Apache-2.0

langwatch

02

langwatch/langwatch

Plugin Claude Code

Records which repository and branch each coding-agent session worked in, and teaches the agent to read its own traces back from LangWatch.

3.5k yesterday A tokens not measured original Apache-2.0

eval-view

03

hidai25/eval-view

Plugin Claude Code

Open-source testing and regression detection framework for AI agents. Golden baseline diffing, CI/CD integration, works with LangGraph, CrewAI, OpenAI, Anthropic Claude, HuggingFace, Ollama, and MCP.

132 8d ago A tokens not measured original Apache-2.0

skill-optimizer

04

fastxyz/skill-optimizer

Plugin Claude Code

Benchmark, evaluate, and optimize skills to ensure reliable performance across all LLMs.

77 3mo ago A tokens not measured original MIT

skill-optimizer

05

fastxyz/skill-optimizer

Plugin Claude Code

Benchmark, evaluate, and optimize skills to ensure reliable performance across all LLMs.

77 3mo ago A tokens not measured original MIT

JuanMarchetto/agent-skills

Plugin Claude Code

App Store and Google Play screenshot creation with exact platform specs, gallery ordering, and preview videos.

5 5mo ago A tokens not measured original MIT

architect

12

JuanMarchetto/agent-skills

Plugin Claude Code

Multidisciplinary idea and project evaluator — dispatches parallel specialist agents, synthesizes unified assessment with scores.

5 5mo ago A tokens not measured original MIT

proofloop

13

sattyamjjain/proofloop

Plugin Claude Code

Plugin marketplace listing 1 plugin: proofloop.

5 2mo ago A tokens not measured original MIT

proofloop

14

sattyamjjain/proofloop

Plugin Claude Code

The first universal judge for Claude Code & Cowork. Auto-evaluates skill and agent execution quality with 7-dimension scoring, configurable rubrics, and dual-mode operation (auto hooks + manual /judge command).

5 2mo ago A tokens not measured original MIT

proofrag

16

unshDee/proofrag

Plugin Claude Code

Plugin marketplace listing 1 plugin: proofrag.

2 22d ago A tokens not measured original MIT

proofrag

17

unshDee/proofrag

Plugin Claude Code

Evaluate a RAG/LLM app: generate a golden set from your docs, run LLM-as-judge + retrieval metrics, and produce a shareable HTML scorecard with a CI gate.

2 22d ago A tokens not measured original MIT

promptup

18

gabikreal1/promptup-plugin

Plugin Claude Code

AI coding skill evaluator — 11-dimension session scoring, decision tracking, PR reports.

2 5mo ago A tokens not measured original MIT

evals-mcp-server

19

cyanheads/evals-mcp-server

Plugin Claude Code

Author verifiable eval records through a draft → review → revise → submit loop with server-enforced graders; compile to JSONL/CSV/Inspect/lm-eval via MCP. STDIO or Streamable HTTP.

1 9d ago A tokens not measured original Apache-2.0

governed-context

21

leonardoeverling-spec/governed-context

Plugin Claude Code

Governed memory (explicit states, fail-closed sensitivity gates, quarantine-only writes) as a dependency-free MCP server, plus a blind-evaluation protocol skill for honest A/B comparisons.

0 29d ago A tokens not measured original MIT

cold-run

23

sina-heidariaan/cold-run

Plugin Claude Code

Cold-run your agent skills to find which ones change nothing, plus 24 skills that survived the same test.

0 9d ago A tokens not measured