llm evaluation plugins

17 tagged llm evaluation, measured the same way as everything else here.

Browse within: python 5

deepeval-plugins

01

confident-ai/deepeval

Plugin Claude Code

DeepEval plugins for LLM evaluation, tracing, and testing in Claude Code.

18k 3d ago A tokens not measured original Apache-2.0

deepeval

02

confident-ai/deepeval

Plugin Claude Code

Skills for adding DeepEval evaluations, tracing, datasets, Confident AI reports, and iterative improvement loops to AI applications.

18k 3d ago A tokens not measured original Apache-2.0

nuguard

03

NuGuardAI/nuguard

Plugin Claude Code

Plugin marketplace listing 1 plugin: nuguard.

36 yesterday A tokens not measured

nuguard

04

NuGuardAI/nuguard

Plugin Claude Code

AI application security for Claude — generate an AI Bill of Materials (AI-SBOM), run static analysis, behavioral validation, and adversarial red-team testing for AI agents and LLM-powered applications.

36 yesterday A tokens not measured

nuguard

05

NuGuardAI/nuguard

Plugin Claude Code

AI Application Security — SBOM generation, static analysis, behavioral testing, and adversarial red-teaming for AI agents and LLM-powered applications.

36 yesterday A tokens not measured

deslop

06

MrZoyo/deslop-GPT

Plugin Claude Code

Claude Code distribution for the deslop Agent Skill.

35 4d ago A tokens not measured original MIT

deslop

07

MrZoyo/deslop-GPT

Plugin Claude Code

Deletion-first audit and cleanup for accumulated test, verification, and fallback bloat.

35 4d ago A tokens not measured original MIT

mcpbr

09

greynewell/mcpbr

Plugin Claude Code

mcpbr - MCP Benchmark Runner plugin marketplace.

10 4mo ago A tokens not measured original MIT

mcpbr

10

greynewell/mcpbr

Plugin Claude Code

Expert benchmark runner for MCP servers using mcpbr. Handles Docker checks, config generation, and result parsing.

10 4mo ago A tokens not measured original MIT

tunelab

11

rchaz/tunelab

Plugin Claude Code

Plugin marketplace listing 1 plugin: tunelab.

6 1mo ago A tokens not measured original MIT

tunelab

12

rchaz/tunelab

Plugin Claude Code

Cut your AI bill without losing accuracy. tunelab helps you move repetitive LLM work — classifying, routing, extracting, drafting — onto small models that run for free on your Mac. It decides by experiment (testing on your own data first), trains locally with MLX, evaluates honestly, and explains every step so you…

6 1mo ago A tokens not measured original MIT

agenda-intelligence

14

vassiliylakhonin/agenda-intelligence-md

Plugin Claude Code

Deterministic evidence-packet linter for claim-backed AI output, with compatibility agenda-analysis skills and MCP tools. Reports packet completeness, not factual truth.

6 2d ago A tokens not measured original MIT

hermeneutic-gate

15

hermes-labs-ai/hermeneutic

Plugin Claude Code

Legacy advisory Stop-hook bundle for the fixed English gate. Current Claude Stop compatibility is not certified in v0.1.7; use the CLI directly. Requires the hermeneutic package.

5 7d ago A tokens not measured original Apache-2.0

eval-coach

16

BayramAnnakov/eval-coach

Plugin Claude Code

AI evaluation strategy design assistant using Evaluation-Driven Development (EDD).

4 7mo ago A tokens not measured original MIT

althing

17

DataViking-Tech/Althing

Plugin Claude Code

Run synthetic focus groups and user research panels using AI personas. Define personas in YAML, design survey instruments, and collect structured qualitative feedback — all from Claude Code.

2 22d ago A tokens not measured original MIT