Plugin Claude Code
DeepEval plugins for LLM evaluation, tracing, and testing in Claude Code.
9 tagged evaluation framework, measured the same way as everything else here.
Plugin Claude Code
DeepEval plugins for LLM evaluation, tracing, and testing in Claude Code.
Plugin Claude Code
Skills for adding DeepEval evaluations, tracing, datasets, Confident AI reports, and iterative improvement loops to AI applications.
Plugin Claude Code
Test whether your Claude Code skills actually trigger, and benchmark AI coding agents against your own tasks.
Plugin Claude Code
Test whether your Claude Code skills actually trigger, and benchmark any coding agent — author, run, and analyze the eval suite that proves it, locally or as a CI gate.
Plugin Claude Code
Evaluate and optimize AI agents.
Plugin Claude Code
Design and review AI agent systems: architecture patterns, workflow design, and plugin quality review.
Plugin Claude Code
AgentV CLI skills for evaluating, optimizing, and governing AI agents.
Plugin Claude Code
Plugin marketplace listing 1 plugin: auto-itera.
Plugin Claude Code
Autonomous experimentation engine for AI engineering decisions. Define a goal + candidates + threshold, get a defensible ship-or-kill verdict in hours — sealed held-out test, pre-registered metric, sprint-and-generalize iteration.