ai evaluation plugins

8 tagged ai evaluation, measured the same way as everything else here.

ifixai-community

01

ifixai-ai/iFixAi

Plugin Claude Code

iFixAi tools for Claude Code. Claude guides you through an independent, open-source audit of your own agent: is it doing the job it's supposed to do, given your business rules and org structure? Test any model (Anthropic, OpenAI, Gemini, Azure, Bedrock, …) graded by the judge(s) of your choice.

12k 3d ago A tokens not measured original Apache-2.0

ifixai

02

ifixai-ai/iFixAi

Plugin Claude Code

Independent auditing for AI agents, open source. Claude investigates your agent, authors an iFixAi fixture from its setup, and audits it against the one question other tools skip: is it doing the job it's supposed to do, given your business rules and org structure? Adversarial depth, assurance discipline. Test any…

12k 3d ago A tokens not measured original Apache-2.0

truesight

04

Goodeye-Labs/truesight-mcp-skills

Plugin Claude Code

MCP server and agent skills for the Truesight AI quality platform. Score inputs, build evaluations, analyze errors, and review results through natural language.

7 5mo ago A tokens not measured original MIT

datapoint

06

impel-intelligence/datapoint-mcp

Plugin Claude Code

Get real human opinions — surveys, preference comparisons, ratings, and rankings — from inside Claude Code, powered by Datapoint AI.

6 2d ago A tokens not measured original MIT

deep-process-skills

07

Deep-Process/deep-process

Plugin Claude Code

Structured workflows that make LLMs think instead of just respond. Skills for verification, exploration, architecture, risk, feasibility, synthesis, and more.

2 5mo ago A tokens not measured

jleonceo/skill-adherencia-reglas

Plugin Claude Code

Skill instalable para medir la adherencia real a las reglas de un agente, con la herramienta y el método en el repositorio adherencia-reglas.

1 1mo ago A tokens not measured original MIT