iris
01MCP server Claude CodeCodexCursor +2
Stop shipping agents on vibes. Score every agent output for quality, safety, and cost. Runs locally from the @iris-eval/mcp-server npm package. Needs 1 environment variable to run.
6 tagged eval, measured the same way as everything else here.
MCP server Claude CodeCodexCursor +2
Stop shipping agents on vibes. Score every agent output for quality, safety, and cost. Runs locally from the @iris-eval/mcp-server npm package. Needs 1 environment variable to run.
MCP server Claude CodeCodexCursor +2
The agent eval standard for MCP. Score every agent output for quality, safety, and cost. Runs locally from the @iris-eval/mcp-server npm package. Needs 3 environment variables to run.
MCP server Claude CodeCodexCursor
Independent, reproducible benchmark harness for agent-memory backends (MemPalace, Mem0, Zep/Graphiti, OpenViking): runs LongMemEval, LoCoMo, contradiction-detection, and a dozen other evals against all four and publishes the raw logs, not vendor-curated numbers. Runs locally from the memtrust-cli Python package.
MCP server Claude CodeCodexCursor +2
MCP server "invarianteval" as configured in AlpharomeroJL/invarianteval. Runs locally from the invarianteval Python package.
MCP server Claude CodeCodexCursor
Sends the same prompt to multiple LLM providers in parallel and returns a divergence score via MCP. Runs locally from the truthroute-cli npm package.
MCP server Claude CodeCodexCursor +2
MCP server "kaigo-gap" as configured in ossudesu-lab/kaigomcp. Runs kaigo_mcp with python.