agent benchmark mcp servers

4 tagged agent benchmark, measured the same way as everything else here.

evalview

01

hidai25/eval-view

MCP server Claude CodeCodexCursor +2

MCP server "evalview" as configured in hidai25/eval-view. Launched with evalview mcp serve. Needs 2 environment variables to run.

133 9d ago A tokens not measured original Apache-2.0

evalview-mcp

02

hidai25/eval-view

MCP server Claude CodeCodexCursor +2

Open-source testing and regression detection framework for AI agents. Golden baseline diffing, CI/CD integration, works with LangGraph, CrewAI, OpenAI, Anthropic Claude, HuggingFace, Ollama, and MCP. Runs locally from the evalview Python package. Needs 1 environment variable to run.

133 9d ago A tokens not measured original Apache-2.0

arcade1v1

03

agustincf/Arcade1v1

MCP server Claude CodeCodexCursor

TypeScript Execute (tsx): Node.js enhanced with esbuild to run TypeScript & ESM files. Runs locally from the tsx npm package. Needs 1 environment variable to run.

0 16d ago A tokens not measured original MIT

arcade1v1

04

agustincf/Arcade1v1

MCP server Claude CodeCodexCursor

Play 1v1 arcade games vs AI agents & humans, ranked by ELO. Replay-verified, on-chain escrow (Base). Runs locally from the @arcade1v1/mcp npm package. Needs 1 environment variable to run.

0 16d ago A tokens not measured original MIT