playwright
01MCP server Claude CodeCodexCursor +2
MCP server "playwright" as configured in langwatch/langwatch. Launched with bash -c d=$PWD; while [ "$d" != / ] && [ ! -f "$d/dev/scripts/playwright-mcp.sh".
30 tagged evaluation, measured the same way as everything else here.
MCP server Claude CodeCodexCursor +2
MCP server "playwright" as configured in langwatch/langwatch. Launched with bash -c d=$PWD; while [ "$d" != / ] && [ ! -f "$d/dev/scripts/playwright-mcp.sh".
MCP server Claude CodeCodexCursor +2
MCP server "playwright-headed" as configured in langwatch/langwatch. Launched with bash -c d=$PWD; while [ "$d" != / ] && [ ! -f "$d/dev/scripts/playwright-mcp.sh".
MCP server Claude CodeCodexCursor +2
MCP server "codebase-memory-mcp" as configured in suyoumo/ClawProBench. Launched with codebase-memory-mcp.
MCP server Claude CodeCodexCursor +2
MCP server "evalview" as configured in hidai25/eval-view. Launched with evalview mcp serve. Needs 2 environment variables to run.
MCP server Claude CodeCodexCursor +2
Open-source testing and regression detection framework for AI agents. Golden baseline diffing, CI/CD integration, works with LangGraph, CrewAI, OpenAI, Anthropic Claude, HuggingFace, Ollama, and MCP. Runs locally from the evalview Python package. Needs 1 environment variable to run.
redhat-community-ai-tools/harness-eval
MCP server Claude CodeCodexCursor +2
Gives the agent read and write access to a set of allowed directories on the local filesystem. Runs locally from the @modelcontextprotocol/server-filesystem npm package.
MCP server Claude CodeCodexCursor +2
The Fastest Way to Audit Your RAG - Generate QA datasets & evaluate RAG systems in Colab, Jupyter, or CLI. Privacy-first, any LLM, visual reports. Runs locally from the ragscore Python package. Needs 2 environment variables to run.
MCP server Claude CodeCodexCursor +2
MCP server "tuneforge" as configured in JustVugg/tuneforge. Launched with tuneforge. Needs 3 environment variables to run.
MCP server Claude CodeCodexCursor +2
The MCP documentation itself, served as an MCP server: lets the agent search the protocol spec and guides. Remote server at modelcontextprotocol.io.
MCP server Claude CodeCodexCursor +2
Governed AI-agent memory, Evidence Ledger traces, evals, and portable context tools. Runs locally from the @lore-context/server npm package. Needs 3 environment variables to run.
MCP server Claude CodeCodexCursor +2
MCP server "agora-memory-toolkit" as configured in DanceNitra/agora. Runs locally from the agora-memory-toolkit Python package.
MCP server Claude CodeCodexCursor
MCP server "promptup" as configured in gabikreal1/promptup-plugin. Runs /Users/gabik/.promptup/plugin/dist/index.js with node.
MCP server Claude CodeCodexCursor
Author verifiable eval records through a draft→review→revise→submit loop with enforced graders. Runs locally from the @cyanheads/evals-mcp-server npm package. Needs 5 environment variables to run.
MCP server Claude CodeCodexCursor
Check whether an AI answer is grounded in its context — deterministic, no LLM judge. Runs locally from the @pharmatools/opengate-mcp npm package.
MCP server Claude CodeCodexCursor
Open, always-on leaderboard + CI gate for autonomous coding agents — every patch sandboxed, every run traced, every regression fails the build. Runs locally from the forgejudge Python package.
MCP server Claude CodeCodexCursor
AgentTrust MCP Server — verify AI agent and MCP server quality from your IDE. Runs locally from the mcp-agenttrust Python package. Needs 3 environment variables to run.
MCP server Claude CodeCodexCursor
Assertions for AI-generated media. The bugs don't throw; rendercheck makes them throw. Runs locally from the rendercheck Python package. Needs 1 environment variable to run.
MCP server Claude CodeCodexCursor
MCP server that packages LLM evaluation gates as reusable CI/CD primitives. Runs locally from the mcp-llm-eval Python package.
MCP server Claude CodeCodexCursor
MCP server "reasonforge" as configured in euuuuuuan/reasonforge-public. Runs locally from the reasonforge npm package.
MCP server Claude CodeCodexCursor
Reusable execution memory and evaluation infrastructure for AI agent frameworks. Runs locally from the engramia Python package. Needs 8 environment variables to run.
MCP server Claude CodeCodexCursor
MCP server "gq-insight-mcp" as configured in yusufdxb/gq-insight-mcp. Runs locally from the gq-insight-mcp Python package.
MCP server Claude CodeCodexCursor
MCP server "evalmine" as configured in hishamalward/evalmine. Runs locally from the evalmine Python package.
MCP server Claude CodeCodexCursor +2
MCP server "posthog-context-mcp" as configured in josephruocco/posthog-mcp-mini. Runs locally from the posthog-context-mcp Python package.
leonardoeverling-spec/governed-context
MCP server Claude CodeCodexCursor +2
MCP server "governed-memory" as configured in leonardoeverling-spec/governed-context. Runs ${CLAUDE_PLUGIN_ROOT}/mcp/server.mjs with node. Needs 3 environment variables to run.