evalview
01MCP server Claude CodeCodexCursor +2
MCP server "evalview" as configured in hidai25/eval-view. Launched with evalview mcp serve. Needs 2 environment variables to run.
18 tagged agent evaluation, measured the same way as everything else here.
MCP server Claude CodeCodexCursor +2
MCP server "evalview" as configured in hidai25/eval-view. Launched with evalview mcp serve. Needs 2 environment variables to run.
MCP server Claude CodeCodexCursor +2
Open-source testing and regression detection framework for AI agents. Golden baseline diffing, CI/CD integration, works with LangGraph, CrewAI, OpenAI, Anthropic Claude, HuggingFace, Ollama, and MCP. Runs locally from the evalview Python package. Needs 1 environment variable to run.
MCP server Claude CodeCodexCursor +2
Stop shipping agents on vibes. Score every agent output for quality, safety, and cost. Runs locally from the @iris-eval/mcp-server npm package. Needs 1 environment variable to run.
MCP server Claude CodeCodexCursor +2
The agent eval standard for MCP. Score every agent output for quality, safety, and cost. Runs locally from the @iris-eval/mcp-server npm package. Needs 3 environment variables to run.
MCP server Claude CodeCodexCursor +2
MCP server "governance-mcp" as configured in cirwel/unitares. Runs locally from the governance-mcp Python package.
MCP server Claude CodeCodexCursor
Certify whether an AI agent is safe to operate your internal web application before production rollout. Author a YAML task suite against your own app, run it with Playwright, and gate CI/CD on a forbidden-action safety score. Runs locally from the deskcert-cli Python package.
MCP server Claude CodeCodexCursor
Test whether AI agents retain critical instructions as conversations grow. Runs locally from the rule-drift npm package.
MCP server Claude CodeCodexCursor
Evaluate, benchmark, and simulate AI agents on the VerifyAX agent-evaluation platform. Runs locally from the @verifyax/mcp-server npm package. Needs 2 environment variables to run.
MCP server Claude CodeCodexCursor
Trust scoring for AI agents. Investigate, verify, and compare agent trustworthiness through MCP. Runs locally from the agentscore-mcp npm package.
whitestone1121-web/signalbrain
MCP server Claude CodeCodexCursor
Trust layer for AI-modified software — receipts, ledger, calibrated autonomy. Runs locally from the signalbrain Python package.
MCP server Claude CodeCodexCursor
Free AI-agent challenges with ID-free daily evaluation, hints, lessons, and batch replay. Remote server at worldorder.club.
ankitkapur1992-hlido/hlido-mcp
MCP server Claude CodeCodexCursor
Independent AI-agent reviews: trust checks, evidence scorecards, incident registry, recommendations. Remote server at hlido.eu.
MCP server Claude CodeCodexCursor
Read-only MCP server for the OPERANT AI operating-agent calibration benchmark. Runs locally from the saagar-operant-mcp npm package.
MCP server Claude CodeCodexCursor
MCP server "mcp-testbed" as configured in anthonys1760/mcp-testbed. Runs locally from the mcp-testbed npm package.
christian140903-sudo/agent-invariants
MCP server Claude CodeCodexCursor +2
MCP server "agent-invariants" as configured in christian140903-sudo/agent-invariants. Runs locally from the agent-invariants npm package.
Oshawott324/datalox-gated-runtime
MCP server Claude CodeCodexCursor +2
MCP server "datalox-gated-runtime" as configured in Oshawott324/datalox-gated-runtime. Runs locally from the datalox-gated-runtime Python package.
MCP server Claude CodeCodexCursor +2
MCP server "tracefact" as configured in Alex0AI/tracefact. Runs locally from the tracefact npm package.
jianghongcheng/radmeasure-agent
MCP server Claude CodeCodexCursor +2
MCP server "radmeasure" as configured in jianghongcheng/radmeasure-agent. Runs locally from the radmeasure Python package.