evalview-mcp

A local testing tool for AI agents that compares their results with saved expected results. It connects to several agent frameworks and AI services, and runs through the EvalView Python package.

In plain words
What is it for?
Use it to save baseline results, compare later runs, and check agent behavior in CI/CD, the automated process that builds and tests code.
Why use it?
It helps detect when an agent starts producing different results after a code change. This makes regressions—new problems caused by changes—easier to find in local work or automated checks.

MCP server for Claude CodeCodexCursorGemini CLIOpenCode

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add mcp/hidai25/eval-view/evalview-mcp
Clone the repo
git clone --depth 1 https://github.com/hidai25/eval-view

Made for: Claude Code, Codex, Cursor, Gemini CLI, OpenCode.

Per session not measured What this adds to a session before it is invoked.
When invoked not measured Not applicable: nothing here is loaded into a session.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Security

Grade A, and why

evalview-mcp scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

server.json · 11 lines

What it actually says

{
  "evalview-mcp": {
    "command": "uvx",
    "args": [
      "evalview"
    ],
    "env": {
      "OPENAI_API_KEY": ""
    }
  }
}
Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 2d ago First seen · 11 lines scan A bed0209b457f

Subscribe to this mod's changes

evalview-mcp is an MCP server published in the GitHub repository hidai25/eval-view (133 stars, last pushed 10d ago), licensed Apache-2.0. Its token cost is not measured: an MCP server costs its tool schemas, not its config file. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.

Related

Other mcp servers, from other repositories

portable-agent-memory

MCP server "portable-agent-memory" as configured in skywaller0/Portable-Agent-Memory. Runs locally from the portable-agent-memory Python package.

skywaller0/Portable-Agent-Memory · not measured

radmeasure

MCP server "radmeasure" as configured in jianghongcheng/radmeasure-agent. Runs locally from the radmeasure Python package.

jianghongcheng/radmeasure-agent · not measured

agent-eval

Statistical regression testing for LLM agents. Get a p-value on whether behavior actually shifted between versions, not just whether one run looked different. Apache 2.0, self-hostable, no SaaS dependency -- a Promptfoo alternative for the statistical testing gap threshold-based eval tools don't cover. Runs locally…

RudrenduPaul/agent-eval · not measured

ai-ticket-triage

MCP server "ai-ticket-triage" as configured in Mohemed-Amine-Chalhy/ai-ticket-triage. Runs locally from the ai-ticket-triage Python package.

Mohemed-Amine-Chalhy/ai-ticket-triage · not measured

engramia

Reusable execution memory and evaluation infrastructure for AI agent frameworks. Runs locally from the engramia Python package. Needs 8 environment variables to run.

engramia/engramia · not measured

network-ai

AI agent orchestration framework for TypeScript/Node.js - 32 adapters (LangChain, AutoGen, CrewAI, OpenAI Assistants, OpenAI Responses, LlamaIndex, Semantic Kernel, Haystack, DSPy, Agno, MCP, OpenClaw, A2A, Codex, MiniMax, NemoClaw, APS, Copilot, LangGrap. Runs locally from the network-ai npm package.

Jovancoding/Network-AI · not measured