Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
git clone --depth 1 https://github.com/cdywolf/llm-eval-mcpWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/mcp/cdywolf/llm-eval-mcp/llm-eval-mcp)<a href="https://agentmods.dev/mcp/cdywolf/llm-eval-mcp/llm-eval-mcp"><img src="https://agentmods.dev/badge/mcp/cdywolf/llm-eval-mcp/llm-eval-mcp.svg" alt="Measured on agentmods" height="20"></a>Grade A, and why
llm-eval-mcp scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 7d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
What it actually says
{ "llm-eval-mcp": { "command": "uvx", "args": [ "llm-eval-mcp" ] } }
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 7d ago First seen · 8 lines scan A a781ea781c96
llm-eval-mcp is an MCP server published in the GitHub repository cdywolf/llm-eval-mcp (1 stars, last pushed 25d ago), licensed MIT. Its token cost is not measured: an MCP server costs its tool schemas, not its config file. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other mcp servers, from other repositories
promptfoo
TypeScript Execute (tsx): Node.js enhanced with esbuild to run TypeScript & ESM files. Runs locally from the tsx npm package.
opengate-mcp
Check whether an AI answer is grounded in its context — deterministic, no LLM judge. Runs locally from the @pharmatools/opengate-mcp npm package.
mcp-failure-lab
A chaos-engineering and resilience-testing toolkit for Model Context Protocol servers. Runs locally from the mcp-failure-lab npm package.
golden-dataset-mcp
MCP server wrapping golden-dataset-studio for version-controlled golden dataset management and RAG evaluation. Runs locally from the golden-dataset-mcp Python package.
perf-mcp
Fact-check and fix AI outputs. Hallucination detection, schema validation, auto-repair. Runs locally from the perf-mcp npm package. Needs 1 environment variable to run.
rag-regression-gate
MCP server "rag-regression-gate" as configured in db-inlee/rag-regression-gate-mcp. Runs app.mcp.server with python.