agent-reliability

A remote service for testing, comparing, and checking autonomous AI agents—software that can act on its own. It provides methods, test harnesses, and evidence for judging agent behaviour.

In plain words
What is it for?
Use it to benchmark agents, audit their behaviour, and build or apply tests for autonomous-agent systems.
Why use it?
It helps replace informal impressions with repeatable tests and documented evidence when you need to know whether an agent works reliably.

MCP server for Claude CodeCodexCursor

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add mcp/citarium/agentreliability-mcp/agent-reliability
Clone the repo
git clone --depth 1 https://github.com/citarium/agentreliability-mcp

Made for: Claude Code, Codex, Cursor.

Per session not measured What this adds to a session before it is invoked.
When invoked not measured Not applicable: nothing here is loaded into a session.
Security scan A 0 findings. Scan, not verified.
Origin unknown No closer match found in the catalogue.
Security

Grade A, and why

agent-reliability scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured yesterday.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

server.json · 6 lines

The source is not reproduced here

Licensed CC-BY-4.0

The repository is licensed CC-BY-4.0, which this catalogue does not treat as permission to reproduce the file. Read it at the source.

Read it on GitHub

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. yesterday First seen · 6 lines scan A 0717d6dd4d33

Subscribe to this mod's changes

agent-reliability is an MCP server published in the GitHub repository citarium/agentreliability-mcp (0 stars, last pushed 2d ago), licensed CC-BY-4.0. Its token cost is not measured: an MCP server costs its tool schemas, not its config file. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.

Related

Other mcp servers, from other repositories

NOUZ-MCP

MCP server for Obsidian — semantic knowledge graph with auto-classification, DAG hierarchy, and cross-domain bridge detection. Runs locally from the nouz-mcp Python package. Needs 8 environment variables to run.

Semiotronika/NOUZ-MCP · not measured

google-knowledge-graph

Search Google's Knowledge Graph for structured information about real-world entities. Runs locally from the @houtini/google-knowledge-graph-mcp npm package. Needs 1 environment variable to run.

houtini-ai/google-knowledge-graph-mcp · not measured

thebrain

MCP server for TheBrain 15: search by meaning, graph traversal and batch writes over the local API. Runs locally from the thebrain-mcp-server npm package. Needs 7 environment variables to run.

yBookoff/thebrain-mcp · not measured

wikidata-mcp-server

Search and fetch Wikidata entities, execute SPARQL queries, and resolve external identifiers. Runs locally from the @cyanheads/wikidata-mcp-server npm package. Needs 4 environment variables to run.

cyanheads/wikidata-mcp-server · not measured

LINZA-MCP

Local-first MCP sidecar for agent workspaces, artifacts, review intents, and context export. Runs locally from the linza-mcp Python package. Needs 9 environment variables to run.

Semiotronika/LINZA-MCP · not measured

june-mcp

MCP server for Junê — give any MCP agent (Claude Desktop, Claude Code, …) a shared, cited knowledge-graph memory backed by a June endpoint. Runs locally from the june-mcp Python package. Needs 13 environment variables to run.

Junemind/june-mcp · not measured