AI-reliability plugins

12 tagged AI-reliability, measured the same way as everything else here.

relai-sdk AGENTS.md

01

relai-ai/relai-sdk

Instructions file CodexOpenCode

Instructions for relai-ai/relai-sdk, covering repository guidelines, project structure & module organization, build, test, and development commands, coding style & naming conventions and testing guidelines.

102 5mo ago A 609 tokens original Apache-2.0

replayd AGENTS.md

02

TaimoorKhan10/replayd

Instructions file CodexOpenCode

Instructions for TaimoorKhan10/replayd, covering replayd — three agent examples, what replayd does, the four-step loop, 1. capture a run and 2. mark it as failed.

18 3mo ago A 4,501 tokens original MIT

hermes-labs-ai/hermeneutic

Instructions file CodexOpenCode

Instructions for hermes-labs-ai/hermeneutic, covering agents.md — using hermeneutic from a coding agent, what this tool does, when to invoke it, programmatic use and calibration.

5 9d ago A 740 tokens original Apache-2.0

hermes-labs-ai/hermeneutic

Instructions file

Instructions for hermes-labs-ai/hermeneutic: Read AGENTS.md first — it is the canonical in-session protocol for coding agents using or modifying this repo, and everything there applies to Claude Code sessions too.

5 9d ago A 351 tokens original Apache-2.0

hermeneutic-gate

05

hermes-labs-ai/hermeneutic

Plugin Claude Code

Legacy advisory Stop-hook bundle for the fixed English gate. Current Claude Stop compatibility is not certified in v0.1.7; use the CLI directly. Requires the hermeneutic package.

5 9d ago A tokens not measured original Apache-2.0

Stop

06

hermes-labs-ai/hermeneutic

Hook

Runs when the agent finishes a response, running python3. From hermes-labs-ai/hermeneutic.

5 9d ago A tokens not measured original Apache-2.0

agent-reliability

07

ByteStack-Labs/claude-plugins

Plugin Claude Code

Claude skills for AI agent and ML reliability: reproduce the eval-to-production gap, catch confidently-wrong outputs, and prove root cause with verified numbers. Start with production-autopsy.

2 2mo ago A tokens not measured original MIT

calibration-guard

08

ByteStack-Labs/claude-plugins

Skill Claude CodeCodex

Detects and quantifies confidently-wrong behavior: where a model, classifier, agent, or LLM judge is highly confident and incorrect, and where the coupling between confidence and correctness breaks down under distribution shift. Measures calibration on the in-distribution set and again on the production distribution…

2 2mo ago A 208 tokens original MIT

production-autopsy

09

ByteStack-Labs/claude-plugins

Skill Claude CodeCodex

Start here. Audits a deployed ML or LLM or agent system that scores well on evaluation but fails, regresses, or behaves unexpectedly in production. Runs a reproducible root-cause "autopsy": frames the eval-to-deployment gap, reproduces the production failure, quantifies it by slice, tests confidence calibration under…

2 2mo ago A 258 tokens original MIT

tool-eval

10

ByteStack-Labs/claude-plugins

Skill Claude CodeCodex

Verifies the tools an agent or orchestrator depends on, by re-deriving a tool evaluation's real accuracy instead of trusting a single pass/fail score. Separates a formatting miss (a correct value scored wrong) from a real failure (a wrong value scored right), recomputes the expected answer from the raw inputs rather…

2 2mo ago A 253 tokens original MIT

At most 3 mods per repository are shown here — the rest are on their repository pages: