Instructions file CodexOpenCode
Instructions for relai-ai/relai-sdk, covering repository guidelines, project structure & module organization, build, test, and development commands, coding style & naming conventions and testing guidelines.
12 tagged AI-reliability, measured the same way as everything else here.
Instructions file CodexOpenCode
Instructions for relai-ai/relai-sdk, covering repository guidelines, project structure & module organization, build, test, and development commands, coding style & naming conventions and testing guidelines.
Instructions file CodexOpenCode
Instructions for TaimoorKhan10/replayd, covering replayd — three agent examples, what replayd does, the four-step loop, 1. capture a run and 2. mark it as failed.
Instructions file CodexOpenCode
Instructions for hermes-labs-ai/hermeneutic, covering agents.md — using hermeneutic from a coding agent, what this tool does, when to invoke it, programmatic use and calibration.
Instructions file
Instructions for hermes-labs-ai/hermeneutic: Read AGENTS.md first — it is the canonical in-session protocol for coding agents using or modifying this repo, and everything there applies to Claude Code sessions too.
Plugin Claude Code
Legacy advisory Stop-hook bundle for the fixed English gate. Current Claude Stop compatibility is not certified in v0.1.7; use the CLI directly. Requires the hermeneutic package.
Hook
Runs when the agent finishes a response, running python3. From hermes-labs-ai/hermeneutic.
Plugin Claude Code
Claude skills for AI agent and ML reliability: reproduce the eval-to-production gap, catch confidently-wrong outputs, and prove root cause with verified numbers. Start with production-autopsy.
Skill Claude CodeCodex
Detects and quantifies confidently-wrong behavior: where a model, classifier, agent, or LLM judge is highly confident and incorrect, and where the coupling between confidence and correctness breaks down under distribution shift. Measures calibration on the in-distribution set and again on the production distribution…
Skill Claude CodeCodex
Start here. Audits a deployed ML or LLM or agent system that scores well on evaluation but fails, regresses, or behaves unexpectedly in production. Runs a reproducible root-cause "autopsy": frames the eval-to-deployment gap, reproduces the production failure, quantifies it by slice, tests confidence calibration under…
Skill Claude CodeCodex
Verifies the tools an agent or orchestrator depends on, by re-deriving a tool evaluation's real accuracy instead of trusting a single pass/fail score. Separates a formatting miss (a correct value scored wrong) from a real failure (a wrong value scored right), recomputes the expected answer from the raw inputs rather…
At most 3 mods per repository are shown here — the rest are on their repository pages: