Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
/plugin marketplace add ByteStack-Labs/claude-pluginsnpx agentmods add plugins/bytestack-labs/claude-plugins/agent-reliabilitygit clone --depth 1 https://github.com/ByteStack-Labs/claude-pluginsGrade A, and why
agent-reliability scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured yesterday.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
What it actually says
{
"name": "agent-reliability",
"description": "Claude skills for AI agent and ML reliability: reproduce the eval-to-production gap, catch confidently-wrong outputs, and prove root cause with verified numbers. Start with production-autopsy.",
"version": "0.5.0",
"author": {
"name": "Jesse Moses (ByteStack Labs)"
},
"homepage": "https://bytestacklabs.com",
"repository": "https://github.com/ByteStack-Labs/claude-plugins",
"license": "MIT",
"keywords": ["agents", "reliability", "evals", "production", "ml", "llm", "calibration", "diagnostics"]
}
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- yesterday First seen · 13 lines scan A 03caf3b065b7
agent-reliability is a plugin published in the GitHub repository ByteStack-Labs/claude-plugins (2 stars, last pushed 2mo ago), licensed MIT. Its token cost is not measured: this kind of file is read by the harness, not the model. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other plugins, from other repositories
skillspec
Plugin marketplace listing 1 plugin: skillspec.
skillspec
SkillSpec is a CLI plus a structured skill plan for making agent skills easier to follow, test, route, and prove.
thumbgate-marketplace
Plugin marketplace listing 1 plugin: thumbgate.
thumbgate
One 👎 becomes a hard rule the agent cannot bypass. Captures thumbs-down feedback, distills it into PreToolUse Pre-Action Checks, enforced across every future Claude Code session.
codex-bridge
Run Codex review, adversarial review, and second-pass handoffs from Claude Code while keeping ThumbGate reliability memory in the loop.
qualixar
Plugin marketplace listing 2 plugins: superlocalmemory, superlocalmemory-codex.