nexus-reliability

nexus-reliability is a skill for Claude Code, Codex from aayushostwal/nexus. It costs 74 tokens per session (1,728 once invoked), scanned A, original, MIT.

A procedure for investigating production incidents, such as outages, error spikes, slow responses, or service degradation. It also includes checks for deciding whether a release is ready.

In plain words
What is it for?
Triage, severity classification, incident response, rollback decisions, release-readiness checks, and preparing evidence for root-cause or post-incident work.
Why use it?
It gives teams a consistent way to stabilize a live service, rebuild the event timeline, and base decisions on evidence instead of guesses.

Skill for Claude CodeCodex

Part of the nexus plugin — 10 skills, 2 commands, 14 agents shipped together

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add skills/aayushostwal/nexus/reliability
Any agent
npx skills add aayushostwal/nexus --skill reliability
Clone the repo
git clone --depth 1 https://github.com/aayushostwal/nexus

Made for: Claude Code, Codex.

Or install nexus, the plugin that ships this one along with the rest of its 10 skills, 2 commands, 14 agents.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for nexus-reliability

README.md
[![agentmods](https://agentmods.dev/badge/skills/aayushostwal/nexus/reliability.svg)](https://agentmods.dev/skills/aayushostwal/nexus/reliability)
Your own site
<a href="https://agentmods.dev/skills/aayushostwal/nexus/reliability"><img src="https://agentmods.dev/badge/skills/aayushostwal/nexus/reliability.svg" alt="Measured on agentmods" height="20"></a>
Per session 74 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 1,728 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00074 $0.01728
Opus 5 $0.00037 $0.00864
Sonnet 5 $0.00015 $0.00346
Haiku 4.5 $0.00007 $0.00173

Measured 3d ago against content hash 79aba6a0a30a, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

nexus-reliability scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 3d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

skills/reliability/SKILL.md · 193 lines

How it starts

The opening of the file, as written. The whole thing — 193 lines — stays where its author put it; the contents beside it link to each section on GitHub.

Nexus Production Incident Investigator

Time-boxed, evidence-driven workflow for investigating production incidents and evaluating release readiness.


Compatibility

  • Sub-skills: release-readiness.md | Materials: checklists/incident-checklist.md, anti-patterns/common-mistakes.md, validation/output-validation.md
  • Required: Read, Bash, Grep | Optional: WebSearch | Hands off to: nexus:debugging (code RCA), the roadmap-planner agent (post-mortem actions)

Core Principle

Stabilize first. Investigate second. Never skip the timeline.


Incident Classification

Severity Definition Response target
P0 Complete outage or data loss affecting all/most users Immediate; page all on-call
P1 Major feature broken or >25% requests failing < 5 min acknowledgement; page on-call lead
P2 Degraded performance or partial failure < 15 min acknowledgement; one engineer
P3 Minor anomaly, no user impact Next business hour

If you cannot classify immediately, default to P1 and downgrade after Phase 1.


Investigation Workflow

Phase 1 — Triage (0–5 min)

  1. Confirm the signal — cross-check alert against at least one independent source (logs, dashboard, manual repro). Never declare from a single data point.
  2. Classify severity using the table above.
  3. Identify blast radius — which services, user segments, and whether data integrity is at risk.
  4. Open incident channel and post: [P0/P1/P2] <service> — <symptom> — investigating
  5. Assign roles (P0/P1): Incident Commander (comms only), Tech Lead (investigation), Scribe (timeline).

Phase 2 — Stabilize (5–15 min)

Apply mitigations before investigating. Stop at first "yes":

Question Action
Deploy in last 2 hours? Roll back; confirm error rate drops
Config/feature-flag change? Revert; confirm error rate drops
One bad instance/pod? Remove from LB, restart
Downstream dependency down? Enable fallback/circuit breaker
DB connection pool exhausted? Restart pooler, reduce pool size
Traffic spike causing saturation? Rate limit, add capacity, shed load

Read the full file on GitHub · 193 lines

Files

What ships with it

3 files beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 3d ago First seen · 193 lines · 74 tokens per session scan A 79aba6a0a30a

Subscribe to this mod's changes

nexus-reliability is a skill published in the GitHub repository aayushostwal/nexus (18 stars, last pushed 24d ago), licensed MIT. It adds 74 tokens to every session and 1,728 once invoked, about $0.0004 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.

Related

Other skills, from other repositories

productize-yourself

当用户在纠结职业/副业/自由职业方向、问「我该做什么才能赚钱/不被替代」、或想找到自己的独特优势时调用。 核心理念: 特殊知识(不可培训、对你像玩对别人像工作) × 产品化(杠杆规模化) = 无可替代的致富定位。 不适用于: 纯求职投递、写简历、已有明确方向的执行细节。 Triggers: 找方向/独特优势/副业/不可替代/productize/special knowledge/moat.

kangarooking/cangjie-skill · 127 tokens

screen-detox

当用户刷手机/短视频/社交媒体上瘾、感觉空虚、想戒断多巴胺零食时调用。 核心理念: 所有屏幕活动与更少幸福相关(作者断言无例外); 屏幕=用长期后果换短期快感的多巴胺零食; 用习惯替换五步戒断。 不适用于: 工作需要屏幕的职业场景(区分工作屏幕与消费屏幕)。 Triggers: 刷手机/上瘾/短视频/社交媒体/多巴胺/戒断/屏幕时间/screen time/dopamine/addiction.

kangarooking/cangjie-skill · 148 tokens

happiness-skill

当用户问「怎么才能更幸福/为什么得到了还不满足/怎么减少焦虑」时调用。 核心理念: 幸福是缺憾感清空的默认状态, 是可训练的技能; 欲望是与自己的契约(得到前不快乐), 同时只留一个重大欲望; 活在当下。 不适用于: 临床抑郁等需要专业治疗的场景(本书方法不能替代医疗)。 Triggers: 幸福/不快乐/欲望/焦虑/知足/活在当下/happiness/desire/anxiety.

kangarooking/cangjie-skill · 136 tokens

long-term-compounding

当用户在选择合作者/生意模式/人生策略、问「要不要长期投入这段关系/这个项目」「如何积累声誉」时调用。 核心理念: 财富、知识、声誉、关系都遵循复利; 只玩长期正和游戏, 与能想象共事一辈子的人合作, 拒绝短期思维交易。 不适用于: 紧急止损、短期现金周转等必须立即决策的场景。 Triggers: 长期/复利/声誉/合作/信任/compounding/long-term/trust.

kangarooking/cangjie-skill · 136 tokens

monkey-mind-meditation

当用户脑子停不下来、焦虑反刍、想学冥想/提升专注力时调用。 核心理念: 内心独白(心猴)是失控程序不是"我"; 用调试模式逐条观察念头, 意识到即失去控制; 冥想=收件箱归零/心灵间歇性禁食。 不适用于: 严重精神疾病急性发作(先就医); 需要立刻处理的实际问题(先做事)。 Triggers: 冥想/焦虑/脑子停不下来/胡思乱想/专注/心猴/monkey mind/meditation/anxiety.

kangarooking/cangjie-skill · 158 tokens

hook-craft

Specializes in chapter openings (hooks) and chapter endings (pulls). Every chapter must start with a reason to keep reading and end with a reason to turn the page. Runs after prose-craft, before chaos-engine. The skill that prevents the reader from putting the book down.

felipelobomotta-blip/book-genesis-v4 · 61 tokens