reliability-audit

reliability-audit is a skill for Claude Code, Codex from Nordic-AI/production-readiness-skills. It costs 101 tokens per session (3,873 once invoked), scanned A, original, Apache-2.0.

A review of how an application handles failures, including errors, timeouts, retries, duplicate requests, service outages, and partial work.

In plain words
What is it for?
Use it to check retry backoff and jitter, idempotency keys, circuit breakers, transaction boundaries, time limits, failure isolation, and graceful degradation.
Why use it?
It helps prevent one failed dependency or repeated request from turning into lost data, inconsistent results, or a wider outage.

Skill for Claude CodeCodex

Written for no agent in particular: nothing here depends on one.

Good fit Use it to check retry backoff and jitter, idempotency keys, circuit breakers, transaction boundaries, time limits, failure isolation, and graceful degradation.

Compare 6 skills from other repositories ↓
Install with agentmods
npx agentmods add skills/nordic-ai/production-readiness-skills/reliability-audit
Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

Any agent
npx skills add Nordic-AI/production-readiness-skills --skill reliability-audit
Clone the repo
git clone --depth 1 https://github.com/Nordic-AI/production-readiness-skills

Made for: Claude Code, Codex.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for reliability-audit

README.md
[![agentmods](https://agentmods.dev/badge/skills/nordic-ai/production-readiness-skills/reliability-audit/github.svg)](https://agentmods.dev/skills/nordic-ai/production-readiness-skills/reliability-audit)
Your own site
<a href="https://agentmods.dev/skills/nordic-ai/production-readiness-skills/reliability-audit"><img src="https://agentmods.dev/badge/skills/nordic-ai/production-readiness-skills/reliability-audit/github.svg" alt="Measured on agentmods" height="20"></a>

Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.

agentmods 80×15 button for reliability-audit

Your own site · 80×15
<a href="https://agentmods.dev/skills/nordic-ai/production-readiness-skills/reliability-audit"><img src="https://agentmods.dev/badge/skills/nordic-ai/production-readiness-skills/reliability-audit.svg" alt="Reviewed on agentmods" width="80" height="20"></a>
Per session 101 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 3,873 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 1 finding. A grade says what 26 rules found in the file — not that it is safe.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5.1 $0.00101 $0.03873
Opus 5 $0.00051 $0.01937
Sonnet 5 $0.00020 $0.00775
Haiku 4.5 $0.00010 $0.00387

Measured 12d ago against content hash 27b8f5f67a04, method: parsed. Prices are Anthropic first-party input rates as of 2026-09-12, from the pricing page.

Security

Grade A, and why

reliability-audit scanned grade A with 1 finding against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 12d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Makes network callslowCapability

Not a fault in itself. Listed so you know the mod talks to something, and to what.

const client = axios.create({
skills/reliability-audit/SKILL.md · 354 lines

How it starts

The opening of the file, as written. The whole thing — 354 lines — stays where its author put it; the contents beside it link to each section on GitHub.

Reliability Audit

You review whether the application fails safely, recovers predictably, and doesn't cascade partial failures into total outages. A reliable system is not a bug-free system — it's one that tolerates the inevitable.

Inputs

From orchestrator: scope_tier, stack_summary, gitnexus_indexed, entry points.

Mode detection

  • Plan mode — produce report with categorized reliability gaps.
  • Edit mode — offer to apply fixes. Things like adding timeouts, adding retries with backoff, and adding idempotency keys are low-risk edits. Things that change transaction boundaries or insert circuit breakers between services need explicit confirmation per change.

Thresholds by tier

Tier Timeouts Retries Idempotency Circuit breakers Graceful degradation
prototype advisory advisory advisory optional optional
team required on all I/O required with backoff on transient failures required on write endpoints recommended required for critical flows
scalable required + bounded end-to-end required + jitter + max attempts + budget required with durable keys required required everywhere user-visible

Review surface

1. Timeouts

Every outbound I/O call must have a timeout. Unbounded waits cause thread / connection / goroutine exhaustion under dependency slowdown — the most common way a small upstream issue turns into a full outage.

Check:

  • HTTP clients: is a client-level timeout set? Per-request override available?
    • Node.js: fetch has no default timeout; check for AbortController usage. axios default is infinity; check timeout config.
    • Go: http.Client{Timeout: ...} set? http.DefaultClient has no timeout — never use it for external calls.
    • Python: requeststimeout=... required (it's missing by default). httpxtimeout set?
    • Java: HttpClient.newBuilder().connectTimeout(...); request-level .timeout(...).
  • Database clients: query timeout, connection timeout, idle timeout set?
  • Message queue clients: publish timeout, consume timeout?
  • External SaaS SDKs (Stripe, Twilio, SendGrid, etc.): default timeouts often too high.
  • gRPC: context.WithTimeout on every outbound call?
  • End-to-end request timeout enforced at the edge (nginx / ALB / API gateway)?

Read the full file on GitHub · 354 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 12d ago First seen · 354 lines · 101 tokens per session scan A 27b8f5f67a04

Subscribe to this mod's changes

reliability-audit is a skill published in the GitHub repository Nordic-AI/production-readiness-skills (2 stars, last pushed 4mo ago), licensed Apache-2.0. It adds 101 tokens to every session and 3,873 once invoked, about $0.0005 per session on Opus 5. A static security scan graded it A with 1 finding (makes network calls). No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.

Related

Other skills, from other repositories

fedramp

Expert guidance for FedRAMP certification and compliance under CR26 (FedRAMP Consolidated Rules for 2026). Use this skill whenever a user asks about FedRAMP authorization, ATO (Authority to Operate), cloud security for federal government, NIST SP 800-53 controls, CSP compliance, or any of the core FedRAMP document…

Sushegaad/Claude-Skills-Governance-Risk-and-Compliance · 236 tokens

gdpr-compliance

Expert GDPR compliance assistant covering all four core workflows: (1) auditing code and systems for GDPR violations, (2) drafting GDPR-compliant documents such as privacy policies, Data Processing Agreements (DPAs), and consent notices, (3) answering GDPR compliance questions with authoritative article citations, and…

Sushegaad/Claude-Skills-Governance-Risk-and-Compliance · 176 tokens

iso42001

Expert ISO 42001 AI Management System (AIMS) compliance advisor. Use this skill whenever a user asks about ISO/IEC 42001:2023, AI governance, AI management systems, AI risk assessment, AI system impact assessment, Annex A controls for AI, Statement of Applicability for AI systems, AI policy, responsible AI, AI…

Sushegaad/Claude-Skills-Governance-Risk-and-Compliance · 173 tokens

eu-cra

Expert EU Cyber Resilience Act (CRA) advisor for Regulation (EU) 2024/2847 — mandatory cybersecurity and vulnerability handling requirements for all products with digital elements (PDEs) sold in the EU. Use this skill for gap analysis, product classification (Default / Class I / Class II), conformity assessment route…

Sushegaad/Claude-Skills-Governance-Risk-and-Compliance · 133 tokens

nist-800-53

NIST SP 800-53 Rev 5 compliance advisor — all 20 control families (AC, AT, AU, CA, CM, CP, IA, IR, MA, MP, PE, PL, PM, PS, PT, RA, SA, SC, SI, SR), Low/Moderate/High baseline selection, FIPS 199/200 system categorization, control tailoring and overlays, privacy controls (PT family), supply chain risk management (SR…

Sushegaad/Claude-Skills-Governance-Risk-and-Compliance · 177 tokens

soc2

Expert SOC 2 compliance assistant covering all five Trust Services Criteria (Security/CC, Availability/A, Confidentiality/C, Processing Integrity/PI, Privacy/P). Use this skill whenever a user mentions SOC 2, Trust Services Criteria, SOC 2 Type 1 or Type 2, audit readiness, compliance gaps, control documentation…

Sushegaad/Claude-Skills-Governance-Risk-and-Compliance · 159 tokens