reliability-improvement-plan

reliability-improvement-plan is a skill for Claude Code, Codex from Bilal140202/the-lord-of-the-skills. It costs 39 tokens per session (2,790 once invoked), scanned A, original, MIT.

A codebase review process for finding single points of failure and assessing how infrastructure recovers from outages.

In plain words
What is it for?
Use it to inspect infrastructure-as-code, scaling, databases, caches, load balancers, queues, storage, and DNS, then produce a prioritized repair plan.
Why use it?
It exposes weak redundancy and recovery arrangements before they cause prolonged downtime.

Skill for Claude CodeCodex

Written for no agent in particular: nothing here depends on one.

Good fit Use it to inspect infrastructure-as-code, scaling, databases, caches, load balancers, queues, storage, and DNS, then produce a prioritized repair plan.

Compare 6 skills from other repositories ↓
Install with agentmods
npx agentmods add skills/bilal140202/the-lord-of-the-skills/aws-samples__sample-well-architected-skills-and-steering
Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

Any agent
npx skills add Bilal140202/the-lord-of-the-skills --skill aws-samples__sample-well-architected-skills-and-steering
Clone the repo
git clone --depth 1 https://github.com/Bilal140202/the-lord-of-the-skills

Made for: Claude Code, Codex.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for reliability-improvement-plan

README.md
[![agentmods](https://agentmods.dev/badge/skills/bilal140202/the-lord-of-the-skills/aws-samples__sample-well-architected-skills-and-steering/github.svg)](https://agentmods.dev/skills/bilal140202/the-lord-of-the-skills/aws-samples__sample-well-architected-skills-and-steering)
Your own site
<a href="https://agentmods.dev/skills/bilal140202/the-lord-of-the-skills/aws-samples__sample-well-architected-skills-and-steering"><img src="https://agentmods.dev/badge/skills/bilal140202/the-lord-of-the-skills/aws-samples__sample-well-architected-skills-and-steering/github.svg" alt="Measured on agentmods" height="20"></a>

Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.

agentmods 80×15 button for reliability-improvement-plan

Your own site · 80×15
<a href="https://agentmods.dev/skills/bilal140202/the-lord-of-the-skills/aws-samples__sample-well-architected-skills-and-steering"><img src="https://agentmods.dev/badge/skills/bilal140202/the-lord-of-the-skills/aws-samples__sample-well-architected-skills-and-steering.svg" alt="Reviewed on agentmods" width="80" height="20"></a>
Per session 39 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 2,790 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. A grade says what 26 rules found in the file — not that it is safe.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5.1 $0.00039 $0.02790
Opus 5 $0.00019 $0.01395
Sonnet 5 $0.00008 $0.00558
Haiku 4.5 $0.00004 $0.00279

Measured 9d ago against content hash b22e710a2fb7, method: parsed. Prices are Anthropic first-party input rates as of 2026-09-12, from the pricing page.

Security

Grade A, and why

reliability-improvement-plan scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 9d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

skills/gondor/claude-code/aws-samples__sample-well-architected-skills-and-steering/SKILL.md · 310 lines

How it starts

The opening of the file, as written. The whole thing — 310 lines — stays where its author put it; the contents beside it link to each section on GitHub.

Reliability Improvement Plan

Step 1: Gather context

Ask the user:

What workload would you like me to assess for reliability? Please share:

  • Workload name and code packages/directories to analyze
  • Availability target (99.9%, 99.95%, 99.99%, etc.)
  • Recovery objectives (RTO and RPO if defined)
  • Past incidents (optional — recent outages or near-misses)

If context is already provided or you are in a codebase with IaC, proceed directly.

Step 2: Fault Tolerance Discovery

Analyze infrastructure for single points of failure.

You MUST examine:

  • Compute deployments (AZ distribution, instance count, ASG configs)
  • Database configurations (Multi-AZ, read replicas, cluster topology)
  • Cache configurations (cluster mode, replica counts, failover)
  • Load balancer configurations (cross-zone, health checks, target groups)
  • NAT Gateway placement (single vs per-AZ)
  • DNS configurations (Route 53 health checks, failover routing)
  • Queue and messaging configs (DLQ, redrive policies)
  • Storage redundancy (S3 replication, EBS snapshots, EFS)

For each component, document:

  • File path and line numbers
  • Current redundancy level (single-AZ, multi-AZ, multi-region)
  • Failure blast radius
  • Failover mechanism (automatic, manual, none)

You MUST flag as HIGH RISK:

  • Single-AZ database deployments for production workloads
  • Compute without auto-scaling (fixed instance count)
  • No health checks on load-balanced targets
  • Single NAT Gateway serving multiple AZs
  • Stateful services without replication
  • Missing DLQ on async invocations (Lambda, SQS, EventBridge)
  • No circuit breaker or timeout on external service calls

Step 3: Recovery Capability Discovery

Analyze backup and recovery configurations.

You MUST examine:

  • AWS Backup plans and rules
  • RDS automated backup settings (retention, PITR)
  • S3 versioning and replication rules
  • DynamoDB PITR and backup settings
  • EBS snapshot configurations
  • Cross-region replication rules
  • Disaster recovery configurations (pilot light, warm standby resources)

Read the full file on GitHub · 310 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 9d ago First seen · 310 lines · 39 tokens per session scan A b22e710a2fb7

Subscribe to this mod's changes

reliability-improvement-plan is a skill published in the GitHub repository Bilal140202/the-lord-of-the-skills (4 stars, last pushed 6d ago), licensed MIT. It adds 39 tokens to every session and 2,790 once invoked, about $0.0002 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-09-03.

Related

Other skills, from other repositories

serpsmith

Publish SEO articles reliably across AI-agent runtimes.

emiliojohann/SERPsmith · 14 tokens

tool-calling-tutor

Use when a tool-calling agent does not call a tool, sends wrong arguments, loops without stopping, or needs a function schema. Guides a four-branch diagnosis and five-step schema repair. Do not use for framework-specific, MCP-server, or production-observability questions.

WenyuChiou/awesome-agentic-ai-zh · 62 tokens

performing-threat-hunting-with-yara-rules

Use YARA pattern-matching rules to hunt for malware, suspicious files, and indicators of compromise across filesystems and memory dumps. Covers rule authoring, yara-python scanning, and integration with threat intel feeds.

adriannoes/awesome-agentic-ai · 53 tokens

hunt-idor

Hunting skill for idor vulnerabilities. Built from 26 public bug bounty reports. Use when hunting idor on any target.

adriannoes/awesome-agentic-ai · 30 tokens

performing-soc2-type2-audit-preparation

Automates SOC 2 Type II audit preparation including gap assessment against AICPA Trust Services Criteria (CC1-CC9), evidence collection from cloud providers and identity systems, control testing validation, remediation tracking, and continuous compliance monitoring. Covers all five TSC categories (Security…

adriannoes/awesome-agentic-ai · 112 tokens

testrail

Sync tests with TestRail. Use when user mentions "testrail", "test management", "test cases", "test run", "sync test cases", "push results to testrail", or "import from testrail".

adriannoes/awesome-agentic-ai · 50 tokens