Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add HoangNguyen0403/agent-skills-standard --skill system-design-resilience-opsgit clone --depth 1 https://github.com/HoangNguyen0403/agent-skills-standardWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/hoangnguyen0403/agent-skills-standard/system-design-resilience-ops)<a href="https://agentmods.dev/skills/hoangnguyen0403/agent-skills-standard/system-design-resilience-ops"><img src="https://agentmods.dev/badge/skills/hoangnguyen0403/agent-skills-standard/system-design-resilience-ops/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/hoangnguyen0403/agent-skills-standard/system-design-resilience-ops"><img src="https://agentmods.dev/badge/skills/hoangnguyen0403/agent-skills-standard/system-design-resilience-ops.svg" alt="Reviewed on agentmods" width="80" height="20"></a>- NVIDIA SkillSpector pass
What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00060 | $0.00810 |
| Opus 5 | $0.00030 | $0.00405 |
| Sonnet 5 | $0.00012 | $0.00162 |
| Haiku 4.5 | $0.00006 | $0.00081 |
Grade A, and why
system-design-resilience-ops scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 13d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 74 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Resilience and Operations
Priority: P1 (HIGH)
A design is not done until its failure and its rollout are designed.
SPOF Elimination
- Walk every component and ask what happens when exactly one instance dies, then when the whole zone dies.
- Any component with one instance, one writer, or one shared config plane is a single point of failure. Name it or remove it.
- Redundancy only helps when failure modes are independent: shared credentials, shared config, and a shared control plane cancel the benefit.
- Blast radius: state which users or flows are affected per component failure, and cap it with cells, bulkheads, or per-tenant quotas.
Failover and Recovery
| Topology | Recovery time | Cost | Fits |
|---|---|---|---|
| Single region, multi-AZ | Minutes, automatic | Low | Most products |
| Active-passive across regions | Minutes to hours, drill-dependent | Medium | Regulated or high-value flows |
| Active-active across regions | Seconds | High | Global low-latency, conflict-tolerant data |
- Set RPO (tolerable data loss) and RTO (tolerable downtime) as numbers before choosing a topology; the numbers pick the topology, not the reverse.
- Untested failover is a hypothesis. Schedule a drill and record the measured RTO against the target.
- Backups need a restore test. A backup that has never been restored is not a backup.
Observability
- Instrument the four signals per service: traffic, error rate, latency percentiles, saturation.
- Alert on user-visible symptoms and on error-budget burn rate, not on raw CPU.
- Propagate a trace and correlation id across every hop, including queue messages.
- Every alert needs an owner, a runbook link, and a defined next action; an alert nobody acts on is noise.
Rollout
| Strategy | Blast radius | Rollback | Cost |
|---|---|---|---|
| Rolling | Grows during the roll | Roll forward or back, slow | Low |
| Blue-green | Full switch at cutover | Instant switch back | Double capacity |
| Canary | Small cohort first | Stop and drain the cohort | Needs routing plus metrics |
| Feature flag | Per user or tenant | Instant, no redeploy | Flag lifecycle debt |
What ships with it
1 file beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 13d ago First seen · 74 lines · 60 tokens per session scan A 0d9f714d908b
system-design-resilience-ops is a skill published in the GitHub repository HoangNguyen0403/agent-skills-standard (565 stars, last pushed 3d ago), licensed MIT. It adds 60 tokens to every session and 810 once invoked, about $0.0003 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other skills, from other repositories
nextjs-app-router
Full end-to-end tRPC setup for Next.js App Router. Covers route handler with fetchRequestHandler (GET + POST exports), TRPCProvider with QueryClientProvider, createTRPCOptionsProxy for RSC prefetching, HydrateClient/HydrationBoundary for hydration, useSuspenseQuery for Suspense, and server-side callers.
nextjs-pages-router
Set up tRPC in Next.js Pages Router with createNextApiHandler, createTRPCNext, withTRPC HOC, SSR via ssr option and ssrPrepass, SSG via createServerSideHelpers with getStaticProps, and server-side helpers for getServerSideProps prefetching.
openapi
Generate OpenAPI 3.1 spec from a tRPC router with @trpc/openapi CLI or programmatic API. Generate typed REST client with @hey-api/openapi-ts and configureTRPCHeyApiClient(). Configure transformers (superjson, EJSON) for generated clients. Alpha status.
server-setup
Initialize tRPC with initTRPC.create(), define routers with t.router(), create procedures with .query()/.mutation()/.subscription(), configure context with createContext(), export AppRouter type, merge routers with t.mergeRouters(), lazy-load routers with lazy().
subscriptions
Set up real-time event streams with async generator subscriptions using .subscription(async function() { yield }). SSE via httpSubscriptionLink is recommended over WebSocket. Use tracked(id, data) from @trpc/server for reconnection recovery with lastEventId. WebSocket via wsLink and createWSClient from @trpc/client…
client-setup
Create a vanilla tRPC client with createTRPCClient (), configure link chain with httpBatchLink/httpLink, dynamic headers for auth, transformer on links (not client constructor). Infer types with inferRouterInputs and inferRouterOutputs. AbortController signal support. TRPCClientError typing.