Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add rules/kok-o/koko-contextos-agents/system-designgit clone --depth 1 https://github.com/kok-o/koko-contextos-agentsWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/rules/kok-o/koko-contextos-agents/system-design)<a href="https://agentmods.dev/rules/kok-o/koko-contextos-agents/system-design"><img src="https://agentmods.dev/badge/rules/kok-o/koko-contextos-agents/system-design.svg" alt="Measured on agentmods" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00000 | $0.03345 |
| Opus 5 | $0.00000 | $0.01673 |
| Sonnet 5 | $0.00000 | $0.00669 |
| Haiku 4.5 | $0.00000 | $0.00334 |
Grade A, and why
system-design scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured today.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 420 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Skill: system-design
system-design
Overview
A brief summary of what the skill does and its core philosophy.
When to Use
Context for when this skill is applicable.
Rules & Patterns
Based on donnemartin/system-design-primer — the most starred system design resource on GitHub.
Core Principle
Everything is a trade-off. Before writing a single line of backend code, reason through the system at scale. A flat monolith that works now fails at 10× load.
Mandatory Pre-Design Checklist
Before architecting any backend system, answer these questions:
- Scale: What is the expected QPS (queries per second)? Peak vs average?
- Data volume: How much data? Growth rate? 1GB? 1TB? 1PB?
- Consistency vs Availability: Can we tolerate eventual consistency? (CAP theorem)
- Read/Write ratio: Is it read-heavy (cache it!) or write-heavy (shard it!)?
- Latency requirements: Real-time (<100ms)? Near-real-time (<1s)? Batch?
- Global distribution: Single region or multi-region?
- Deployment model: Traditional servers, Serverless, or Edge functions?
Core Architecture Patterns
Load Balancing
Clients → Load Balancer → [App Server 1, App Server 2, App Server N]
- Use Round Robin for stateless services
- Use Least Connections for varying request times
- Use IP Hash for session affinity (or move sessions to Redis)
- Always add health checks — remove unhealthy nodes automatically
Rule: Any service expecting > 1000 RPS needs a load balancer. No exceptions.
Caching Strategy
App → [Cache Layer: Redis/Memcached] → Database
Cache decision ladder (check in order):
- Is it read > write? → Cache it
- Is it expensive to compute? → Cache it
- Is it user-specific? → Cache with user key
- Is it global? → Shared cache, shorter TTL
Cache patterns:
- Cache-aside (lazy loading): check cache → miss → load DB → write cache
- Write-through: write to DB AND cache simultaneously (consistency > performance)
- Write-behind: write to cache → async flush to DB (performance > consistency)
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- today Changed · +38 lines c4e1da091d23
- 5d ago First seen · 382 lines · 0 tokens per session scan A b8dad5f3e59f
system-design is a cursor rule published in the GitHub repository kok-o/koko-contextos-agents (2 stars, last pushed yesterday), licensed MIT. It costs nothing until one of its globs matches a file; then it loads 3,345 tokens. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other cursor rules, from other repositories
craft-skill
Create new bigpowers skills with proper structure, progressive disclosure, and bundled resources. Use when user wants to create, write, or build a new skill for the bigpowers lifecycle.
grill-me
Interactive assumption-surfacing Q&A that stress-tests a plan through relentless questioning until every decision is resolved. Use when user wants to challenge a plan, validate decisions from conversation/context, or mentions "grill me". For doc-grounded variant, use grill-with-docs.
run-planning
DISCOVER-PHASE ADVANCER — Drive the discover-phase checklist (specs/planning-status.yaml) through survey-context → scope-work → research-first → elaborate-spec → plan-release → slice-tasks. NOT a duplicate of plan-work or the planning spine; it orchestrates the pre-coding discover phase only.
terse-mode
Fallback ultra-compressed communication mode. Cuts token usage 75% by dropping filler, articles, and pleasantries while keeping full technical accuracy. Use ONLY when context is critically long and compressing output is necessary to continue. Not a strategy — token discipline comes from code shape (small functions…
hatch3r-dynamic-stack-verification
Runtime-evidence gates for stacks with no static type layer (plain-JS Vue, dynamic persistence) — smoke-mount changed components, dormant-collection-read ratchet, runtime smoke evidence on module-boundary changes; a green lint+build+mock-suite proves a layer, not the system.
hatch3r-spec-currency
Spec-currency mandate — any diff changing user-observable behavior covered by a docs/specs/ file updates that spec in the same delivery; staleness definition, review-gate drift finding with a named owner, and a release-readiness sweep.