Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add Jkudjo/oh-my-cursor --skill chaosgit clone --depth 1 https://github.com/Jkudjo/oh-my-cursorWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/jkudjo/oh-my-cursor/chaos)<a href="https://agentmods.dev/skills/jkudjo/oh-my-cursor/chaos"><img src="https://agentmods.dev/badge/skills/jkudjo/oh-my-cursor/chaos/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/jkudjo/oh-my-cursor/chaos"><img src="https://agentmods.dev/badge/skills/jkudjo/oh-my-cursor/chaos.svg" alt="Reviewed on agentmods" width="80" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00014 | $0.00216 |
| Opus 5 | $0.00007 | $0.00108 |
| Sonnet 5 | $0.00003 | $0.00043 |
| Haiku 4.5 | $0.00001 | $0.00022 |
Grade A, and why
chaos scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 10d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
What it actually says
chaos
Tests whether an implementation degrades gracefully under failure conditions — not just correctness under happy path.
What it injects
- Service unavailability (external APIs, Redis, DB)
- Intermittent failures (30% failure rate)
- Slow dependencies (10x normal latency)
- Null/empty/boundary inputs
- Concurrent access and duplicate requests
- Auth bypass and privilege escalation attempts
- Input injection (SQL, HTML, command)
Output
- Per-scenario verdict: RESILIENT | DEGRADES GRACEFULLY | FAILS UNGRACEFULLY | CATASTROPHIC
- Resilience score (0–100)
- Hard blocks on CATASTROPHIC findings (data loss, security bypass, silent corruption)
Example
@chaos "the new checkout flow"
@chaos "the JWT authentication middleware"
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 10d ago First seen · 30 lines · 14 tokens per session scan A 94c7c8425225
chaos is a skill published in the GitHub repository Jkudjo/oh-my-cursor (1 stars, last pushed 5mo ago), licensed MIT. It adds 14 tokens to every session and 216 once invoked, about $0.0001 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other skills, from other repositories
testing-pyramid
Test Distribution Analysis Command 🧪.
android-preflight
Final verification checklist to run before declaring Android work finished — build, both themes, string resources, lifecycle and leak risks, registered permissions and components, resource parity between values and values-night, and honest reporting of what was and was not verified. Use at the end of any feature, fix…
ai-evals-builder
Build AI evals using the Husain-Shankar framework (error analysis, open/axial coding, LLM-as-judge). Use when a user needs to create, improve, or debug evals for an AI product — including defining failure modes, building LLM judges, or setting up production monitoring for an LLM application.
test-driven-development
Use when implementing any feature, bugfix, or code change - write tests first, watch them fail, then implement.
character-animation-qa
Review local character animation with schema checks, Playwright browser previews, frame sampling, and FFmpeg/ffprobe final output checks.
patrol-e2e-testing
Use when writing E2E/integration tests, testing native interactions like permissions or system dialogs, capturing UI regressions, or validating cross-platform behavior (Patrol 4.x).