Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add skills/huzaifa525/claude-code-optimizer/modenpx skills add huzaifa525/claude-code-optimizer --skill modegit clone --depth 1 https://github.com/huzaifa525/claude-code-optimizerWhat it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00029 | $0.00653 |
| Opus 5 | $0.00015 | $0.00327 |
| Sonnet 5 | $0.00006 | $0.00131 |
| Haiku 4.5 | $0.00003 | $0.00065 |
Grade A, and why
mode scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 81 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Switch optimization mode to: $ARGUMENTS
Available Modes
Aggressive Mode
Goal: Maximum token savings. Use for routine tasks, simple edits, well-understood codebases.
Rules:
MAX_THINKING_TOKENS=4000— minimal thinking for simple operations- Fork ALL exploration — never explore in main context
- Read ONLY the specific lines needed (always use offset/limit)
- Skip reading files you can infer from context
- Use Haiku for all subagent tasks
- Compact after every 5 tool calls
- Never read more than 50 lines from any file at once
- Prefer Grep → targeted Read over full file reads
- One-sentence responses unless the user asks for detail
Balanced Mode (Default)
Goal: Good savings with reasonable context. Use for general development work.
Rules:
MAX_THINKING_TOKENS=10000— standard thinking budget- Fork exploration and review tasks
- Read full files when needed, but use offset/limit for large files
- Use Sonnet for most tasks, Opus for complex ones
- Compact after reading 10+ files
- Normal response length
Thorough Mode
Goal: Maximum context and quality. Use for complex debugging, architecture decisions, security audits.
Rules:
MAX_THINKING_TOKENS=32000— deep thinking for complex problems- Read related files fully to understand context
- Use Opus for implementation, Sonnet for simple subagents
- Don't compact aggressively — keep context for cross-referencing
- Explore broadly before narrowing down
- Detailed explanations and reasoning in responses
- Run extra verification passes
How to Apply
When the user invokes /mode [name]:
- Acknowledge the mode switch
- State the key behavior changes
- Apply the rules for the rest of this session
Example response:
Switched to **aggressive** mode.
- Minimal thinking, maximum token savings
- All exploration forked to subagents
- Targeted file reads only
- Compact responses
If no argument given, show current mode and the three options.
Mode Comparison
| Aspect | Aggressive | Balanced | Thorough |
|---|---|---|---|
| Thinking budget | 4K | 10K | 32K |
| Exploration | Always forked | Fork when large | Direct when useful |
| File reading | Targeted lines | Full when needed | Full + related |
| Model selection | Haiku subagents | Sonnet default | Opus default |
| Response style | Terse | Normal | Detailed |
| Best for | Routine edits | General dev | Complex problems |
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 2d ago First seen · 81 lines · 29 tokens per session scan A 95fe7b4d747a
mode is a skill published in the GitHub repository huzaifa525/claude-code-optimizer (9 stars, last pushed 5mo ago), licensed MIT. It adds 29 tokens to every session and 653 once invoked, about $0.0001 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other skills, from other repositories
throughline
Use when the user asks to use Throughline from Codex, continue or restore Throughline memory, export read-only handoff context for a local launcher, prepare a new Codex thread handoff, summarize a captured Codex session, or check whether the Throughline Codex Stop hook captured the current session. Hide long…
critical-review
Use when you want structured feedback on a plan or document before implementation.
save
Save all session progress to status tracking files. Use when you want to checkpoint work mid-session or before ending.
code-review
Use when you want a thorough code review of files, changes, or the entire project before shipping.
implement-batch
Use when you want to implement the next batch of a plan. Handles module implementation, testing, and validation.
audit-orchestrator
Universal Pre-Scan → Analysis → Optimization → Report orchestrator for ANY project type — web apps (Astro/SvelteKit/Next), infrastructure/homelab repos, CLI tools, libraries, backend services, monorepos, data/ML projects, docs. Self-detects project type and runs the matching analysis track. Session state lives in…