Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add skills/phuonghx/aim-cli/context-engineeringnpx skills add phuonghx/aim-cli --skill context-engineeringgit clone --depth 1 https://github.com/phuonghx/aim-cliWhat it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00084 | $0.01362 |
| Opus 5 | $0.00042 | $0.00681 |
| Sonnet 5 | $0.00017 | $0.00272 |
| Haiku 4.5 | $0.00008 | $0.00136 |
Grade A, and why
context-engineering scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured yesterday.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 110 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Context Engineering
The context window is a budget, not a bucket. Every token must earn its place.
Why it matters
Models do not get smarter with more context — past a point they get worse. Irrelevant, stale, or redundant tokens cause context rot (degraded recall) and distraction (the model fixates on noise). Curate aggressively: the goal is the smallest set of high-signal tokens that lets the model succeed.
Context types
| Type | What it is | Sourcing strategy |
|---|---|---|
| Instructions | System prompt, role, rules, output format | Preload (stable, always needed) |
| Knowledge | Retrieved docs, facts, code | Just-in-time (fetch per query) |
| Tools | Tool/function schemas available | Scope to the task; don't expose all |
| Memory | Persistent cross-session facts | Ranked recall (importance + recency) |
| History | Prior turns in this session | Compress as it grows |
Token budgeting
- Set an explicit budget per section before assembling context (e.g. instructions 10%, retrieval 50%, history 25%, headroom 15%).
- Leave headroom for the model's output — a full window leaves no room to answer.
- Measure, don't guess: count tokens of each section; log the breakdown.
- When over budget, cut the lowest-signal section first (usually old history or low-ranked retrieval), never the instructions.
Retrieval and ranking
Getting the right knowledge in matters more than getting more in.
- Retrieve a candidate pool, then rerank and keep only the top few — precision over recall.
- Deduplicate near-identical chunks before they reach the window.
- Filter by metadata (recency, source, permissions) before semantic ranking.
- Attach provenance (source id/URL) to each chunk so the model can cite and you can debug.
- Tune chunk size to the content: too small loses context, too large wastes budget. Test it.
Just-in-time vs preloading
| Preload (put it in now) | Just-in-time (fetch when needed) |
|---|---|
| Small, stable, always-relevant | Large, conditional, or rarely needed |
| System rules, output schema, key conventions | Document bodies, search results, file contents |
| Core tool schemas | Niche tools gated behind a router |
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- yesterday First seen · 110 lines · 84 tokens per session scan A 945b65e8299e
context-engineering is a skill published in the GitHub repository phuonghx/aim-cli (1 stars, last pushed 2mo ago), licensed MIT. It adds 84 tokens to every session and 1,362 once invoked, about $0.0004 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other skills, from other repositories
hs-release
Cut a core Hindsight release (vX.Y.Z) and open the changelog + blog PR. Use when asked to cut/start a release, bump the version, or publish a new Hindsight version.
hindsight-local
Store user preferences, learnings from tasks, and procedure outcomes. Use to remember what works and recall context before new tasks. (user).
research-repository
Build a repository that makes findings findable, reusable, and cumulative across teams. Use when the same research keeps getting redone. For synthesising one study, use affinity-diagram.
design-negotiation
Advocate for design quality, scope, and timeline with partners and leadership using evidence and shared goals. Use in the conversation itself. For the commercial vocabulary behind it, use business-design (ux-strategy).
user-persona
Build research-grounded personas with goals, frustrations, and behavioural patterns. Use when decisions need a consistent user reference. For one session's emotional snapshot use empathy-map; for motivation framing use jobs-to-be-done.
version-control-strategy
Define version control for design files, components, and libraries — branching, naming, and release. Use when file history is chaotic. For design system contribution rules, use design-system-governance (design-systems).