Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add skills/phuonghx/aim-cli/context-compressionnpx skills add phuonghx/aim-cli --skill context-compressiongit clone --depth 1 https://github.com/phuonghx/aim-cliWhat it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00067 | $0.01084 |
| Opus 5 | $0.00034 | $0.00542 |
| Sonnet 5 | $0.00013 | $0.00217 |
| Haiku 4.5 | $0.00007 | $0.00108 |
Grade A, and why
context-compression scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured yesterday.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 142 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Context Compression for Long Sessions
Over a long session the context window fills with finished work, and the assistant starts repeating itself or losing the thread. The fix is to summarize what's done — preserving the decisions, discarding the transcript — so attention stays on what's still active.
Overview
Extended sessions (roughly 30+ turns) degrade: earlier work fades, suggestions repeat, decisions get forgotten. Compressing completed phases as you go keeps the window focused on live work.
Payoff: reclaiming on the order of 5,000–15,000 tokens in a long session by swapping bulky tool output for tight semantic summaries.
When to Compress
| Signal | Response |
|---|---|
| Past ~20 turns | Consider compressing proactively |
| Suggestions start repeating | Window is saturated — compress now |
| User notes "we covered this already" | Compress right away |
| Moving into a new phase | Summarize the phase you're leaving |
| A tool dumps 500+ lines | Compact that output on the spot |
Three Levels
Level 1 — Compact a Tool Output
Shrink a single noisy result down to its meaning.
Before — raw search dump (~200 lines, ~4,000 tokens):
src/auth/jwt.ts:15: import { verify } from 'jsonwebtoken'
src/auth/jwt.ts:23: export function validateToken(token: string) {
src/auth/jwt.ts:24: try {
... (and so on)
After — the gist (~5 lines, ~100 tokens):
Searched "jwt": 8 files, 42 hits. Core: src/auth/jwt.ts (JWT logic),
src/middleware/auth.ts (guard), src/api/login.ts (issues tokens).
Validation lives at jwt.ts:23-40, error handling at 42-55, secret read from env at line 8.
Level 2 — Summarize a Phase
Collapse a whole stretch of exploration into its conclusions.
Before — full research trail (~3,000 tokens):
[turn 1] read package.json
[turn 2] read src/index.ts
[turn 3] searched "auth"
... (a dozen more exploration turns)
After — phase summary (~200 tokens):
## Research done
- Stack: Next.js 15 app, JWT-based auth
- Auth code: 8 files across src/auth, src/middleware, src/api
- Flow: login -> mint JWT -> httpOnly cookie -> verified in middleware
- Defect: src/auth/jwt.ts:45 — expiry test uses `<` where it needs `<=`
- Plan: fix the operator, add a boundary test
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- yesterday First seen · 142 lines · 67 tokens per session scan A 2d2110a82c8b
context-compression is a skill published in the GitHub repository phuonghx/aim-cli (1 stars, last pushed 2mo ago), licensed MIT. It adds 67 tokens to every session and 1,084 once invoked, about $0.0003 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other skills, from other repositories
hs-release
Cut a core Hindsight release (vX.Y.Z) and open the changelog + blog PR. Use when asked to cut/start a release, bump the version, or publish a new Hindsight version.
hindsight-local
Store user preferences, learnings from tasks, and procedure outcomes. Use to remember what works and recall context before new tasks. (user).
research-repository
Build a repository that makes findings findable, reusable, and cumulative across teams. Use when the same research keeps getting redone. For synthesising one study, use affinity-diagram.
design-negotiation
Advocate for design quality, scope, and timeline with partners and leadership using evidence and shared goals. Use in the conversation itself. For the commercial vocabulary behind it, use business-design (ux-strategy).
user-persona
Build research-grounded personas with goals, frustrations, and behavioural patterns. Use when decisions need a consistent user reference. For one session's emotional snapshot use empathy-map; for motivation framing use jobs-to-be-done.
version-control-strategy
Define version control for design files, components, and libraries — branching, naming, and release. Use when file history is chaotic. For design system contribution rules, use design-system-governance (design-systems).