AI Agent: Design Principles and Engineering Practice is an open-source book that explains how AI agents combine language models, context, and tools, with accompanying experiments and code. It is intended for readers studying the principles and engineering of AI agents, from fundamentals through production use. The catalogue skills support coding-agent work related to the book's subject matter.
Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add bojieli/ai-agent-book --skill context-compressiongit clone --depth 1 https://github.com/bojieli/ai-agent-bookWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/bojieli/ai-agent-book/context-compression)<a href="https://agentmods.dev/skills/bojieli/ai-agent-book/context-compression"><img src="https://agentmods.dev/badge/skills/bojieli/ai-agent-book/context-compression/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/bojieli/ai-agent-book/context-compression"><img src="https://agentmods.dev/badge/skills/bojieli/ai-agent-book/context-compression.svg" alt="Reviewed on agentmods" width="80" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00092 | $0.02975 |
| Opus 5.5 | $0.00037 | $0.01190 |
| Sonnet 5.5 | $0.00018 | $0.00595 |
| Haiku 4.5 | $0.00009 | $0.00298 |
Grade A, and why
context-compression scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 14d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 109 lines — stays where its author put it; the contents beside it link to each section on GitHub.
上下文压缩策略
压缩解决的是相反方向的问题:如何为上下文做减法——什么时候压缩、怎么压缩、为什么即使上下文没满也应该压缩。核心认知:上下文学习本质上是检索而非推理,压缩就是把需要思考才能得到的结论,变成可以直接检索的知识。
何时使用
- 工具调用结果动辄数万字符,几轮交互就撑满 128K 窗口
- Agent 在长任务中「明明窗口没满,却找不到关键信息」或反复纠结已解决的问题
- 模型在任务尚未完成前提前收尾、草率下结论
- 设计压缩策略(何时触发、压缩什么、保留什么)或排查压缩副作用
- 评估「把大任务委派给子 Agent」与「主 Agent 自己做完再压缩」的取舍
核心原则
- 三个动机,层次不同:① 长度与成本约束(窗口有限、token 越贵、延迟越高);② 提升思考质量——总结后的知识比原始形式更利于模型使用,十几轮搜索的原始结果散落各处,模型每次决策都要在数万 token 中反复检索;③ 缓解上下文焦虑(Context Anxiety)——模型认为窗口即将耗尽时可能在任务完成前提前收尾,在未接近耗尽时就提前压缩可能提升决策质量。
- 上下文窗口是一台只有一半的检索引擎:检索这一半极强(相当于每次前向传播都内置了 RAG),但缺「提炼层」——上下文里的东西从来不会被自动数一遍、建索引或就地总结。任何「关于这些内容的结论」都要从原始记录现算一遍,代价随内容量 N 上涨。
- 压缩与状态栏是同一枚硬币的两面:状态栏把算好的结论加进上下文(由代码确定性维护),压缩把臃肿的原始记录换成算好的结论(多用一次 LLM 调用蒸馏)。
- 上下文腐化(Context Rot)≠ 溢出:溢出是「装不下了」,腐化是「装得下但找不到了」——后者更隐蔽,Agent 表面正常工作,决策质量悄然下降。注意力权重被分散到更多 token 上,无关内容一旦占大头,决策质量明显下滑。
- 主动显式提炼,不要期望模型自动学习:与其让模型被动在海量信息中检索,不如主动提供经过提炼的高密度结构化知识。
- 压缩最容易丢失:早期的架构决策、约束背后的理由、失败的路径。因此 Agent 需要定期把进展记录到文档,而不是把所有信息零散堆在执行历史里。
- 隔离优于压缩:压缩是信息已进入上下文后的事后有损补救;隔离让大体积中间信息根本不进入主上下文。
实践模式
1. 六种压缩策略实测对比(实验 2-10)
任务:识别并追踪 OpenAI 联合创始人的职业状态;Kimi K3(原生约 1M 窗口,实验刻意限制在 128K 预算以触发压缩)。
| 策略 | 做法 | 迭代 | Token | 压缩率 | 问题 |
|---|---|---|---|---|---|
| 无压缩 | 完整保留原始结果 | 5 次即溢出失败 | ~165,000 | — | 7 次搜索累计约 367,000 字符(平均 52,000/次),数次搜索即耗尽 128K |
| 个体摘要 | 每个结果独立生成 2-3 段摘要 | 12 | 276,608 | 10.9% | 信息碎片化,多页重复描述同一事件 |
| 组合摘要 | 所有结果合并后一份综合摘要 | 10 | 93,449 | 4.3% | 输入超长必须截断,可能丢失末尾信息 |
| 上下文感知 | 压缩提示中带入查询意图与已积累信息 | 7 | 40,157 | ~3.0% | 最优平衡点 |
| 带引用的上下文感知 | 智能压缩 + 每条事实附带来源 URL | 7 | — | — | 内容有损、索引无损,可回溯原文 |
| 自适应窗口化 | 阈值触发 + 批量压缩 + 防重复 | — | 174,601 | — | 初期保留完整原始信息,灵活性最大 |
压缩率定义为「压缩后体积 / 原文体积」,数值越小表示压得越狠。上下文感知压缩将 token 使用量减少 75% 以上。
上下文感知的提示词写法:在压缩提示中指定 Given the search query: {query} 和 Current context: {context},引导模型生成针对性摘要。实测把约 150K 字符压到 2K 字符时,仍保留创始人姓名与职位变动等后续任务需要的关键信息。
2. 自适应窗口化三机制(推荐默认)
- 阈值触发:持续监控上下文使用率,prompt token 超过窗口 80% 时才激活压缩(如 102,400 / 128K)。
- 批量压缩:触发时一次性压缩所有未标记的工具结果,而不是每轮零碎压。
- 防重复保护:添加
[COMPRESSED]标记,确保已压缩内容永不被重复处理。
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 14d ago First seen · 109 lines · 92 tokens per session scan A b68d34311763
context-compression is a skill published in the GitHub repository bojieli/ai-agent-book (52,744 stars, last pushed today), licensed Apache-2.0. It adds 92 tokens to every session and 2,975 once invoked, about $0.0004 per session on Opus 5.5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-09-24.
Other skills, from other repositories
remem-mcp
Long-term memory for coding agents. Auto-applies at the start of any coding task — recall past context before answering, capture decisions/learnings/fixes after work, use CodeGraph instead of grep for symbol lookup. Invoke when you see [remem-mcp] in your context or when starting any non-trivial coding work.
acontext-installer
Install Acontext, Login & Init Acontext Project, Add Skill Memory to Agent.
agent-memory-discipline
Rules for when an agent should recall from long-term memory before acting and when it should save decisions, corrections and failures afterwards. Works with any memory backend.
kayba-ace
This skill ships learnfromtraces.py, a script that reads OpenClaw session transcripts, feeds them through the ACE learning pipeline, and writes an updated skillbook to disk.
skill-creator
Create or update AgentSkills. Use when designing, structuring, or packaging skills with scripts, references, and assets.
reminder
Set reminders and manage todos with natural language. Uses built-in cron scheduling, no external service needed.