Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add shawnwei512/doc-distillation-mcp --skill skillgit clone --depth 1 https://github.com/shawnwei512/doc-distillation-mcpWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/shawnwei512/doc-distillation-mcp/skill)<a href="https://agentmods.dev/skills/shawnwei512/doc-distillation-mcp/skill"><img src="https://agentmods.dev/badge/skills/shawnwei512/doc-distillation-mcp/skill/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/shawnwei512/doc-distillation-mcp/skill"><img src="https://agentmods.dev/badge/skills/shawnwei512/doc-distillation-mcp/skill.svg" alt="Reviewed on agentmods" width="80" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00172 | $0.10894 |
| Opus 5 | $0.00086 | $0.05447 |
| Sonnet 5 | $0.00034 | $0.02179 |
| Haiku 4.5 | $0.00017 | $0.01089 |
Grade A, and why
doc-distillation scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 5d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 673 lines — stays where its author put it; the contents beside it link to each section on GitHub.
资料蒸馏技能 (Document Distillation)
将各类来源(飞书文档、网页、PDF等)的原始资料,通过整理、分析、结构化、重组,生成两种高质量输出:HTML蒸馏文稿 和 Obsidian笔记。
合规使用边界
[!important] 本技能仅用于个人学习目的的知识整理
适用场景(用户需确保对内容拥有合法访问权):
- 用户自己创建的飞书文档/笔记
- 用户被授权访问的团队文档
- 公开免费的网页/视频/播客内容
- 用户已购买的课程(仅限个人学习笔记,不公开发布)
不适用场景:
- 未经授权访问付费/受限内容
- 批量复制他人版权内容用于分发
- 绕过任何平台的技术保护措施
蒸馏原则:蒸馏是知识提炼(用自己的话概括核心观点),不是原文复制。输出内容应大幅压缩原文,保留信息密度而非原文照搬。
触发条件
当用户出现以下意图时激活:
- 提供文档链接(飞书/网页等)并要求"提取"、"蒸馏"、"整理"、"总结"
- 提供视频链接(YouTube/B站/抖音/快手/小红书/视频号等)并要求"提取文案"、"蒸馏"、"整理"
- 提供播客链接或本地音视频文件并要求"转录"、"蒸馏"
- 要求将文档/视频内容转化为笔记或文稿
- 提及"蒸馏"一词与文档/视频处理相关
- 要求从资料中提炼核心知识和结构
核心原则
- 双输出:始终生成 HTML蒸馏文稿 + Obsidian笔记,缺一不可
- 原文尊重:蒸馏是提炼而非改写,保留原文核心观点和表达
- 结构优先:先理解整体结构,再填充细节
- 图片完整:提取所有图片,保持图文关联
- 图片内容蒸馏(关键!):不仅要提取图片本身,还必须对包含文字信息的图表、框架图、流程图等做内容蒸馏——用文字完整提取图片中的所有文字、结构和关系,融入Obsidian笔记中。这是确保笔记"可独立阅读"的核心要求。
- 独立成篇:每篇文档独立生成一套输出,不合并多篇
- 图片智能过滤(效率优先):不盲目下载所有图片,通过三层过滤机制(自动规则→上下文预判→选择性蒸馏),只下载和蒸馏含关键信息的图片,节约token和时间
- 完整性优先于效率(关键!):节约token的前提是不遗漏关键信息。压缩表达方式,不压缩信息量。通过"结构骨架法+关键要素标记+双向校验"三重机制,在蒸馏前建立完整性清单,在蒸馏后执行缺口检查,确保零遗漏。
token节约与完整性的平衡原则
[!important] 核心准则:压缩表达方式,不压缩信息量 蒸馏的目标是用更少的字数传达同样的信息密度,而非用更少的字数传达更少的信息。
| 区域 | 策略 | 说明 |
|---|---|---|
| 高信息密度区(不可压缩) | 原样保留或结构化重组 | 核心论点、关键数据(数字/比例/金额)、操作步骤、公式/模板、框架/模型、清单/表格、避坑提醒 |
| 低信息密度区(可压缩) | 提炼为一句话或删除 | 过渡性描述、重复论述、冗长案例、情感修辞、背景铺垫 |
| 图片过滤区 | 三层过滤+安全网 | 装饰图跳过,内容图必须蒸馏,边界图存疑时保留 |
用户偏好(务必遵守)
- 沟通语言:中文
- HTML蒸馏文稿路径:
~/Documents/蒸馏文稿/ - Obsidian笔记路径:
~/Documents/obsidian/(具体子目录由用户每次指定) - 图片存储:Obsidian笔记目录下的
assets/子目录 - 不需要:小红书帖子等其他格式输出
蒸馏工作流程
阶段1:来源识别与内容提取
1.1 识别来源类型
- 飞书文档(feishu.cn / larkoffice.com / yitang.top/fs-doc):优先尝试 lark-doc 技能,权限不足时用浏览器提取
- 普通网页:直接用浏览器提取
- PDF/PPT:用 pdfplumber/PyMuPDF 提取文本,PPT型PDF需渲染页面为图片后Read蒸馏
- 视频/播客(YouTube/B站/抖音/快手/小红书/视频号/播客):用
scripts/video_transcript.py提取文案,详见references/video-transcript-guide.md - 本地音视频文件(mp3/m4a/mp4/wav等):直接用
scripts/video_transcript.py转录 - 用户粘贴文本:直接进入阶段2
What ships with it
9 files beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
- examples/视频文案蒸馏示例.md 799 B
- examples/飞书文档蒸馏示例.md 930 B
- references/feishu-image-extraction-guide.md 8.4 KB
- references/html-template.md 7.7 KB
- references/obsidian-template.md 5.1 KB
- references/quality-checklist.md 8.9 KB
- references/video-transcript-guide.md 8.0 KB
- scripts/image_preprocessor.py 6.6 KB runs code
- scripts/video_transcript.py 16 KB runs code
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 5d ago Changed · +34 lines · +14 tokens per session c1fe3eebb6ac
- 9d ago First seen · 639 lines · 158 tokens per session scan A 214aeffdea36
doc-distillation is a skill published in the GitHub repository shawnwei512/doc-distillation-mcp (1 stars, last pushed 6d ago), licensed MIT. It adds 172 tokens to every session and 10,894 once invoked, about $0.0009 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-09-04.
Other skills, from other repositories
baoyu-youtube-transcript
A tool for downloading the written captions, subtitles, chapter information, speaker labels, and cover image from a YouTube video using its URL or ID.
orbit-notion
Open Orbit briefing skill — selected by the Orbit pipeline when Notion is the user's only connected connector, or when the user explicitly scopes their daily digest to Notion. Pulls the past 24 hours of document edits, comments, mentions, and database row changes from the user's authenticated Notion connection and…
instrument-data-to-allotrope
Convert laboratory instrument output files (PDF, CSV, Excel, TXT) to Allotrope Simple Model (ASM) JSON format or flattened 2D CSV. Use this skill when scientists need to standardize instrument data for LIMS systems, data lakes, or downstream analysis. Supports auto-detection of instrument types. Outputs include full…
feishu
Work with Feishu or Lark bots, docs, sheets, bitables, approval flows, and OpenAPI/MCP setup without hardcoding credentials.
read
Reads URLs and PDFs by fetching source content, defaulting to concise summaries for plain read requests and clean Markdown when asked to convert, save, quote, cite, or feed downstream work. Use when users ask in any language to read, fetch, check, summarize, quote, cite, convert, or save a URL or PDF. Not for local…
overleaf-sync
A two-way connection between a local paper folder and Overleaf, a web-based LaTeX editor for writing research papers. It lets you move changes between the local files and the shared Overleaf project.