Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add handsomeng/Hskill-chatcut --skill edit-talking-head-videosgit clone --depth 1 https://github.com/handsomeng/Hskill-chatcutWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/handsomeng/hskill-chatcut/edit-talking-head-videos)<a href="https://agentmods.dev/skills/handsomeng/hskill-chatcut/edit-talking-head-videos"><img src="https://agentmods.dev/badge/skills/handsomeng/hskill-chatcut/edit-talking-head-videos/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/handsomeng/hskill-chatcut/edit-talking-head-videos"><img src="https://agentmods.dev/badge/skills/handsomeng/hskill-chatcut/edit-talking-head-videos.svg" alt="Reviewed on agentmods" width="80" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00099 | $0.01691 |
| Opus 5 | $0.00049 | $0.00846 |
| Sonnet 5 | $0.00020 | $0.00338 |
| Haiku 4.5 | $0.00010 | $0.00169 |
Grade A, and why
edit-talking-head-videos scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 10d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 126 lines — stays where its author put it; the contents beside it link to each section on GitHub.
口播视频剪辑
使用 ChatCut 把原始口播整理成连贯、易懂、可发布的视频。根据内容、受众、发布平台和用户品牌决定视觉方案,不套用固定行业风格。
前置要求
- 确认 ChatCut 插件可用,并按任务需要读取对应 ChatCut Skill。
- 优先复用已经存在的项目和时间线。重新接管项目时先读取当前状态,再做增量修改。
- 用户未指定画幅时沿用源视频。面向竖屏短视频且素材允许时,默认使用 9:16、1080×1920、30fps。
- 用户未要求导出时,先交付可预览版本。
决策顺序
按以下优先级做决定:
- 用户明确要求
- 用户已有品牌规范或参考图
- 视频内容、受众和发布平台
- 本 Skill 的克制默认值
遇到视觉偏好不明确时,先做一版干净、真实、易读的方案。不得自行套用科技、手账、综艺、新闻或其他固定模板。
标准流程
1. 理解内容
- 导入素材并完成转录。
- 阅读完整转录,提取主题、目标受众、核心观点、论据、案例、转折和结论。
- 标记错误重录、重复表达、无意义填充、长停顿、离题内容和不可用画面。
- 先拟定剪辑结构,再修改时间线。
2. 连贯型粗剪
- 删除明确错误的重录、完整重复、无意义口头填充和过长空白。
- 同一内容存在多次录制时,保留表达最完整、语气最自然的一次。
- 检查每个剪切点的主语、指代、因果、转折和结论。
- 保留必要气口和自然停顿,避免为了节奏频繁跳切。
- 不改变讲述者观点,不凭空补充口播事实。
- 删除一句话中的片段前,先确认剩余部分仍是完整句子。
3. 字幕
默认添加与口播同语言的字幕。只有用户要求双语字幕或目标受众需要时,才添加第二语言。
字幕遵守以下规则:
- 每屏展示一个完整句子或能够独立成立的语义单元。
- 允许一句话在同一字幕事件中自然换行,不因换行拆成多个事件。
- 长句只在自然分句点跨屏,不拆开专有名词、数字与单位、主谓结构或因果结构。
- 删除跨屏残字、重复词和没有语义作用的识别尾巴。
- 核对人名、公司名、品牌名、模型名、数字、日期、价格和专业术语。
- 字幕位于安全区,避免遮住人物眼睛、嘴部、产品和关键画面。
- 字号根据画布和每屏字数自适应。先保证手机端可读,再控制遮挡。
- 默认每种语言不超过 2 行。内容过长时优先优化断句,不盲目缩小字号。
启用双语字幕时:
- 为每屏原文单独改写译文,保持主语、动作、因果、转折、语气、数字和专有名词一致。
- 两种语言使用相同起止时间。
- 使用自然表达,不直接采用未经核验的机器翻译。
- 逐屏检查原文和译文的语义对应关系。
4. 内容结构提示
根据内容决定是否增加章节进度和重点总结。结构简单或视频很短时可以省略。
使用章节进度时:
- 按真实语义拆分模块,不固定模块数量。
- 模块标题简短、互不重复,并从观众视角概括该段价值。
- 在顶部显示完整内容地图,让观众知道前后结构。
- 使用从左向右连续推进的进度线。当前模块高亮,其他模块降低明度。
- 保持元素紧凑,避免遮挡人物或字幕。
使用重点总结时:
- 只在真正影响理解和记忆的位置出现。
- 直接显示总结文案,不添加“重点总结”“核心观点”等固定标签。
- 每次显示约 2 至 4 秒,控制在 1 至 3 行。
- 文案必须忠于原口播,不添加未经讲述的结论。
5. 封面
用户需要封面时,根据视频内容和品牌生成方案:
- 从视频选择眼睛睁开、表情自然、脸部清楚、构图稳定的真实帧,或使用用户提供的形象素材。
- 根据视频核心观点生成 2 至 5 个标题候选。标题必须来自视频内容,不夸大事实。
- 根据受众和品牌选择字体、颜色、背景和排版。优先保证手机缩略图可读。
- 核对标题的每个字、行数、对齐和安全区。
- 没有品牌规范时,默认使用真实人物画面、克制背景和大号高对比标题。
封面嵌入视频时只占第 1 帧。正文从第 2 帧直接开始,不添加停留、淡入淡出或转场。若平台支持独立上传封面,优先保留独立封面文件。
6. 音频与画面
- 人声优先,做必要的响度平衡、降噪和去爆音。
- 用户未要求时不默认添加背景音乐。
- 添加音乐时选择与内容和品牌一致的无人声音乐,并对口播做 ducking。
- 不默认使用高频缩放、卡点、强音效、密集花字和频繁转场。
- 必要的画面重构、B-roll 和动效必须服务于理解。
- 涉及驾驶、医疗、金融或其他高风险场景时,根据画面证据、发布地区和用户要求添加合适声明,禁止写死固定文案。
7. 验证与导出
交付前完整读取 references/quality-checks.md,执行其中的全片检查和截图审核。
What ships with it
1 file beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 10d ago First seen · 126 lines · 99 tokens per session scan A 9fc2a8ccfd86
edit-talking-head-videos is a skill published in the GitHub repository handsomeng/Hskill-chatcut (6 stars, last pushed 1mo ago), licensed MIT. It adds 99 tokens to every session and 1,691 once invoked, about $0.0005 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other skills, from other repositories
voice
Text-to-Speech (TTS), voice cloning, voiceover, narration placement/sync, and custom sound effects (SFX) generator. Use when the user wants generated speech from text, wants to clone a consented voice from uploaded reference audio, wants to add/replace/align narration or voiceover for an existing video/timeline, wants…
create-motion-graphics
Use whenever an ACP or local CLI agent in ChatCut Desktop needs to add, create, hand-author, patch, or place Motion Graphic JSX assets in a project. Covers direct inline JSX authoring, visual language, editable properties, asset binding, timeline placement, and local verification. Not for the built-in ChatCut Agent.
talking-head-guide
A guide for editing videos where spoken delivery or conversation drives the structure, such as interviews, podcasts, lectures, tutorials, and courses.
video-gen
A video-generation guide for creating or changing clips from text, images, frames, or reference material.
multicam-sync
Synchronize footage from a multi-camera / multi-recorder shoot — several cameras plus separate audio recorders covering one session, imported as loose clips — and optionally turn all or part of it into speaker-follow footage for a larger edit. Use when a user drops in multiple clips from the same recording and wants…
asset-import
Use in Claude Code when acquiring or importing media into a ChatCut project, including local or attached videos, readable user-provided paths, files already uploaded in the editor, public media URLs, Browser-pane injection, transcription readiness, and upload fallback decisions.