edit-talking-head-videos

edit-talking-head-videos is a skill for Claude Code, Codex from handsomeng/Hskill-chatcut. It costs 99 tokens per session (1,691 once invoked), scanned A, original, MIT.

A set of instructions for editing talking-head videos in ChatCut, using the spoken transcript to guide the cut, subtitles, pacing, and optional visual summaries. Talking-head videos are recordings where a person speaks directly to the camera.

In plain words
What is it for?
It is for making publishable talking-head videos with transcript-led rough cuts, complete-sentence subtitles, optional bilingual subtitles, chapter progress, summary overlays, and an optional cover frame.
Why use it?
It helps turn raw speech into a coherent, readable video by removing mistakes, repetition, filler words, and long pauses while preserving the speaker’s meaning. It also provides checks for subtitle accuracy, continuity, and export quality.

Skill for Claude CodeCodex

Written for no agent in particular: nothing here depends on one.

Good fit It is for making publishable talking-head videos with transcript-led rough cuts, complete-sentence subtitles, optional bilingual subtitles, chapter progress, summary overlays, and an optional cover frame.

Compare 6 skills from other repositories ↓
Install with agentmods
npx agentmods add skills/handsomeng/hskill-chatcut/edit-talking-head-videos
Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

Any agent
npx skills add handsomeng/Hskill-chatcut --skill edit-talking-head-videos
Clone the repo
git clone --depth 1 https://github.com/handsomeng/Hskill-chatcut

Made for: Claude Code, Codex.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for edit-talking-head-videos

README.md
[![agentmods](https://agentmods.dev/badge/skills/handsomeng/hskill-chatcut/edit-talking-head-videos/github.svg)](https://agentmods.dev/skills/handsomeng/hskill-chatcut/edit-talking-head-videos)
Your own site
<a href="https://agentmods.dev/skills/handsomeng/hskill-chatcut/edit-talking-head-videos"><img src="https://agentmods.dev/badge/skills/handsomeng/hskill-chatcut/edit-talking-head-videos/github.svg" alt="Measured on agentmods" height="20"></a>

Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.

agentmods 80×15 button for edit-talking-head-videos

Your own site · 80×15
<a href="https://agentmods.dev/skills/handsomeng/hskill-chatcut/edit-talking-head-videos"><img src="https://agentmods.dev/badge/skills/handsomeng/hskill-chatcut/edit-talking-head-videos.svg" alt="Reviewed on agentmods" width="80" height="20"></a>
Per session 99 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 1,691 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. A grade says what 26 rules found in the file — not that it is safe.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5.1 $0.00099 $0.01691
Opus 5 $0.00049 $0.00846
Sonnet 5 $0.00020 $0.00338
Haiku 4.5 $0.00010 $0.00169

Measured 10d ago against content hash 9fc2a8ccfd86, method: parsed. Prices are Anthropic first-party input rates as of 2026-09-10, from the pricing page.

Security

Grade A, and why

edit-talking-head-videos scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 10d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

skills/edit-talking-head-videos/SKILL.md · 126 lines

How it starts

The opening of the file, as written. The whole thing — 126 lines — stays where its author put it; the contents beside it link to each section on GitHub.

口播视频剪辑

使用 ChatCut 把原始口播整理成连贯、易懂、可发布的视频。根据内容、受众、发布平台和用户品牌决定视觉方案,不套用固定行业风格。

前置要求

  • 确认 ChatCut 插件可用,并按任务需要读取对应 ChatCut Skill。
  • 优先复用已经存在的项目和时间线。重新接管项目时先读取当前状态,再做增量修改。
  • 用户未指定画幅时沿用源视频。面向竖屏短视频且素材允许时,默认使用 9:16、1080×1920、30fps。
  • 用户未要求导出时,先交付可预览版本。

决策顺序

按以下优先级做决定:

  1. 用户明确要求
  2. 用户已有品牌规范或参考图
  3. 视频内容、受众和发布平台
  4. 本 Skill 的克制默认值

遇到视觉偏好不明确时,先做一版干净、真实、易读的方案。不得自行套用科技、手账、综艺、新闻或其他固定模板。

标准流程

1. 理解内容

  1. 导入素材并完成转录。
  2. 阅读完整转录,提取主题、目标受众、核心观点、论据、案例、转折和结论。
  3. 标记错误重录、重复表达、无意义填充、长停顿、离题内容和不可用画面。
  4. 先拟定剪辑结构,再修改时间线。

2. 连贯型粗剪

  1. 删除明确错误的重录、完整重复、无意义口头填充和过长空白。
  2. 同一内容存在多次录制时,保留表达最完整、语气最自然的一次。
  3. 检查每个剪切点的主语、指代、因果、转折和结论。
  4. 保留必要气口和自然停顿,避免为了节奏频繁跳切。
  5. 不改变讲述者观点,不凭空补充口播事实。
  6. 删除一句话中的片段前,先确认剩余部分仍是完整句子。

3. 字幕

默认添加与口播同语言的字幕。只有用户要求双语字幕或目标受众需要时,才添加第二语言。

字幕遵守以下规则:

  • 每屏展示一个完整句子或能够独立成立的语义单元。
  • 允许一句话在同一字幕事件中自然换行,不因换行拆成多个事件。
  • 长句只在自然分句点跨屏,不拆开专有名词、数字与单位、主谓结构或因果结构。
  • 删除跨屏残字、重复词和没有语义作用的识别尾巴。
  • 核对人名、公司名、品牌名、模型名、数字、日期、价格和专业术语。
  • 字幕位于安全区,避免遮住人物眼睛、嘴部、产品和关键画面。
  • 字号根据画布和每屏字数自适应。先保证手机端可读,再控制遮挡。
  • 默认每种语言不超过 2 行。内容过长时优先优化断句,不盲目缩小字号。

启用双语字幕时:

  • 为每屏原文单独改写译文,保持主语、动作、因果、转折、语气、数字和专有名词一致。
  • 两种语言使用相同起止时间。
  • 使用自然表达,不直接采用未经核验的机器翻译。
  • 逐屏检查原文和译文的语义对应关系。

4. 内容结构提示

根据内容决定是否增加章节进度和重点总结。结构简单或视频很短时可以省略。

使用章节进度时:

  • 按真实语义拆分模块,不固定模块数量。
  • 模块标题简短、互不重复,并从观众视角概括该段价值。
  • 在顶部显示完整内容地图,让观众知道前后结构。
  • 使用从左向右连续推进的进度线。当前模块高亮,其他模块降低明度。
  • 保持元素紧凑,避免遮挡人物或字幕。

使用重点总结时:

  • 只在真正影响理解和记忆的位置出现。
  • 直接显示总结文案,不添加“重点总结”“核心观点”等固定标签。
  • 每次显示约 2 至 4 秒,控制在 1 至 3 行。
  • 文案必须忠于原口播,不添加未经讲述的结论。

5. 封面

用户需要封面时,根据视频内容和品牌生成方案:

  1. 从视频选择眼睛睁开、表情自然、脸部清楚、构图稳定的真实帧,或使用用户提供的形象素材。
  2. 根据视频核心观点生成 2 至 5 个标题候选。标题必须来自视频内容,不夸大事实。
  3. 根据受众和品牌选择字体、颜色、背景和排版。优先保证手机缩略图可读。
  4. 核对标题的每个字、行数、对齐和安全区。
  5. 没有品牌规范时,默认使用真实人物画面、克制背景和大号高对比标题。

封面嵌入视频时只占第 1 帧。正文从第 2 帧直接开始,不添加停留、淡入淡出或转场。若平台支持独立上传封面,优先保留独立封面文件。

6. 音频与画面

  • 人声优先,做必要的响度平衡、降噪和去爆音。
  • 用户未要求时不默认添加背景音乐。
  • 添加音乐时选择与内容和品牌一致的无人声音乐,并对口播做 ducking。
  • 不默认使用高频缩放、卡点、强音效、密集花字和频繁转场。
  • 必要的画面重构、B-roll 和动效必须服务于理解。
  • 涉及驾驶、医疗、金融或其他高风险场景时,根据画面证据、发布地区和用户要求添加合适声明,禁止写死固定文案。

7. 验证与导出

交付前完整读取 references/quality-checks.md,执行其中的全片检查和截图审核。

Read the full file on GitHub · 126 lines

Files

What ships with it

1 file beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 10d ago First seen · 126 lines · 99 tokens per session scan A 9fc2a8ccfd86

Subscribe to this mod's changes

edit-talking-head-videos is a skill published in the GitHub repository handsomeng/Hskill-chatcut (6 stars, last pushed 1mo ago), licensed MIT. It adds 99 tokens to every session and 1,691 once invoked, about $0.0005 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.

Related

Other skills, from other repositories

voice

Text-to-Speech (TTS), voice cloning, voiceover, narration placement/sync, and custom sound effects (SFX) generator. Use when the user wants generated speech from text, wants to clone a consented voice from uploaded reference audio, wants to add/replace/align narration or voiceover for an existing video/timeline, wants…

nigedazhima/dsh-plugin-chatcut · 109 tokens

create-motion-graphics

Use whenever an ACP or local CLI agent in ChatCut Desktop needs to add, create, hand-author, patch, or place Motion Graphic JSX assets in a project. Covers direct inline JSX authoring, visual language, editable properties, asset binding, timeline placement, and local verification. Not for the built-in ChatCut Agent.

nigedazhima/dsh-plugin-chatcut · 70 tokens

talking-head-guide

A guide for editing videos where spoken delivery or conversation drives the structure, such as interviews, podcasts, lectures, tutorials, and courses.

nigedazhima/dsh-plugin-chatcut · 178 tokens

video-gen

A video-generation guide for creating or changing clips from text, images, frames, or reference material.

nigedazhima/dsh-plugin-chatcut · 90 tokens

multicam-sync

Synchronize footage from a multi-camera / multi-recorder shoot — several cameras plus separate audio recorders covering one session, imported as loose clips — and optionally turn all or part of it into speaker-follow footage for a larger edit. Use when a user drops in multiple clips from the same recording and wants…

nigedazhima/dsh-plugin-chatcut · 114 tokens

asset-import

Use in Claude Code when acquiring or importing media into a ChatCut project, including local or attached videos, readable user-provided paths, files already uploaded in the editor, public media URLs, Browser-pane injection, transcription readiness, and upload fallback decisions.

nigedazhima/dsh-plugin-chatcut · 53 tokens