video-script

video-script is a skill for Codex from bozhouDev/video-skills-toolkit. It costs 197 tokens per session (4,225 once invoked), scanned A, original, MIT.

A video planning guide that turns a finalized voice recording and time-coded subtitles into a shot-by-shot visual plan. It can account for recordings, screenshots, real footage, supporting footage, and animation.

In plain words
What is it for?
Use it to plan scenes, supporting footage, evidence, transitions, and motion before building a video in a production tool.
Why use it?
It keeps the visuals aligned with the exact spoken words and timing. It prevents production planning from being based on guessed timing or an unfinished script.

Skill for Codex

Written for Codex: agents/openai.yaml present.

Good fit Use it to plan scenes, supporting footage, evidence, transitions, and motion before building a video in a production tool.

Compare 6 skills from other repositories ↓
Install with agentmods
npx agentmods add skills/bozhoudev/video-skills-toolkit/video-script
Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

Any agent
npx skills add bozhouDev/video-skills-toolkit --skill video-script
Clone the repo
git clone --depth 1 https://github.com/bozhouDev/video-skills-toolkit

Made for: Codex.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for video-script

README.md
[![agentmods](https://agentmods.dev/badge/skills/bozhoudev/video-skills-toolkit/video-script.svg)](https://agentmods.dev/skills/bozhoudev/video-skills-toolkit/video-script)
Your own site
<a href="https://agentmods.dev/skills/bozhoudev/video-skills-toolkit/video-script"><img src="https://agentmods.dev/badge/skills/bozhoudev/video-skills-toolkit/video-script.svg" alt="Measured on agentmods" height="20"></a>
Per session 197 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 4,225 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. A grade says what 26 rules found in the file — not that it is safe.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5.1 $0.00197 $0.04225
Opus 5 $0.00098 $0.02112
Sonnet 5 $0.00039 $0.00845
Haiku 4.5 $0.00020 $0.00422

Measured 7d ago against content hash 5e4b7d812ce7, method: parsed. Prices are Anthropic first-party input rates as of 2026-09-07, from the pricing page.

Security

Grade A, and why

video-script scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 7d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

skills/video-script/SKILL.md · 184 lines

How it starts

The opening of the file, as written. The whole thing — 184 lines — stays where its author put it; the contents beside it link to each section on GitHub.

Video Script · Director Layer(短视频留存增强)

把已锁定的完整声音和字幕时间轴转成全片一致、可确认、可还原的导演处理稿。先回答“观众经历什么、画面世界如何演进”,再回答“每一段真实口播时间里发生什么”。

本 Skill 是逐字稿与实现管线之间的导演层。它不默认把所有内容做成 MG,也不把“运动优先”误解为“所有东西一直动”。真人、真实录屏、截图、B-roll、图形动画和留白都可以成为正确答案;关键是它们是否共同服务同一条叙事与注意力路径。

对短视频,本 Skill 可作为“爆款潜力导演”:它负责让镜头继续履行上游稿件的停滑、留存和传播意图,但不承诺流量结果,也不用画面替上游修稿。

职责门禁

  • 正式视频导演稿必须同时具备:已确认逐字稿、已锁定完整音频、主字幕文件和校准后的带时间码字幕数据。只有选题、素材、大纲、逐字稿或 8–15s 声音试听时,先完成上游,不得进入正式分镜。
  • 用户同时要求“写逐字稿+做分镜”时,先完成当前工作区可用的上游逐字稿创作/审稿流程,再完成声音试听、完整配音、字幕生成和校准,最后把同一版本交给本 Skill。不要一边写正文、一边生成声音、一边预设画面。
  • 不改写原稿的事实、观点、顺序、句子或连接词。发现口语、逻辑、事实或段间衔接阻塞制作时,引用具体相邻句,标为“上游逐字稿阻塞”,转回当前可用的逐字稿修订流程;若未安装专用 Skill,则明确告知用户需先修稿。用户明确要求按草稿原样继续时才做导演处理。
  • 不输出 HTML、JSX、Remotion 组件、HyperFrames DOM 或 ChatCut MG 实体。实现属于下游 Skill。

执行职责

  • 导演规划者完成导演诊断、素材职责、导演稿和制作规格;角色和能力要求以当前运行时与工作区指令为准,不在 Skill 中写死模型版本。
  • 本阶段不得启动任何镜头执行者,也不得提前写场景代码。只有导演稿、制作规格、素材计划和 motion contract 锁定后,才交给已选管线的单一执行者;本工具包的 HyperFrames 执行者是 hyperframes-scene-animator

单一导演来源

锁定后的 导演契约 + beat graph 是下游的叙事与视觉单一事实源。

  • 下游可以补坐标、组件、帧、选择器、缓动和渲染细节。
  • 下游不得擅自更换核心视觉命题、贯穿物、证据顺序、运动语法或转场关系。
  • 实现不可行时,记录冲突与保守替代方案,返回导演阶段确认;不要静默重新导演。

默认保持管线无关

  1. 用户未指定工具时,不选择默认实现工具,也不写 pipeline
  2. 使用真人、录屏、截图、B-roll、仿真 UI、流程关系、图形动画、空间和镜头等通用语言,不在导演契约中混入平台属性、坐标、帧号或组件名。
  3. 用户明确指定 Remotion、ChatCut、HyperFrames、剪映或其他管线时,保留同一份导演契约,只追加最少必要的应用说明。
  4. 能用真实证据直接说明的内容,不用抽象隐喻替代;抽象逻辑才优先寻找可演进的视觉隐喻。

按需读取

  • 每次正式生成导演稿时读取 references/director-treatment.md
  • 导演方向锁定后读取 references/beat-graph.md
  • 完成导演稿前读取 references/anti-ppt-gate.md;写制作规格时在实现前再检查一次。低清 proof 后的语义运动、转场中点、AI 味硬拒绝项和联系表连续性由下游执行者检查;HyperFrames 管线使用 hyperframes-scene-animator
  • 正式交付时读取 references/output-template.md
  • 目标是短视频、平台分发型视频,或用户明确要求完播、留存、爆款/传播潜力时,读取 references/short-video-retention-directing.md;只提取稿件已有意图,不调用上游选题或片头改写流程。
  • 中长视频、复杂剪辑判断或节奏诊断时读取 references/editing-methodology.md
  • 命中既有平台风险或用户要求复盘经验时读取 references/lessons-learned.md,只应用相关规则。
  • 分镜已确认且用户要求继续制作时读取 references/production-spec.md。制作规格必须使用已经锁定的最终声音和带时间码字幕,并在代码动画管线中同步产出 work/motion-contract.json
  • 用户明确要求 HyperFrames 制作时:若尚无固定舞台工程,先读取 talking-head-hyperframes 生成模板与素材交接;交接状态为 READY_FOR_EXECUTION 后,一律读取 hyperframes-scene-animator 实现所有镜头。模板 Skill 不制作场景。
  • 使用 Studio 暖白 Remotion 模板、沿用 J-space 视觉或用户锁定同类背景时,在给方向卡前用 view_image 查看 Studio 暖白蓝金网格高清参考图,并按 references/director-treatment.md 把背景当作导演约束,而不是实现阶段再补的装饰。
  • 需要实际生成图片时读取 imagegen;导演阶段默认只写资产职责和规格。

Read the full file on GitHub · 184 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 7d ago First seen · 184 lines · 197 tokens per session scan A 5e4b7d812ce7

Subscribe to this mod's changes

video-script is a skill published in the GitHub repository bozhouDev/video-skills-toolkit (137 stars, last pushed 1mo ago), licensed MIT. It adds 197 tokens to every session and 4,225 once invoked, about $0.0010 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.

Related

Other skills, from other repositories

video-hyperframes

Hyperframes / Remotion-compatible continuous frame animation with autoplay support.

nexu-io/open-design · 18 tokens

video-hyperframes

A web-based sequence of video frames designed for Hyperframes or Remotion, with each frame presenting one visual idea.

nexu-io/html-anything · 22 tokens

remocn

Build Remotion videos with remocn — copy-paste animation components and timeline-driven UI primitives from a shadcn registry. Use when composing a video or scene in a Remotion project, adding a single animation, transition, background, or UI-block sim, or reaching for a video-ready UI primitive (button, dialog…

Remocn/remocn · 88 tokens

minimax-cli

Nested swiss-knife reference for the MiniMax mmx CLI and the canonical MiniMax CLI procedure shipped with the TUI: install mmx-cli, discover the correct TUI-managed MiniMax preset/key slot without leaking secrets, match mainland vs international regions, and route image/video/music/TTS generation or one-shot shell…

Lingtai-AI/lingtai · 72 tokens

videoagent-audio-studio

Tired of juggling multiple audio APIs? This skill gives you one-command access to TTS, music generation, sound effects, and voice cloning. Use when you want to generate any audio without managing multiple API keys.

pexoai/pexo-skills · 50 tokens

multimodal-llm

Vision, audio, video generation, and multimodal LLM integration patterns. Use when processing images, transcribing audio, generating speech, generating AI video (Kling v3, Sora 2, Veo 3.1 std/lite/fast, Runway Gen-4.5 via gen4turbo), or building multimodal AI pipelines.

yonatangross/orchestkit · 82 tokens