make-ai-video

make-ai-video is a skill for Codex from wanghui2323/ai-video-maker. It costs 127 tokens per session (2,067 once invoked), scanned A, original, MIT.

An AI video production assistant that helps turn ideas, articles, outlines, spoken material, source files, or audio into reviewable video packages.

In plain words
What is it for?
Use it to prepare video briefs, scripts, storyboards, voice work, subtitles, production files, and rendering or status checks.
Why use it?
It separates content, voice, visuals, subtitles, rendering, and publishing checks, making it clear what is ready and what still needs human approval.

Skill for Codex

Written for Codex: agents/openai.yaml present.

Good fit Use it to prepare video briefs, scripts, storyboards, voice work, subtitles, production files, and rendering or status checks.

Compare 6 skills from other repositories ↓
Install with agentmods
npx agentmods add skills/wanghui2323/ai-video-maker/make-ai-video
Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

Any agent
npx skills add wanghui2323/ai-video-maker --skill make-ai-video
Clone the repo
git clone --depth 1 https://github.com/wanghui2323/ai-video-maker

Made for: Codex.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for make-ai-video

README.md
[![agentmods](https://agentmods.dev/badge/skills/wanghui2323/ai-video-maker/make-ai-video/github.svg)](https://agentmods.dev/skills/wanghui2323/ai-video-maker/make-ai-video)
Your own site
<a href="https://agentmods.dev/skills/wanghui2323/ai-video-maker/make-ai-video"><img src="https://agentmods.dev/badge/skills/wanghui2323/ai-video-maker/make-ai-video/github.svg" alt="Measured on agentmods" height="20"></a>

Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.

agentmods 80×15 button for make-ai-video

Your own site · 80×15
<a href="https://agentmods.dev/skills/wanghui2323/ai-video-maker/make-ai-video"><img src="https://agentmods.dev/badge/skills/wanghui2323/ai-video-maker/make-ai-video.svg" alt="Reviewed on agentmods" width="80" height="20"></a>
Per session 127 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 2,067 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. A grade says what 26 rules found in the file — not that it is safe. Third-party audits
  • NVIDIA SkillSpector pass 7 Sept 2026
How audits are shown
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5.1 $0.00127 $0.02067
Opus 5 $0.00063 $0.01033
Sonnet 5 $0.00025 $0.00413
Haiku 4.5 $0.00013 $0.00207

Measured 10d ago against content hash c1aa7edba9e6, method: parsed. Prices are Anthropic first-party input rates as of 2026-09-10, from the pricing page.

Security

Grade A, and why

make-ai-video scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 10d ago.

The scan reads SKILL.md. This mod also ships 8 executable files (scripts/create-package.mjs, scripts/doctor.mjs, scripts/providers/qwen3-tts-mlx.py, …), listed below but not scanned — reading those needs a real analyzer, not pattern matching.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

make-ai-video/SKILL.md · 143 lines

How it starts

The opening of the file, as written. The whole thing — 143 lines — stays where its author put it; the contents beside it link to each section on GitHub.

AI 视频制作助手

从用户已有的想法或素材开始,先判断输入和当前阶段,再协助完成内容、声音、画面与审核。不要把文章当作默认入口,也不要把首次声音建档、本次正式配音和整片发布合并成一次生成。

启动即执行

触发 Skill 后,立即进入工作流并执行当前可安全完成的步骤,不要只复述流程、列菜单或让用户自己拼命令。

  1. 先运行 node scripts/doctor.mjs --json 检查基础能力。定位或创建本次生产包,检查已有文件、设备、运行时和当前状态;没有生产包时运行 node scripts/create-package.mjs --dir <目录> --input-mode <类型> --summary <摘要>,不要临时拼接一套不可复用目录。
  2. 把用户现有输入写入 video-brief.json,明确缺失项,并继续执行到下一个真实人工门禁。
  3. 可逆的本地检查、建目录、生成合同对象和校验应主动执行。安装依赖、下载大模型或需要额外系统权限时,说明体积、目录和用途后发起所需批准;获准后继续,不要退回成教程。
  4. 只有授权/权利不明、核心事实待核验、口播确认、声音所有者选声、整片审核和发布授权等人工门禁才暂停。
  5. 每次暂停或交付都报告:当前阶段 / 已完成动作 / 产物路径 / 校验结果 / 需要用户做的一个决定 / 批准后下一步

用户说“帮我做成视频”代表持续推进到下一个人工门禁;用户明确要求本人克隆声音且没有可用 Profile 时,按下文时钟 A 先执行本地声音建档,然后自动回到时钟 B。不要把“给一句开始话术”当作完成。

读取所需合同

  1. 先读 references/input-routing.md,把用户输入整理为 VideoBrief
  2. 设计内容时读 references/content-contract.md
  3. 选择画幅和分镜时读 references/visual-routing.md
  4. 使用声音、渲染或报告状态前读 references/production-gates.md
  5. 只有用户要求克隆或复用克隆声音时,才读 references/voice-cloning.md
  6. 进入画面与渲染阶段时读 references/rendering-adapter.md。如果当前项目没有渲染适配器,继续交付内容、声音、字幕和 video-unit.json,但不得声称已经可以生成正式 MP4。
  7. assets/example-package/ 只用于理解完整对象和测试,不作为新项目直接改写;需要首次本地声纹建档时,再复制 assets/voice-clone-starter/

使用两条时钟

时钟 A:低频声音能力建立

仅在用户明确选择克隆声音且没有可用 VoiceProfile 时执行:

授权与私有范围
→ 参考录音和准确逐字稿
→ 三个校准候选
→ 机器 QA
→ 声音所有者选择
→ production-pilot VoiceProfile

这条链通常只在首次建档、参考或模型漂移、授权范围变化时重跑。它先于完整视频生产,但不生成本次正式口播。

时钟 B:每条视频自己的生产

每个视频项目都执行:

VideoBrief
→ ContentDecision
→ VideoContentPlan
→ 口播人工确认
→ 本次 VoiceRun 或其他配音
→ 正式时序
→ 视觉预检
→ 渲染与技术检查
→ 整片人工审核
→ 发布候选

如果已有可用声音档案,在项目开始时做一次 preflight,口播确认后再生成本次三个候选。不要在口播未冻结时提前生成正式音频。

执行工作流

1. 建立 video-brief.json

识别 ideaarticleoutlinescriptsource-packaudio 输入。记录受众、期望变化、来源、权利、证据成熟度、时长/画幅偏好和声音意图。

用户只有想法时,协助展开方向,但把假设与待核验事实写入 verificationNeeds;不要编造个人经历、数据或案例。

2. 检查声音依赖

  • nonehuman 或普通 synthetic:按相应合同继续。
  • cloned 且已有 Profile:立即 preflight;漂移则阻断。
  • cloned 且没有 Profile:读取 references/voice-cloning.md,先做设备与本地模型预检;缺少模型时按该参考执行下载与锁定,随后走时钟 A,再自动进入完整生产。

Read the full file on GitHub · 143 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 10d ago First seen · 143 lines · 127 tokens per session scan A c1aa7edba9e6

Subscribe to this mod's changes

make-ai-video is a skill published in the GitHub repository wanghui2323/ai-video-maker (103 stars, last pushed 25d ago), licensed MIT. It adds 127 tokens to every session and 2,067 once invoked, about $0.0006 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.

Related

Other skills, from other repositories

webgl-holographic-foil

A self-contained WebGL2 hero: thin-film interference over a crushed-foil surface whose palette shifts with the viewing angle; move the cursor to tilt the film.

nexu-io/open-design · 41 tokens

general-video

Author or edit a custom HyperFrames composition when no specialized workflow fits, or when BRIEF.md sets flow: companion. Use for longer or multi-scene pieces, brand and sizzle reels, montages, static loops, static title cards, footage remixes, and freeform builds. Use motion-graphics instead for a short unnarrated…

heygen-com/hyperframes · 92 tokens

html-ppt-hermes-cyber-terminal

OpenDesign + BYOK: choosing and wiring your own model, hands-on — cost, quality, and the routing decision. Built as a decision-grade AI literacy deck for engineers, IT, applied-AI teams.

nexu-io/open-design · 53 tokens

html-ppt-taste-brutalist

16:9 HTML deck in tactical-telemetry / CRT-terminal taste. Deactivated-CRT charcoal slides, white-phosphor monospace, hazard-red accent, scanline overlay, ASCII syntax, density over decoration. Distilled from Leonxlnx/taste-skill brutalist-skill (Tactical Telemetry mode).

nexu-io/open-design · 78 tokens

diagnostic-stem-delivery

Audio production with diagnostic analysis, timecode parsing from documents, and verified export workflow.

HKUDS/OpenSpace · 23 tokens

chengfeng-check-updates

An environment manager for a video-editing system. It checks whether its skills and runtime—the software needed to run them—are installed and compatible.

Agentchengfeng/chengfeng-videocut-skills · 120 tokens