Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add commands/coding-ax/docvideoer/docvideogit clone --depth 1 https://github.com/coding-ax/docvideoerWhat it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00017 | $0.00388 |
| Opus 5 | $0.00009 | $0.00194 |
| Sonnet 5 | $0.00003 | $0.00078 |
| Haiku 4.5 | $0.00002 | $0.00039 |
Grade A, and why
docvideo scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured yesterday.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
What it actually says
/docvideo — 文档转视频
使用方法
/docvideo <文档URL或文本内容>
参数
<url>— 文档的 URL 地址(支持任何公开网页)<text>— 直接粘贴的文档文本内容
示例
/docvideo https://example.com/article
/docvideo 这是一段关于 AI 技术的介绍文章……
执行流程
- 获取内容:如果提供 URL,抓取网页正文;如果提供文本,直接使用
- 分析分镜:提取要点,生成分镜脚本(展示给用户确认)
- 语音合成:使用
scripts/tts.mjs为每段文案生成语音 - 视频制作:搭建 Remotion 项目,生成场景组件
- 渲染输出:渲染 MP4 视频文件
- 结果展示:显示视频路径、描述文案、渲染统计
用户需要确认的事项
- 分镜脚本(可以修改场景顺序、文案、时长等)
- 视频风格偏好(音色、是否需要背景音乐等)
首次使用 TTS
默认使用 Edge TTS,推荐先安装:
pip3 install edge-tts
也可以复制 .env.example 为 .env,按需调整 TTS_BACKEND、TTS_VOICE、TTS_RATE。如果 Edge TTS 不可用,macOS 可回退到系统 say。
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- yesterday First seen · 50 lines · 17 tokens per session scan A 3ae0cb5cf051
docvideo is a command published in the GitHub repository coding-ax/docvideoer (5 stars, last pushed 3mo ago), licensed MIT. It adds 17 tokens to every session and 388 once invoked, about $0.0001 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other commands, from other repositories
stt
Transcribe a local audio file or remote audio URL into text.
audition-voices
Generate voice audition samples for a character using Venice TTS.
status
Show 3d-design team status and recent activity.
music-suno-prompt
Grounded Suno prompt synthesis from local knowledge corpus + persona canon + label canon. No vibes-prompting.
develop-image-prompt.eval
Generates a detailed image generation prompt from a document or content description. Good output: a prompt that is specific, visual, non-abstract, includes style/composition/lighting guidance, and is calibrated to the specified dimensions and style options.
frontend-3d
You are an expert in 3D web development using Three.js, React Three Fiber, WebGL, and WebGPU. You create immersive 3D experiences for the web.