Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add chubbyguan/chubbyskills --skill youtube-transcribegit clone --depth 1 https://github.com/chubbyguan/chubbyskillsWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/chubbyguan/chubbyskills/youtube-transcribe)<a href="https://agentmods.dev/skills/chubbyguan/chubbyskills/youtube-transcribe"><img src="https://agentmods.dev/badge/skills/chubbyguan/chubbyskills/youtube-transcribe/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/chubbyguan/chubbyskills/youtube-transcribe"><img src="https://agentmods.dev/badge/skills/chubbyguan/chubbyskills/youtube-transcribe.svg" alt="Reviewed on agentmods" width="80" height="20"></a>- NVIDIA SkillSpector warn
SkillSpector: 1 finding, up to medium
These are SkillSpector’s own severities. On a checked sample its high-severity flags on skills were ~96% false positives — a documented command, a public API, a “never do X” rule — so we show them as a caution to read, not a verdict. Why →
- medium Privilege Escalation · line 40 Commands invoke sudo or root privileges. Verify this elevated access is necessary and justified.Fix: Avoid sudo/root unless strictly required. Prefer least-privilege patterns. If elevation is needed, document the justification and scope.
What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00041 | $0.01143 |
| Opus 5 | $0.00020 | $0.00571 |
| Sonnet 5 | $0.00008 | $0.00229 |
| Haiku 4.5 | $0.00004 | $0.00114 |
Grade B, and why
youtube-transcribe scanned grade B with 1 finding against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 11d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Asks for rootmediumPrivilege escalation
A mod that escalates privileges can change anything on the machine, not only the project.
# Ubuntu: sudo apt install ffmpeg && pip install yt-dlp How it starts
The opening of the file, as written. The whole thing — 145 lines — stays where its author put it; the contents beside it link to each section on GitHub.
YouTube 视频转录 + 翻译 Skill
字幕优先:先抓官方/自动字幕(秒级、免 GPU),抓不到再下载音频用 SenseVoice-Small 转录;英文内容自动翻译成中文,输出中英对照 Markdown。加 --no-subtitle 可强制走音频转录。
为什么需要这个
YouTube 有大量优质英文内容,但很多人没时间看完或语言不通。这个 skill:
- 转录视频为文字(支持中英文)
- 英文内容自动翻译成高质量中文
- 输出中英对照,方便学习
环境要求
# Python 3.9+
python -m venv .venv
source .venv/bin/activate
# 依赖
pip install funasr modelscope torch torchaudio
# 系统依赖
# macOS: brew install ffmpeg yt-dlp
# Ubuntu: sudo apt install ffmpeg && pip install yt-dlp
# 翻译功能(可选,英文视频需要)
export DEEPSEEK_API_KEY="your-api-key"
使用方法
python scripts/transcribe.py "https://www.youtube.com/watch?v=xxxxx"
python scripts/transcribe.py "https://www.youtube.com/watch?v=xxxxx" --output ./output
python scripts/transcribe.py "https://www.youtube.com/watch?v=xxxxx" --no-translate # 不翻译
python scripts/transcribe.py "https://www.youtube.com/watch?v=xxxxx" --no-subtitle # 强制音频转录
python scripts/batch_transcribe.py ../../examples/youtube-urls.txt -o ./output --no-translate
流程
Step 1: yt-dlp 下载音频
yt-dlp --extract-audio --audio-format mp3 --audio-quality 128K \
"https://www.youtube.com/watch?v=xxxxx"
Step 2: SenseVoice-Small 转录
from funasr import AutoModel
model = AutoModel(model="iic/SenseVoiceSmall", ...)
result = model.generate(input=audio_path, language="auto", use_itn=True)
Step 3: 语言检测 + 翻译
- 中文内容 → 直接输出
- 英文内容 → LLM 翻译成中文,输出中英对照
Step 4: 生成 Markdown
---
title: 视频标题
type: note
tags: [YouTube]
language: en
translated: true
---
# 视频标题
> 转录引擎:SenseVoice-Small | 翻译:DeepSeek
## 中文翻译
翻译内容...
---
## English Original
原文内容...
批量模式
scripts/batch_transcribe.py 支持 URL 列表批处理,默认单条失败后继续处理下一条;需要严格模式时加 --stop-on-error。
翻译配置
默认使用 DeepSeek API 翻译。需要设置环境变量:
export DEEPSEEK_API_KEY="your-api-key"
支持的 LLM:
- DeepSeek(推荐,性价比高)
- OpenAI
- 任何兼容 OpenAI API 格式的服务
性能数据
| 视频时长 | 下载 | 转录 | 翻译 | 总耗时 |
|---|---|---|---|---|
| 5 min | ~10s | ~5s | ~10s | ~30s |
| 10 min | ~15s | ~10s | ~20s | ~50s |
| 30 min | ~30s | ~30s | ~60s | ~2min |
What ships with it
3 files beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 11d ago First seen · 145 lines · 41 tokens per session scan B f25baf21e3d0
youtube-transcribe is a skill published in the GitHub repository chubbyguan/chubbyskills (667 stars, last pushed 23d ago), licensed MIT. It adds 41 tokens to every session and 1,143 once invoked, about $0.0002 per session on Opus 5. A static security scan graded it B with 1 finding (asks for root). No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other skills, from other repositories
youtube-transcribe
Transcribe YouTube videos and playlists. Extract audio to text with visual context, generate summaries and detailed notes.
fanyi
A command-line tool for translating between Chinese and English and looking up words or short phrases. It combines dictionary results with a language-model translation, including meanings, pronunciation, related words and examples.
multimodal-llm
Vision, audio, video generation, and multimodal LLM integration patterns. Use when processing images, transcribing audio, generating speech, generating AI video (Kling v3, Sora 2, Veo 3.1 std/lite/fast, Runway Gen-4.5 via gen4turbo), or building multimodal AI pipelines.
en-zh-translation-polish
An English-to-Chinese translation and editing process focused on natural Chinese rather than word-for-word wording.
citedy-content-ingestion
Turn any URL into structured content — YouTube videos (via Gemini Video API), web articles, PDFs, and audio files. Extract transcripts, summaries, and metadata for use in any LLM pipeline. Powered by Citedy.
translation-accuracy-reviewer
Compares a source text against its translation, flags potential inaccuracies and mistranslations, and explains each issue so an editor can make an informed correction.