Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add chubbyguan/chubbyskills --skill bilibili-transcribegit clone --depth 1 https://github.com/chubbyguan/chubbyskillsWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/chubbyguan/chubbyskills/bilibili-transcribe)<a href="https://agentmods.dev/skills/chubbyguan/chubbyskills/bilibili-transcribe"><img src="https://agentmods.dev/badge/skills/chubbyguan/chubbyskills/bilibili-transcribe/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/chubbyguan/chubbyskills/bilibili-transcribe"><img src="https://agentmods.dev/badge/skills/chubbyguan/chubbyskills/bilibili-transcribe.svg" alt="Reviewed on agentmods" width="80" height="20"></a>- NVIDIA SkillSpector warn
SkillSpector: 1 finding, up to medium
These are SkillSpector’s own severities. On a checked sample its high-severity flags on skills were ~96% false positives — a documented command, a public API, a “never do X” rule — so we show them as a caution to read, not a verdict. Why →
- medium Privilege Escalation · line 31 Commands invoke sudo or root privileges. Verify this elevated access is necessary and justified.Fix: Avoid sudo/root unless strictly required. Prefer least-privilege patterns. If elevation is needed, document the justification and scope.
What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00042 | $0.00921 |
| Opus 5 | $0.00021 | $0.00461 |
| Sonnet 5 | $0.00008 | $0.00184 |
| Haiku 4.5 | $0.00004 | $0.00092 |
Grade B, and why
bilibili-transcribe scanned grade B with 1 finding against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 12d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Asks for rootmediumPrivilege escalation
A mod that escalates privileges can change anything on the machine, not only the project.
# Ubuntu: sudo apt install ffmpeg && pip install yt-dlp How it starts
The opening of the file, as written. The whole thing — 106 lines — stays where its author put it; the contents beside it link to each section on GitHub.
B 站视频转录 Skill
字幕优先:先尝试抓取官方/自动字幕(秒级、免 GPU、免 funasr);抓不到再下载音频用 SenseVoice-Small 转录,存为 Markdown 文件。加 --no-subtitle 可强制走音频转录。
环境要求
# Python 3.9+
python -m venv .venv
source .venv/bin/activate
# 依赖
pip install funasr modelscope torch torchaudio
# 系统依赖
# macOS: brew install ffmpeg yt-dlp
# Ubuntu: sudo apt install ffmpeg && pip install yt-dlp
使用方法
python scripts/transcribe.py "https://www.bilibili.com/video/BV1rrQGBeEen/"
python scripts/transcribe.py "BV1rrQGBeEen"
python scripts/transcribe.py "BV1rrQGBeEen" --no-subtitle # 强制音频转录
# 批量处理:每行一个 B站 URL 或 BV 号
python scripts/batch_transcribe.py ../../examples/bilibili-urls.txt -o ./output
流程
Step 1: yt-dlp 下载音频
B 站有反爬机制,需要加 User-Agent 和 Referer:
yt-dlp --extract-audio --audio-format mp3 --audio-quality 128K \
--user-agent "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) ..." \
--referer "https://www.bilibili.com" \
"https://www.bilibili.com/video/BV1rrQGBeEen/"
Step 2: SenseVoice-Small 转录
from funasr import AutoModel
from funasr.utils.postprocess_utils import rich_transcription_postprocess
model = AutoModel(
model="iic/SenseVoiceSmall",
trust_remote_code=True,
vad_model="fsmn-vad",
vad_kwargs={"max_single_segment_time": 30000},
device="cpu",
)
result = model.generate(input=audio_path, language="zh", use_itn=True, batch_size_s=60)
Step 3: 生成 Markdown
自动创建带 frontmatter 的 Markdown 文件。
批量模式
scripts/batch_transcribe.py 会逐条调用单篇转录脚本,默认单条失败后继续处理下一条;需要严格模式时加 --stop-on-error。
性能数据
| 视频时长 | 转录耗时 |
|---|---|
| 5 min | ~10s |
| 10 min | ~20s |
| 30 min | ~60s |
已知限制
- B 站 4K/1080P60 需要大会员,普通用户可下载 720P 以下
- SenseVoice-Small 模型首次下载约 893MB
- 不支持说话人分离
- 合集视频会下载第一个分 P
参考项目
- yt-dlp/yt-dlp - 视频下载工具
- FunAudioLLM/SenseVoice - 语音识别模型
⚖️ 合规声明
仅供个人学习与研究使用。请遵守目标平台的服务条款(ToS)与 robots 规则,控制请求频率,不要用于批量抓取、商用爬取或侵犯他人权益的场景。下载内容的版权归原作者所有。
What ships with it
3 files beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 12d ago First seen · 106 lines · 42 tokens per session scan B cf3a19308a39
bilibili-transcribe is a skill published in the GitHub repository chubbyguan/chubbyskills (667 stars, last pushed 23d ago), licensed MIT. It adds 42 tokens to every session and 921 once invoked, about $0.0002 per session on Opus 5. A static security scan graded it B with 1 finding (asks for root). No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other skills, from other repositories
multimodal-llm
Vision, audio, video generation, and multimodal LLM integration patterns. Use when processing images, transcribing audio, generating speech, generating AI video (Kling v3, Sora 2, Veo 3.1 std/lite/fast, Runway Gen-4.5 via gen4turbo), or building multimodal AI pipelines.
Transcription Automation
Automate audio/video transcription, meeting notes, subtitle generation, and content processing.
media-metadata
Extract and display metadata from images, audio, and video files.
audio-transcriber
Speech-to-text transcription using Whisper API or local engine.
mk:multimodal
Process images, video, audio, PDFs with Gemini API. Generate images (Nano Banana 2), videos (Veo 3), speech (MiniMax TTS), music (MiniMax). Convert documents to Markdown. Multi-provider fallback (Gemini → MiniMax → OpenRouter). Activate when task references media files, asks to…
youtube-transcribe
Transcribe YouTube videos and playlists. Extract audio to text with visual context, generate summaries and detailed notes.