Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add Knuckles-Team/media-downloader --skill media-watchgit clone --depth 1 https://github.com/Knuckles-Team/media-downloaderWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/knuckles-team/media-downloader/media-watch)<a href="https://agentmods.dev/skills/knuckles-team/media-downloader/media-watch"><img src="https://agentmods.dev/badge/skills/knuckles-team/media-downloader/media-watch/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/knuckles-team/media-downloader/media-watch"><img src="https://agentmods.dev/badge/skills/knuckles-team/media-downloader/media-watch.svg" alt="Reviewed on agentmods" width="80" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00135 | $0.02411 |
| Opus 5.5 | $0.00054 | $0.00964 |
| Sonnet 5.5 | $0.00027 | $0.00482 |
| Haiku 4.5 | $0.00014 | $0.00241 |
Grade A, and why
media-watch scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured yesterday.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 203 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Media Watch
Turn a video into something an agent can work from: the transcript plus what was actually on screen.
When to use
- Learn a method from a tutorial, demo, conference talk, or screen recording.
- Capture what a video shows that the narration never says out loud — a command typed on screen, a menu path, a config value, an error message.
- Turn a video into a reusable skill, checklist, or runbook.
When NOT to use
- Just archiving a video →
media-download. - Only the audio track as MP3 →
media-audio. - A media file already on disk →
audio-transcriber-transcription. - Querying media already in the KG → query the
:MediaAssetnodes directly.
Prerequisites & environment
Connect via the mcp-client skill against the media-downloader MCP server.
ffmpeg and ffprobe must be on PATH for key frames; without them the run
still returns captions and reports frames.status: "unavailable".
| Variable | Required | Notes |
|---|---|---|
MEDIA_DOWNLOADER_OUTPUT_ROOT |
optional | Root the bundle is written beneath |
GRAPH_SERVICE_ENDPOINTS |
optional | Engine endpoint for native KG ingestion |
MCP_TOOL_MODE (condensed|verbose|both) selects the tool surface.
Tools & actions
| Tool | Purpose |
|---|---|
watch_media |
Download a URL with captions + key frames; returns the bundle manifest |
list_watch_skills |
What video-built skills exist and which videos each already has |
build_watch_skill |
Write a new skill from a bundle, or extend an existing one |
transcribe_audio |
(audio-transcriber package) fallback when captions are missing |
Key parameters
video_url— the media URL (required).download_directory— where the bundle is written (default.).max_frames— cap on extracted key frames (default 24).scene_threshold— ffmpeg scene sensitivity, 0-1 (default 0.3); lower finds more.- Frames also carry a
sharpnessscore; below ~3 usually means motion blur. subtitle_languages— comma-separated caption codes (defaulten,en-orig).
What ships with it
1 file beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- yesterday First seen · 203 lines · 135 tokens per session scan A d2a1511af368
media-watch is a skill published in the GitHub repository Knuckles-Team/media-downloader (4 stars, last pushed 2d ago), licensed MIT. It adds 135 tokens to every session and 2,411 once invoked, about $0.0005 per session on Opus 5.5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-09-29.
Other skills, from other repositories
video-transcriber
Fetch video transcripts (subtitles) and metadata from online video sources and output them as Markdown, JSON, or plain text. Currently supported: YouTube. Prefers manually created subtitles, falls back to auto-generated ones.
youtube-learn
Phân tích video (YouTube, LinkedIn, Facebook, X, TikTok) theo hướng Belief Archaeology, kết hợp Multimodal phân tích hình ảnh và thế giới quan của người nói.
youtube-summarizer
Summarize YouTube videos by extracting transcripts and generating structured summaries with key points, timestamps, and topic segmentation.
youtube-transcribe
Transcribe YouTube videos and playlists. Extract audio to text with visual context, generate summaries and detailed notes.
youtube
Busqueda en YouTube, extraccion de transcripciones y analisis/interpretacion de contenido de video usando herramientas locales (yt-dlp, youtube-transcript-api) y modelos AI (Gemini gratis / Ollama local). Use when: (1) Buscar videos en YouTube por tema o query, (2) Extraer transcripcion o subtitulos de un video, (3)…
mediago
A tool for downloading videos from direct links, m3u8 or HLS streams, and Bilibili through a running MediaGo service. HLS is a common format for streamed video.