Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add 1596941391qq/anything-to-md --skill skillgit clone --depth 1 https://github.com/1596941391qq/anything-to-mdWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/1596941391qq/anything-to-md/skill)<a href="https://agentmods.dev/skills/1596941391qq/anything-to-md/skill"><img src="https://agentmods.dev/badge/skills/1596941391qq/anything-to-md/skill/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/1596941391qq/anything-to-md/skill"><img src="https://agentmods.dev/badge/skills/1596941391qq/anything-to-md/skill.svg" alt="Reviewed on agentmods" width="80" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00265 | $0.02049 |
| Opus 5 | $0.00133 | $0.01025 |
| Sonnet 5 | $0.00053 | $0.00410 |
| Haiku 4.5 | $0.00026 | $0.00205 |
Grade A, and why
anything-to-md scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 10d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
This is a copy
100% identical to anything-to-md — 600 lines differ, which has more behind it and is treated as the original. This page carries a canonical link to it rather than competing with it.
How it starts
The opening of the file, as written. The whole thing — 301 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Anything-to-MD Skill
Convert any file to clean, LLM-ready Markdown. Supports 50+ file formats including documents, images (OCR), audio/video, and YouTube URLs.
Quick Start
# Convert single file
anything-to-md file document.pdf -o ./output
# Convert entire directory
anything-to-md dir ./my-docs ./my-mds --report
# Extract YouTube transcript
anything-to-md youtube "https://youtube.com/watch?v=xxx"
Capabilities
File Types Supported
| Category | Formats |
|---|---|
| Documents | PDF, DOCX, XLSX, PPTX, ODT, ODS, ODP |
| Web | HTML, HTM, XHTML |
| eBooks | EPUB, MOBI |
| Data | CSV, TSV, JSON, XML |
| Images | PNG, JPG, GIF, BMP, TIFF, WEBP (with OCR) |
| Audio | MP3, WAV, M4A, FLAC, OGG, AAC |
| Video | MP4, MKV, AVI, MOV, WEBM |
| URLs | YouTube, Wikipedia, RSS feeds |
Video Intelligent Routing (NEW)
Videos are processed through a smart 4-phase pipeline:
PROBE → DECIDE → EXTRACT → FUSE
Phase 1: PROBE (< 5 seconds)
- ffprobe: Detect embedded subtitle tracks, audio streams
- Sidecar detection: Check for .srt/.vtt files
- Sample frame OCR: Extract 5 frames, run quick OCR to detect on-screen text
Phase 2: DECIDE - Choose optimal strategy
| Video Type | Strategy | Description |
|---|---|---|
| Has embedded subtitles | embedded_subtitle |
Extract with ffmpeg, parse SRT |
| Has sidecar .srt/.vtt | sidecar_subtitle |
Parse external subtitle file |
| Audio + on-screen text | hybrid |
faster-whisper + frame OCR |
| Pure audio (podcast) | audio_transcribe |
faster-whisper transcription |
| PPT recording / tutorial | visual_ocr |
Scene detection + keyframe OCR |
| Unknown / mixed | full_pipeline |
Run all extraction methods |
Phase 3: EXTRACT
- Subtitles: ffmpeg extraction → SRT parsing
- Audio: faster-whisper (4x faster than OpenAI Whisper)
- Frames: PySceneDetect for scene changes + perceptual hash deduplication
- OCR: RapidOCR (PaddleOCR models + ONNX runtime)
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 10d ago First seen · 301 lines · 265 tokens per session scan A a3476135a8f4
anything-to-md is a skill published in the GitHub repository 1596941391qq/anything-to-md (18 stars, last pushed 4mo ago), licensed MIT. It adds 265 tokens to every session and 2,049 once invoked, about $0.0013 per session on Opus 5. A static security scan graded it A with 0 findings. It is 100% identical to anything-to-md, differing in 600 lines, and is treated as a copy.
Other skills, from other repositories
youtube-fetcher
Retrieve YouTube transcripts and subtitles, summarize or analyze what was said, or save an Obsidian-ready Markdown knowledge-base note with captions, creator metadata, chapters, language, and source provenance. Use for a YouTube URL or video ID when the request needs spoken content or an archival note. A bare YouTube…
design-drawing-svg-md
Use when converting vector PDF engineering drawings (steel-structure shop drawings, A0 drawings, linework-only PDFs without a text layer) into LLM-consumable dual-carrier records — semantic Markdown reading thread plus layered SVG geometry archive linked by shared IDs and crosswalk.
docutranslate
Use when translating documents locally via LLM — PDF, Word, Excel, Markdown, SRT subtitles with format preservation. DocuTranslate: LLM-powered multi-format local file translation tool with MCP server support.
mineru
An AI-Native skill for parsing PDF / Office / image files into clean Markdown with MinerU — a fast, zero-config document parser for AI agents. Works with NO token via the lightweight Agent API and auto-upgrades to the Standard API (token) for large files, batches, and DOCX/HTML/LaTeX export. Use when: (1) Converting…
bilibili-to-doc
A workflow for turning a Bilibili video—a video-hosting site popular in China—into a structured Markdown document using its Chinese subtitles.
policy-drift-audit
Run the policy/process drift audit. Surfaces stale, superseded, duplicate, non-canonical, and conflicting policy documents using DocGraph's built-in drift engine. Presents findings and optional remediation suggestions without writing governance decisions.