Automatically fetch YouTube video transcripts, generate structured summaries, and send full transcripts to messaging platforms. Detects YouTube URLs and provides metadata, key insights, and downloadable transcripts.
Analyze, summarize, and extract insights from DeLive transcription sessions. Use when: user mentions DeLive, transcription, meeting transcripts, live captions, audio transcription, AI correction, corrected transcript, or transcript analysis; user wants to search, retrieve, summarize, correct, or process recorded…
Use when turning a YouTube URL or local presentation, explainer, interview, podcast, or product-demo video into complete screenshot-led notes, faithful light-polished text, HTML, PDF, or a shareable ZIP.
Watch a video from YouTube, Instagram, X/Twitter, Vimeo, TikTok or any of 1800 yt-dlp sites (or a local path). Downloads with yt-dlp, extracts auto-scaled frames with ffmpeg, pulls the transcript from captions (or local mlx-whisper fallback, no API key), and hands the result to Claude so it can answer questions about…
Open or rebuild the video-lens gallery index — your personal library of saved video summaries. Use this whenever the user wants to browse, open, or search saved video reports: "show my gallery", "open video library", "browse saved videos", "build gallery", "what videos have I saved", "show my video notes", "my video…
Fetch a YouTube transcript and generate an executive summary, key points, and timestamped topic list as a polished HTML report. Activate on YouTube URLs or requests like "summarize this video", "what's this about", "give me the highlights", "TL;DR this", "digest this video", "watch this for me", "I watched this and…
Use when writing an article, newsletter, blog post or long social post that must genuinely be the user's own writing. Triggers on "voiceprint", "write this in my voice", "turn my recording into an article", "cut my transcript down", "make this actually mine", "write from my video". Turns the user's own spoken words…
A workflow that turns a public Xiaoyuzhou FM podcast episode into structured Markdown notes and a saved transcript. Xiaoyuzhou FM is a Chinese podcast platform.
A tool for reading web pages, videos, podcasts, links, and local audio recordings, then producing structured Chinese explanations and answering questions about the source material.
A video analysis tool that indexes speech and visible screen text in recordings from YouTube, Loom, Zoom, Kinescope, or local MP4 files. It links claims about the screen to exact frames and time codes, while reporting sections that were not visually covered.
A video-analysis skill that indexes speech and screen content in videos, including online recordings and local files, and links findings to timestamps and frames.
Transcribes a local video or audio file into a traceable Chinese or mixed Chinese-English package with TXT, optional timestamps and SRT/VTT, glossary handling, and quality evidence. Use for 视频转文字, 音频转写, 逐字稿, 字幕, word timestamps, 本地模型选择, 转写环境配置, 专业术语保留, StepFun/阶跃/StepAudio cloud ASR, or Volcengine/豆包/Seed ASR. Not for…
Download videos from 1000+ websites (YouTube, Bilibili, Twitter/X, TikTok, Vimeo, Instagram, Twitch, etc.) using yt-dlp. Use this skill whenever a user shares a video URL, asks to save or download a video, wants to extract audio from an online video, needs a specific quality like 1080p or 4K, or mentions downloading a…
Generate professional voiceover narration for a video with audio-video sync using Azure TTS by default, or Gemini 3.1 Flash TTS when configured. Use this skill whenever the user wants to add narration, voiceover, commentary, or voice dubbing to any video file — even if they just say "add audio to this video" or "make…
Extract transcript or subtitles from a local video file. Use this skill whenever the user asks to transcribe a video, extract speech-to-text, get subtitles, or wants a text version of what's said in a video. Also trigger on "提取字幕", "视频转文字", "语音转文字", "transcribe", "extract audio text", or when the user references…
Use when a task requires reading, transcribing, or extracting content from audio or voice — WhatsApp PTT/voice notes (.opus/.ogg), meeting recordings, .mp3/.m4a/.wav/.mp4 files, or audio attachments inside a WhatsApp chat-export zip.
Local speech-to-text using faster-whisper. 4-6x faster than OpenAI Whisper with identical accuracy; GPU acceleration enables 20x realtime transcription. SRT/VTT/TTML/CSV subtitles, speaker diarization, URL/YouTube input, batch processing with ETA, transcript search, chapter detection, per-file language map.
Build, run, and drive the WhatsApp group collector — the Baileys collector daemon, the Next.js panel, the MLX transcriber service, and the MCP server. Use to run/start/build/screenshot the project or to call its MCP tools (listargrupos, lermensagens, buscar, resumododia, verimagem, vervideo, lerdocumento, transcrever…
MOLDE (ainda não preenchido) do perfil de escrita do operador deste coletor, usado ao redigir qualquer mensagem que vá pro WhatsApp. Preencha os campos {{...}} com os números do SEU histórico rodando extrair-voz.py — veja o README desta pasta. Enquanto estiver com placeholders, esta skill não descreve ninguém: escreva…
Use when working with audio or video content, or when a user pastes a URL and asks what was said. Provides workflows for Augent MCP tools — transcription, search, notes, highlights, speaker ID, visual context, and more. Activate when the user mentions audio, video, podcasts, transcription, or URLs to media content.…