speech-to-text skills

45 tagged speech-to-text, measured the same way as everything else here.

Browse within: Apple Silicon 16text-to-speech 16asr 15bun 12coreml 12openclaw 12transcription 12audio 8whisper 8voice 7mlx 5

XimilalaXiang/DeLive

Skill Claude CodeCodex

Analyze, summarize, and extract insights from DeLive transcription sessions. Use when: user mentions DeLive, transcription, meeting transcripts, live captions, audio transcription, AI correction, corrected transcript, or transcript analysis; user wants to search, retrieve, summarize, correct, or process recorded…

264 +3 1mo ago A 98 tokens original Apache-2.0

release-cli

02

drakulavich/kesha-voice-kit

Skill Claude CodeCodex

Cuts a STABLE CLI release (vX.Y.Z-cli marker tag; not for beta or alpha markers, which this lane silently skips while burning the tag) — the 🚀 Release (CLI) lane builds the Linux packages, publishes the marker release, and dispatches npm publish with provenance. Covers version alignment across package.json and…

73 3d ago A 117 tokens original MIT

ticket-team

03

drakulavich/kesha-voice-kit

Skill Claude CodeCodex

Run one ticket from plan to merged pull request with two agents — a team lead that judges and owns the ticket, and an implementer that builds it. Use when handing a ticket to the team rather than doing it yourself.

73 3d ago A 48 tokens original MIT

kesha-voice-kit

04

drakulavich/kesha-voice-kit

Skill Claude CodeCodex

Local multilingual voice toolkit — speech-to-text (STT), text-to-speech (TTS), speaker diarization, and language detection, over a CLI or an MCP server. Runs entirely offline on Apple Silicon, Linux, and Windows. No API keys, no cloud. NVIDIA Parakeet TDT for STT across 25 European languages, Kokoro-82M + Vosk-TTS for…

73 3d ago A 111 tokens original MIT

KingJing1/podcast-transcript-txt-skill

Skill Claude CodeCodex

A workflow for finding and exporting cleaned podcast or video transcripts as TXT files. It accepts sources such as YouTube, episode webpages, podcast searches, social-media links, audio URLs, and episode titles.

37 1mo ago A 78 tokens original MIT

yw-transcribe

06

yuwen-cool/yw-transcribe

Skill Claude CodeCodex

A workflow for transcribing one local audio or video file into a traceable Chinese or mixed Chinese-English transcript, with optional timestamps and subtitle files.

22 1mo ago A 131 tokens original MIT

whisper-skill

07

Mobiss11/Whisper-Skill

Skill Claude CodeCodex

A local Whisper-based tool for converting audio or video into text, creating voice input, and adding subtitles to MP4 videos.

15 25d ago A 237 tokens original MIT

faster-whisper

08

ThePlasmak/faster-whisper

Skill Claude CodeCodex

Local speech-to-text using faster-whisper. 4-6x faster than OpenAI Whisper with identical accuracy; GPU acceleration enables 20x realtime transcription. SRT/VTT/TTML/CSV subtitles, speaker diarization, URL/YouTube input, batch processing with ETA, transcript search, chapter detection, per-file language map.

11 +1 6mo ago A 74 tokens original MIT

noisy/noisy-coding

Skill Claude CodeCodex

Part of noisy-coding

What to tell the user right after publishing a noisy-coding release — derive the minimal refresh steps (image? plugin? per-session reloads?) from what actually changed and present them as a short spoken summary plus a bulleted console checklist. Use every time a release/tag is pushed, when the user asks "what do I…

6 4d ago A 84 tokens original MIT

creating-releases

10

noisy/noisy-coding

Skill Claude CodeCodex

Part of noisy-coding

How to cut a noisy-coding release — version bump, tag, GitHub release via gh, and above all HOW TO WRITE the release notes (agent-quotable Highlights, pain-first framing, upgrade notes derived from what changed). Repo-local skill for maintainers; use whenever asked to release, publish a version, or write release notes.

6 4d ago A 73 tokens original MIT

local-dev-setup

11

noisy/noisy-coding

Skill Claude CodeCodex

Part of noisy-coding

Set up or fix the side-by-side LOCAL DEV instance of noisy-coding in this repo — dev daemon on port 7765, noisy-coding-dev MCP, project-scoped hooks. Use when asked to prepare the local development environment, when the dev daemon is down, or when a session in this repo should talk to the dev instance instead of…

6 4d ago A 77 tokens original MIT

augent

12

AugentDevs/augent

Skill Claude CodeCodex

The audio & video layer for agents. 22 local MCP tools. No cloud, no API keys.

5 3d ago A 24 tokens original MIT

augent

13

AugentDevs/augent

Skill Claude CodeCodex

Use when working with audio or video content, or when a user pastes a URL and asks what was said. Provides workflows for Augent MCP tools — transcription, search, notes, highlights, speaker ID, visual context, and more. Activate when the user mentions audio, video, podcasts, transcription, or URLs to media content.…

5 3d ago A 91 tokens original MIT

vidscribe

14

XFWang522/vidscribe

Skill Claude CodeCodex

Transcribe video to text from any platform (Bilibili, YouTube, Douyin, Twitter/X, TikTok, Vimeo, etc.). Use when the user provides a video URL and asks for transcript, subtitles, text content, or speech-to-text conversion.

4 5mo ago A 56 tokens original MIT

chicogong/ffvoice-engine

Skill Claude CodeCodex

Offline speech-to-text and speaker diarization with the ffvoice engine. Use when the user wants to transcribe an audio file, generate subtitles (SRT/VTT/JSON), identify who spoke when (speaker diarization), caption or transcribe live microphone input, or list audio input devices — all fully on-device, with no cloud…

3 3mo ago A 78 tokens original MIT

voxpip

16

EzekielOgunkunle/voxpip

Skill Claude CodeCodex

Part of voxpip

Watch any video and get a transcript — even without an API key. Extends /watch with a local speech-to-text fallback.

2 4mo ago C 30 tokens original MIT

deyo

17

casatwy/deyo-skill

Skill Claude CodeCodex

Use only when the current user explicitly asks to use Deyo to transcribe one provided URL or one exact local audio/video file path, or explicitly asks for Deyo install, status, or troubleshooting. Do not trigger from a mere Deyo mention, ambient context, an implicit attachment, directory browsing, a glob, stdin, a…

2 1mo ago B 86 tokens

vcr

18

WahabF/vcr-skill

Skill Claude CodeCodex

Process and reconcile new Just Press Record iPhone recordings into a configured Git-backed Markdown vault. Use when the user invokes $vcr or asks to process the configured Just Press Record capture source; do not use for generic transcription or unrelated voice apps.

2 3d ago A 52 tokens original MIT

jump-in

19

michael-L-i/cadence-code

Skill Claude CodeCodex

Part of cadence-code

Interrupt an active Cadence Code conversation and add fresh spoken guidance to the current Codex or Antigravity task. Use only when the user explicitly invokes $jump-in or /jump-in after stopping the current host turn.

2 5d ago A 47 tokens original MIT

start-talking

20

michael-L-i/cadence-code

Skill Claude CodeCodex

Part of cadence-code

Start and run an explicit, interactive Cadence Code conversation with Codex or Antigravity using fully local speech input and output. Use only when the user explicitly invokes $start-talking, /start-talking, or asks to start talking with Cadence Code.

2 5d ago A 57 tokens original MIT

voice-settings

21

michael-L-i/cadence-code

Skill Claude CodeCodex

Part of cadence-code

Choose Cadence Code's local speech and transcription models from the Codex or Antigravity UI. Use only when the user explicitly invokes $voice-settings, /voice-settings, or asks to open Cadence Code settings.

2 5d ago A 47 tokens original MIT

Knuckles-Team/audio-transcriber

Skill Claude CodeCodex

Speech-to-text on the audio-transcriber MCP server — run Whisper (faster-whisper, falling back to openai-whisper) over a local audio/video file or a microphone recording, and export txt/srt/vtt/json captions. Use when the agent must transcribe or translate spoken audio, generate subtitle/caption files, or pick a…

2 5d ago A 119 tokens original MIT

session-context

23

sandraschi/speech-mcp

Skill Claude CodeCodex

Part of speech-mcp

You have access to a multi-provider speech gateway with TTS, STT, wake word, RAG, and fleet voice command bus.

2 14d ago A 0 tokens original MIT

speech-expert

24

sandraschi/speech-mcp

Skill Claude CodeCodex

Part of speech-mcp

Expert guide for speech-mcp - TTS, local/streaming STT, barge-in, wake word, voice command bus, and provider selection.

2 14d ago A 36 tokens original MIT

At most 3 mods per repository are shown here, and a mod shipped inside a plugin is left to that plugin's page — the rest are on their repository pages: