speech-to-text

speech-to-text is a skill for Claude Code, Codex from thewolffish/wolffish-app. It costs 68 tokens per session (2,932 once invoked), scanned B, original, MIT.

A local audio transcription tool that converts recordings into written text using faster-whisper. It supports 99 languages and runs offline after setup, with language detection available as an option.

In plain words
What is it for?
Transcribing separate audio files, voice memos, or recordings, and optionally detecting their language.
Why use it?
It lets users transcribe audio without cloud services, API keys, or a separate ffmpeg installation. The chosen speech model balances speed against accuracy.

Skill for Claude CodeCodex

Written for no agent in particular: nothing here depends on one.

Not installable: its command points at a path on the author’s own machine, so it runs nowhere else. The line is /Users/me/recording.wav.

Good fit Transcribing separate audio files, voice memos, or recordings, and optionally detecting their language.

Compare 6 skills from other repositories ↓
Install

Getting it into your agent

There is no command for this one: it runs only inside a plugin, and the catalogue could not identify which plugin ships it. The source is linked below.

Made for: Claude Code, Codex.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for speech-to-text

README.md
[![agentmods](https://agentmods.dev/badge/skills/thewolffish/wolffish-app/speech-to-text/github.svg)](https://agentmods.dev/skills/thewolffish/wolffish-app/speech-to-text)
Your own site
<a href="https://agentmods.dev/skills/thewolffish/wolffish-app/speech-to-text"><img src="https://agentmods.dev/badge/skills/thewolffish/wolffish-app/speech-to-text/github.svg" alt="Measured on agentmods" height="20"></a>

Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.

agentmods 80×15 button for speech-to-text

Your own site · 80×15
<a href="https://agentmods.dev/skills/thewolffish/wolffish-app/speech-to-text"><img src="https://agentmods.dev/badge/skills/thewolffish/wolffish-app/speech-to-text.svg" alt="Reviewed on agentmods" width="80" height="20"></a>
Per session 68 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 2,932 The whole file, excluding the scripts and references it only reads on demand.
Security scan B 1 finding. A grade says what 26 rules found in the file — not that it is safe.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5.1 $0.00068 $0.02932
Opus 5 $0.00034 $0.01466
Sonnet 5 $0.00014 $0.00586
Haiku 4.5 $0.00007 $0.00293

Measured 5d ago against content hash 6d43de87264f, method: parsed. Prices are Anthropic first-party input rates as of 2026-09-10, from the pricing page.

Security

Grade B, and why

speech-to-text scanned grade B with 1 finding against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 5d ago.

The scan reads SKILL.md. This mod also ships 2 executable files (plugin/index.mjs, plugin/transcribe.py), listed below but not scanned — reading those needs a real analyzer, not pattern matching.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Asks for rootmediumPrivilege escalation

A mod that escalates privileges can change anything on the machine, not only the project.

- Everything is provisioned managed-first with no admin rights. Do NOT propose `sudo apt` / `brew` / `winget` installs for speech-to-text.
src/defaults/workspace/brain/cerebellum/speech-to-text/SKILL.md · 254 lines

How it starts

The opening of the file, as written. The whole thing — 254 lines — stays where its author put it; the contents beside it link to each section on GitHub.

Speech-to-text

Interface

  • Tools: stt_transcribe, stt_transcribe_upload, stt_transcribe_voice_memo, stt_detect_language
  • Engine: faster-whisper (CTranslate2 + PyAV) — free, open-source, 100% local, and far lighter than reference Whisper (no PyTorch, no external ffmpeg).
  • Models: tiny → base → small → medium → large (accuracy vs. speed tradeoff)
  • Setup: fully automatic — the python capability provisions a hermetic Python runtime and installs faster-whisper into an isolated venv on first use (no host Python and no ffmpeg required). The chosen model is downloaded on first use.

When to use each tool

Never transcribe the user's own voice note. When the user's message is tagged <voice_note>, it was ALREADY transcribed before the agent ran — the visible message text IS the transcript. The audio attached to that message is only the source of that transcript. Do NOT call any stt_* tool on it; just respond to the text (per the <voice_prompts> rules in your prompt). These tools are for a separate audio file the user hands you to transcribe — not for their own spoken message.

  • "transcribe this", "what does this audio say?", "what did they say?" with an uploaded audio file → stt_transcribe_upload using the uploaded file's name. The <attachments> block in the user message lists every uploaded filename and type.

  • "transcribe the last voice memo", "transcribe what you just said"stt_transcribe_voice_memo with the filename from the most recent text-to-speech tool result.

  • User provides an absolute path: "transcribe /Users/me/recording.wav"stt_transcribe with that path.

  • "what language is this?", "what language are they speaking?"stt_detect_language returns the language without a full transcription.

  • "my voice notes come out in the wrong language", "transcribe in Arabic from now on", "use a more accurate model" → the settings tools below. This is yours to fix, not something to send the user to Settings for.

Read the full file on GitHub · 254 lines

Files

What ships with it

2 files beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 5d ago Changed · +62 lines 6d43de87264f
  2. 10d ago First seen · 192 lines · 68 tokens per session scan B 892b017a52ce

Subscribe to this mod's changes

speech-to-text is a skill published in the GitHub repository thewolffish/wolffish-app (5 stars, last pushed 4d ago), licensed MIT. It adds 68 tokens to every session and 2,932 once invoked, about $0.0003 per session on Opus 5. A static security scan graded it B with 1 finding (asks for root). No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.

Related

Other skills, from other repositories

agent-computer-use

REQUIRED for any task that involves operating a desktop application — opening apps, clicking buttons, typing into fields, pressing keys, scrolling, dragging, reading what's on screen, moving or resizing windows, or verifying state after an action. Always use the agent-cu CLI commands (open, snapshot, click, type, key…

kortix-ai/agent-computer-use · 216 tokens

claude-api

Build, debug, and optimize Claude API / Anthropic SDK apps. Apps built with this skill should include prompt caching. Also handles migrating existing Claude API code between Claude model versions (4.5 → 4.6, 4.6 → 4.7, retired-model replacements). TRIGGER when: code imports anthropic/@anthropic-ai/sdk; user asks for…

warpdotdev/warp · 193 tokens

create-skill

Create new skills, modify and improve existing skills, and measure skill performance. Use when users want to create a skill from scratch, edit, or optimize an existing skill, run evals to test a skill, benchmark skill performance with variance analysis, or optimize a skill's description for better triggering accuracy.

warpdotdev/warp · 64 tokens

figma-create-design-system-rules

Generates custom design system rules for the user's codebase. Use when user says "create design system rules", "generate rules for my project", "set up design rules", "customize design system guidelines", or wants to establish project-specific conventions for Figma-to-code workflows. Requires Figma MCP server…

warpdotdev/warp · 71 tokens

figma-generate-design

Use this skill alongside figma-use when the task involves translating an application page, view, or multi-section layout into Figma. Triggers: 'write to Figma', 'create in Figma from code', 'push page to Figma', 'take this app/page and build it in Figma', 'create a screen', 'build a landing page in Figma', 'update the…

warpdotdev/warp · 160 tokens

figma-generate-library

Build or update a professional-grade design system in Figma from a codebase. Use when the user wants to create variables/tokens, build component libraries, set up theming (light/dark modes), document foundations, or reconcile gaps between code and Figma. This skill teaches WHAT to build and in WHAT ORDER — it…

warpdotdev/warp · 95 tokens