Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add commands/moore-developers/grok-cli/sttgit clone --depth 1 https://github.com/Moore-developers/grok-cliWhat it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00000 | $0.00648 |
| Opus 5 | $0.00000 | $0.00324 |
| Sonnet 5 | $0.00000 | $0.00130 |
| Haiku 4.5 | $0.00000 | $0.00065 |
Grade A, and why
stt scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 84 lines — stays where its author put it; the contents beside it link to each section on GitHub.
grok-cli stt
Purpose
Transcribe a local audio file or remote audio URL into text.
Common Usage
grok-cli stt ./sample.wav
Specify language:
grok-cli stt ./sample.mp3 --language zh
Transcribe remote audio:
grok-cli stt --url https://example.com/sample.wav --language auto
Advanced transcription parameters:
grok-cli stt ./meeting.wav --diarize --keyterm Grok --keyterm xAI --filler-words
Script or skill usage:
grok-cli stt --json --file ./sample.wav
Parameters
PATH: positional audio file to transcribe.--file <PATH>: explicit script-friendly file path.--url <URL>: transcribe a remote audio URL; cannot be used together withPATHor--file.--json: use the standard JSON envelope.--auth-file <PATH>: override the OAuth state file path.--model <MODEL>: command-level model override, mainly for compatibility and usage tagging.--language <LANG>: language code, defaulten.--format <true|false>: whether to request formatted transcription text, defaulttrue.--audio-format <FORMAT>: explicitly declare the raw audio format when container metadata is unavailable.--sample-rate <HZ>: raw audio sample rate.--multichannel: treat audio as multichannel.--channels <CHANNELS>: choose channels such as0,1.--diarize: enable speaker diarization.--keyterm <TERM>: repeatable keyword hint.--filler-words: keep filler words.--timeout <SECONDS>: request timeout, default120; increase it for large files.
Behavior
- One of
PATH,--file, or--urlis required. --urlcannot be combined with local file input.- Local files must exist, otherwise
invalid_argsis returned. - Multipart requests send
fileorurland includeformat,language,audio_format,sample_rate,multichannel,channels,diarize,keyterm, andfiller_wordsas needed. - Access token expiry is checked before the request. If needed, the command refreshes first.
- Responses read
textortranscript. - If the upstream returns
language,duration,words, orchannels,--jsonpreserves those structured fields. - Successful calls are written to the local usage SQLite database under audio usage.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 2d ago First seen · 84 lines · 0 tokens per session scan A 97859ef8601f
stt is a command published in the GitHub repository Moore-developers/grok-cli (52 stars, last pushed 19d ago), licensed MIT. It costs nothing until one of its globs matches a file; then it loads 648 tokens. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other commands, from other repositories
audition-voices
Generate voice audition samples for a character using Venice TTS.
status
Show 3d-design team status and recent activity.
music-suno-prompt
Grounded Suno prompt synthesis from local knowledge corpus + persona canon + label canon. No vibes-prompting.
develop-image-prompt.eval
Generates a detailed image generation prompt from a document or content description. Good output: a prompt that is specific, visual, non-abstract, includes style/composition/lighting guidance, and is calibrated to the specified dimensions and style options.
frontend-3d
You are an expert in 3D web development using Three.js, React Three Fiber, WebGL, and WebGPU. You create immersive 3D experiences for the web.
diagram-create
Generate a diagram. Pass a prompt describing what to visualize.