Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add commands/moore-developers/grok-cli/ttsgit clone --depth 1 https://github.com/Moore-developers/grok-cliWhat it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00000 | $0.00668 |
| Opus 5 | $0.00000 | $0.00334 |
| Sonnet 5 | $0.00000 | $0.00134 |
| Haiku 4.5 | $0.00000 | $0.00067 |
Grade A, and why
tts scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 88 lines — stays where its author put it; the contents beside it link to each section on GitHub.
grok-cli tts
Purpose
Convert text to speech and save the audio locally.
Common Usage
grok-cli tts "Hello from Grok"
Choose voice, language, and output file:
grok-cli tts "Hello, I am Grok" --voice-id eve --language zh --output ./out/grok.mp3
Script or skill usage:
grok-cli tts --json --text "Hello from Grok"
List available voices:
grok-cli tts --list-voices --json
Control output format explicitly:
grok-cli tts "Hello" --output-format mp3 --sample-rate 24000 --bit-rate 128000
Parameters
TEXT: positional input text.--text <TEXT>: explicit script-friendly text.--list-voices: list available TTS voices without synthesizing audio.--json: use the standard JSON envelope.--auth-file <PATH>: override the OAuth state file path.--voice-id <VOICE>: voice id, defaulteve.--language <LANG>: language code, defaulten;autois allowed.--output <PATH>: output audio path.--output-format <FORMAT>: output format, such asmp3orwav.--sample-rate <HZ>: output sample rate.--bit-rate <BPS>: output bit rate.--optimize-streaming-latency <MODE>: pass through xAI TTS streaming latency optimization mode.--text-normalization <MODE>: pass through xAI TTS text normalization mode.--model <MODEL>: command-level model override, mainly for compatibility and usage tagging.--timeout <SECONDS>: request timeout, default120.
Behavior
- Default output path lives under
~/.hermes/cache/audio/audio_cache/. - If the output file extension is
.wav, the request sendsoutput_format=wav. - If
--output-formatis passed explicitly, it must match the output file extension. - The xAI TTS request body sends
text,voice_id,language, and optionallyoutput_format,optimize_streaming_latency, andtext_normalization. --list-voicescallsGET /v1/tts/voicesand returns a voice list without requiring text.- Access token expiry is checked before the request. If needed, the command refreshes first.
- Successful calls are written to the local usage SQLite database under audio usage.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 2d ago First seen · 88 lines · 0 tokens per session scan A bd25089d77e7
tts is a command published in the GitHub repository Moore-developers/grok-cli (52 stars, last pushed 19d ago), licensed MIT. It costs nothing until one of its globs matches a file; then it loads 668 tokens. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other commands, from other repositories
audition-voices
Generate voice audition samples for a character using Venice TTS.
status
Show 3d-design team status and recent activity.
music-suno-prompt
Grounded Suno prompt synthesis from local knowledge corpus + persona canon + label canon. No vibes-prompting.
develop-image-prompt.eval
Generates a detailed image generation prompt from a document or content description. Good output: a prompt that is specific, visual, non-abstract, includes style/composition/lighting guidance, and is calibrated to the specified dimensions and style options.
frontend-3d
You are an expert in 3D web development using Three.js, React Three Fiber, WebGL, and WebGPU. You create immersive 3D experiences for the web.
diagram-create
Generate a diagram. Pass a prompt describing what to visualize.