Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add calesthio/generative-media-skills --skill cartesia-sonicgit clone --depth 1 https://github.com/calesthio/generative-media-skillsWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/calesthio/generative-media-skills/cartesia-sonic)<a href="https://agentmods.dev/skills/calesthio/generative-media-skills/cartesia-sonic"><img src="https://agentmods.dev/badge/skills/calesthio/generative-media-skills/cartesia-sonic/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/calesthio/generative-media-skills/cartesia-sonic"><img src="https://agentmods.dev/badge/skills/calesthio/generative-media-skills/cartesia-sonic.svg" alt="Reviewed on agentmods" width="80" height="20"></a>- NVIDIA SkillSpector warn
SkillSpector: 3 findings, up to high
These are SkillSpector’s own severities. On a checked sample its high-severity flags on skills were ~96% false positives — a documented command, a public API, a “never do X” rule — so we show them as a caution to read, not a verdict. Why →
- high Memory Poisoning · line 178 Skill manipulates agent memory, state, or stored context. Memory corruption can alter personality, override safety rules, or cause unpredictable behavior.Fix: Protect agent memory and state from modification by untrusted content. Use read-only memory for critical instructions and validate all state changes.
- high Privilege Escalation · line 235 Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.Fix: Remove references to credential paths. Use environment variables or secrets managers. For docs, use placeholder paths (e.g., /path/to/config). Never load .env or token files in production code paths.
- high Privilege Escalation · line 395 Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.Fix: Remove references to credential paths. Use environment variables or secrets managers. For docs, use placeholder paths (e.g., /path/to/config). Never load .env or token files in production code paths.
What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00080 | $0.06236 |
| Opus 5 | $0.00040 | $0.03118 |
| Sonnet 5 | $0.00016 | $0.01247 |
| Haiku 4.5 | $0.00008 | $0.00624 |
Grade A, and why
cartesia-sonic scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 9d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 473 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Cartesia Sonic production skill
Use this skill when a media-production agent is choosing, prompting, integrating, or reviewing Cartesia Sonic for speech output. Treat Sonic as a low-latency text-to-speech provider first; use adjacent Cartesia voice APIs only when the production need explicitly requires them.
All volatile facts below were checked against official Cartesia documentation or legal pages on 2026-07-10. Re-check before committing budgets, compliance claims, model IDs, pricing, concurrency, supported languages, or API behavior.
Provider boundary
Documented facts:
- Sonic is Cartesia's text-to-speech model family. The current production default is
sonic-3.5; API references also listsonic-3andsonic-latest. - Sonic takes text and returns generated speech. It is suitable for realtime voice agents, narration, dubbing, avatars, notifications, ads, and localization when the workflow starts from a transcript.
- Cartesia also documents Ink speech-to-text, Line hosted voice agents, voice cloning, voice localization, infill, and voice changer. Do not present all of these as "Sonic TTS"; name the exact Cartesia surface being used.
- Voice changer is not TTS: it takes an input speech clip and returns speech with the same intonation in a different target voice.
- Speech-to-speech or voice conversion claims should be limited to the documented
voice-changerendpoints unless Cartesia's docs explicitly add another supported path.
Production heuristics:
- Prefer Cartesia Sonic when low time-to-first-byte, conversational pacing, multilingual speech, and API streaming matter more than offline local rendering or maximal post-production editability.
- Prefer non-realtime batch TTS providers or local TTS when the project requires air-gapped processing, predictable no-network rendering, or a license/security posture not satisfied by Cartesia's current account settings.
- For a media pipeline, record the selected endpoint, model ID, voice ID, language, output format, Cartesia-Version, generation_config, pronunciation dictionary ID, and source transcript checksum in the asset manifest.
What ships with it
1 file beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 9d ago First seen · 473 lines · 80 tokens per session scan A 40e88c19c232
cartesia-sonic is a skill published in the GitHub repository calesthio/generative-media-skills (170 stars, last pushed 2mo ago), licensed MIT. It adds 80 tokens to every session and 6,236 once invoked, about $0.0004 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-09-03.
Other skills, from other repositories
cliptalk-caption-layout-director
Lays out subtitles for an existing ClipTalk timeline, including top/bottom placement, safe margins, readable wrapping, bilingual order, and review rendering.
video-translate
Translate and dub existing videos into multiple languages using HeyGen. Use when: (1) Translating a video into another language, (2) Dubbing video content with lip-sync, (3) Creating multi-language versions of existing videos, (4) Audio-only translation without lip-sync, (5) Working with HeyGen's /v2/videotranslate…
translation
A translation tool for converting current text, conversations, or explicitly provided files into a target language while preserving their meaning and key information.
Stable Diffusion
State-of-the-art text-to-image generation with Stable Diffusion models via HuggingFace Diffusers. Use when generating images from text prompts, performing image-to-image translation, inpainting, or building custom diffusion pipelines.
cliptalk-cover-director
Produces evidence-backed cover candidates and reviewable cover variants for a ClipTalk video. Use when the user asks for a cover, poster frame, thumbnail, or multiple cover directions; do not use for timeline editing or social-video reframing.
cliptalk-smart-reframe
Creates a subject-aware, time-varying crop track and a review-only social-format preview from an accepted ClipTalk cut. Use for automatic vertical, square, or portrait reframing; do not use for a fixed manual crop or before content editing is accepted.