Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add calesthio/generative-media-skills --skill qwen3-ttsgit clone --depth 1 https://github.com/calesthio/generative-media-skillsWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/calesthio/generative-media-skills/qwen3-tts)<a href="https://agentmods.dev/skills/calesthio/generative-media-skills/qwen3-tts"><img src="https://agentmods.dev/badge/skills/calesthio/generative-media-skills/qwen3-tts/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/calesthio/generative-media-skills/qwen3-tts"><img src="https://agentmods.dev/badge/skills/calesthio/generative-media-skills/qwen3-tts.svg" alt="Reviewed on agentmods" width="80" height="20"></a>- NVIDIA SkillSpector warn
SkillSpector: 3 findings, up to medium
These are SkillSpector’s own severities. On a checked sample its high-severity flags on skills were ~96% false positives — a documented command, a public API, a “never do X” rule — so we show them as a caution to read, not a verdict. Why →
- medium Data Exfiltration · line 64 Data is being sent to an external URL. This could be legitimate telemetry or data exfiltration. Manual review is recommended.Fix: Verify the destination URL is trusted and necessary. Remove or replace with documented APIs. Ensure no secrets, tokens, or PII are transmitted.
- medium Excessive Agency · line 158 Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.Fix: Add human-in-the-loop confirmation for destructive, irreversible, or high-impact operations. Never auto-execute commands that modify files, send data, or alter system state.
- medium Excessive Agency · line 182 Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.Fix: Add human-in-the-loop confirmation for destructive, irreversible, or high-impact operations. Never auto-execute commands that modify files, send data, or alter system state.
What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00110 | $0.05439 |
| Opus 5 | $0.00055 | $0.02720 |
| Sonnet 5 | $0.00022 | $0.01088 |
| Haiku 4.5 | $0.00011 | $0.00544 |
Grade A, and why
qwen3-tts scanned grade A with 1 finding against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 9d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Makes network callslowCapability
Not a fault in itself. Listed so you know the mod talks to something, and to what.
curl -X POST "https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/api/v1/services/aigc/multimodal-generation/generation" \ How it starts
The opening of the file, as written. The whole thing — 256 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Qwen3-TTS production guidance
Use this skill to plan, invoke, review, and document Qwen3-TTS speech generation. Treat every model ID, price, region, quota, voice list, and limit below as volatile; the facts marked "verified 2026-07-10" were checked against official Alibaba/Qwen sources on that date.
What Qwen3-TTS is, and which boundary matters
Documented facts:
- Qwen3-TTS is Alibaba/Qwen's multilingual text-to-speech model family for expressive, controllable, streaming-capable speech generation, with hosted DashScope/Model Studio routes and open-weight research/deployment checkpoints. Qwen's technical report describes 12Hz and 25Hz model variants, 0.6B and 1.7B sizes, voice cloning, voice design, multilingual generation, streaming, and Apache-2.0 model/tokenizer release intent. Source: https://arxiv.org/html/2601.15621v1 and https://github.com/QwenLM/Qwen3-TTS (verified 2026-07-10).
- Alibaba Cloud Model Studio exposes Qwen-TTS/Qwen3-TTS via HTTP/SSE non-real-time synthesis and WebSocket real-time synthesis. The non-real-time user guide says it is suited to latency-tolerant audiobook, e-learning, and content-production work, while real-time synthesis is designed for low-latency assistants, audiobook streaming, and customer service. Sources: https://help.aliyun.com/en/model-studio/non-realtime-tts-user-guide and https://help.aliyun.com/zh/model-studio/realtime-tts-user-guide (verified 2026-07-10).
- Hosted Qwen3-TTS is not one universal endpoint. It has different model families for built-in voices, instruction control, voice design, voice cloning, and realtime variants. Verify the chosen region and model list immediately before production. Source: https://help.aliyun.com/en/model-studio/non-realtime-tts-user-guide (verified 2026-07-10).
Production heuristic:
- Use hosted DashScope when you need managed infrastructure, built-in voices, quick custom voice creation, or production traceability. Use open weights only when you have GPU capacity, need local/private inference, need to inspect or adapt model behavior, or cannot send script/audio data to a hosted provider.
What ships with it
1 file beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 9d ago First seen · 256 lines · 110 tokens per session scan A e99b8ee47491
qwen3-tts is a skill published in the GitHub repository calesthio/generative-media-skills (171 stars, last pushed 2mo ago), licensed MIT. It adds 110 tokens to every session and 5,439 once invoked, about $0.0006 per session on Opus 5. A static security scan graded it A with 1 finding (makes network calls). No closer match exists in the catalogue, so it is treated as the original; first seen 2026-09-03.
Other skills, from other repositories
cliptalk-caption-layout-director
Lays out subtitles for an existing ClipTalk timeline, including top/bottom placement, safe margins, readable wrapping, bilingual order, and review rendering.
video-translate
Translate and dub existing videos into multiple languages using HeyGen. Use when: (1) Translating a video into another language, (2) Dubbing video content with lip-sync, (3) Creating multi-language versions of existing videos, (4) Audio-only translation without lip-sync, (5) Working with HeyGen's /v2/videotranslate…
translation
A translation tool for converting current text, conversations, or explicitly provided files into a target language while preserving their meaning and key information.
Stable Diffusion
State-of-the-art text-to-image generation with Stable Diffusion models via HuggingFace Diffusers. Use when generating images from text prompts, performing image-to-image translation, inpainting, or building custom diffusion pipelines.
cliptalk-cover-director
Produces evidence-backed cover candidates and reviewable cover variants for a ClipTalk video. Use when the user asks for a cover, poster frame, thumbnail, or multiple cover directions; do not use for timeline editing or social-video reframing.
cliptalk-smart-reframe
Creates a subject-aware, time-varying crop track and a review-only social-format preview from an accepted ClipTalk cut. Use for automatic vertical, square, or portrait reframing; do not use for a fixed manual crop or before content editing is accepted.