scenario-audio

scenario-audio is a skill for Claude Code from scenario-labs/skills. It costs 110 tokens per session (1,615 once invoked), scanned A, original, MIT.

An audio-generation workflow for Scenario, a service for creating and processing digital media. It helps find models for music, sound effects, voice and speech, and soundtrack creation, then runs them and saves the resulting audio.

In plain words
What is it for?
Use it to create songs, background music, sound effects, ambience, voiceovers, narration, speech, or soundtracks, and to listen to or download the results.
Why use it?
It brings different audio tasks into one process and handles model settings, long-running jobs, playback, and downloads. It is intended for generated or existing audio work rather than image creation.

Skill for Claude Code

Written for Claude Code: shipped in a Claude Code plugin.

Part of the 05.-video-and-audio plugin — 10 skills shipped together

Good fit Use it to create songs, background music, sound effects, ambience, voiceovers, narration, speech, or soundtracks, and to listen to or download the results.

Compare 6 skills from other repositories ↓
Install with agentmods
npx agentmods add skills/scenario-labs/skills/scenario-audio
Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

Any agent
npx skills add scenario-labs/skills --skill scenario-audio
Clone the repo
git clone --depth 1 https://github.com/scenario-labs/skills

Made for: Claude Code.

Or install 05.-video-and-audio, the plugin that ships this one along with the rest of its 10 skills.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for scenario-audio

README.md
[![agentmods](https://agentmods.dev/badge/skills/scenario-labs/skills/scenario-audio/github.svg)](https://agentmods.dev/skills/scenario-labs/skills/scenario-audio)
Your own site
<a href="https://agentmods.dev/skills/scenario-labs/skills/scenario-audio"><img src="https://agentmods.dev/badge/skills/scenario-labs/skills/scenario-audio/github.svg" alt="Measured on agentmods" height="20"></a>

Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.

agentmods 80×15 button for scenario-audio

Your own site · 80×15
<a href="https://agentmods.dev/skills/scenario-labs/skills/scenario-audio"><img src="https://agentmods.dev/badge/skills/scenario-labs/skills/scenario-audio.svg" alt="Reviewed on agentmods" width="80" height="20"></a>
Per session 110 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 1,615 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 1 finding. A grade says what 26 rules found in the file — not that it is safe.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5.1 $0.00110 $0.01615
Opus 5 $0.00055 $0.00807
Sonnet 5 $0.00022 $0.00323
Haiku 4.5 $0.00011 $0.00161

Measured 7d ago against content hash ffbec3578d5b, method: parsed. Prices are Anthropic first-party input rates as of 2026-09-10, from the pricing page.

Security

Grade A, and why

scenario-audio scanned grade A with 1 finding against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 7d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Makes network callslowCapability

Not a fault in itself. Listed so you know the mod talks to something, and to what.

| Save | `asset_download` | returns a download URL: `curl -L -o out.mp3 "<url>"` |
skills/scenario-audio/SKILL.md · 72 lines

How it starts

The opening of the file, as written. The whole thing — 72 lines — stays where its author put it; the contents beside it link to each section on GitHub.

Scenario Audio Generation

Overview

Scenario generates audio through the same loop as images. The live catalog covers three generation lanes (music, sound effects, voice/speech) plus video-to-audio soundtrack models and audio utilities. Connection and the core generation loop: see the scenario skill. If a sibling skill named here is missing from your available skills, ask the user to install it (npx skills add scenario-labs/skills --skill <name>); unattended, proceed from tool schemas and flag the gap.

Quick reference

Step Tool Notes
Find a model recommend with the need in the user's own words; search only for a member known by name capability="txt2audio" covers music, SFX, and TTS; optional, inferred from the prompt when omitted
Inspect inputs model_schema_get audio schemas vary widely: durations, lyrics, voices, looping
Generate model_run schema-conformant parameters; wait=false for long jobs
Wait jobs_wait blocks server-side; on timeout re-call with pending_job_ids
Listen asset_display renders an inline audio player
Save asset_download returns a download URL: curl -L -o out.mp3 "<url>"

Read the full file on GitHub · 72 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 7d ago Changed ffbec3578d5b
  2. 10d ago First seen · 72 lines · 110 tokens per session scan A f0341c9890c3

Subscribe to this mod's changes

scenario-audio is a skill published in the GitHub repository scenario-labs/skills (11 stars, last pushed 7d ago), licensed MIT. It adds 110 tokens to every session and 1,615 once invoked, about $0.0006 per session on Opus 5. A static security scan graded it A with 1 finding (makes network calls). No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.

Related

Other skills, from other repositories

Video Generation

Implement AI-powered video generation capabilities using the z-ai-web-dev-sdk. Use this skill when the user needs to generate videos from text prompts or images, create video content programmatically, or build applications that produce video outputs. Supports asynchronous task management with status polling and result…

jjyaoao/HelloAgents · 58 tokens

video-understand

Implement specialized video understanding capabilities using the z-ai-web-dev-sdk. Use this skill when the user needs to analyze video content, understand motion and temporal sequences, extract information from video frames, describe video scenes, or perform video-based AI analysis. Optimized for MP4, AVI, MOV, and…

jjyaoao/HelloAgents · 68 tokens

TTS

Implement text-to-speech (TTS) capabilities using the z-ai-web-dev-sdk. Use this skill when the user needs to convert text into natural-sounding speech, create audio content, build voice-enabled applications, or generate spoken audio files. Supports multiple voices, adjustable speed, and various audio formats.

jjyaoao/HelloAgents · 64 tokens

image-generation

Implement AI image generation capabilities using the z-ai-web-dev-sdk. Use this skill when the user needs to create images from text descriptions, generate visual content, create artwork, design assets, or build applications with AI-powered image creation. Supports multiple image sizes and returns base64 encoded…

jjyaoao/HelloAgents · 69 tokens

Podcast Generate

Generate podcast episodes from user-provided content or by searching the web for specified topics. If user uploads a text file/article, creates a dual-host dialogue podcast (or single-host upon request). If no content is provided, searches the web for information about the user-specified topic and generates a podcast.…

jjyaoao/HelloAgents · 113 tokens

image_generate

An image-generation skill whose detailed work and output rules come from the Lime project.

limecloud/lime · 44 tokens