Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add calesthio/generative-media-skills --skill audio-mixing-masteringgit clone --depth 1 https://github.com/calesthio/generative-media-skillsWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/calesthio/generative-media-skills/audio-mixing-mastering)<a href="https://agentmods.dev/skills/calesthio/generative-media-skills/audio-mixing-mastering"><img src="https://agentmods.dev/badge/skills/calesthio/generative-media-skills/audio-mixing-mastering/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/calesthio/generative-media-skills/audio-mixing-mastering"><img src="https://agentmods.dev/badge/skills/calesthio/generative-media-skills/audio-mixing-mastering.svg" alt="Reviewed on agentmods" width="80" height="20"></a>- NVIDIA SkillSpector pass
What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00100 | $0.06133 |
| Opus 5 | $0.00050 | $0.03067 |
| Sonnet 5 | $0.00020 | $0.01227 |
| Haiku 4.5 | $0.00010 | $0.00613 |
Grade A, and why
audio-mixing-mastering scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 12d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 336 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Audio Mixing and Mastering Direction
Use this skill when audio must survive real playback: phone speakers, earbuds, laptops, TV soundbars, cinema-style trailers, podcasts, client review links, broadcast deliveries, and social feeds. Treat the mix as a production decision, not as a last-minute normalization step.
The job is to make the audience understand the foreground, feel the intended energy, avoid fatigue or distortion, and pass the declared delivery spec.
Start with the audio contract
Before touching levels, write a short audio contract for the project:
- Delivery context: social post, YouTube upload, podcast RSS, streaming ad, broadcast, OTT, internal review, theatrical-style trailer, music video, documentary, localization, accessibility version.
- Foreground hierarchy: dialogue/VO first, performance/music first, sound-design impact first, or a changing hierarchy by scene.
- Required deliverables: full mix, dialogue stem, music stem, effects stem, M&E, narration-only, clean captions, audio-described version, alternate language mix, stereo fold-down, 5.1/Atmos printmaster, client preview.
- Target spec: use the client/platform/broadcaster spec if supplied. If no spec is supplied, choose a conservative target and label it as a heuristic.
- Monitoring assumption: headphones-only, nearfield speakers, phone/laptop check, calibrated room, or unavailable.
- Source risk: AI voice artifacts, room noise, inconsistent clips, clipping, music licensing, generated SFX harshness, missing room tone, mono/stereo mismatch, translated VO timing, or stem bleed.
Do not master blindly to the loudest reference. Streaming and broadcast systems may normalize loudness, and over-limiting can reduce clarity while gaining little or nothing at playback.
Evidence categories
Documented facts:
- ITU-R BS.1770-5 specifies algorithms for programme loudness and true-peak signal level measurement, including K-weighting, channel weighting, gating, and true-peak guidance. Verified 2026-07-10: https://www.itu.int/dms_pubrec/itu-r/rec/bs/R-REC-BS.1770-5-202311-I!!PDF-E.pdf
- EBU R 128 version 5.0 recommends normalising programme loudness to -23.0 LUFS; where attaining target level is not practically achievable, a tolerance of +/-1.0 LU is permitted, and QC workflows may allow +/-0.2 LU for measurement error. The maximum true peak level during production linear audio should not exceed -1 dBTP. Verified 2026-07-10: https://tech.ebu.ch/files/live/sites/tech/files/shared/r/r128.pdf
- ATSC A/85:2026-07 quick reference lists -24 LKFS as the target for delivery/exchange without metadata where no prior arrangement exists, -2 dBTP maximum true peak, and a -23 to -27 LKFS range for streaming delivery services unless parties arrange otherwise. Verified 2026-07-10: https://www.atsc.org/wp-content/uploads/2026/07/A85-2026-07-Annex-M.pdf
- AES TD1008 recommends, for internet audio streaming/on-demand distribution, maximum true peak not exceeding -1 dBTP at the lossy codec input, with examples such as speech/assorted content around -18 LUFS, track-normalized music at -16 LUFS, album loudest track at -14 LUFS, interstitials at -18 LUFS, and format examples from -16 to -18 LUFS. Verified 2026-07-10: https://aes.org/wp-content/uploads/2024/01/20210924_TD1008_v3.13.pdf
- Spotify for Artists recommends targeting -14 dB integrated LUFS and keeping true peak below -1 dBTP; if a master is louder than -14 LUFS, Spotify recommends keeping true peak below -2 dBTP. Verified 2026-07-10: https://support.spotify.com/us/artists/article/loudness-normalization/
- Apple Podcasts recommends overall loudness around -16 dB LKFS with +/-1 dB tolerance and true peak not exceeding -1 dB FS, calculated according to ITU-R BS.1770-5. Verified 2026-07-10: https://podcasters.apple.com/support/893-audio-requirements
- YouTube's official upload encoding help lists recommended upload audio bitrates: mono 128 kbps, stereo 384 kbps, 5.1 512 kbps, and immersive audio at 128 kbps per channel. The same official page does not state a public loudness target. Verified 2026-07-10: https://support.google.com/youtube/answer/1722171
- Netflix branded sound mix specifications require -27 LKFS +/-2 LU dialog-gated loudness, true peaks not exceeding -2 dB True Peak, and 48 kHz/24-bit for original language mix or M&E mix masters. Verified 2026-07-10: https://partnerhelp.netflixstudios.com/hc/en-us/articles/360001794307-Netflix-Sound-Mix-Specifications-Best-Practices-v1-6
- WCAG 2.2 understanding guidance for low/no background audio says background sounds should be at least 20 dB lower than foreground speech, except brief sounds, to support users who are hard of hearing. Verified 2026-07-10: https://www.w3.org/WAI/WCAG22/Understanding/low-or-no-background-audio.html
- FFmpeg's
loudnormfilter implements EBU R128 loudness normalization, supports single- and double-pass operation, can target integrated loudness, loudness range, and true peak, and upsamples to 192 kHz in dynamic mode for true-peak detection. Verified 2026-07-10: https://ffmpeg.org/ffmpeg-filters.html#loudnorm
What ships with it
3 files beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 12d ago First seen · 336 lines · 100 tokens per session scan A 19e8d97bef7c
audio-mixing-mastering is a skill published in the GitHub repository calesthio/generative-media-skills (170 stars, last pushed 1mo ago), licensed MIT. It adds 100 tokens to every session and 6,133 once invoked, about $0.0005 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other skills, from other repositories
cliptalk-audio-polish-mixer
Produces a task-local audio-polished preview or output from an existing cut, applying loudness normalization, optional noise reduction, fades, muting, and voice-first mix policy.
cliptalk-broll-overlay-editor
Adds relevant B-roll or cutaway overlays to an existing timeline while preserving the primary audio and making the result reviewable before final export.
cliptalk-cover-intro-composer
Creates task-local cover candidates and composes the confirmed cover into the beginning of the current output as a short intro, keeping cover selection separate from video editing.
feature-demo-recording
Record a demo video of a web feature from a real browser. Two modes -- a NARRATED film where measured voiceover drives the timeline (designed slides, subtitles, punch-in camera, rendered from an HTML timeline), and a SILENT evidence clip for a PR or a QA pass. Use when the user asks to record a video, demo, or screen…
image-authoring
Author images and diagrams as code — SVG, Pillow, Excalidraw, mermaid. Load when asked to draw, illustrate, or make an image, icon, logo, poster, or diagram.
bento-slides
Create and edit Bento presentations — self-contained .bento.html decks whose document is JSON. Use whenever the user wants a slide deck or presentation: from scratch, from source material, or by improving an existing file.