Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add modu-ai/moai-cowork --skill media-audio-gengit clone --depth 1 https://github.com/modu-ai/moai-coworkWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/modu-ai/moai-cowork/media-audio-gen)<a href="https://agentmods.dev/skills/modu-ai/moai-cowork/media-audio-gen"><img src="https://agentmods.dev/badge/skills/modu-ai/moai-cowork/media-audio-gen/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/modu-ai/moai-cowork/media-audio-gen"><img src="https://agentmods.dev/badge/skills/modu-ai/moai-cowork/media-audio-gen.svg" alt="Reviewed on agentmods" width="80" height="20"></a>- NVIDIA SkillSpector pass
What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00099 | $0.05614 |
| Opus 5 | $0.00049 | $0.02807 |
| Sonnet 5 | $0.00020 | $0.01123 |
| Haiku 4.5 | $0.00010 | $0.00561 |
Grade A, and why
media-audio-gen scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 8d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 355 lines — stays where its author put it; the contents beside it link to each section on GitHub.
media-audio-gen
개요
AI 기반 오디오 생성을 위한 통합 스킬입니다. ElevenLabs의 고성능 TTS 엔진을 활용하여 32개국어(한국 포함)의 자연스러운 음성을 생성하고, 단 1분의 샘플로 개인/브랜드 보이스를 클로닝합니다. 또한 비디오 다국어 더빙(립싱크 자동), 효과음 생성, 실시간 대화형 AI 보이스 에이전트 구축을 지원합니다.
주요 기능
- TTS (Text-to-Speech): 32개국어 자연스러운 음성 생성 (한국어 최적화)
- 보이스 클로닝: 1분 샘플로 개인/브랜드 보이스 복제
- 다국어 더빙: 한국어 비디오 → 영어/일본어 등 (립싱크 자동 조정)
- 효과음 생성: 영화, 게임, 콘텐츠용 사운드 이펙트
- ConvAI: 실시간 대화형 보이스 에이전트 (챗봇, AI 상담원)
트리거 키워드
다음 요청 시 이 스킬을 사용하세요:
- 음성 생성 관련: "목소리 만들어줘", "TTS", "음성 합성", "나레이션 녹음", "AI 성우"
- 보이스 클로닝: "내 목소리 복제", "브랜드 보이스 만들기", "보이스 클로닝", "샘플 음성으로 학습"
- 더빙 관련: "영어 더빙", "일본어 번역+녹음", "외국어 자막+음성", "다국어 버전"
- 효과음: "효과음 생성", "사운드 이펙트", "배경음", "비디오 소리"
- 대화형 AI: "AI 상담원", "보이스봇", "실시간 음성 대화", "전화 자동 응답"
워크플로우
1. 기본 TTS (Text-to-Speech)
1. 텍스트 입력 (한국/영어/32개국어)
2. 보이스 프리셋 선택 또는 커스텀 보이스 ID 지정
3. 모델 선택 (eleven_multilingual_v2 기본)
4. 합성 전 견적 확인 ← 아래 §유료 합성 견적 게이트
5. 오디오 생성 (MP3/WAV)
1-1. 유료 합성 견적 게이트 (TTS·더빙·효과음 공통)
ElevenLabs 할당량은 문자 수로 소진되고 되돌아오지 않는다. 장문 나레이션 한 번이 무료 플랜 월 한도(10,000자)를 통째로 쓸 수 있다. 그래서 합성 전에 얼마가 나가는지 보여주고 승인을 받는다.
- [HARD] 크레딧이 나가는 합성 전에는 견적을 보여주고 승인을 받는다. 보이스 목록 조회·잔여 할당량 조회 같은 무료 호출은 대상이 아니다.
- [HARD] 요약하지 말고 실제로 넘어가는 것을 그대로 보여준다:
| 보여줄 것 | TTS | 더빙 | 효과음 |
|---|---|---|---|
| 합성될 원문 전문과 문자 수 | ✓ | ✓ (원문 + 번역문 양쪽) | 프롬프트 |
| 보이스 이름·ID (클로닝 보이스면 그 사실도) | ✓ | ✓ | — |
모델 (eleven_multilingual_v2 등) |
✓ | ✓ | — |
| 대상 언어 목록 | — | ✓ | — |
| 출력 개수 · 길이 | ✓ | ✓ | ✓ |
| 소진되는 문자 수 / 할당량과 잔여분 | ✓ | ✓ | ✓ |
- [HARD] 대상 언어가 여럿이면 언어별 소진분과 합계를 함께 보여준다. 합계만 보여주면 언어 하나를 빼는 판단을 할 수 없다. 더빙은 언어 수만큼 배수로 나간다.
- [HARD] 한국어 나레이션은 합성 전에 ⟨한국어 감사 3단⟩을 통과시킨다. 음성으로 굳은 뒤 고치면 할당량을 다시 쓴다.
moai-coworker:ai-slop-reviewer 1차 일반 슬롭 정리
→ moai-writer:korean-spell-check 2차 맞춤법 — 제안 수집 (미공개 정보가 섞였으면 건너뜀)
→ moai-writer:korean-humanize 3차 정밀 윤문 + 맞춤법 반영 + Phase 6 최종 검수
감사 뒤에 문자 수를 다시 센다 — 감사가 문장을 고치므로 이전 견적은 무효다.
승인 선택지는 이렇게 구성한다:
| 선택지 | 뜻 |
|---|---|
| 이대로 합성 (권장) | 보여준 견적 그대로 실행 |
| 고쳐 쓰기 / 언어·개수 줄이기 | 원문을 다듬거나 범위를 줄여 견적을 다시 |
| 취소 | 합성하지 않고 종료. 할당량 소진 없음 |
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 8d ago First seen · 355 lines · 99 tokens per session scan A fbcbe71ffcfb
media-audio-gen is a skill published in the GitHub repository modu-ai/moai-cowork (300 stars, last pushed 9d ago), licensed Apache-2.0. It adds 99 tokens to every session and 5,614 once invoked, about $0.0005 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-09-03.
Other skills, from other repositories
market-intelligence-report
Produce a Market Intelligence Report — YouTube competitive research, channel analysis, content gap discovery, idea generation, daily scanning, and AI trend scouting — then render it as a BenAI-branded HTML dashboard in the instant-ui design language. Use this skill whenever the user says "market intelligence report"…
linkedin-writer-vault
Vault-aware LinkedIn writer. Same step-by-step LinkedIn post process as linkedin-writer, but ICP, voice, and offer context come from the vault's Context/ folder instead of being bundled inside the skill. Update one file in the vault and every skill pointing to it inherits the change. TRIGGERS: LinkedIn post, LinkedIn…
marketing-os-carousel
Build an image-first social carousel from an asset already filed in the Marketing OS, export it as a PDF, and record it back as a real channel asset. Brand palette, typography, the logo pointer and the never-black-background rule all resolve from Context/brand/brand-kit.md. Source is a filed newsletter edition…
operator
Build and schedule a personalized Operator prompt that runs a Baalda vault as a second brain on a recurring cadence. Run it from inside the vault: it reads Context/ and CLAUDE.md first to infer org, team, brand voice and paths, then asks only the gaps (cadence, connectors, DM recipient, budgets, signature), writes the…
crm-prospect-mining
Mine high-value prospects from CRM pipeline stages (Lost, No Show, Churned, Stalled) by cross-referencing records with LinkedIn company data and comms history. Connects to any CRM, pulls records from target stages, filters out personal email domains, finds company LinkedIn pages via web research, bulk-scrapes company…
seo-hreflang
Hreflang and international SEO audit, validation, and generation. Detects common mistakes, validates language/region codes, and generates correct hreflang implementations for HTML, HTTP headers, and XML sitemaps. Use when user says "hreflang", "i18n SEO", "international SEO", "multi-language", "multi-region"…