voice-optimization

voice-optimization is a skill for Claude Code, Codex from kangarooking/system-prompt-skills. It costs 84 tokens per session (1,668 once invoked), scanned A, original, MIT.

A set of guidelines for making voice-based assistant conversations brief, natural, and safe. It covers spoken wording, supported languages, response length, and boundaries such as not identifying or imitating speakers.

In plain words
What is it for?
Use it when designing or adapting a voice assistant, writing voice-specific system instructions, or deciding how a chatbot should respond through audio.
Why use it?
Voice users cannot scan a screen, so long introductions and visual formatting are less useful. The rules also address privacy and safety issues that arise when software handles spoken audio.

Skill for Claude CodeCodex

Written for no agent in particular: nothing here depends on one.

Good fit Use it when designing or adapting a voice assistant, writing voice-specific system instructions, or deciding how a chatbot should respond through audio.

Compare 6 skills from other repositories ↓
Install with agentmods
npx agentmods add skills/kangarooking/system-prompt-skills/voice-optimization
Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

Any agent
npx skills add kangarooking/system-prompt-skills --skill voice-optimization
Clone the repo
git clone --depth 1 https://github.com/kangarooking/system-prompt-skills

Made for: Claude Code, Codex.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for voice-optimization

README.md
[![agentmods](https://agentmods.dev/badge/skills/kangarooking/system-prompt-skills/voice-optimization/github.svg)](https://agentmods.dev/skills/kangarooking/system-prompt-skills/voice-optimization)
Your own site
<a href="https://agentmods.dev/skills/kangarooking/system-prompt-skills/voice-optimization"><img src="https://agentmods.dev/badge/skills/kangarooking/system-prompt-skills/voice-optimization/github.svg" alt="Measured on agentmods" height="20"></a>

Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.

agentmods 80×15 button for voice-optimization

Your own site · 80×15
<a href="https://agentmods.dev/skills/kangarooking/system-prompt-skills/voice-optimization"><img src="https://agentmods.dev/badge/skills/kangarooking/system-prompt-skills/voice-optimization.svg" alt="Reviewed on agentmods" width="80" height="20"></a>
Per session 84 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 1,668 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. A grade says what 26 rules found in the file — not that it is safe.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5.1 $0.00084 $0.01668
Opus 5 $0.00042 $0.00834
Sonnet 5 $0.00017 $0.00334
Haiku 4.5 $0.00008 $0.00167

Measured 9d ago against content hash 70dc18eec11d, method: parsed. Prices are Anthropic first-party input rates as of 2026-09-08, from the pricing page.

Security

Grade A, and why

voice-optimization scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 9d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

voice-optimization/SKILL.md · 75 lines

How it starts

The opening of the file, as written. The whole thing — 75 lines — stays where its author put it; the contents beside it link to each section on GitHub.

语音场景优化

R — 原文 (Reading)

Perplexity Voice 要求"请快速说话"、仅支持英语、禁止说话人识别、禁止唱歌哼唱、禁止模仿;Claude Mobile 强调"始终先给答案、无前言"、列表在小屏幕上更易扫描;Sesame AI Maya 专为语音优化的对话模式。核心模式:简洁优先、口语化适配、音频安全约束、语言限制、去除视觉格式。

I — 方法论骨架 (Interpretation)

  1. 简洁优先原则:语音场景中用户注意力窗口极短,回答必须开门见山,禁止"好的,让我来回答您的问题"等前言式表达。
  2. 口语化转换:将书面语转换为自然口语——使用短句、主动语态、日常词汇,避免从句嵌套和术语堆砌。
  3. 格式降级:去除 Markdown 表格、代码块、嵌套列表等视觉格式,改用自然语言描述或简短列举。
  4. 音频安全约束:禁止识别特定说话人身份、禁止模仿真实人物声线特征、禁止唱歌或哼唱旋律。
  5. 语言限制声明:明确支持的语种范围,超出范围时引导用户调整设置而非强行处理。
  6. 长度自适应:根据问题复杂度动态调整回答长度——简单问题一两句话,复杂问题控制在合理时长内。

A1 — 案例分析 (Past Application)

案例: Perplexity Voice 的音频安全边界

  • 问题: 语音交互中用户可能要求模仿名人声音、识别通话对象、或要求 AI 唱歌,这些行为涉及隐私、版权和安全风险。
  • 设计模式的使用: Perplexity Voice 在系统提示中设置明确禁区——不对语音输入进行说话人识别("No speaker identification from voice"),不执行唱歌或哼唱请求,不进行人物模仿。同时限定仅支持英语,超出能力范围时引导用户修改设置。
  • 结论: 语音场景有独特的安全边界(声纹、模仿、演唱),这些在纯文本场景中不存在,需要专项防护。

案例: Claude Mobile 的回答优先策略

  • 问题: 移动端语音回答中,用户听到冗长的开场白会快速失去耐心,尤其在驾驶、行走等场景下。
  • 设计模式的使用: Claude Mobile 明确指令"Always lead with answer. No preamble.",将答案前置,解释后置。对于不同复杂度的问题设定长度层级——简单问题 1-2 句,操作指南用短列表,实质性问题 2-3 段。
  • 结论: 语音场景对延迟感知极度敏感,去除前言可显著提升用户满意度和信息获取效率。

A2 — 触发场景 (Future Trigger) ★

用户在什么情境下需要?

  1. 设计语音助手(如智能音箱、车载助手)的系统提示
  2. 为现有文本聊天机器人添加语音交互模式
  3. 构建电话客服 AI 的对话脚本
  4. 优化播客生成或有声内容合成中的口语表达

语言信号

  • "语音场景下的输出优化"
  • "需要口语化回答"
  • "用户通过语音提问"
  • "回答会被朗读出来"
  • "如何让 AI 说话更自然"

与相邻 skill 的区分

  • mobile-adaptation 区别:移动适配关注屏幕尺寸约束,语音优化关注听觉通道约束;但两者都强调简洁优先,常联合使用
  • citation-system 区别:引用系统在语音场景中需要特殊处理(无法使用视觉标记),但语音优化不涉及引用格式设计本身

E — 可执行步骤 (Execution)

  1. 步骤 1:设定回答长度层级 - 完成标准:为问题复杂度定义 3-4 个层级(简单/操作/中等/复杂),每个层级规定最大句数或预估朗读时长,并在系统提示中以示例说明。
  2. 步骤 2:编写前言禁令与答案前置规则 - 完成标准:明确声明"禁止在回答开头添加确认性前言",提供正确和错误的示例对比(如错误:"好的,让我为您解答...",正确:直接给出答案)。
  3. 步骤 3:定义格式降级规则 - 完成标准:列出需降级的视觉格式(表格→自然语言描述、嵌套列表→扁平列举、代码块→口语化步骤说明),并给出每种降级的示例。
  4. 步骤 4:设定音频安全边界 - 完成标准:明确禁止的行为清单(说话人识别、声线模仿、唱歌哼唱、人物扮演),定义超出能力范围时的标准回退话术。
  5. 步骤 5:添加口语化转换指南 - 完成标准:列出书面语到口语的转换规则(从句→短句、被动→主动、术语→日常词汇),提供 3 个以上转换示例。

B — 边界 (Boundary) ★

不要在以下情况使用

  • 纯文本聊天界面,用户通过键盘输入和屏幕阅读
  • 语音合成(TTS)引擎的技术选型或参数调优
  • 音频信号处理(降噪、回声消除等)
  • 多模态场景中语音仅为辅助通道(如视频会议中的字幕场景)

Read the full file on GitHub · 75 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 9d ago First seen · 75 lines · 84 tokens per session scan A 70dc18eec11d

Subscribe to this mod's changes

voice-optimization is a skill published in the GitHub repository kangarooking/system-prompt-skills (182 stars, last pushed 4mo ago), licensed MIT. It adds 84 tokens to every session and 1,668 once invoked, about $0.0004 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.