voice-agents

voice-agents is a skill for Claude Code, Codex from davila7/claude-code-templates. It costs 103 tokens per session (522 once invoked), scanned A, original, MIT.

A guide to building AI agents that people can talk to naturally by voice. It covers speech recognition, language-model responses, speech synthesis, and turn-taking.

In plain words
What is it for?
Use it to choose between direct speech-to-speech and separate speech-to-text, AI, and text-to-speech pipelines, then design voice interactions.
Why use it?
It helps address awkward pauses, interruptions, background noise, and the added delay caused by processing speech in several stages.

Skill for Claude CodeCodex

Written for no agent in particular: nothing here depends on one.

Good fit Use it to choose between direct speech-to-speech and separate speech-to-text, AI, and text-to-speech pipelines, then design voice interactions.

Compare 6 skills from other repositories ↓
Install with agentmods
npx agentmods add skills/davila7/claude-code-templates/voice-agents
About the project

Claude Code Templates is a command-line tool and catalogue for configuring Anthropic’s Claude Code with agents, commands, settings, hooks, integrations, skills, and project templates. Developers use it to browse and install reusable components for their coding workflows. The catalogue includes many of these Claude Code components.

davila7/claude-code-templates · 30,576 stars · on GitHub · aitmpl.com

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

Any agent
npx skills add davila7/claude-code-templates --skill voice-agents
Clone the repo
git clone --depth 1 https://github.com/davila7/claude-code-templates

Made for: Claude Code, Codex.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for voice-agents

README.md
[![agentmods](https://agentmods.dev/badge/skills/davila7/claude-code-templates/voice-agents/github.svg)](https://agentmods.dev/skills/davila7/claude-code-templates/voice-agents)
Your own site
<a href="https://agentmods.dev/skills/davila7/claude-code-templates/voice-agents"><img src="https://agentmods.dev/badge/skills/davila7/claude-code-templates/voice-agents/github.svg" alt="Measured on agentmods" height="20"></a>

Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.

agentmods 80×15 button for voice-agents

Your own site · 80×15
<a href="https://agentmods.dev/skills/davila7/claude-code-templates/voice-agents"><img src="https://agentmods.dev/badge/skills/davila7/claude-code-templates/voice-agents.svg" alt="Reviewed on agentmods" width="80" height="20"></a>
Per session 103 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 522 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. A grade says what 26 rules found in the file — not that it is safe. Third-party audits
  • Socket pass 18 Mar 2026
  • Snyk pass 15 Feb 2026
  • NVIDIA SkillSpector pass 7 Sept 2026
How audits are shown
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5.1 $0.00103 $0.00522
Opus 5 $0.00051 $0.00261
Sonnet 5 $0.00021 $0.00104
Haiku 4.5 $0.00010 $0.00052

Measured 6d ago against content hash c31bb21df6e1, method: parsed. Prices are Anthropic first-party input rates as of 2026-09-10, from the pricing page.

Security

Grade A, and why

voice-agents scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 6d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

cli-tool/components/skills/ai-research/voice-agents/SKILL.md · 69 lines

What it actually says

Voice Agents

You are a voice AI architect who has shipped production voice agents handling millions of calls. You understand the physics of latency - every component adds milliseconds, and the sum determines whether conversations feel natural or awkward.

Your core insight: Two architectures exist. Speech-to-speech (S2S) models like OpenAI Realtime API preserve emotion and achieve lowest latency but are less controllable. Pipeline architectures (STT→LLM→TTS) give you control at each step but add latency. Mos

Capabilities

  • voice-agents
  • speech-to-speech
  • speech-to-text
  • text-to-speech
  • conversational-ai
  • voice-activity-detection
  • turn-taking
  • barge-in-detection
  • voice-interfaces

Patterns

Speech-to-Speech Architecture

Direct audio-to-audio processing for lowest latency

Pipeline Architecture

Separate STT → LLM → TTS for maximum control

Voice Activity Detection Pattern

Detect when user starts/stops speaking

Anti-Patterns

❌ Ignoring Latency Budget

❌ Silence-Only Turn Detection

❌ Long Responses

⚠️ Sharp Edges

Issue Severity Solution
Issue critical # Measure and budget latency for each component:
Issue high # Target jitter metrics:
Issue high # Use semantic VAD:
Issue high # Implement barge-in detection:
Issue medium # Constrain response length in prompts:
Issue medium # Prompt for spoken format:
Issue medium # Implement noise handling:
Issue medium # Mitigate STT errors:

Works well with: agent-tool-builder, multi-agent-orchestration, llm-architect, backend

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 6d ago First seen · 69 lines · 103 tokens per session scan A c31bb21df6e1

Subscribe to this mod's changes

voice-agents is a skill published in the GitHub repository davila7/claude-code-templates (30,576 stars, last pushed yesterday), licensed MIT. It adds 103 tokens to every session and 522 once invoked, about $0.0005 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-09-03.

Related

Other skills, from other repositories

ccc-prompt-fix

Fix and sharpen a prompt. Diagnoses it against the 6 prompt-quality patterns, returns a tightened rewrite with the reasoning, and suggests the right library prompt for your task.

KevinZai/commander · 41 tokens

fabric-mlv

Use for Fabric Materialized Lake Views (MLVs) — CREATE MATERIALIZED LAKE VIEW Spark SQL (GA March 2026) + still-preview @fmlv.materializedlakeview PySpark decorator on a schema-enabled lakehouse (Runtime 1.3). Covers CREATE / SHOW / ALTER RENAME / DROP / REFRESH FULL syntax, CONSTRAINT ... CHECK ... ON MISMATCH…

wardawgmalvicious/agent-config · 245 tokens

ccc-data

For large datasets and data files, the Files API can ingest CSVs, JSON, Parquet, and other formats directly — avoiding token limits for bulk data analysis. Use data-ingestion from ccc-research for document-scale inputs.

KevinZai/commander · 35 tokens

fabric-ai-functions

Use for Microsoft Fabric AI Functions (Data Science) — one-line LLM transformations on pandas and PySpark DataFrames in Fabric notebooks: ai.analyzesentiment, ai.classify, ai.extract, ai.embed, ai.summarize, ai.translate, ai.fixgrammar, ai.generateresponse, ai.similarity. Covers the two import paths (synapse.ml.aifunc…

wardawgmalvicious/agent-config · 284 tokens

fabric-data-agent

Use when configuring Microsoft Fabric Data Agents (GA March 2026) — conversational Q&A over Lakehouse / Warehouse / KQL / Semantic Model / Fabric SQL DB / Mirrored DB / Ontology / MS Graph (≤5 sources per agent), consumed in-product or via the agent's MCP endpoint (Assistants API and Copilot-in-Power-BI paths retired…

wardawgmalvicious/agent-config · 230 tokens

fabric-semantic-model-ai-instructions

Use when configuring AI instructions on a Power BI semantic model — the 10,000-character blob attached via Prep data for AI → Add AI instructions in Desktop or the service. Applies everywhere Copilot uses the model (reports, Q&A, Copilot pane). Covers what belongs in the blob (business context, terminology, date…

wardawgmalvicious/agent-config · 154 tokens