Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add calesthio/generative-media-skills --skill hume-evigit clone --depth 1 https://github.com/calesthio/generative-media-skillsWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/calesthio/generative-media-skills/hume-evi)<a href="https://agentmods.dev/skills/calesthio/generative-media-skills/hume-evi"><img src="https://agentmods.dev/badge/skills/calesthio/generative-media-skills/hume-evi/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/calesthio/generative-media-skills/hume-evi"><img src="https://agentmods.dev/badge/skills/calesthio/generative-media-skills/hume-evi.svg" alt="Reviewed on agentmods" width="80" height="20"></a>- NVIDIA SkillSpector warn
SkillSpector: 1 finding, up to medium
These are SkillSpector’s own severities. On a checked sample its high-severity flags on skills were ~96% false positives — a documented command, a public API, a “never do X” rule — so we show them as a caution to read, not a verdict. Why →
- medium analysis-evasion · line 1 Suspicious Unicode normalization or mixed-script contentFix: Review the flagged content for security risks. Ensure no credentials, secrets, or sensitive data are exposed.
What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00085 | $0.06802 |
| Opus 5 | $0.00043 | $0.03401 |
| Sonnet 5 | $0.00017 | $0.01360 |
| Haiku 4.5 | $0.00009 | $0.00680 |
Grade A, and why
hume-evi scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 9d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 479 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Hume EVI production guide
Use this skill when the task is to design, implement, troubleshoot, or review a Hume Empathic Voice Interface (EVI) realtime voice agent. Do not use it for offline-only Hume Text-to-Speech/Octave work unless the user is explicitly configuring EVI voices or comparing EVI to TTS.
Hume's EVI is a realtime speech-to-speech agent interface. It streams user audio, measures expressive vocal modulation, generates a language response, and produces expressive assistant speech. Treat it as a live conversation system, not as a batch transcription or batch TTS API. Documented facts below were verified on 2026-07-10 unless a source date is explicitly noted.
Primary official sources:
- Hume EVI overview: https://dev.hume.ai/docs/speech-to-speech-evi/overview
- EVI version guide: https://dev.hume.ai/docs/speech-to-speech-evi/configuration/evi-version
- Configuration guide: https://dev.hume.ai/docs/speech-to-speech-evi/configuration/build-a-configuration
- Chat WebSocket API reference: https://dev.hume.ai/reference/speech-to-speech-evi/chat
- Audio guide: https://dev.hume.ai/docs/speech-to-speech-evi/guides/audio
- Tool use guide: https://dev.hume.ai/docs/speech-to-speech-evi/features/tool-use
- Prompting guide: https://dev.hume.ai/docs/speech-to-speech-evi/guides/prompting
- Privacy controls: https://dev.hume.ai/docs/resources/privacy
- Pricing page: https://www.hume.ai/pricing
- Acceptable Use Policy: https://www.hume.ai/acceptable-use-policy
Capability boundaries
Documented facts:
- EVI is for realtime voice interaction. The Chat WebSocket accepts streamed
audio_input,session_settings,user_input,assistant_input, and tool response messages; audio input must be streamed continuously in small chunks, not sent as a whole prerecorded file. Hume recommends roughly 20 ms audio buffers generally or 100 ms for web apps. Source: https://dev.hume.ai/reference/speech-to-speech-evi/chat - Direct browser or app integration is typically WebSocket-based. Hume's React SDK abstracts microphone capture, playback, and WebSocket connection management; the TypeScript and Python SDKs expose lower-level integration paths. Source: https://dev.hume.ai/docs/speech-to-speech-evi/quickstart/nextjs and https://dev.hume.ai/docs/speech-to-speech-evi/quickstart/typescript
- Use the Python SDK primarily for CLIs and desktop apps where the Python process can access the user's audio device. For hosted web apps, put audio capture in the browser; server-side Python cannot directly access the user's microphone and routing audio through a backend adds latency. Source: https://dev.hume.ai/docs/speech-to-speech-evi/quickstart/python
- EVI can use Hume Voice Library voices and account-private Custom Voices. The EVI voice can be specified in a persistent config or overridden for a session with a voice ID. Custom voice cloning requires rights/consent to the voice sample. Sources: https://dev.hume.ai/docs/speech-to-speech-evi/configuration/voice and https://dev.hume.ai/docs/voice/voice-cloning
- EVI supports tools/function calling through supported supplemental LLMs, not through every native EVI-only configuration. Hume documents tool use with Claude, GPT, Gemini, Moonshot AI, and custom language models that follow OpenAI function-calling conventions. Source: https://dev.hume.ai/docs/speech-to-speech-evi/features/tool-use
- Expression measurements are available from audio user messages, not text-only
user_input, because the prosody model relies on audio input. Source: https://dev.hume.ai/reference/speech-to-speech-evi/chat - Chat history and reconstructed audio are retrievable only when data retention is enabled. If data retention is disabled, resume chats, chat history retrieval, and audio reconstruction are not supported. Sources: https://dev.hume.ai/docs/speech-to-speech-evi/features/resume-chats and https://dev.hume.ai/docs/speech-to-speech-evi/features/audio-reconstruction
What ships with it
1 file beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 9d ago First seen · 479 lines · 85 tokens per session scan A f09ce699bce0
hume-evi is a skill published in the GitHub repository calesthio/generative-media-skills (171 stars, last pushed 2mo ago), licensed MIT. It adds 85 tokens to every session and 6,802 once invoked, about $0.0004 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-09-03.
Other skills, from other repositories
cliptalk-cover-director
Produces evidence-backed cover candidates and reviewable cover variants for a ClipTalk video. Use when the user asks for a cover, poster frame, thumbnail, or multiple cover directions; do not use for timeline editing or social-video reframing.
cliptalk-smart-reframe
Creates a subject-aware, time-varying crop track and a review-only social-format preview from an accepted ClipTalk cut. Use for automatic vertical, square, or portrait reframing; do not use for a fixed manual crop or before content editing is accepted.
cliptalk-content-extractor
Locates and assembles source passages matching a semantic request. Use for extracting explanations, topics, quotes, demonstrations, or other specifically described content.
cliptalk-interview-editor
Produces a coherent interview edit by combining speaker discovery, topic selection, dialogue context, cleanup, subtitles, and preview. Use for interviews, podcasts, testimonials, or question-and-answer recordings.
cliptalk-shortform-hook-director
Finds and assembles a reviewable short-form cut with a strong opening hook. Use for Shorts, Reels, social clips, talking-head cutdowns, or requests for a punchier opening.
cliptalk-social-reframe-exporter
Creates a review-only 9:16, 4:5, 1:1, or 16:9 version from an existing accepted ClipTalk cut, then checks the rendered preview. Use only when a cut already exists and the user asks to adapt it for Shorts, Reels, Douyin, Xiaohongshu, WeChat Channels, or square feeds; do not use when the user still needs content found…