Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add instructions/sandraschi/speech-mcp/gemini-mdgit clone --depth 1 https://github.com/sandraschi/speech-mcpWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/instructions/sandraschi/speech-mcp/gemini-md)<a href="https://agentmods.dev/instructions/sandraschi/speech-mcp/gemini-md"><img src="https://agentmods.dev/badge/instructions/sandraschi/speech-mcp/gemini-md.svg" alt="Measured on agentmods" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00532 | $0.00532 |
| Opus 5 | $0.00266 | $0.00266 |
| Sonnet 5 | $0.00106 | $0.00106 |
| Haiku 4.5 | $0.00053 | $0.00053 |
Grade A, and why
speech-mcp GEMINI.md scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 5d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
What it actually says
Speech-MCP: System Context
This document provides grounding for the AI agent. Refer to this before performing architectural changes, debugging, or tool development.
🛠 Tech Stack
- Backend: Python 3.12 + FastMCP + FastAPI
- Frontend: React + Vite + Tailwind v4 + Biome (Linting)
- Data Layer:
- RAG: LanceDB + FastEmbed (
bge-small-en-v1.5) - Storage: JSON-based state in
data/
- RAG: LanceDB + FastEmbed (
- Voice Stack:
- TTS: Gemini 3.1 Flash (Priority), Hume Octave, ElevenLabs, Windows SAPI5
- STT: Gemini 3.1 Multimodal Live
- Wake Word: openWakeWord (Offline ONNX Models)
📡 Registry & Ports
- Backend API:
http://localhost:10909 - Frontend UI:
http://localhost:10908 - WebSocket Gateway:
ws://localhost:10909/ws/stream - Telemetry Stream:
ws://localhost:10909/ws/logs(JSON Payload)
🏗 Directory Structure
/src/speech_mcp: Core Python logic and MCP tool definitionsserver.py: API and MCP entry pointstreaming.py: WebSocket audio handlerstools/: Modular tool implementationswake_word.py: openWakeWord engine integrationspeech.py: TTS / Dialogue synthesisrag.py: LanceDB knowledge base
/web: React frontend/docs: Detailed technical and user documentation/scripts/demos: Standalone demo scripts for provider testing
🧠 Development Standards
- Neutral Branding: Component labels and logs should use technical descriptors (e.g., "Wake-Word Detection"). Avoid redundant status markers like "Alpha" in functional names.
- Zero-Diagnostic Policy: Maintain zero Ruff/Biome warnings. Always run
just fixbefore finishing a task. - Semantic HTML: All interactive elements must have unique, descriptive IDs for automated testing.
- Tool Integrity: Preserve existing docstrings and FastMCP decorators.
Reference global patterns in D:\Dev\repos\mcp-central-docs\patterns
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 5d ago First seen · 42 lines · 532 tokens per session scan A 20dc5d66745c
speech-mcp GEMINI.md is an instructions file published in the GitHub repository sandraschi/speech-mcp (2 stars, last pushed 5d ago), licensed MIT. It adds 532 tokens to every session, about $0.0027 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other instructions, from other repositories
sayna CLAUDE.md
Instructions for SaynaAI/sayna, covering claude.md, project overview, development commands, feature flags and high-level architecture.
cadence-code AGENTS.md
AGENTS.md instructions for michael-L-i/cadence-code, covering agent guide, project overview, plugin layout, important paths and local commands.
cadence-code CLAUDE.md
Claude Code instructions for michael-L-i/cadence-code, covering claude code guide, project summary, claude-specific flow and working expectations.
noisy-coding CLAUDE.md
Claude Code instructions for noisy/noisy-coding, covering noisy-coding — agent notes, local development setup, key docs, releasing and restarting the daemon.
AI_Secretary_System CLAUDE.md
Claude Code instructions for ShaerWare/AI_Secretary_System, covering claude.md, project overview, commands, build & run and docker (recommended).
cc-hooks CLAUDE.md
Instructions for husniadil/cc-hooks, covering claude.md, documentation structure, best practices, pep 723 for standalone scripts and /// script.