OmniRoute is an AI gateway that gives coding agents one endpoint for accessing models from many providers, with quota-aware fallback between them. It is for people using Claude Code, Codex, Cursor, OpenCode, Cline, Copilot, and similar tools who want a single route to their available AI models.
Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add skills/diegosouzapw/omniroute/omni-settingsnpx skills add diegosouzapw/OmniRoute --skill omni-settingsgit clone --depth 1 https://github.com/diegosouzapw/OmniRouteWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/diegosouzapw/omniroute/omni-settings)<a href="https://agentmods.dev/skills/diegosouzapw/omniroute/omni-settings"><img src="https://agentmods.dev/badge/skills/diegosouzapw/omniroute/omni-settings.svg" alt="Measured on agentmods" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00031 | $0.08803 |
| Opus 5 | $0.00015 | $0.04401 |
| Sonnet 5 | $0.00006 | $0.01761 |
| Haiku 4.5 | $0.00003 | $0.00880 |
Grade A, and why
omni-settings scanned grade A with 1 finding against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Makes network callslowCapability
Not a fault in itself. Listed so you know the mod talks to something, and to what.
curl https://localhost:20128/api/settings/memory \ Copies of this mod
1 near-identical copy found in the catalogue:
- omni-settings — 94% identical, 1,034 lines differ
How it starts
The opening of the file, as written. The whole thing — 1,391 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Overview
Read and update global application settings: system prompts, thinking budget, IP filters, payload rules, combo defaults, and require-login configuration.
Authentication
All requests require a valid Bearer token or session cookie. Obtain a token via POST /api/auth/login or configure REQUIRE_API_KEY=false for local development.
Endpoints
GET /api/settings/memory
Get memory settings
Returns the extended memory settings including 7 new fields added in plan 21 (embeddingSource, embeddingProviderModel, transformersEnabled, staticEnabled, rerankEnabled, rerankProviderModel, vectorStore).
curl https://localhost:20128/api/settings/memory \
-H "Authorization: Bearer $OMNIROUTE_TOKEN"
PUT /api/settings/memory
Update memory settings
Update any subset of the extended memory settings. All fields are optional; only provided fields are updated. Schema: MemorySettingsExtendedSchema.
curl -X PUT https://localhost:20128/api/settings/memory \
-H "Authorization: Bearer $OMNIROUTE_TOKEN" \
-H "Content-Type: application/json" \
-d '{}'
GET /api/settings/qdrant
Get Qdrant settings
Returns current Qdrant configuration. The apiKey field is never returned raw — use hasApiKey / apiKeyMasked instead.
curl https://localhost:20128/api/settings/qdrant \
-H "Authorization: Bearer $OMNIROUTE_TOKEN"
PUT /api/settings/qdrant
Update Qdrant settings
Update Qdrant configuration. Pass apiKey: "" to remove the stored key. Schema: QdrantSettingsUpdateSchema.
curl -X PUT https://localhost:20128/api/settings/qdrant \
-H "Authorization: Bearer $OMNIROUTE_TOKEN" \
-H "Content-Type: application/json" \
-d '{}'
GET /api/settings/qdrant/health
Qdrant health probe
Performs a liveness check against the configured Qdrant instance. Returns latency and any connection error (sanitized — no stack traces).
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 2d ago Changed · +975 lines 338ca2023928
- 5d ago First seen · 416 lines · 31 tokens per session scan A 769874473ebd
omni-settings is a skill published in the GitHub repository diegosouzapw/OmniRoute (59,841 stars, last pushed 2d ago), licensed MIT. It adds 31 tokens to every session and 8,803 once invoked, about $0.0002 per session on Opus 5. A static security scan graded it A with 1 finding (makes network calls). No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other skills, from other repositories
workflow-ai-coding
Edit, validate, debug, publish, and inspect ReachAI Workflow drafts through the Workflow AI Coding REST API. Use when asked to create or modify a workflow graph, add/update/delete nodes or edges, validate GraphSpec, dry-run or debug-run a workflow, inspect trace/run output, check release readiness, publish a validated…
reachai-onboarding
Integrate Java business systems with ReachAI SDK registration, SDK instance heartbeat, gateway/embed access, and optional API Management handoff. Use when asked to connect a Spring Boot service to ReachAI, add reachai-capability-sdk or reachai-spring-boot2-starter, configure…
9router-web-fetch
Fetch URL → markdown / text / HTML via 9Router /v1/web/fetch using Ollama Cloud / Firecrawl / Jina Reader / Tavily Extract / Exa Contents. Use when the user wants to scrape a webpage, extract URL content, read article, or convert a URL to markdown.
9router-web-search
Web and X search via 9Router /v1/search using Tavily / Exa / Brave / Serper / SearXNG / Google PSE / Linkup / SearchAPI / You.com / Perplexity / Xquik. Use when the user wants to search the web, find articles, or search public X posts.
9router-embeddings
Generate vector embeddings via 9Router /v1/embeddings using OpenAI / Gemini / Mistral / Voyage / Nvidia / GitHub embedding models for RAG, semantic search, similarity. Use when the user wants embeddings, vectors, RAG, semantic search, or to embed text.
9router-stt
Speech-to-text via 9Router /v1/audio/transcriptions using OpenAI Whisper / Groq / Gemini / Deepgram / AssemblyAI / NVIDIA / HuggingFace models. Use when the user wants to transcribe audio, convert speech to text, or get subtitles from audio files.