Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add kayzaa/k.i.t.-bot --skill model-failovergit clone --depth 1 https://github.com/kayzaa/k.i.t.-botWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/kayzaa/k.i.t.-bot/model-failover)<a href="https://agentmods.dev/skills/kayzaa/k.i.t.-bot/model-failover"><img src="https://agentmods.dev/badge/skills/kayzaa/k.i.t.-bot/model-failover/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/kayzaa/k.i.t.-bot/model-failover"><img src="https://agentmods.dev/badge/skills/kayzaa/k.i.t.-bot/model-failover.svg" alt="Reviewed on agentmods" width="80" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00000 | $0.01581 |
| Opus 5 | $0.00000 | $0.00790 |
| Sonnet 5 | $0.00000 | $0.00316 |
| Haiku 4.5 | $0.00000 | $0.00158 |
Grade A, and why
model-failover scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 7d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
This is a copy
100% identical to model-failover — 0 lines differ, which has more behind it and is treated as the original. This page carries a canonical link to it rather than competing with it.
How it starts
The opening of the file, as written. The whole thing — 220 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Skill #85: Model Failover Manager
Enterprise-grade AI provider rotation, cooldowns, and automatic failover for K.I.T.'s trading decisions.
Why Model Failover?
AI providers have rate limits, outages, and billing issues. K.I.T. needs:
- 24/7 uptime for autonomous trading
- Automatic recovery from provider failures
- Cost optimization across providers
- Quality maintenance when switching models
Features
Multi-Provider Support
| Provider | Models | Rate Limit Handling |
|---|---|---|
| Anthropic | Claude Opus, Sonnet, Haiku | Per-minute, per-day |
| OpenAI | GPT-4o, GPT-4-turbo, o1 | TPM, RPM |
| Gemini 2.0, 1.5 Pro | Per-minute | |
| xAI | Grok 2 | Per-minute |
| DeepSeek | DeepSeek V3, R1 | Per-minute |
| Groq | Llama, Mixtral | Per-minute, free tier |
| OpenRouter | All models | Aggregated |
| Local | Ollama, vLLM | No limits |
Failover Strategies
1. Round-Robin Rotation Distributes load across providers:
Request 1 → Anthropic
Request 2 → OpenAI
Request 3 → Google
Request 4 → Anthropic (back to start)
2. Priority Cascade Falls back through priority order:
Primary: Claude Opus 4
Fallback1: GPT-4o
Fallback2: Gemini 2.0
Fallback3: Local Ollama (never fails)
3. Cost-Optimized Routes to cheapest available provider:
Simple queries → Haiku ($0.25/1M)
Complex analysis → Sonnet ($3/1M)
Critical decisions → Opus ($15/1M)
4. Latency-Optimized Tracks response times and routes to fastest:
Groq: ~200ms (when available)
Anthropic: ~800ms
OpenAI: ~1200ms
Cooldown System
Exponential backoff for failures:
1st failure: 1 minute cooldown
2nd failure: 5 minutes
3rd failure: 25 minutes
4th+ failure: 1 hour (cap)
Separate handling for:
- Rate limits: Standard cooldown + retry
- Auth errors: Immediate failover, longer cooldown
- Billing issues: 5-hour initial backoff, doubles each time
- Timeouts: Shorter cooldown (30 seconds)
What ships with it
1 file beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 7d ago First seen · 220 lines · 0 tokens per session scan A 8bbf24e894c1
model-failover is a skill published in the GitHub repository kayzaa/k.i.t.-bot (5 stars, last pushed 6mo ago), licensed MIT. It costs nothing until one of its globs matches a file; then it loads 1,581 tokens. A static security scan graded it A with 0 findings. It is 100% identical to model-failover, differing in 0 lines, and is treated as a copy.
Other skills, from other repositories
embeddings
Vector embeddings configuration and semantic search.
BytesAgain Crypto Toolkit — 200+ Technical Indicators, Real-Time Market Data
Use when you need real-time crypto prices, technical indicators (RSI, MACD, Bollinger, 50+), market rankings, on-chain data, or trading signals. Zero API key required.
add-prompt
Scaffold a new MCP prompt template. Use when the user asks to add a prompt, create a reusable message template, or define a prompt for LLM interactions.
public
For agents: This document explains how to integrate with Clodds APIs.
feeds
Real-time market data feeds from 8 prediction market platforms.
integrations
External data sources, connectors, and custom data streams.