horizon-router

horizon-router is a skill for Claude Code, Codex from HorizonRobotics/OE-Skills. It costs 46 tokens per session (5,815 once invoked), scanned A, original, Apache-2.0.

A routing skill for OpenExplorer and Horizon tools used to quantize, compile, deploy, and evaluate machine-learning models on Horizon hardware. It selects the relevant sub-skill and defines rules for complete deployment workflows.

In plain words
What is it for?
Use it for PTQ or QAT quantization, model compilation, board deployment, performance or accuracy evaluation, and workflows involving ONNX, BC, HBM, or PT model files.
Why use it?
It helps send each model task to the right tool while keeping settings and deliverables consistent across quantization, compilation, and deployment.

Skill for Claude CodeCodex

Written for no agent in particular: nothing here depends on one. Also seen: mentions subagents.

Good fit Use it for PTQ or QAT quantization, model compilation, board deployment, performance or accuracy evaluation, and workflows involving ONNX, BC, HBM, or PT model files.

Compare 6 skills from other repositories ↓
Install with agentmods
npx agentmods add skills/horizonrobotics/oe-skills/horizon-router
Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

Any agent
npx skills add HorizonRobotics/OE-Skills --skill horizon-router
Clone the repo
git clone --depth 1 https://github.com/HorizonRobotics/OE-Skills

Made for: Claude Code, Codex.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for horizon-router

README.md
[![agentmods](https://agentmods.dev/badge/skills/horizonrobotics/oe-skills/horizon-router/github.svg)](https://agentmods.dev/skills/horizonrobotics/oe-skills/horizon-router)
Your own site
<a href="https://agentmods.dev/skills/horizonrobotics/oe-skills/horizon-router"><img src="https://agentmods.dev/badge/skills/horizonrobotics/oe-skills/horizon-router/github.svg" alt="Measured on agentmods" height="20"></a>

Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.

agentmods 80×15 button for horizon-router

Your own site · 80×15
<a href="https://agentmods.dev/skills/horizonrobotics/oe-skills/horizon-router"><img src="https://agentmods.dev/badge/skills/horizonrobotics/oe-skills/horizon-router.svg" alt="Reviewed on agentmods" width="80" height="20"></a>
Per session 46 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 5,815 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. A grade says what 26 rules found in the file — not that it is safe.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5.1 $0.00046 $0.05815
Opus 5 $0.00023 $0.02908
Sonnet 5 $0.00009 $0.01163
Haiku 4.5 $0.00005 $0.00581

Measured 12d ago against content hash fc16f8d4ae06, method: parsed. Prices are Anthropic first-party input rates as of 2026-09-11, from the pricing page.

Security

Grade A, and why

horizon-router scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 12d ago.

The scan reads SKILL.md. This mod also ships 2 executable files (oe-llm-package-install/install.sh, oe-package-install/install.sh), listed below but not scanned — reading those needs a real analyzer, not pattern matching.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

horizon/skills/horizon-router/SKILL.md · 329 lines

How it starts

The opening of the file, as written. The whole thing — 329 lines — stays where its author put it; the contents beside it link to each section on GitHub.

Horizon Router

OpenExplorer / Horizon 工具链的顶层路由入口。

当请求涉及量化、编译、部署、板端推理、性能/精度评估,或用户提供了 .onnx.bc.hbm.pt 等模型文件时,从本 Skill 进入。

路由前必须先读取 .horizon/skill-index.json,通过其中每个 Skill 的 description 字段理解能力边界,再决定路由到哪个子 Skill。不要凭名字猜测。


⛔ 关键规则(最高优先级)

全链路部署规范优先于子 Skill

当用户需求涉及「量化 → 编译 → 部署」完整链路时,references/deployment-workflow.md 中的全链路部署规范是最高权威。如果子 Skill 的默认行为与全链路规范冲突,以全链路规范为准。常见冲突场景:

全链路规范要求 子 Skill 默认值 处理方式
calibration_type: histogram HMCT 默认 max 按全链路规范,使用 histogram
all_node_type: float16(nash-p) HMCT 默认 int8 按全链路规范,使用 fp16 + conv int8
remove_node_type: [Quantize, Dequantize] hbdk-compile 默认 [Quantize] 按全链路规范,同时删除两者
部署交付物 = UCP 推理代码 hbm_infer SDK 即可完成验证 按全链路规范,UCP 代码才是部署交付物

原则:子 Skill 服务于单步操作,全链路规范服务于端到端目标。端到端任务中,单步的"合理默认"可能不符合全链路要求。

量化配置默认原则

除非用户明确要求混合精度调优,或当前任务已通过评测确认全 int8 精度不达标,否则涉及量化配置的任务(QAT 适配、导出、全流程代码生成)应默认使用全 int8 配置。

  • 禁止在没有精度不达标证据的情况下,主动将算子升高到 int16 或 fp16
  • 如果全 int8 精度不达标,应先路由到 j6-plugin-precision-tuning,按敏感度分析结果决定哪些算子需要升高精度,而不是凭经验预设混合精度
  • 用户明确说"用混合精度"或"int8 不够"时,才跳过全 int8 默认

长时间任务的等待策略

当 agent 启动了耗时较长的后台任务(敏感度分析、HBDK 编译、QAT 训练等,通常 >3 分钟)时,必须使用以下策略之一等待完成,禁止反复轮询:

策略 A:后台启动 + 等待通知(推荐)
1. 编写包含完整处理逻辑的脚本(脚本自身负责生成最终结果文件)
2. 使用 Bash run_in_background: true 启动脚本
3. 回复"任务已在后台运行,等待完成通知",然后停止——不执行任何进度检查
4. 收到系统自动通知后,读取脚本输出的结果文件
策略 B:单次超时等待

如果必须用 wait 或前台运行,设置足够长的 timeout(如 600000ms),一次等到结束

# 一次等到完成,不中途检查
python3 long_running_script.py  # timeout: 600000
⛔ 禁止行为(会导致 API 崩溃)
# ❌ 以下模式会触发重复调用检测,导致 400 错误终止:
Bash: tail -c 1000 tuning_run.log   # 第1次
Bash: tail -c 1000 tuning_run.log   # 第2次
Bash: tail -c 1000 tuning_run.log   # 第3次 → 崩溃!

# ❌ 即使参数微调也会被检测:
Bash: tail -n 5 tuning_run.log
Bash: tail -n 10 tuning_run.log
Bash: wc -l tuning_run.log          # 仍然可能触发

原因:API 层面会检测短时间内相似的工具调用。连续 3 次语义相近的命令即可能触发保护机制。后台任务完成时系统会自动通知,无需主动检查。

连续失败时的策略切换

当同一操作连续失败 2 次时,必须切换策略,禁止继续重试同一方法:

Read the full file on GitHub · 329 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 12d ago First seen · 329 lines · 46 tokens per session scan A fc16f8d4ae06

Subscribe to this mod's changes

horizon-router is a skill published in the GitHub repository HorizonRobotics/OE-Skills (19 stars, last pushed 2mo ago), licensed Apache-2.0. It adds 46 tokens to every session and 5,815 once invoked, about $0.0002 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.

Related

Other skills, from other repositories

spark-environment-setup

Set up a working ML training/inference environment on NVIDIA DGX Spark (GB10, aarch64, CUDA 13). Use when installing PyTorch/Unsloth/TRL/vLLM on DGX Spark, hitting libcudart or wheel-ABI errors on aarch64, or choosing between NGC containers and bare pip installs.

wshobson/agents · 76 tokens

spark-memory-thermal-ops

Manage unified memory and thermals during long-running ML jobs on NVIDIA DGX Spark. Use when planning memory headroom for a training run on GB10, when a job OOMs on unified memory, or when monitoring temperature and power during multi-hour training.

wshobson/agents · 59 tokens

spark-training-gotchas

Preflight and diagnose the ten known failure modes for ML training on NVIDIA DGX Spark. Use when a training run on DGX Spark fails to start, OOMs below the 128GB limit, slows down mid-run, or before any multi-hour training job on GB10.

wshobson/agents · 63 tokens

llama-cpp

Runs LLM inference on CPU, Apple Silicon, and consumer GPUs without NVIDIA hardware. Use for edge deployment, M1/M2/M3 Macs, AMD/Intel GPUs, or when CUDA is unavailable. Supports GGUF quantization (1.5-8 bit) for reduced memory and 4-10× speedup vs PyTorch on CPU.

davila7/claude-code-templates · 76 tokens

minicpm5-deploy-vllm-ascend

Deploy MiniCPM5-2B with vLLM on Huawei Ascend NPU using vLLM-Ascend. Use when the user mentions vLLM-Ascend, Ascend NPU, Huawei Ascend, CANN, torchnpu, davinci devices, or wants an OpenAI-compatible MiniCPM5 server on Ascend hardware.

OpenBMB/MiniCPM · 87 tokens

minicpm5-deploy-litert

Run MiniCPM5-2B or MiniCPM5-1B on-device with Google's LiteRT-LM runtime — the litert-lm CLI or its OpenAI-compatible server on a desktop, the Kotlin API or the AI Edge Gallery app on Android, the same .litertlm bundle on CPU or GPU. Use when the user says "LiteRT", "LiteRT-LM", "litertlm", ".litertlm", "Android"…

OpenBMB/MiniCPM · 121 tokens