A programmable Mixture-of-Models router for heterogeneous LLM inference
vLLM Semantic Router is a programmable routing layer that chooses or combines language models for each request in a system using multiple models and types of computing infrastructure. It helps teams route inference by signals such as user preferences, application policies, quality, cost, latency, privacy, and safety requirements. The catalogue skills and instructions support configuring and operating this model-routing system.
Latest release v0.3.0 — Release v0.3.0 · 5 Jun 2026
These files are vllm-project/semantic-router's own configuration. They tell GitHub Copilot, Codex and OpenCode how to work on this repository, so they are not mods to install elsewhere. Copy one as a starting point and replace the parts that are about this project.
.github/copilot-instructions.md A 304 tok AGENTS.md A 2,046 tok