vllm-project/semantic-router

A programmable Mixture-of-Models router for heterogeneous LLM inference

About the project

vLLM Semantic Router is a programmable routing layer that chooses or combines language models for each request in a system using multiple models and types of computing infrastructure. It helps teams route inference by signals such as user preferences, application policies, quality, cost, latency, privacy, and safety requirements. The catalogue skills and instructions support configuring and operating this model-routing system.

Latest release v0.3.0 — Release v0.3.0 · 5 Jun 2026

These files are vllm-project/semantic-router's own configuration. They tell GitHub Copilot, Codex and OpenCode how to work on this repository, so they are not mods to install elsewhere. Copy one as a starting point and replace the parts that are about this project.

5,600Stars on the repository
2Files it configures its agents with
2,350Tokens loaded in every session
3Agents configured

Instructions