vllm-project

21 mods across 2 repositories, 5.5k stars between them.

vllm-project/semantic-router

Instructions file CodexOpenCode

Instructions for vllm-project/semantic-router, covering vllm semantic router agent entry, read first, task routing, repository map and supported environments.

5.4k 2d ago A 1,993 tokens original Apache-2.0

openclaw-vsr-bridge

03

vllm-project/semantic-router

Skill Claude CodeCodex

Install vLLM Semantic Router in agent-safe mode, import supported OpenClaw model providers into canonical VSR config, and rewrite OpenClaw to target VSR.

5.4k 2d ago C 42 tokens original Apache-2.0

vllm-project/semantic-router

Skill Claude CodeCodex

Synchronizes config representations across router config, Python CLI schema, and dashboard config UI. Use when adding or changing a config concept that spans those surfaces or addressing config representation debt before Kubernetes-facing translation.

5.4k 2d ago A 43 tokens original Apache-2.0

vllm-project/semantic-router

Skill Claude CodeCodex

Modifies the repository's agent contract including AGENTS.md, docs index, manifests, validation scripts, and contributor-facing harness wrappers. Use when updating agent documentation, changing repo manifests, editing validation scripts, modifying CI/workflow classification, or updating contributor-facing guides like…

5.4k 2d ago A 69 tokens original Apache-2.0

vllm-project/semantic-router

Skill Claude CodeCodex

Manages GitHub issue and pull-request lifecycle including creation, updates, triage labelling, and closeout metadata using canonical templates and repository taxonomy. Use when a maintainer asks to create, update, close, or triage GitHub issues or PRs, or when issue creation requires codebase analysis for scope…

5.4k 2d ago A 77 tokens original Apache-2.0

vllm-project/semantic-router

Skill Claude CodeCodex

Maintainer release and milestone operating workflow. Use when a maintainer wants to plan a release, assess milestone health, coordinate release blockers, or generate a release-focused review brief.

5.4k 2d ago A 41 tokens original Apache-2.0

vllm-project/semantic-router

Skill Claude CodeCodex

Calibrates routing changes against a live router endpoint with executable probes, local DSL validation, versioned deploys, and structured failure review. Use when tuning signals, projections, decisions, or maintained route examples against a real apiserver.

5.4k 2d ago A 53 tokens original Apache-2.0

plugin-end-to-end

09

vllm-project/semantic-router

Skill Claude CodeCodex

Implements end-to-end plugin changes spanning router config, post-decision processing, optional CLI/UI exposure, and E2E test coverage. Use when adding a new plugin type, changing plugin config schema or execution semantics, updating plugin chain behavior, or modifying plugin-exposed metadata across surfaces.

5.4k 2d ago A 63 tokens original Apache-2.0

project-change

10

vllm-project/semantic-router

Skill Claude CodeCodex

Handles a focused repository change when no specialized primary skill applies. Use when changed-file routing selects this fallback for a feature, fix, refactor, documentation update, or subsystem-local task.

5.4k 2d ago A 40 tokens original Apache-2.0

vllm-project/semantic-router

Skill Claude CodeCodex

Modifies routing policy after signal extraction, including matched-decision logic, candidate-model selection, and downstream looper behavior. Use when changing decision predicates, thresholds, priorities, model ranking, cost or latency routing, or other post-signal routing policy.

5.4k 2d ago A 54 tokens original Apache-2.0

signal-end-to-end

12

vllm-project/semantic-router

Skill Claude CodeCodex

Implements end-to-end signal changes spanning router config, signal extraction, CLI schema, optional bindings, router-owned metadata headers, and E2E test coverage. Use when adding a new signal type, changing signal configuration or extraction logic, updating CLI schema for signal parameters, or modifying router-owned…

5.4k 2d ago A 67 tokens original Apache-2.0

vllm-project/semantic-router

Skill Claude CodeCodex

Modifies the local startup chain including image build, container serve/bootstrap logic, and canonical smoke test behavior. Use when changing vllm-sr serve behavior, image selection or pull policy, container startup sequences, local Docker/Make bootstrap, or canonical smoke config.

5.4k 2d ago A 59 tokens original Apache-2.0

vllm-skills

14

vllm-project/vllm-skills

Plugin Claude Code

Plugin marketplace listing 1 plugin: vllm-skills.

95 5mo ago A tokens not measured original Apache-2.0

vllm-project/vllm-skills

Skill Claude CodeCodex

Run vLLM performance benchmark using synthetic random data to measure throughput, TTFT (Time to First Token), TPOT (Time per Output Token), and other key performance metrics. Use when the user wants to quickly test vLLM serving performance without downloading external datasets.

95 5mo ago A 64 tokens original Apache-2.0

vllm-bench-serve

17

vllm-project/vllm-skills

Skill Claude CodeCodex

Benchmark vLLM or OpenAI-compatible serving endpoints using vllm bench serve. Supports multiple datasets (random, sharegpt, sonnet, HF), backends (openai, openai-chat, vllm-pooling, embeddings), throughput/latency testing with request-rate control, and result saving. Use when benchmarking LLM serving performance…

95 5mo ago A 93 tokens original Apache-2.0

vllm-deploy-docker

18

vllm-project/vllm-skills

Skill Claude CodeCodex

Deploy vLLM using Docker (pre-built images or build-from-source) with NVIDIA GPU support and run the OpenAI-compatible server.

95 5mo ago B 35 tokens original Apache-2.0

vllm-deploy-k8s

19

vllm-project/vllm-skills

Skill Claude CodeCodex

Deploy vLLM to Kubernetes (K8s) with GPU support, health probes, and OpenAI-compatible API endpoint. Use this skill whenever the user wants to deploy, run, or serve vLLM on a Kubernetes cluster, including creating deployments, services, checking existing deployments, or managing vLLM on K8s.

95 5mo ago A 76 tokens original Apache-2.0

vllm-deploy-simple

20

vllm-project/vllm-skills

Skill Claude CodeCodex

Quick install and deploy vLLM, start serving with a simple LLM, and test OpenAI API.

95 5mo ago A 29 tokens original Apache-2.0

vllm-project/vllm-skills

Skill Claude CodeCodex

This is a skill for benchmarking the efficiency of automatic prefix caching in vLLM using fixed prompts, real-world datasets, or synthetic prefix/suffix patterns. Use when the user asks to benchmark prefix caching hit rate, caching efficiency, or repeated-prompt performance in vLLM.

95 5mo ago A 64 tokens original Apache-2.0