vllm skills

43 tagged vllm, measured the same way as everything else here.

Browse within: ai-gateway 11kubernetes 11llmrouter 11mixture-of-models 11semantic-router 11primary 7SGLang 6

mooncake-api

01

kvcache-ai/Mooncake

Skill Claude CodeCodex

Help users work with the Mooncake Python APIs for distributed storage and high-performance data transfer. Use when working with Mooncake Store (distributed KV cache), Transfer Engine (RDMA/TCP transfers), service setup (master, metadata server), PyTorch tensors in the Store, zero-copy/buffer management, batch…

6.4k 2d ago A 122 tokens original Apache-2.0

mooncake-ci-local

02

kvcache-ai/Mooncake

Skill Claude CodeCodex

Run Mooncake pre-PR local validation through scripts/runcitest.sh. Use this skill whenever the user wants to validate a branch before opening or submitting a PR, run local CI, run ci test, check changes before PR, reproduce GitHub Actions locally, or force a full pre-submit verification. Trigger on phrases like "提交 PR…

6.4k 2d ago A 108 tokens original Apache-2.0

kvcache-ai/Mooncake

Skill Claude CodeCodex

Automatically diagnose Mooncake deployment and runtime issues. Checks services (mooncakemaster, metadata server), RDMA devices, environment variables, connectivity, memory limits, object integrity, and analyzes logs for common error patterns. Use when Mooncake deployment fails, services won't start, connections fail…

6.4k 2d ago A 134 tokens original Apache-2.0

vllm-project/semantic-router

Skill Claude CodeCodex

Manages GitHub issue and pull-request lifecycle including creation, updates, triage labelling, and closeout metadata using canonical templates and repository taxonomy. Use when a maintainer asks to create, update, close, or triage GitHub issues or PRs, or when issue creation requires codebase analysis for scope…

5.4k 2d ago A 77 tokens original Apache-2.0

vllm-project/semantic-router

Skill Claude CodeCodex

Maintainer release and milestone operating workflow. Use when a maintainer wants to plan a release, assess milestone health, coordinate release blockers, or generate a release-focused review brief.

5.4k 2d ago A 41 tokens original Apache-2.0

vllm-project/semantic-router

Skill Claude CodeCodex

Calibrates routing changes against a live router endpoint with executable probes, local DSL validation, versioned deploys, and structured failure review. Use when tuning signals, projections, decisions, or maintained route examples against a real apiserver.

5.4k 2d ago A 53 tokens original Apache-2.0

add-model-bundle

07

Tencent-Hunyuan/UniRL

Skill Claude CodeCodex

Add or update UniRL model package support. Use when adding diffusion or autoregressive model pipelines, model config dataclasses, Bundle/Pipeline/Stage/Conditions implementations, LoRA targets, FSDP wrapping hints, Sample/Part plumbing, or multimodal text/image/video/audio conditioning.

921 3d ago A 62 tokens

pr-workflow

08

Tencent-Hunyuan/UniRL

Skill Claude CodeCodex

Create and repair UniRL pull requests. Use when creating a PR, editing a PR body, responding to PR Body or Semantic Pull Request CI failures, running gh pr create, or when the user mentions PR body, pull request template, body check, title check, or pull request formatting.

921 3d ago A 62 tokens

code-standards

09

Tencent-Hunyuan/UniRL

Skill Claude CodeCodex

A development workflow and coding-standard skill for Python, PyTorch, deep-learning code, shell scripts, and configuration files. It requires understanding the project, proposing a plan for approval, writing tests, simplifying code, and reviewing the result.

921 3d ago A 238 tokens

mudler/vllm.cpp

Skill Claude CodeCodex

Apply the repository writing rules to commits, pull requests, branches, changelogs, and release notes.

367 yesterday A 29 tokens original Apache-2.0

mudler/vllm.cpp

Skill Claude CodeCodex

Apply the repository technical-English rules to repository documents and all session prose.

367 yesterday A 20 tokens original Apache-2.0

add-model

12

guoqingbao/xinfer

Skill Claude CodeCodexCursor

Adapt and port new LLM model architectures to this xinfer project. Use when the user asks to add, port, support, or adapt a new model (e.g. Llama, Gemma, Qwen, GPT-OSS, DeepSeek, or any HuggingFace architecture) including safetensors and GGUF formats, Dense and MoE architectures, and quantization formats (MXFP4…

309 5d ago B 95 tokens original MIT

check-model

13

guoqingbao/xinfer

Skill Claude CodeCodexCursor

Check model compatibility with xinfer before loading. Validates config.json, weight tensor shapes and naming, quantization format correctness, and multi-rank (tensor-parallel) divisibility. Use when the user asks to check, validate, audit, or verify a model will load correctly — from a HuggingFace URL/config, local…

309 5d ago A 76 tokens original MIT

test-model

14

guoqingbao/xinfer

Skill Claude CodeCodexCursor

Test LLM models served by xinfer for correctness, output quality, and performance. Use when the user asks to test, benchmark, validate, or verify models — either from a local folder path or HuggingFace model IDs. Supports all xinfer-compatible formats: BF16, FP8, MXFP4, NVFP4, GGUF, GPTQ, AWQ, ISQ, Dense, MoE, and…

309 5d ago A 93 tokens original MIT

shen-shanshan/vllm-dev-skills

Skill Claude CodeCodex

Analyze contribution opportunities in the vllm-project/vllm repository for community developers. Given a module, feature, or model area, this skill collects information from open issues, recent PRs, GitHub discussions, code TODOs/FIXMEs, roadmap labels, and maintainer activity to generate a structured Markdown report…

16 6d ago A 226 tokens original Apache-2.0

vllm-rocm-pr-review

16

shen-shanshan/vllm-dev-skills

Skill Claude CodeCodex

Review AMD/ROCm-related pull requests from the vllm-project/vllm GitHub repository and produce a concise Chinese review report covering motivation, code-change summary, severity-sorted and type-categorized review findings, existing discussion, and a verdict. Use when the user provides a vllm PR number or link and asks…

16 6d ago A 137 tokens original Apache-2.0

shen-shanshan/vllm-dev-skills

Skill Claude CodeCodex

Write or complete Chinese vLLM technical blog posts in the author's established Zhihu style. Use when the user provides a vLLM feature, model, architecture, optimization, or other topic and asks for a full blog post, or provides an existing Markdown outline/draft plus references and asks to research current…

16 6d ago A 104 tokens original Apache-2.0