Estimates max concurrency under TTFT/TPOT SLO for arbitrary vLLM-Ascend models (Qwen any size, GLM, DeepSeek, MiniMax, …) by analyzing vllm-ascend attention dispatch (usemla/usesparse) and msmodeling profilingdatabase. Requires local clones of both repos. Use as sub-skill of serving-parallel-strategy-tuning.
Extract and compare configuration switches between vLLM and vLLM-Ascend repos. Invoke when user needs to audit, compare, or document config options across vLLM and vLLM-Ascend.
A guide to tuning vLLM-Ascend, a system for running language models on Ascend hardware. It covers standalone use, pipeline use, and quantization tuning, which reduces model precision to change performance or resource use.
★not rated 2 22d agoB147 tokens
At most 3 mods per repository are shown here, and a mod shipped inside a plugin is left to that plugin's page — the rest are on their repository pages: