A high-performance inference engine for LLM, VLM, DiT and REC models, optimized for diverse AI accelerators. It is hosted in OpenAtom Foundation.
xLLM is an inference engine, meaning software that runs trained AI models to produce outputs from inputs, for large language, vision-language, diffusion, and recommendation models on different AI accelerators. Organizations use it to deploy these models with high-throughput and low-latency inference. The catalogue entries provide skills and instructions for working with xLLM.
Latest release v0.10.1 · 14 Jul 2026
These files are xLLM-AI/xllm's own configuration. They tell Claude Code, Codex and OpenCode how to work on this repository, so they are not mods to install elsewhere. Copy one as a starting point and replace the parts that are about this project.
AGENTS.md A 805 tok CLAUDE.md A 5 tok .agents/skills/add-unit-test/SKILL.md A 76 tok .agents/skills/code-review/SKILL.md A 71 tok .agents/skills/git-workflow/SKILL.md A 52 tok .agents/skills/tilelang-ascend-kernel/SKILL.md A 60 tok