xLLM-AI/xllm

A high-performance inference engine for LLM, VLM, DiT and REC models, optimized for diverse AI accelerators. It is hosted in OpenAtom Foundation.

About the project

xLLM is an inference engine, meaning software that runs trained AI models to produce outputs from inputs, for large language, vision-language, diffusion, and recommendation models on different AI accelerators. Organizations use it to deploy these models with high-throughput and low-latency inference. The catalogue entries provide skills and instructions for working with xLLM.

This repository also configures its own agents. See what xllm tells them →

1.6kStars on the repository
7Mods indexed here, across every type
2d agoLast push, which is what freshness is scored on
Apache-2.0Licence, which decides whether bodies are shown

clang-format

01

xLLM-AI/xllm

Cursor rule Cursor

Always clang-format C++ edits before finishing.

not rated 1.6k +13 2d ago A 113 tokens original Apache-2.0