flashinfer-ai/flashinfer

FlashInfer: Kernel Library for LLM Serving

About the project

FlashInfer is a library and kernel generator that supplies GPU operations used to run large language model inference, including attention, matrix multiplication, and mixture-of-experts computations. It helps engineers build and optimize LLM serving systems across supported GPU hardware and backend implementations. Its catalogue add-ons provide skills and instructions for working with FlashInfer.

Latest release v0.6.18.post1 — Release v0.6.18.post1 · 5 Sept 2026

These files are flashinfer-ai/flashinfer's own configuration. They tell Claude Code, Codex and OpenCode how to work on this repository, so they are not mods to install elsewhere. Copy one as a starting point and replace the parts that are about this project.

6,340Stars on the repository
5Files it configures its agents with
11,398Tokens loaded in every session
3Agents configured

Instructions

Skills