FlashInfer: Kernel Library for LLM Serving
FlashInfer is a library and kernel generator that supplies GPU operations used to run large language model inference, including attention, matrix multiplication, and mixture-of-experts computations. It helps engineers build and optimize LLM serving systems across supported GPU hardware and backend implementations. Its catalogue add-ons provide skills and instructions for working with FlashInfer.
Latest release v0.6.18.post1 — Release v0.6.18.post1 · 5 Sept 2026
These files are flashinfer-ai/flashinfer's own configuration. They tell Claude Code, Codex and OpenCode how to work on this repository, so they are not mods to install elsewhere. Copy one as a starting point and replace the parts that are about this project.
AGENTS.md A 45 tok CLAUDE.md E 11,353 tok .claude/skills/add-cuda-kernel/SKILL.md A 18 tok .claude/skills/benchmark-kernel/SKILL.md B 14 tok .claude/skills/debug-cuda-crash/SKILL.md A 14 tok