NVIDIA/TensorRT-LLM

TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT LLM also contains components to create Python and C++ runtimes that orchestrate the inference execution in a performant way.

Latest release v1.2.1 · 20 Apr 2026

These files are NVIDIA/TensorRT-LLM's own configuration. They tell Claude Code, Codex and OpenCode how to work on this repository, so they are not mods to install elsewhere. Copy one as a starting point and replace the parts that are about this project.

14,556Stars on the repository
42Files it configures its agents with
2,974Tokens loaded in every session
3Agents configured

Instructions

Skills

Agents