A unified library of SOTA model optimization techniques like quantization, distillation, pruning, neural architecture search, speculative decoding, etc. It compresses deep learning models for downstream deployment frameworks like TensorRT-LLM, TensorRT, vLLM, etc. to optimize inference speed.
NVIDIA Model Optimizer is a library that changes deep-learning models through techniques such as quantization, pruning, distillation, neural architecture search, speculative decoding, and sparsity so they can run more efficiently. Developers use it to optimize Hugging Face, PyTorch, or ONNX models and export checkpoints for inference frameworks such as SGLang, TensorRT-LLM, TensorRT, or vLLM. The catalogue entries provide skills, instructions, and a plugin for using its workflows.
Latest release 0.46.0 — ModelOpt 0.46.0 Release · 18 Aug 2026
These files are NVIDIA/Model-Optimizer's own configuration. They tell Codex, OpenCode and Claude Code how to work on this repository, so they are not mods to install elsewhere. Copy one as a starting point and replace the parts that are about this project.
AGENTS.md A 920 tok CLAUDE.md A 3 tok