NVIDIA/Model-Optimizer

A unified library of SOTA model optimization techniques like quantization, distillation, pruning, neural architecture search, speculative decoding, etc. It compresses deep learning models for downstream deployment frameworks like TensorRT-LLM, TensorRT, vLLM, etc. to optimize inference speed.

3.6kStars on the repository
22Mods indexed here, across every type
yesterdayLast push, which is what freshness is scored on
Apache-2.0Licence, which decides whether bodies are shown

NVIDIA/Model-Optimizer

Instructions file CodexOpenCode

Instructions for NVIDIA/Model-Optimizer, covering agent instructions for modelopt, repository orientation, coding guidelines, iterative development and contributing and pr readiness.

3.6k yesterday A 920 tokens original Apache-2.0

NVIDIA/Model-Optimizer

Instructions file

Instructions for NVIDIA/Model-Optimizer, a project described as: A unified library of SOTA model optimization techniques like quantization, distillation, pruning, neural architecture search, speculative decoding, etc. It compresses deep learning models for downstream deployment frameworks like TensorRT-LLM, TensorRT…

3.6k yesterday A 3 tokens copy · 100% Apache-2.0