NVIDIA Model Optimizer (ModelOpt)
- Type
- code
- Venue
- GitHub
- Year
- 2026
- Source
- github
- Access
- public
- Language
- en
- Added
- 2026-09-29
- Verified
- 2026-09-29
Summary
NVIDIA's unified open-source library of state-of-the-art model optimization techniques — post-training quantization, quantization-aware training/distillation, pruning, distillation, neural architecture search, speculative decoding, and sparsity. Takes Hugging Face, PyTorch, or ONNX models and exports optimized quantized checkpoints ready for TensorRT-LLM, TensorRT, vLLM, or SGLang deployment.
Keywords
quantization · pruning · distillation · inference-optimization · NVFP4 · deployment
Topics
quantization, pruning, distillation, inference-optimization, NVFP4
Research notes
- Discovery: Shared in #random-papers as a bare GitHub link.
- Method: Composable Python APIs for PTQ, QAT/QAD, pruning (incl. Minitron), distillation, NAS (Puzzletron), speculative decoding, and sparsity; integrated with Megatron-Bridge, Megatron-LM, and Hugging Face Accelerate; unified HF export API for transformers and diffusers.
- Key findings: Production use cases reported in the README: NVFP4 W4A4 PTQ + quantization-aware distillation reaching 1.30x vLLM throughput over BF16 on Qwen3.6-35B-A3B; Nemotron 3 Ultra (550B) quantized to NVFP4 with up to 5.9x higher decode-heavy throughput vs GLM-5.1 754B FP4 at matched BF16 accuracy; Domyn compressed Colosseum-355B to 260B via Minitron pruning + distillation. Repo: 5,075 stars, 1,299 commits, 50 releases, Apache-2.0.
- Limitations: Capability claims are NVIDIA's own; no independent benchmark verification performed for this catalog entry.
- Rebranded from 'NVIDIA TensorRT Model Optimizer' in December 2025. Docs at nvidia.github.io; pip package nvidia-modelopt.