← Back to explorer

NVIDIA Model Optimizer (ModelOpt)

Type
code
Venue
GitHub
Year
2026
Source
github
Access
public
Language
en
Added
2026-09-29
Verified
2026-09-29

Summary

NVIDIA's unified open-source library of state-of-the-art model optimization techniques — post-training quantization, quantization-aware training/distillation, pruning, distillation, neural architecture search, speculative decoding, and sparsity. Takes Hugging Face, PyTorch, or ONNX models and exports optimized quantized checkpoints ready for TensorRT-LLM, TensorRT, vLLM, or SGLang deployment.

Keywords

quantization · pruning · distillation · inference-optimization · NVFP4 · deployment

Topics

quantization, pruning, distillation, inference-optimization, NVFP4

Research notes

  • Discovery: Shared in #random-papers as a bare GitHub link.
  • Method: Composable Python APIs for PTQ, QAT/QAD, pruning (incl. Minitron), distillation, NAS (Puzzletron), speculative decoding, and sparsity; integrated with Megatron-Bridge, Megatron-LM, and Hugging Face Accelerate; unified HF export API for transformers and diffusers.
  • Key findings: Production use cases reported in the README: NVFP4 W4A4 PTQ + quantization-aware distillation reaching 1.30x vLLM throughput over BF16 on Qwen3.6-35B-A3B; Nemotron 3 Ultra (550B) quantized to NVFP4 with up to 5.9x higher decode-heavy throughput vs GLM-5.1 754B FP4 at matched BF16 accuracy; Domyn compressed Colosseum-355B to 260B via Minitron pruning + distillation. Repo: 5,075 stars, 1,299 commits, 50 releases, Apache-2.0.
  • Limitations: Capability claims are NVIDIA's own; no independent benchmark verification performed for this catalog entry.
  • Rebranded from 'NVIDIA TensorRT Model Optimizer' in December 2025. Docs at nvidia.github.io; pip package nvidia-modelopt.