Optimizing Large Language Model Training Using FP4 Quantization
- Type
- other
- Venue
- arXiv / Microsoft Research Asia / University of Science and Technology of China
Summary
First from-scratch FP4 LLM pretraining: Differentiable Gradient Estimator (DGE, k=5) corrects E2M1 weight grads vs STE, and Outlier Clamping and Compensation (OCC, α=0.99 plus a ~2% sparse residual) keeps activations from collapsing. Mixed-precision W4A4 GeMM simulated on H100 FP8 cores with token-wise activation and channel-wise weight scaling. LLaMA-2 1.3B/7B/13B on 100B DCLM tokens: train loss 2.55/2.17/1.97 vs BF16 2.49/2.07/1.88; zero-shot avg 53.13/54.42/54.95 vs 53.23/53.87/54.44. Direct W4A4 diverges. ICML 2025. Framework https://aka.ms/MS.AMP (Azure/MS-AMP).
Keywords
fp4 · quantization · dge · occ · mixed-precision · llama2 · dclm · icml · microsoft · e2m1
Topics
quantization, FP4 training, mixed precision
Research notes
- Primary: arxiv abs (cs.LG; also cs.CL). License: arXiv.org perpetual non-exclusive on HTML at check. Comment: ICML 2025; PMLR 267:62937-62957. Wang USTC/Microsoft SIGMA; Gong/Liu/Zhao/Yang/Cheng Microsoft Research Asia / SIGMA; Guo MSRA; Zha USTC. Correspondence yegong@microsoft.com, pengc@microsoft.com. Framework https://aka.ms/MS.AMP → https://github.com/Azure/MS-AMP (637 stars at check; MIT). HF paper page 36 upvotes; unofficial linked model w-hy21/lcqat_qwen3_1.7B not copied into hf_* fields. Discord posted abs. Trains on public DCLM rather than a new hosted corpus, so no datasets_local row. License field left blank per catalog convention.