← Back to explorer

Optimizing Large Language Model Training Using FP4 Quantization

Type
other
Venue
arXiv / Microsoft Research Asia / University of Science and Technology of China

Summary

First from-scratch FP4 LLM pretraining: Differentiable Gradient Estimator (DGE, k=5) corrects E2M1 weight grads vs STE, and Outlier Clamping and Compensation (OCC, α=0.99 plus a ~2% sparse residual) keeps activations from collapsing. Mixed-precision W4A4 GeMM simulated on H100 FP8 cores with token-wise activation and channel-wise weight scaling. LLaMA-2 1.3B/7B/13B on 100B DCLM tokens: train loss 2.55/2.17/1.97 vs BF16 2.49/2.07/1.88; zero-shot avg 53.13/54.42/54.95 vs 53.23/53.87/54.44. Direct W4A4 diverges. ICML 2025. Framework https://aka.ms/MS.AMP (Azure/MS-AMP).

Keywords

fp4 · quantization · dge · occ · mixed-precision · llama2 · dclm · icml · microsoft · e2m1

Topics

quantization, FP4 training, mixed precision

Research notes

  • Primary: arxiv abs (cs.LG; also cs.CL). License: arXiv.org perpetual non-exclusive on HTML at check. Comment: ICML 2025; PMLR 267:62937-62957. Wang USTC/Microsoft SIGMA; Gong/Liu/Zhao/Yang/Cheng Microsoft Research Asia / SIGMA; Guo MSRA; Zha USTC. Correspondence yegong@microsoft.com, pengc@microsoft.com. Framework https://aka.ms/MS.AMP → https://github.com/Azure/MS-AMP (637 stars at check; MIT). HF paper page 36 upvotes; unofficial linked model w-hy21/lcqat_qwen3_1.7B not copied into hf_* fields. Discord posted abs. Trains on public DCLM rather than a new hosted corpus, so no datasets_local row. License field left blank per catalog convention.