Q-Sparse: All Large Language Models can be Fully Sparsely-Activated
- Type
- other
- Venue
- arXiv / Microsoft Research / University of Chinese Academy of Sciences
Summary
Q-Sparse applies top-K (absolute-value) masks on activations of all linear layers, STE so non-selected grads are not zeroed, optional 8-bit quantized top-K, squared-ReLU/ReLU²GLU in FFNs, and Block Q-Sparse (N:M, e.g. 16:32) for batched GPU kernels. Inference-optimal sparsity ~45.58% full-precision (1.84× params at matched activated N_a) and ~61.25% for 1.58-bit. From-scratch 40% sparsity matches dense loss at the same size/tokens (RedPajama 50B). Continue-train Mistral-7B: 3.8B activated avg 63.7 vs dense 64.6 (ARC/HS/MMLU/WG/TQA); 2.9B 61.7, beating ReLUfication 60.8 and dReLU 61.0. Works with BitNet b1.58. No official code on abs.
Keywords
q-sparse · activation-sparsity · top-k · ste · bitnet · block-sparsity · microsoft · inference-optimal-scaling
Topics
activation sparsity, LLM inference, quantization
Research notes
- Primary: arxiv abs (cs.CL; also cs.LG). License: arXiv.org perpetual non-exclusive on HTML at check. Equal contrib Wang/Ma; corresponding Shuming Ma and Furu Wei (Microsoft Research). H. Wang and R. Wang at University of Chinese Academy of Sciences. No official code URL on abs (OpenReview version: “The code will be released”). Later github.com/ustcwhy/Q-Sparse not treated as abs-linked. HF paper page 23 upvotes; no linked models/datasets. Discord posted PDF. Trains on public RedPajama/FineWeb-Edu/Open-Orca rather than a new hosted corpus, so no datasets_local row. License field left blank per catalog convention.