← Back to explorer

BitNet a4.8: 4-bit Activations for 1-bit LLMs

Type
other
Venue
arXiv / Microsoft Research / University of Chinese Academy of Sciences

Summary

Hybrid recipe: INT4 activations into attention/FFN, plus sparsify then 8-bit on outlier-heavy intermediate states. Two-stage continue-training from BitNet b1.58 (W1.58A8 then W1.58A4; last 5B tokens for the 4-bit stage). Matches b1.58 at equal train cost while enabling 4-bit kernels; ~55% of parameters activated and 3-bit KV cache. 7B: 44.5% overall sparsity / 3.4B active vs b1.58 6.0B; a 2B model on 2T tokens stays at parity with b1.58. No official code on abs.

Keywords

bitnet · quantization · 1-bit · int4 · sparsity · microsoft · activation-quantization

Topics

quantization, 1-bit LLMs, efficient inference

Research notes

  • Primary: arxiv abs (cs.CL; also cs.LG). License: arXiv.org perpetual non-exclusive on HTML at check. Comment: Work in progress. Equal contrib Wang/Ma; corresponding S. Ma and F. Wei (Microsoft Research). H. Wang UCAS. https://aka.ms/GeneralAI. No official code on abs. HF paper page 70 upvotes; no githubRepo; unofficial linked dataset DavidRy/bitnet-distill-logits (1 like / 14 downloads at HF API check) not copied into hf_* fields and not substantial, so no datasets_local row. Discord posted HTML. License field left blank per catalog convention.