BitNet b1.58 2B4T
- Type
- model
- Venue
- Microsoft Research
- Year
- 2026
- Source
- huggingface
- Access
- free
- Language
- en
- Added
- 2026-08-14T18:28:00Z
- Verified
- 2026-08-14T18:28:00Z
Summary
Packed ternary BitNet weights (~2B) trained from scratch on 4T tokens (SmolLM-Corpus, dclm-baseline-1.0, open-web-math) then SFT+DPO. Architecture: BitLinear, RoPE, ReLU² FFN, subln, no biases; LLaMA 3 tokenizer (128,256); ctx 4096. Card: non-emb memory 0.4GB, CPU decode 29ms, energy 0.028J vs 1–2B dense open models; average 54.19 vs Qwen2.5-1.5B 55.23 (GSM8K 58.38). Efficiency needs microsoft/BitNet (bitnet.cpp), not stock transformers. Siblings: bf16 master weights and GGUF.
Keywords
bitnet · 1-bit · ternary · microsoft · bitnet-cpp · w1.58a8 · llm
Topics
1-bit LLMs, quantization, efficient inference
Research notes
- Primary: HF model card (MIT, transformers, pipeline_tag text-generation, downloads=19771 likes=1491 checked 2026-08-14). Technical report arXiv 2504.12285. Official inference https://github.com/microsoft/BitNet. Discord posted the HF model pointer. Weights/checkpoint item, not a new hosted pretraining corpus, so no datasets_local row. Training used public SmolLM-Corpus / dclm-baseline-1.0 / open-web-math.