← Back to explorer

BitNet b1.58 2B4T

Type
model
Venue
Microsoft Research
Year
2026
Source
huggingface
Access
free
Language
en
Added
2026-08-14T18:28:00Z
Verified
2026-08-14T18:28:00Z

Summary

Packed ternary BitNet weights (~2B) trained from scratch on 4T tokens (SmolLM-Corpus, dclm-baseline-1.0, open-web-math) then SFT+DPO. Architecture: BitLinear, RoPE, ReLU² FFN, subln, no biases; LLaMA 3 tokenizer (128,256); ctx 4096. Card: non-emb memory 0.4GB, CPU decode 29ms, energy 0.028J vs 1–2B dense open models; average 54.19 vs Qwen2.5-1.5B 55.23 (GSM8K 58.38). Efficiency needs microsoft/BitNet (bitnet.cpp), not stock transformers. Siblings: bf16 master weights and GGUF.

Keywords

bitnet · 1-bit · ternary · microsoft · bitnet-cpp · w1.58a8 · llm

Topics

1-bit LLMs, quantization, efficient inference

Research notes

  • Primary: HF model card (MIT, transformers, pipeline_tag text-generation, downloads=19771 likes=1491 checked 2026-08-14). Technical report arXiv 2504.12285. Official inference https://github.com/microsoft/BitNet. Discord posted the HF model pointer. Weights/checkpoint item, not a new hosted pretraining corpus, so no datasets_local row. Training used public SmolLM-Corpus / dclm-baseline-1.0 / open-web-math.