Byte Latent Transformer: Patches Scale Better Than Tokens
- Type
- other
- Venue
- arXiv / FAIR Meta / University of Washington / University of Chicago
Summary
BLT maps bytes to variable patches via next-byte entropy, then a light local encoder, a large latent Transformer, and a local decoder. First flop-controlled byte-level scaling study to 8B / 4T bytes: matches Llama 3 at equal train FLOPs and can cut inference FLOPs ~50% with patch size 6–8. For a fixed inference budget, growing patch size and model size together beats BPE. Robustness: noisy HellaSwag 64.3 vs Llama 3 56.9; CUTE 54.1 vs 27.5. Code https://github.com/facebookresearch/blt.
Keywords
blt · byte-level · entropy-patching · tokenization · fair · meta · llama3 · cute
Topics
tokenization, byte-level LMs, efficient transformers
Research notes
- Primary: arxiv abs (cs.CL). License not stated on abs/HTML at check. FAIR at Meta / UW / UChicago. Correspondence artidoro at cs.washington.edu, sviyer at meta.com. Code https://github.com/facebookresearch/blt (2,059 stars at check). HF paper page 109 upvotes; githubRepo linked; official linked models facebook/blt, facebook/blt-7b, facebook/blt-1b, facebook/blt-entropy not copied into hf_* fields. Discord posted abs. Pretraining mix is not a new public corpus, so no datasets_local row. License field left blank per catalog convention.