SuperBPE: Space Travel for Language Models
- Type
- other
- Venue
- arXiv / University of Washington / NVIDIA / Allen Institute for AI
Summary
Two-stage BPE curriculum: subwords first, then merge across whitespace. At vocab 200k, encodes a fixed text with up to 33% fewer tokens than BPE. 8B transformers trained from scratch at matched params/vocab/FLOPs (~330B tokens): SuperBPE (t=180k) beats BPE on 25/30 tasks, +4.0% absolute average and +8.2% on MMLU, with 27% less inference compute. Superwords often match multi-word expressions. COLM 2025 camera-ready. No official code on abs.
Keywords
superbpe · tokenization · superword · bpe · olmo · colm · uw · nvidia · ai2 · mmlu
Topics
tokenization, BPE, pretraining efficiency
Research notes
- Primary: arxiv abs (cs.CL; also cs.LG). License: arXiv.org perpetual non-exclusive on HTML at check. Comment: COLM 2025 camera-ready. Equal contrib Liu/Hayase. UW / NVIDIA / Ai2. No official code on abs. HF paper page 14 upvotes; official-looking linked models UW/OLMo2-8B-SuperBPE-t180k, UW/OLMo2-8B-SuperBPE-t160k, UW/OLMo2-11B-SuperBPE-t180k (plus unofficial UniversalComputingResearch/Limen0.2B) not copied into hf_* fields. Discord posted abs. Pretraining mix is not a new public corpus, so no datasets_local row. License field left blank per catalog convention.