Decoupling the Benefits of Subword Tokenization for Language Model Training via Byte-level Simulation
- Type
- other
- Venue
- arXiv / Nous Research
Summary
Trains a 1.7B LLaMA-3 byte LM on FineWeb-Edu UTF-8 and injects one subword effect at a time. Biggest gains: 4× isoFLOP sample throughput (chunk-4 embeddings) and subword boundary priors/inductive biases (end-boundaries leak future bytes; start-boundaries remain useful when removed at val). Embedding-table scaling, subword-distance RoPE, per-subword CE, and next-subword MTP are weak or harmful at this scale. Interventions often run 50k steps then revert. No official code on abs.
Keywords
tokenization · byte-level · bpe · fineweb-edu · nous-research · sample-throughput · inductive-bias
Topics
tokenization, byte-level LMs, pretraining ablations
Research notes
- Primary: arxiv abs (cs.CL). CC BY 4.0 on HTML. Nous Research; correspondence theo/bloc/emozilla@nousresearch.com. No official code on abs. HF paper page 11 upvotes, org NousResearch; no linked models/datasets. Discord posted abs. Uses public FineWeb-Edu; no new corpus, so no datasets_local row. License field left blank per catalog convention.