← Back to explorer

Decoupling the Benefits of Subword Tokenization for Language Model Training via Byte-level Simulation

Type
other
Venue
arXiv / Nous Research

Summary

Trains a 1.7B LLaMA-3 byte LM on FineWeb-Edu UTF-8 and injects one subword effect at a time. Biggest gains: 4× isoFLOP sample throughput (chunk-4 embeddings) and subword boundary priors/inductive biases (end-boundaries leak future bytes; start-boundaries remain useful when removed at val). Embedding-table scaling, subword-distance RoPE, per-subword CE, and next-subword MTP are weak or harmful at this scale. Interventions often run 50k steps then revert. No official code on abs.

Keywords

tokenization · byte-level · bpe · fineweb-edu · nous-research · sample-throughput · inductive-bias

Topics

tokenization, byte-level LMs, pretraining ablations

Research notes

  • Primary: arxiv abs (cs.CL). CC BY 4.0 on HTML. Nous Research; correspondence theo/bloc/emozilla@nousresearch.com. No official code on abs. HF paper page 11 upvotes, org NousResearch; no linked models/datasets. Discord posted abs. Uses public FineWeb-Edu; no new corpus, so no datasets_local row. License field left blank per catalog convention.