Efficient Pre-Training with Token Superposition
- Type
- other
- Venue
- arXiv / Nous Research
Summary
Token-Superposition Training (TST) averages embeddings of non-overlapping bags of s contiguous tokens and trains with multi-hot cross-entropy, then a recovery phase reverts to standard next-token CE. Drop-in: no change to parallelism, optimizer, tokenizer, data, or architecture. Evaluated at 270M/600M (SmolLM2-shaped Llama3, untied embeddings) and validated at 3B and a 10B A1B MoE on DCLM. Consistently beats baseline loss and downstream metrics; equal-loss wall-clock up to 2.5× faster at the 10B A1B scale. Project https://nousresearch.com/token-superposition.
Keywords
tst · token-superposition · pretraining · multi-hot-ce · dclm · smollm · nous-research · moe
Topics
LLM pretraining, data efficiency, token packing
Research notes
- Primary: arxiv abs (cs.CL). CC BY 4.0 on HTML. Nous Research; Peng/Gigant equal contrib. Correspondence bloc/theo/emozilla@nousresearch.com. No official code on abs. HF paper page 48 upvotes, org NousResearch; linked models are unofficial (GestaltLabs/joelhenwang), not recorded in hf_* fields. Discord posted abs. Uses public DCLM; no new corpus, so no datasets_local row. License field left blank per catalog convention.