← Back to explorer

Efficient Pre-Training with Token Superposition

Type
other
Venue
arXiv / Nous Research

Summary

Token-Superposition Training (TST) averages embeddings of non-overlapping bags of s contiguous tokens and trains with multi-hot cross-entropy, then a recovery phase reverts to standard next-token CE. Drop-in: no change to parallelism, optimizer, tokenizer, data, or architecture. Evaluated at 270M/600M (SmolLM2-shaped Llama3, untied embeddings) and validated at 3B and a 10B A1B MoE on DCLM. Consistently beats baseline loss and downstream metrics; equal-loss wall-clock up to 2.5× faster at the 10B A1B scale. Project https://nousresearch.com/token-superposition.

Keywords

tst · token-superposition · pretraining · multi-hot-ce · dclm · smollm · nous-research · moe

Topics

LLM pretraining, data efficiency, token packing

Research notes

  • Primary: arxiv abs (cs.CL). CC BY 4.0 on HTML. Nous Research; Peng/Gigant equal contrib. Correspondence bloc/theo/emozilla@nousresearch.com. No official code on abs. HF paper page 48 upvotes, org NousResearch; linked models are unofficial (GestaltLabs/joelhenwang), not recorded in hf_* fields. Discord posted abs. Uses public DCLM; no new corpus, so no datasets_local row. License field left blank per catalog convention.