Late-to-Early Training: LET LLMs Learn Earlier, So Faster and Better
- Type
- paper
- Venue
- arXiv / HKUST-GZ / ByteDance Seed
- Year
- 2026
- Source
- arxiv
- Access
- free
- Language
- en
- Added
- 2026-08-14T17:02:26Z
- Verified
- 2026-08-14T17:02:26Z
Summary
LET adds a decaying cosine-alignment loss so early layers of a larger target, early in training, match the last-layer hidden states of a small already-trained teacher (late-to-early-layer and late-to-early-step). On 1.4B LLaMA-like models on The Pile (~20B tokens) it reports up to 1.6x faster downstream improvement and ~5% higher average one-shot accuracy vs standard CLM, even with a 10x smaller teacher (SmolLM-135M); 7B also beats baseline, reverse KD, and SALT. Ablations: last-to-early (L2E) alignment is best; lambda=0.1; S_stop=1500. No official code on the abs page.
Keywords
let · late-to-early · distillation · pretraining · bytedance · hkust · pile · smollm
Topics
LLM pretraining, knowledge distillation
Research notes
- Primary: arxiv abs (default nonexclusive-distrib 1.0, cs.LG). HKUST (Guangzhou) / ByteDance Seed; correspondence zekexie@hkust-gz.edu.cn. Discord posted AlphaXiv 2602.05393; canonical abs recorded. HF paper page 9 upvotes, org ByteDance-Seed; no linked models/datasets. No official code on abs. Uses existing Pile/OPT/Pythia/SmolLM; no new corpus, so no datasets_local row.