Beyond Random Sampling: Efficient Language Model Pretraining via Curriculum Learning
- Type
- paper
- Venue
- arXiv / Ecole Polytechnique / MBZUAI
- Year
- 2026
- Source
- arxiv
- Access
- free
- Language
- en
- Added
- 2026-08-14T16:54:04Z
- Verified
- 2026-08-14T16:54:04Z
Summary
First systematic curriculum-learning study for LLM pretraining: more than 200 models, up to 100B tokens, three strategies (vanilla sort, pacing, interleaved) times six difficulty metrics (compression ratio, fertility, Flesch, MTLD, n_tokens, perplexity) on CulturaX English. 0.5B LLaMA3.2-like, plus 1B/3B. CL reaches the random baseline 18-45 percent faster in early/mid training. Best warmup (CL then random) gains up to 3.5 percent (0.5B) and about 3.1 percent at 100B tokens. Strongest metrics: compression ratio, MTLD, Flesch; high-perplexity tails are often noisy. Orthogonal to data selection. English decoder-only only; static precomputed scores. No official code on the abs page.
Keywords
curriculum-learning · pretraining · data-ordering · culturax · mtld · flesch · mbzuai
Topics
LLM pretraining, curriculum learning, data ordering
Research notes
- Primary: arxiv abs (CC BY 4.0, cs.CL). Correspondence yang.zhang@polytechnique.edu, guokan.shang@mbzuai.ac.ae. No official code on abs. HF paper page 2 upvotes. Uses existing CulturaX; no new dataset release, so no datasets_local row. Discord posted AlphaXiv 2506.11300.