← Back to explorer

Beyond Random Sampling: Efficient Language Model Pretraining via Curriculum Learning

Type
paper
Venue
arXiv / Ecole Polytechnique / MBZUAI
Year
2026
Source
arxiv
Access
free
Language
en
Added
2026-08-14T16:54:04Z
Verified
2026-08-14T16:54:04Z

Summary

First systematic curriculum-learning study for LLM pretraining: more than 200 models, up to 100B tokens, three strategies (vanilla sort, pacing, interleaved) times six difficulty metrics (compression ratio, fertility, Flesch, MTLD, n_tokens, perplexity) on CulturaX English. 0.5B LLaMA3.2-like, plus 1B/3B. CL reaches the random baseline 18-45 percent faster in early/mid training. Best warmup (CL then random) gains up to 3.5 percent (0.5B) and about 3.1 percent at 100B tokens. Strongest metrics: compression ratio, MTLD, Flesch; high-perplexity tails are often noisy. Orthogonal to data selection. English decoder-only only; static precomputed scores. No official code on the abs page.

Keywords

curriculum-learning · pretraining · data-ordering · culturax · mtld · flesch · mbzuai

Topics

LLM pretraining, curriculum learning, data ordering

Research notes

  • Primary: arxiv abs (CC BY 4.0, cs.CL). Correspondence yang.zhang@polytechnique.edu, guokan.shang@mbzuai.ac.ae. No official code on abs. HF paper page 2 upvotes. Uses existing CulturaX; no new dataset release, so no datasets_local row. Discord posted AlphaXiv 2506.11300.