Synthetic Pre-pretraining Survives Scale, but Not as a Grammatical Prior
- Type
- paper
- Venue
- arXiv:2609.39827 (cs.CL), submitted 30 Sep 2026
- Year
- 2026
- Source
- arxiv
- Access
- free
- Language
- English
- Added
- 2026-10-01
- Verified
- 2026-10-01
Summary
Tests whether pre-pretraining (PPT) on synthetic non-natural language data — previously shown to improve token efficiency in pretraining (PT), and attributed to a learned grammatical prior — survives at realistic scale. Across five PPT tasks, four PT data mixtures, four parameter scales (500M to 7B), and PT budgets up to 100B tokens, the downstream performance and token-efficiency gains persist at scale, saving at least 21B PT tokens at the 3B scale. But there is no consistent evidence the gains come from a grammatical prior: downstream performance does not consistently align with grammatical acceptability across model sizes, and instead the gains arise from PPT tasks that improve long-range retrieval. Gains are robust to PT mixture composition and diminish only when web text is absent, making PPT a low-cost addition to PT — with future task design aimed at long-range retrieval rather than grammar.
Keywords
pre-pretraining · PPT · synthetic data · token efficiency · grammatical prior · inductive bias · long-range retrieval · pretraining data mixtures
Topics
pre-pretraining, synthetic data, language model pretraining, token efficiency, grammatical prior, long-range retrieval
Research notes
- Discovery: shared directly in chat 2026-10-01 (arXiv link)
- Prior work (models at most 1B params, PT budgets below 2B tokens, predominantly web text) attributed PPT gains to a grammatical prior — a structural inductive bias from synthetic-data PPT that transfers to natural language grammar
- This study: five PPT tasks x four PT data mixtures x four parameter scales (500M-7B) x PT budgets up to 100B tokens; gains in downstream performance and token efficiency persist at scale, saving at least 21B PT tokens at 3B scale
- Against the grammatical-prior account: downstream performance does not consistently align with grammatical acceptability across model sizes; instead downstream gains arise from PPT tasks that improve long-range retrieval
- Gains robust to how PT data mixtures are composed (incl. code and math), diminish only when web text is absent; PPT framed as a low-cost addition to PT; recommendation: future PPT task design should target long-range retrieval rather than natural language grammar
- License: not stated on arXiv abstract page