Understanding Reasoning from Pretraining to Post-Training
- Type
- paper
- Venue
- arXiv (cs.LG), v2 revised 2026-08-09
- Year
- 2026
- Source
- arxiv
- Access
- public
- Language
- en
- Added
- 2026-09-29
- Verified
- 2026-09-29
Summary
Studies how pretraining choices shape the returns to RL post-training and what RL actually does to the model, using chess as a controlled testbed. Language models from 5M to 1B parameters are pretrained on human chess games, supervised fine-tuned on synthetic reasoning traces, then RL-trained on chess puzzles with verifiable rewards. Post-RL performance at a given RL compute level is well-predicted from pretraining loss, and RL reward-curve slopes improve approximately linearly with pretraining tokens.
Keywords
reasoning · pretraining · RL post-training · scaling laws · chess · verifiable rewards
Topics
reasoning, pretraining, RL post-training, scaling laws, chess
Research notes
- Method: Full pretraining-to-post-training pipeline replicated in a controlled domain: pretrain LMs (5M-1B) on human chess games, SFT on synthetic reasoning traces, RL on chess puzzles with verifiable rewards; transfer check by training a 1B LM on math-domain text.
- Key findings: Post-RL performance at fixed RL compute is well-predicted from pretraining loss; Slope of RL reward curves improves approximately linearly with pretraining tokens; RL does not merely sharpen the SFT policy: on easy puzzles it amplifies already-preferred correct moves, on hard puzzles it surfaces correct moves nearly absent under SFT; Same predictive pattern transfers to a 1B language model trained on math text
- Limitations: Chess is a narrow, fully-verifiable domain; transfer beyond the math-domain check is not demonstrated. Compute sweeps are still expensive outside the testbed.
- Original Discord link was the /pdf/ form; canonical abs page used.