← Back to explorer

Understanding Reasoning from Pretraining to Post-Training

Type
paper
Venue
arXiv (cs.LG), v2 revised 2026-08-09
Year
2026
Source
arxiv
Access
public
Language
en
Added
2026-09-29
Verified
2026-09-29

Summary

Studies how pretraining choices shape the returns to RL post-training and what RL actually does to the model, using chess as a controlled testbed. Language models from 5M to 1B parameters are pretrained on human chess games, supervised fine-tuned on synthetic reasoning traces, then RL-trained on chess puzzles with verifiable rewards. Post-RL performance at a given RL compute level is well-predicted from pretraining loss, and RL reward-curve slopes improve approximately linearly with pretraining tokens.

Keywords

reasoning · pretraining · RL post-training · scaling laws · chess · verifiable rewards

Topics

reasoning, pretraining, RL post-training, scaling laws, chess

Research notes

  • Method: Full pretraining-to-post-training pipeline replicated in a controlled domain: pretrain LMs (5M-1B) on human chess games, SFT on synthetic reasoning traces, RL on chess puzzles with verifiable rewards; transfer check by training a 1B LM on math-domain text.
  • Key findings: Post-RL performance at fixed RL compute is well-predicted from pretraining loss; Slope of RL reward curves improves approximately linearly with pretraining tokens; RL does not merely sharpen the SFT policy: on easy puzzles it amplifies already-preferred correct moves, on hard puzzles it surfaces correct moves nearly absent under SFT; Same predictive pattern transfers to a 1B language model trained on math text
  • Limitations: Chess is a narrow, fully-verifiable domain; transfer beyond the math-domain check is not demonstrated. Compute sweeps are still expensive outside the testbed.
  • Original Discord link was the /pdf/ form; canonical abs page used.