Finetuning with Sampling: SFT Learns Better Than You Think
- Type
- paper
- Venue
- arXiv:2610.02140 (cs.LG), submitted 1 Oct 2026; to be presented at NeurIPS
- Year
- 2026
- Source
- arxiv
- Access
- free
- Language
- English
- Added
- 2026-10-02
- Verified
- 2026-10-02
Summary
Conventional wisdom: RL generalizes well on new tasks without losing existing capabilities, while SFT suffers weak generalization and catastrophic forgetting — but SFT can learn from off-policy expert data, whereas RL must find successful trajectories by repeated sampling. Instead of modifying the learning objective for off-policy data, this work tailors the data distribution to the learner: an MCMC (Metropolis-Hastings) sampling algorithm progressively transforms off-policy traces to be more on-policy given a reference model. The target distribution is the KL-closest distribution to the base model consistent with a privileged information constraint C; the information constraint is absorbed into the MCMC proposer (used in-context to generate candidates) rather than as a verifier, and the data processing inequality guarantees the trajectory distribution moves closer to the base model as MCMC progresses. Across scientific skill acquisition, mathematical reasoning, and open-ended expertise, sampling-augmented SFT rivals prevailing post-training techniques — often generalizing better and forgetting less than strong on-policy baselines like RL and on-policy distillation — and the pass@k curve can exceed both base model and on-policy learning, i.e., learning genuinely new abilities rather than just sharpening. Sampling is framed as a model-native operator shaping data for learnability, with broader utility as a general primitive across the post-training stack (RL, OPD, self-play).
Keywords
Finetuning with Sampling · SFT · MCMC · Metropolis-Hastings · off-policy · on-policy · post-training · catastrophic forgetting · learnability
Topics
SFT, post-training, MCMC, Metropolis-Hastings, off-policy data, on-policy learning, generalization, forgetting
Research notes
- Discovery: @aakaran31 (Aayush Karan) 11-part X thread 2026-10-02 (https://x.com/aakaran31/status/2106037829059133903)
- Blog: https://aakaran.github.io/finetuning_with_sampling/
- Code: https://github.com/aakaran/finetuning-with-sampling
- Method: sample from the distribution closest in KL divergence to the base model that outputs trajectories consistent with privileged information constraint C; Metropolis-Hastings with the constraint absorbed into the proposer (generate in-context candidates, refine by base model likelihoods)
- Follow-up to the author's prior work on reasoning with sampling; advised by @sitanch and @du_yilun
- License: CC BY 4.0