Extrapolative Weight Averaging Reveals Correctness-Efficiency Frontiers in Code RL
- Type
- paper
- Venue
- arXiv
- Year
- 2026
- Source
- arxiv
- Access
- public
- Language
- en
- Added
- 2026-09-29
- Verified
- 2026-09-29
Summary
Studies RL for competitive programming where hidden unit tests enforce both correctness and efficiency. Training checkpoints under nested unit-test coverage reveals a correctness-efficiency frontier: higher-coverage rewards reduce optimization failures but increase correctness failures. Interpolation between low- and high-coverage checkpoints recovers the frontier; extrapolation extends it beyond trained endpoints, across pure reasoning, tool-use, and agentic coding settings and 32B/7B scales. Ensembles with extrapolative weight averaging broaden coverage and improve pass@250 on LCB/hard by 3.3% over the best single checkpoint.
Keywords
weight averaging · RL · code · correctness-efficiency frontier · extrapolation · inference scaling
Topics
reinforcement learning, code generation, weight averaging
Research notes
- Discovery: shared by Pierre Chambon (@PierreChambon6, FAIR/Meta AI & INRIA) in an X thread on 2026-09-29 presenting 3 papers on code optimization: https://x.com/PierreChambon6/status/2104966972043657560
- Led by Kunhao Zheng (per announcement thread).
- Method: nested unit-test coverage sweeps; interpolative and extrapolative weight averaging of RL checkpoints; ensembling for inference-time scaling.
- Submitted 2026-05-27.