← Back to explorer

Extrapolative Weight Averaging Reveals Correctness-Efficiency Frontiers in Code RL

Type
paper
Venue
arXiv
Year
2026
Source
arxiv
Access
public
Language
en
Added
2026-09-29
Verified
2026-09-29

Summary

Studies RL for competitive programming where hidden unit tests enforce both correctness and efficiency. Training checkpoints under nested unit-test coverage reveals a correctness-efficiency frontier: higher-coverage rewards reduce optimization failures but increase correctness failures. Interpolation between low- and high-coverage checkpoints recovers the frontier; extrapolation extends it beyond trained endpoints, across pure reasoning, tool-use, and agentic coding settings and 32B/7B scales. Ensembles with extrapolative weight averaging broaden coverage and improve pass@250 on LCB/hard by 3.3% over the best single checkpoint.

Keywords

weight averaging · RL · code · correctness-efficiency frontier · extrapolation · inference scaling

Topics

reinforcement learning, code generation, weight averaging

Research notes

  • Discovery: shared by Pierre Chambon (@PierreChambon6, FAIR/Meta AI & INRIA) in an X thread on 2026-09-29 presenting 3 papers on code optimization: https://x.com/PierreChambon6/status/2104966972043657560
  • Led by Kunhao Zheng (per announcement thread).
  • Method: nested unit-test coverage sweeps; interpolative and extrapolative weight averaging of RL checkpoints; ensembling for inference-time scaling.
  • Submitted 2026-05-27.