← Back to explorer

Reinforcement Learning for Code Optimization

Type
paper
Venue
arXiv
Year
2026
Source
arxiv
Access
public
Language
en
Added
2026-09-29
Verified
2026-09-29

Summary

Makes execution time learnable for RL on code optimization via three stages: (1) DMC-Optim benchmark with large optimization tests and a calibrated timing sandbox; (2) composing correctness and speed into the reward using an offline simulator to pick configurations; (3) adapting GRPO and evaluation to the sparser, noisier timed-execution setting. Strongest configurations improve strict top-50% pass@1 from 18.0% to 31.3% on Qwen 2.5 7B and 30.7% to 50.4% on CWM 32B, with further gains at stricter percentiles (125% relative improvement for CWM 32B at top-30%) while preserving pure-correctness scores.

Keywords

RL · code optimization · GRPO · DMC-Optim · execution time reward

Topics

reinforcement learning, code optimization

Research notes

  • Discovery: shared by Pierre Chambon (@PierreChambon6, FAIR/Meta AI & INRIA) in an X thread on 2026-09-29 presenting 3 papers on code optimization: https://x.com/PierreChambon6/status/2104966972043657560
  • Method: DMC-Optim benchmark + calibrated sandbox; correctness-speed reward composition via offline simulator; GRPO adapted to sparse noisy timed rewards.
  • Submitted 2026-07-28.