EasyPPO: Stabilizing the Critic Is Key
- Type
- paper
- Venue
- arXiv
- Year
- 2026
- Source
- arxiv
- Access
- public
- Language
- en
- Added
- 2026-09-29
- Verified
- 2026-09-29
Summary
Argues PPO's learned critic is a major source of instability in RL for LLMs, and identifies two critic failure modes. First, filtering truncated rollouts from both actor and critic shifts the policy objective to reward conditioned on completion, letting truncation grow even as conditional reward improves. Second, heterogeneous return noise lets high-variance prompts dominate critic updates in finite batches. Introduces EasyPPO: actor-only overlong filtering (critic trains on returns from completed and truncated rollouts), noise-normalized critic regression (each prompt's critic loss weighted by the inverse standard deviation of its sampled returns, balancing noise contributions across prompts), and moderately smaller critic minibatches to confine outlier influence during gradient clipping. No new actor loss, no new policy algorithm; actor update unchanged. Across continuous-reward coding on FrontierCS, binary-reward math reasoning on AIME24, and multi-turn search on Search-R1, using Qwen3.5-9B with rollout batch sizes 512-1024, EasyPPO stays stable through the full training horizon with zero training collapse and beats vanilla PPO, VAPO, and HL-Gauss PPO; best validation scores show relative gains of 14.89%, 2.28%, and 9.47% over PPO respectively.
Keywords
PPO · reinforcement learning · post-training · critic · training stability · value function · GRPO
Topics
reinforcement learning, PPO, post-training, critic
Research notes
- Discovery: announced by Qiuyang Mang (@MangQiuyang) on 2026-09-29 in an X thread: https://x.com/MangQiuyang/status/2105138096899764337, quote-tweeted by Wenhao Chai (@wenhaocha1, Princeton/prev Google DeepMind): https://x.com/wenhaocha1/status/2105141024309829970
- Chai's motivation: "we brought PPO back just for one reason: get way more out of each rollout" - GRPO is hard to scale where rollouts are expensive (e.g. a training run costing thousands of dollars, or biology where identical embryos cannot be run in parallel); early work on keeping PPO training stable.
- Submitted to arXiv 2026-09-29, cs.LG.
- Team includes Xuanyi Zhou (project lead), Qiuyang Mang, Huanzhi Mao, Dacheng Li, Wenhao Chai, Mayank Mishra, Yichuan Wang, Karthik Narasimhan, Alvin Cheung, Joseph E. Gonzalez.
- Code: https://github.com/EasyPPO/EasyPPO
- Website: https://easyppo.github.io
- Complements the collection's PPO/GRPO/RL-for-post-training entries.