← Back to explorer

EasyPPO: Stabilizing the Critic Is Key

Type
paper
Venue
arXiv
Year
2026
Source
arxiv
Access
public
Language
en
Added
2026-09-29
Verified
2026-09-29

Summary

Argues PPO's learned critic is a major source of instability in RL for LLMs, and identifies two critic failure modes. First, filtering truncated rollouts from both actor and critic shifts the policy objective to reward conditioned on completion, letting truncation grow even as conditional reward improves. Second, heterogeneous return noise lets high-variance prompts dominate critic updates in finite batches. Introduces EasyPPO: actor-only overlong filtering (critic trains on returns from completed and truncated rollouts), noise-normalized critic regression (each prompt's critic loss weighted by the inverse standard deviation of its sampled returns, balancing noise contributions across prompts), and moderately smaller critic minibatches to confine outlier influence during gradient clipping. No new actor loss, no new policy algorithm; actor update unchanged. Across continuous-reward coding on FrontierCS, binary-reward math reasoning on AIME24, and multi-turn search on Search-R1, using Qwen3.5-9B with rollout batch sizes 512-1024, EasyPPO stays stable through the full training horizon with zero training collapse and beats vanilla PPO, VAPO, and HL-Gauss PPO; best validation scores show relative gains of 14.89%, 2.28%, and 9.47% over PPO respectively.

Keywords

PPO · reinforcement learning · post-training · critic · training stability · value function · GRPO

Topics

reinforcement learning, PPO, post-training, critic

Research notes

  • Discovery: announced by Qiuyang Mang (@MangQiuyang) on 2026-09-29 in an X thread: https://x.com/MangQiuyang/status/2105138096899764337, quote-tweeted by Wenhao Chai (@wenhaocha1, Princeton/prev Google DeepMind): https://x.com/wenhaocha1/status/2105141024309829970
  • Chai's motivation: "we brought PPO back just for one reason: get way more out of each rollout" - GRPO is hard to scale where rollouts are expensive (e.g. a training run costing thousands of dollars, or biology where identical embryos cannot be run in parallel); early work on keeping PPO training stable.
  • Submitted to arXiv 2026-09-29, cs.LG.
  • Team includes Xuanyi Zhou (project lead), Qiuyang Mang, Huanzhi Mao, Dacheng Li, Wenhao Chai, Mayank Mishra, Yichuan Wang, Karthik Narasimhan, Alvin Cheung, Joseph E. Gonzalez.
  • Code: https://github.com/EasyPPO/EasyPPO
  • Website: https://easyppo.github.io
  • Complements the collection's PPO/GRPO/RL-for-post-training entries.