Trust the Critic More
- Type
- paper
- Venue
- arXiv:2609.39247 (cs.LG), submitted 30 Sep 2026 (v1), revised 1 Oct 2026 (v2)
- Year
- 2026
- Source
- arxiv
- Access
- free
- Language
- English
- Added
- 2026-10-02
- Verified
- 2026-10-02
Summary
Standard LLM RL algorithms credit every token of a long rollout with the same advantage from the terminal reward; actor-critic methods could give finer-grained credit, but learned critics are considered too inaccurate to trust, so existing critics are used only for baseline estimation and every trajectory must be rolled to completion. AC2 (Actor-Critic with Action Chunking) removes the need to roll every trajectory to completion: it assigns credit to action chunks — short continuations of prefixes of past trajectories — where a learned critic scores the state reached at the end of each chunk, allowing policy updates without observing a terminal reward. Reliability comes from three design choices: (1) local readiness — critic-based updates on a problem only when the critic is sufficiently accurate on that particular problem; (2) when available, the critic gets a reference solution from a previous successful rollout; (3) credit over action chunks of 10k tokens rather than individual tokens. Training Qwen3-4B on FineProofs-RL and evaluating on IMO-ProofBench, AC2 exceeds GRPO's peak validation score of 18.5% using 2.5x fewer decoding FLOPs (25% fewer steps plus fewer tokens per step since trajectories aren't continued to completion). The critic is the only source of advantages, and it works even on hard tasks like Lean theorem proving. AC2 is a drop-in GRPO replacement: same objective and implementation details (e.g., clip higher), no PPO tricks.
Keywords
Trust the Critic More · AC2 · actor-critic · action chunking · GRPO · RLHF · LLM RL · critic pretraining · FineProofs-RL · IMO-ProofBench
Topics
actor-critic, GRPO, RLHF, action chunking, critic pretraining, LLM RL, credit assignment, theorem proving
Research notes
- Discovery: @wen_kaiyue (Kaiyue Wen) X thread 2026-10-02 (https://x.com/wen_kaiyue/status/2106052620507091244)
- Code: https://github.com/WhenWen/AC2
- Extra ingredients: rollout M chunks (e.g., 2-4), critic pretrained by reusing the first few GRPO training steps with an MSE loss; chunk-level updates gated by problem-level readiness (critic MAE low for most problems) and sample-level readiness (all chunk values agree on sign)
- Result: Qwen3-4B on FineProofs-RL, evaluated on IMO-ProofBench: matches GRPO's peak with 2.5x fewer decoding FLOPs
- License: CC BY 4.0