← Back to explorer

Trust the Critic More

Type
paper
Venue
arXiv:2609.39247 (cs.LG), submitted 30 Sep 2026 (v1), revised 1 Oct 2026 (v2)
Year
2026
Source
arxiv
Access
free
Language
English
Added
2026-10-02
Verified
2026-10-02

Summary

Standard LLM RL algorithms credit every token of a long rollout with the same advantage from the terminal reward; actor-critic methods could give finer-grained credit, but learned critics are considered too inaccurate to trust, so existing critics are used only for baseline estimation and every trajectory must be rolled to completion. AC2 (Actor-Critic with Action Chunking) removes the need to roll every trajectory to completion: it assigns credit to action chunks — short continuations of prefixes of past trajectories — where a learned critic scores the state reached at the end of each chunk, allowing policy updates without observing a terminal reward. Reliability comes from three design choices: (1) local readiness — critic-based updates on a problem only when the critic is sufficiently accurate on that particular problem; (2) when available, the critic gets a reference solution from a previous successful rollout; (3) credit over action chunks of 10k tokens rather than individual tokens. Training Qwen3-4B on FineProofs-RL and evaluating on IMO-ProofBench, AC2 exceeds GRPO's peak validation score of 18.5% using 2.5x fewer decoding FLOPs (25% fewer steps plus fewer tokens per step since trajectories aren't continued to completion). The critic is the only source of advantages, and it works even on hard tasks like Lean theorem proving. AC2 is a drop-in GRPO replacement: same objective and implementation details (e.g., clip higher), no PPO tricks.

Keywords

Trust the Critic More · AC2 · actor-critic · action chunking · GRPO · RLHF · LLM RL · critic pretraining · FineProofs-RL · IMO-ProofBench

Topics

actor-critic, GRPO, RLHF, action chunking, critic pretraining, LLM RL, credit assignment, theorem proving

Research notes

  • Discovery: @wen_kaiyue (Kaiyue Wen) X thread 2026-10-02 (https://x.com/wen_kaiyue/status/2106052620507091244)
  • Code: https://github.com/WhenWen/AC2
  • Extra ingredients: rollout M chunks (e.g., 2-4), critic pretrained by reusing the first few GRPO training steps with an MSE loss; chunk-level updates gated by problem-level readiness (critic MAE low for most problems) and sample-level readiness (all chunk values agree on sign)
  • Result: Qwen3-4B on FineProofs-RL, evaluated on IMO-ProofBench: matches GRPO's peak with 2.5x fewer decoding FLOPs
  • License: CC BY 4.0