Learning More from Less: Reinforcement Learning from Hindsight
- Type
- paper
- Venue
- arXiv / MIT
- Year
- 2026
- Source
- arxiv
- Access
- free
- Language
- en
- Added
- 2026-08-14T16:58:34Z
- Verified
- 2026-08-14T16:58:34Z
Summary
LfH applies hindsight relabeling to GRPO post-training of VLAs: a VLM proposes a hindsight instruction for a low-reward group and scores each rollout against it (0/0.5/1), then the policy trains jointly on original and relabeled groups with an importance correction from commanded to hindsight instruction. On OOD LIBERO-PRO task perturbations, matches standard GRPO final success in about 5 vs 30 steps (~5x sample efficiency) and beats a RoboMETER dense progress-reward baseline by keeping ~70-80% of groups usable vs 20-40%. Gains hold on π0.5, GR00T, and OpenVLA-OFT. On a Franka FR3 held-out pick-and-place from zero SFT success, 56% vs GRPO 22% at 160 rollouts. Relabeler is Qwen3-VL-235B-A22B-Thinking-FP8. No official code on the abs page.
Keywords
lfh · hindsight · vla · grpo · libero-pro · robotics · mit
Topics
VLA post-training, hindsight relabeling, robotics
Research notes
- Primary: arxiv abs (CC BY 4.0, cs.LG). Affiliations MIT, MIT-IBM Computing Research Lab, Stanford, UC San Diego. Correspondence irisxu@mit.edu. Implemented in RLinf; relabeler Qwen3-VL-235B-A22B-Thinking-FP8. No official code on abs. Discord posted abs link. Paper not a dataset; no datasets_local row.