← Back to explorer

Learning More from Less: Reinforcement Learning from Hindsight

Type
paper
Venue
arXiv / MIT
Year
2026
Source
arxiv
Access
free
Language
en
Added
2026-08-14T16:58:34Z
Verified
2026-08-14T16:58:34Z

Summary

LfH applies hindsight relabeling to GRPO post-training of VLAs: a VLM proposes a hindsight instruction for a low-reward group and scores each rollout against it (0/0.5/1), then the policy trains jointly on original and relabeled groups with an importance correction from commanded to hindsight instruction. On OOD LIBERO-PRO task perturbations, matches standard GRPO final success in about 5 vs 30 steps (~5x sample efficiency) and beats a RoboMETER dense progress-reward baseline by keeping ~70-80% of groups usable vs 20-40%. Gains hold on π0.5, GR00T, and OpenVLA-OFT. On a Franka FR3 held-out pick-and-place from zero SFT success, 56% vs GRPO 22% at 160 rollouts. Relabeler is Qwen3-VL-235B-A22B-Thinking-FP8. No official code on the abs page.

Keywords

lfh · hindsight · vla · grpo · libero-pro · robotics · mit

Topics

VLA post-training, hindsight relabeling, robotics

Research notes

  • Primary: arxiv abs (CC BY 4.0, cs.LG). Affiliations MIT, MIT-IBM Computing Research Lab, Stanford, UC San Diego. Correspondence irisxu@mit.edu. Implemented in RLinf; relabeler Qwen3-VL-235B-A22B-Thinking-FP8. No official code on abs. Discord posted abs link. Paper not a dataset; no datasets_local row.