PivotOPD: Learning to Recover from Pivotal Mistakes in Multi-Turn Agents
- Type
- paper
- Venue
- NVIDIA LPR lab (arXiv technical report, 2026-09-30; CC BY 4.0)
- Year
- 2026
- Source
- paper
- Access
- public
- Language
- en
- Added
- 2026-10-01
- Verified
- 2026-10-01
Summary
On-policy distillation (OPD) for training language agents with dense teacher supervision on student trajectories. In multi-turn interaction an incorrect action changes the states the student encounters later, so errors compound. Across three Qwen3 models (8B-235B), more than half of failed rollouts contain an early "pivotal mistake" -- an action that moves the agent farther from task completion -- and these are often recoverable: a few guided turns after the pivotal turn can restore task success. PivotOPD jointly trains the student to prevent pivotal mistakes and to recover from the states they create: at each pivotal mistake the teacher provides a gold action plus recovery actions for the next few turns. Preventive distillation uses the gold action with reverse KL to steer the student away from the mistake; recovery distillation uses the recovery actions with forward KL to transfer recovery behaviors the student rarely samples. Strongest average performance vs 13 baselines on ALFWorld, WebShop, and Search-based QA for Qwen3-1.7B and Qwen3-8B students (+5.5% over the strongest baseline on ALFWorld with the 1.7B student); gains transfer to SWE-Bench Verified, raising a Nemotron-3.5 student's resolve rate by +3.2%.
Keywords
agents · on-policy distillation · multi-turn interaction · error recovery · language agents
Topics
agents, on-policy distillation, multi-turn interaction, error recovery, language agents
Research notes
- Discovery: shared directly in chat (2026-10-01)
- Project page: https://research.nvidia.com/labs/lpr/pivotopd/
- No code repo, dataset, or model checkpoint linked on the arXiv page.
- Connects to the collection's agent-training/distillation entries.