← Back to explorer

PivotOPD: Learning to Recover from Pivotal Mistakes in Multi-Turn Agents

Type
paper
Venue
NVIDIA LPR lab (arXiv technical report, 2026-09-30; CC BY 4.0)
Year
2026
Source
paper
Access
public
Language
en
Added
2026-10-01
Verified
2026-10-01

Summary

On-policy distillation (OPD) for training language agents with dense teacher supervision on student trajectories. In multi-turn interaction an incorrect action changes the states the student encounters later, so errors compound. Across three Qwen3 models (8B-235B), more than half of failed rollouts contain an early "pivotal mistake" -- an action that moves the agent farther from task completion -- and these are often recoverable: a few guided turns after the pivotal turn can restore task success. PivotOPD jointly trains the student to prevent pivotal mistakes and to recover from the states they create: at each pivotal mistake the teacher provides a gold action plus recovery actions for the next few turns. Preventive distillation uses the gold action with reverse KL to steer the student away from the mistake; recovery distillation uses the recovery actions with forward KL to transfer recovery behaviors the student rarely samples. Strongest average performance vs 13 baselines on ALFWorld, WebShop, and Search-based QA for Qwen3-1.7B and Qwen3-8B students (+5.5% over the strongest baseline on ALFWorld with the 1.7B student); gains transfer to SWE-Bench Verified, raising a Nemotron-3.5 student's resolve rate by +3.2%.

Keywords

agents · on-policy distillation · multi-turn interaction · error recovery · language agents

Topics

agents, on-policy distillation, multi-turn interaction, error recovery, language agents

Research notes

  • Discovery: shared directly in chat (2026-10-01)
  • Project page: https://research.nvidia.com/labs/lpr/pivotopd/
  • No code repo, dataset, or model checkpoint linked on the arXiv page.
  • Connects to the collection's agent-training/distillation entries.