← Back to explorer

Shockingly Simple Self-retrospection Improves Agentic Models Without RL

Type
paper
Venue
arXiv
Year
2026
Source
arxiv
Access
public
Language
en
Added
2026-09-29
Verified
2026-09-29

Summary

Proposes Retrospection-Only Fine-Tuning (ROFT): post-train agentic models using only the agent's own self-generated explanations, with no external teacher and no reward-based policy update.

Keywords

agentic models · self-retrospection · SWE-bench · post-training

Topics

agentic models, self-retrospection, SWE-bench, post-training

Research notes

  • Discovery: Shared in #random-papers as an arXiv link.
  • Method: ROFT: collect the agent's self-generated retrospective explanations of its actions and fine-tune on those explanations only.
  • Key findings: On held-out SWE-bench Verified/Pro with Qwen3.5-4B: 49.2%/26.8% after 20 updates, vs GRPO 48.0%/25.3% after 40 updates in the evaluated runs.
  • Shares authors with ProgramDistill (arXiv:2609.18805); both are Microsoft Research agent papers from the same week.