Shockingly Simple Self-retrospection Improves Agentic Models Without RL
- Type
- paper
- Venue
- arXiv
- Year
- 2026
- Source
- arxiv
- Access
- public
- Language
- en
- Added
- 2026-09-29
- Verified
- 2026-09-29
Summary
Proposes Retrospection-Only Fine-Tuning (ROFT): post-train agentic models using only the agent's own self-generated explanations, with no external teacher and no reward-based policy update.
Keywords
agentic models · self-retrospection · SWE-bench · post-training
Topics
agentic models, self-retrospection, SWE-bench, post-training
Research notes
- Discovery: Shared in #random-papers as an arXiv link.
- Method: ROFT: collect the agent's self-generated retrospective explanations of its actions and fine-tune on those explanations only.
- Key findings: On held-out SWE-bench Verified/Pro with Qwen3.5-4B: 49.2%/26.8% after 20 updates, vs GRPO 48.0%/25.3% after 40 updates in the evaluated runs.
- Shares authors with ProgramDistill (arXiv:2609.18805); both are Microsoft Research agent papers from the same week.