Shockingly Simple Self-retrospection Improves Agentic Models Without RL
- Type
- paper
- Venue
- arXiv
- Year
- 2026
- Source
- arxiv
- Access
- public
- Language
- en
- Added
- 2026-09-29
- Verified
- 2026-09-29
Summary
Introduces Retrospection-Only Fine-Tuning (ROFT): a minimal online procedure where an LM agent attempts a task, observes feedback, generates a retrospective explanation, and is fine-tuned with next-token prediction loss on the explanation tokens alone - no teacher, no verifier, no RL updates. In software-engineering experiments with Qwen3.5-4B, ROFT reaches 49.2% and 26.8% solve rates on SWE-bench Verified and Pro after 20 updates, compared with GRPO 48.0% and 25.3% after 40 updates, with faster early progress. Shows that learning to explain can improve learning to do, establishing self-generated retrospections as useful training targets.
Keywords
retrospection · ROFT · self-improvement · agentic models · SWE-bench · GRPO
Topics
reinforcement learning, self-improvement, LLM agents, code generation
Research notes
- Discovery: announced via X thread by @lightetal (Jonathan Li) on 2026-09-29: https://x.com/lightetal/status/2105007134773768234
- Method: Retrospection-Only Fine-Tuning (ROFT): agent attempts a task, observes feedback, generates a retrospective explanation, and is fine-tuned with next-token loss on the explanation tokens alone. No external teacher, no reward-based policy update.
- Key findings: On SWE-bench Verified and Pro, ROFT with Qwen3.5-4B reaches 49.2% and 26.8% solve rates after 20 updates without a verifier, vs GRPO 48.0% and 25.3% after 40 updates; faster early progress; learns tasks where all 64 sampled base attempts failed. ROFT indirectly assigns credit to actions.
- Paper: 62 pages, 18 figures, 5 tables. Submitted 2026-09-28. Subjects: cs.AI, cs.CL.