← Back to explorer

Efficient Agentic Reasoning Through Self-Regulated Simulative Planning

Type
paper
Venue
arXiv / IFM / Carnegie Mellon University
Year
2026
Source
arxiv
Access
free
Language
en
Added
2026-08-14T19:20:00Z
Verified
2026-08-14T19:20:00Z

Summary

SR2AM (Self-Regulated Simulative Reasoning Agentic LLM) decomposes agent deliberation into reactive execution (System I), simulative planning with the LLM as world model (System II), and a configurator that chooses whether to plan, continue, or skip (System III). Two instantiations: v0.1 records a multi-module prompted teacher (o4-mini); v1.0 reconstructs structured plans from DeepSeek-V3.2 traces. Both use SFT then GRPO. SR2AM-v0.1-8B (from Qwen3-8B) overall Pass@1 57.0, competitive with 120-355B tool-using systems; v1.0-30B (from Qwen3-30B-A3B-Thinking-2507) 71.3 Pass@1 with 25.8-95.3% fewer reasoning tokens than similar-scale agentic LLMs. RL lengthens plan horizon +22.8% while planning frequency grows only +2.0 pp. Code sailing-lab/sr2am.

Keywords

sr2am · agentic · world-model · planning · configurator · qwen3 · cmu · ifm

Topics

agentic LLMs, world models, planning

Research notes

  • Primary: arxiv abs 2605.22138. Discord/X https://x.com/mdeng34/status/2057846036367110483 via fxtwitter (reply in a thread). Related SiRA paper https://arxiv.org/abs/2507.23773. Project https://sailing-lab.github.io/sr2am-self-regulated-planning. Models https://huggingface.co/sailing-lab/SR2AM-v0.1-8B and https://huggingface.co/sailing-lab/SR2AM-v1.0-30B. Code Apache-2.0. Trains on public math/science/web sets (Guru, MegaScience, HotpotQA, etc.), not a new hosted corpus, so no datasets_local row. Co-first authors Deng/Hou/Sá Neves.