← Back to explorer

Dream-RSI: Recursive Self-Improvement through Evolving Worlds

Type
paper
Venue
arXiv:2609.14858
Year
2026
Source
github
Access
public
Language
en
Added
2026-09-29
Verified
2026-09-29

Summary

Dream-RSI is a framework for recursive self-improvement of exploration in AI-driven discovery (Tong Zheng et al., Google/DeepMind/UMD/UVA). Its key insight: a finished discovery run's decision tree is an exact replay simulator, so candidate exploration policies can be evaluated and improved offline at zero execution cost, then redeployed — closing a self-improving loop that cut discovery cost substantially on algorithm engineering, math optimization, and GPU kernel tasks.

Keywords

recursive-self-improvement · ai-scientist · exploration · agents

Topics

recursive-self-improvement, ai-scientist, exploration, agents

Research notes

  • Method: Deploy exploration policy π_t online to grow a discovery tree; convert the tree into a replay simulator; an offline policy-development agent writes M code revisions of the orchestration policy (branching/parallelism/stopping); score each by replay over history at zero cost; redeploy the winner, growing the pool of 'worlds'. Repeat.
  • Key findings: A completed discovery run's tree is an exact replay simulator: alternative exploration policies scored at zero executions ('Nothing is predicted: the simulator is exact over the search space that was realized'); Claimed: 1.7× fewer discovery-agent calls vs fixed exploration; up to 162× vs SimpleTES on Lasso; 2.43× fewer generations on KernelBench VGG16; 2.09× higher score on ConvDiv; Winner-takes-all selection guarantees the redeployed policy is never worse than the incumbent; Only the orchestration policy code changes — the coding agent, evaluator, and model weights stay fixed
  • Limitations: Results are self-reported with no independent replication; code unreleased (paper + project page only). Replay simulators only cover already-realized search space — policies can't be 'dreamt' where history never went.
  • Google / Google DeepMind / UMD / UVA. 12 pages, submitted 2026-09-14. Project page: https://dream-rsi.com. Community replications (e.g., dream-rsi-dsh plugin) report matching a circle-packing world record.