AIDE²: The First Evidence of Recursive Self-Improvement
- Type
- other
- Venue
- Weco AI Blog
- Year
- 2026
- Source
- web
- Access
- free
- Language
- en
- Added
- 2026-08-14T19:35:00Z
- Verified
- 2026-08-14T19:35:00Z
Summary
Weco frames RSI as bi-level optimization: AIDEhuman (Claude Opus 4.7) rewrites a simplified AIDE0 inner agent (Gemini 3 Flash) under a fixed dollar budget, keeping a rewrite only if private held-out scores improve. 100 unattended outer steps over eight days yielded seven successive keepers; AIDE47 (best @50) and AIDE85 (best @100) beat the two-year hand-tuned AIDEhuman on held-out MLE-Bench Lite, ALE-Bench Lite, and out-of-distribution WeatherBench 2. Emergent anti-hacking: KernelBench reward-hack rate 63%→34% via prompt guards plus hardcoded checks (statistical layer later found buggy). Claims Level 1 (net-positive) on Weco’s RSI ladder; ignition test with AIDE47 as outer loop mixed, no Level 2 claim. PDF/AIDE85 release promised later.
Keywords
aide2 · rsi · weco · autoresearch · mle-bench · weatherbench · reward-hacking · blog · x
Topics
recursive self-improvement, autonomous research agents
Research notes
- Primary: Weco blog (cite weco2026aide2). Discord/X https://x.com/zhengyaojiang/status/2077079778793042425 via fxtwitter (1/7 thread; unroll links the blog). Companion RSI ladder https://www.weco.ai/blog/4-levels-of-recursive-self-improvement. Original AIDE arXiv 2502.13138; MLE-Bench 2410.07095. WecoAI/AIDE GitHub 404 at check. PDF technical report and AIDE85 code not yet released. Method/blog item, not a hosted corpus, so no datasets_local row.