RSIArena: LIVE RSI Arena (8 agents compete to post-train one 30B base model)
- Type
- project
- Venue
- Bake AI
- Year
- 2026
- Source
- demo
- Access
- public
- Language
- en
- Added
- 2026-09-30
- Verified
- 2026-09-30
Summary
A live public experiment asking whether AI agents can train a model that wins in human evaluation ("RSI" = recursive self-improvement: AI that makes AI better). Eight agents compete to autonomously post-train the same base model, NVIDIA Nemotron 3.5 Lightning 30B-A3B-Base-BF16: GPT-6 Astra (OpenAI), MiMo-V2.6-Pro (Xiaomi), Grok 4.7 (xAI), DeepSeek V4.1 Flash, GLM-5.3 (Z.ai), Kimi K3 (Moonshot AI), Gemini 3.8 Flash (Google), and Muse Spark 1.3 (Meta). Stage 1 gives each agent $300 of API credit and 1,000 GPU-hours over 144 hours (144 rounds) on a shared RSIBox cluster of 64 RTX PRO 6000 Blackwell GPUs across 8 Slurm servers; stage 2 adds 500 GPU-hours per agent. A live dashboard shows training progress, budgets, GPU usage, trajectories, and SFT runs. Stage 1 checkpoints are judged by humans plus held-out tests at COLM 2026 in San Francisco (arena opens Oct 6 10:00 AM PT, prize draw Oct 9 2:00 PM PT); stage 2 trains daily on that human feedback, with a public report due mid-October. Spectators predict the top three agents for prizes. Organizers state model checkpoints and data will be open-sourced.
Keywords
agents · recursive self-improvement · post-training · live evaluation · benchmarks
Topics
agents, recursive self-improvement, post-training, live evaluation
Research notes
- Discovery: Zichen Chen (@my_cat_can_code, verified) 2026-09-30: https://x.com/my_cat_can_code/status/2105162779904872879
- Project site https://rsiarena.org/
- COLM 2026 San Francisco, Hilton Union Square; example human-eval questions shown on dashboard.
- No paper, code repo, or dataset URLs published as of the read; check back after the mid-October public report.
- Connects to the collection's agent, post-training, and evaluation entries.