← Back to explorer

RSIArena: LIVE RSI Arena (8 agents compete to post-train one 30B base model)

Type
project
Venue
Bake AI
Year
2026
Source
demo
Access
public
Language
en
Added
2026-09-30
Verified
2026-09-30

Summary

A live public experiment asking whether AI agents can train a model that wins in human evaluation ("RSI" = recursive self-improvement: AI that makes AI better). Eight agents compete to autonomously post-train the same base model, NVIDIA Nemotron 3.5 Lightning 30B-A3B-Base-BF16: GPT-6 Astra (OpenAI), MiMo-V2.6-Pro (Xiaomi), Grok 4.7 (xAI), DeepSeek V4.1 Flash, GLM-5.3 (Z.ai), Kimi K3 (Moonshot AI), Gemini 3.8 Flash (Google), and Muse Spark 1.3 (Meta). Stage 1 gives each agent $300 of API credit and 1,000 GPU-hours over 144 hours (144 rounds) on a shared RSIBox cluster of 64 RTX PRO 6000 Blackwell GPUs across 8 Slurm servers; stage 2 adds 500 GPU-hours per agent. A live dashboard shows training progress, budgets, GPU usage, trajectories, and SFT runs. Stage 1 checkpoints are judged by humans plus held-out tests at COLM 2026 in San Francisco (arena opens Oct 6 10:00 AM PT, prize draw Oct 9 2:00 PM PT); stage 2 trains daily on that human feedback, with a public report due mid-October. Spectators predict the top three agents for prizes. Organizers state model checkpoints and data will be open-sourced.

Keywords

agents · recursive self-improvement · post-training · live evaluation · benchmarks

Topics

agents, recursive self-improvement, post-training, live evaluation

Research notes

  • Discovery: Zichen Chen (@my_cat_can_code, verified) 2026-09-30: https://x.com/my_cat_can_code/status/2105162779904872879
  • Project site https://rsiarena.org/
  • COLM 2026 San Francisco, Hilton Union Square; example human-eval questions shown on dashboard.
  • No paper, code repo, or dataset URLs published as of the read; check back after the mid-October public report.
  • Connects to the collection's agent, post-training, and evaluation entries.