AI4AI-Bench: Benchmarking LLM Agents in Algorithmic Design for Recursive Self-Improvement
- Type
- paper
- Venue
- arXiv (cs.AI)
- Year
- 2026
- Source
- arxiv
- Access
- public
- Language
- en
- Added
- 2026-09-29
- Verified
- 2026-09-29
Summary
A benchmark isolating whether LLM agents can design training algorithms — the core capability recursive self-improvement turns on. Ten frozen research repositories span ten training-algorithm families; in each task an agent gets 4 hours on one B300 to rewrite the training algorithm, and its code is rerun from scratch for up to 12 hours and scored by a fixed evaluator hidden from the agent against the repo's original algorithm. Scores are normalized so 0 is an uninformative model, 0.1 is the shipped algorithm, and 1.0 is the task optimum.
Keywords
benchmark · agents · recursive self-improvement · training algorithms · evaluation
Topics
benchmark, agents, recursive self-improvement, training algorithms, evaluation
Research notes
- Method: 10 frozen repos x 10 algorithm families; agent explores up to 4h, then only a source patch crosses into a fresh formal environment for up to 12h retraining; fixed hidden evaluator scores reruns; all 10 incommensurable metrics mapped to a common 0-1 scale.
- Key findings: Mean score 0.166 across 29 configurations of 6 systems; best system reaches 0.250 — under a fifth of the distance from shipped algorithm to optimum; Most submissions never change how the model learns at all; the minority that do average 0.226 vs 0.126 for the rest; More reasoning effort raises the share of submissions that change the learning rule from 8% to 64%, and mean score from 0.094 to 0.196
- Limitations: Single-GPU (B300), fixed time budgets, and ten repositories may not generalize to full-scale RSI; scoring depends on the hidden evaluator's fidelity.
- Original Discord link was the /pdf/ form. Official code released at github.com/Einsia/AI4AI-Bench (Apache-2.0); task suite, evaluators, and all scored submissions released. Also cataloged separately as a repo item.