← Back to explorer

AI4AI-Bench: Benchmarking LLM Agents in Algorithmic Design for Recursive Self-Improvement

Type
paper
Venue
arXiv (cs.AI)
Year
2026
Source
arxiv
Access
public
Language
en
Added
2026-09-29
Verified
2026-09-29

Summary

A benchmark isolating whether LLM agents can design training algorithms — the core capability recursive self-improvement turns on. Ten frozen research repositories span ten training-algorithm families; in each task an agent gets 4 hours on one B300 to rewrite the training algorithm, and its code is rerun from scratch for up to 12 hours and scored by a fixed evaluator hidden from the agent against the repo's original algorithm. Scores are normalized so 0 is an uninformative model, 0.1 is the shipped algorithm, and 1.0 is the task optimum.

Keywords

benchmark · agents · recursive self-improvement · training algorithms · evaluation

Topics

benchmark, agents, recursive self-improvement, training algorithms, evaluation

Research notes

  • Method: 10 frozen repos x 10 algorithm families; agent explores up to 4h, then only a source patch crosses into a fresh formal environment for up to 12h retraining; fixed hidden evaluator scores reruns; all 10 incommensurable metrics mapped to a common 0-1 scale.
  • Key findings: Mean score 0.166 across 29 configurations of 6 systems; best system reaches 0.250 — under a fifth of the distance from shipped algorithm to optimum; Most submissions never change how the model learns at all; the minority that do average 0.226 vs 0.126 for the rest; More reasoning effort raises the share of submissions that change the learning rule from 8% to 64%, and mean score from 0.094 to 0.196
  • Limitations: Single-GPU (B300), fixed time budgets, and ten repositories may not generalize to full-scale RSI; scoring depends on the hidden evaluator's fidelity.
  • Original Discord link was the /pdf/ form. Official code released at github.com/Einsia/AI4AI-Bench (Apache-2.0); task suite, evaluators, and all scored submissions released. Also cataloged separately as a repo item.