← Back to explorer

AI4AI-Bench

Type
code
Venue
GitHub
Year
2026
Source
github
Access
public
Language
en
Added
2026-09-29
Verified
2026-09-29

Summary

Official implementation of the AI4AI-Bench benchmark for recursive self-improvement via training-algorithm design. Ten tasks span generation, alignment, reasoning, unlearning, pruning, RL, reward modeling, and model merging. Each run separates open-ended exploration (up to 4h, agent produces a source patch) from reproducible measurement (fresh formal retraining up to 12h, up to 3 checkpoints, frozen validation and final evaluation). Apache-2.0; 41 stars, 3 forks as of late September 2026.

Keywords

benchmark · agents · recursive self-improvement · open source

Topics

benchmark, agents, recursive self-improvement, open source

Research notes

  • Discovery: Official code release for arXiv:2608.20318 (AI4AI-Bench); also shared as a standalone Discord item.
  • Limitations: Requires Linux amd64, Python 3.10+, Docker with NVIDIA Container Toolkit and an NVIDIA GPU (official runs use one B300); no blind evaluation service yet — public defaults label receipts as non-official local results.
  • Companion paper arXiv:2608.20318 cataloged as a separate paper item.