AI4AI-Bench
- Type
- code
- Venue
- GitHub
- Year
- 2026
- Source
- github
- Access
- public
- Language
- en
- Added
- 2026-09-29
- Verified
- 2026-09-29
Summary
Official implementation of the AI4AI-Bench benchmark for recursive self-improvement via training-algorithm design. Ten tasks span generation, alignment, reasoning, unlearning, pruning, RL, reward modeling, and model merging. Each run separates open-ended exploration (up to 4h, agent produces a source patch) from reproducible measurement (fresh formal retraining up to 12h, up to 3 checkpoints, frozen validation and final evaluation). Apache-2.0; 41 stars, 3 forks as of late September 2026.
Keywords
benchmark · agents · recursive self-improvement · open source
Topics
benchmark, agents, recursive self-improvement, open source
Research notes
- Discovery: Official code release for arXiv:2608.20318 (AI4AI-Bench); also shared as a standalone Discord item.
- Limitations: Requires Linux amd64, Python 3.10+, Docker with NVIDIA Container Toolkit and an NVIDIA GPU (official runs use one B300); no blind evaluation service yet — public defaults label receipts as non-official local results.
- Companion paper arXiv:2608.20318 cataloged as a separate paper item.