(auto)²-research: SoTA on Karpathy's NanoChat Benchmark
- Type
- blog
- Venue
- rekursiv.ai blog
- Year
- 2026
- Source
- web
- Access
- public
- Language
- en
- Added
- 2026-09-29
- Verified
- 2026-09-29
Summary
rekursiv.ai's '(auto)²-research' blog describes a swarm of AI agents that improved on the previous NanoChat state of the art in three days, reaching 0.887791 mean bits-per-byte with a 5-minute training budget on one B200. The system ran 6,164 experiments and, notably, also improved its own research process — revising instructions and handoffs between waves — using open-source tooling (Trackinizer, Priml, Configgle).
Keywords
ai-scientist · automated-research · nanochat · agents
Topics
ai-scientist, automated-research, nanochat, agents
Research notes
- Method: A meta-harness runs waves of AI-scientist campaigns on NanoChat; each wave proposes a new harness. Teams share findings via Trackinizer (open-source graph DB of beliefs/experiments), run composable A/B tests in Priml, and inherit configs via Configgle.
- Key findings: Final recipe: 0.887791 mean bits-per-byte across 10 seeds on Karpathy's NanoChat benchmark, within a 5-minute training budget on a single B200; 6,164 experiments launched in 3 days across waves of 2-hour research campaigns; Agents improved not only the recipe but the research process itself: revising instructions, repairing handoffs, changing testing strategies
- Limitations: Self-reported results on a single small benchmark (NanoChat); no independent replication reported. Final evaluation required B200 hardware.
- Follows rekursiv.ai's earlier ARC-AGI and Sudoku campaigns. Emphasis on 'speedrun science' — handoffs must be runnable, not just readable.