← Back to explorer

(auto)²-research: SoTA on Karpathy's NanoChat Benchmark

Type
blog
Venue
rekursiv.ai blog
Year
2026
Source
web
Access
public
Language
en
Added
2026-09-29
Verified
2026-09-29

Summary

rekursiv.ai's '(auto)²-research' blog describes a swarm of AI agents that improved on the previous NanoChat state of the art in three days, reaching 0.887791 mean bits-per-byte with a 5-minute training budget on one B200. The system ran 6,164 experiments and, notably, also improved its own research process — revising instructions and handoffs between waves — using open-source tooling (Trackinizer, Priml, Configgle).

Keywords

ai-scientist · automated-research · nanochat · agents

Topics

ai-scientist, automated-research, nanochat, agents

Research notes

  • Method: A meta-harness runs waves of AI-scientist campaigns on NanoChat; each wave proposes a new harness. Teams share findings via Trackinizer (open-source graph DB of beliefs/experiments), run composable A/B tests in Priml, and inherit configs via Configgle.
  • Key findings: Final recipe: 0.887791 mean bits-per-byte across 10 seeds on Karpathy's NanoChat benchmark, within a 5-minute training budget on a single B200; 6,164 experiments launched in 3 days across waves of 2-hour research campaigns; Agents improved not only the recipe but the research process itself: revising instructions, repairing handoffs, changing testing strategies
  • Limitations: Self-reported results on a single small benchmark (NanoChat); no independent replication reported. Final evaluation required B200 hardware.
  • Follows rekursiv.ai's earlier ARC-AGI and Sudoku campaigns. Emphasis on 'speedrun science' — handoffs must be runnable, not just readable.