How I got the highest score on ARC-AGI again swapping Python for English
- Type
- other
- Venue
- Jeremy’s Substack
- Year
- 2026
- Source
- web
- Access
- free
- Language
- en
- Added
- 2026-08-14T18:56:35Z
- Verified
- 2026-08-14T18:56:35Z
Summary
Sep 2025 writeup of Berman second ARC-AGI-Pub run. Same evolutionary test-time architecture as Dec 2024 v1 (53.6 percent on ARC-AGI-1 with Sonnet 3.5) but evolves natural-language instructions rather than Python transforms because v2 grids are too awkward to code. Grok-4 generates up to 30 instructions; subagents apply them to training grids for a cell-accuracy fitness score; top-5 get individual then pooled revisions (worst case 40 candidates per task). Reports 79.6 percent on ARC-AGI-1 at $8.42/task and 29.4 percent SoTA on ARC-AGI-2 (prior 25 percent). Code already cataloged as papers_local id 616.
Keywords
blog · arc-agi · grok-4 · evolutionary-search · test-time-compute · multi-agent · jeremy-berman
Topics
ARC-AGI, test-time compute, multi-agent
Research notes
- Primary: Substack post. Code already cataloged as papers_local id 616 https://github.com/jerber/arc-lang-public. v1 post Dec 6 2024. Leaderboard arcprize.org. Blog/method writeup, not a new hosted corpus, so no datasets_local row.