Huxley-Gödel Machine: Human-Level Coding Agent Development by an Approximation of the Optimal Self-Improving Machine
- Type
- other
- Venue
- arXiv / KAUST
Summary
Identifies a Metaproductivity–Performance Mismatch: high SWE scores need not predict productive descendants. Clade-metaproductivity (CMP) aggregates descendant outcomes; under coding-agent assumptions a CMP oracle implements a Gödel Machine (Theorem 1). HGM estimates CMP and Thompson-samples expansion vs evaluation asynchronously. SWE-Verified-60: 56.7% vs DGM 53.3 / SICA 50.0 with 2.38× fewer CPU-hours (517 vs 1231). Full SWE-Verified 61.4% (GPT-5-mini). Agent transfers to SWE-Lite+GPT-5 at 57%, matching the best officially checked human-engineered agents. Code https://github.com/metauto-ai/HGM.
Keywords
hgm · godel-machine · self-improvement · swe-bench · coding-agents · kaust · dgm
Topics
coding agents, self-improvement, Gödel machines
Research notes
- Primary: arxiv abs (cs.AI). License not stated on abs/HTML at check. KAUST (Center of Excellence for Generative AI, award 5940). Equal contrib Wang/Piękos. Correspondence via {wenyi.wang, piotr.piekos, ...}@kaust.edu.sa. Code https://github.com/metauto-ai/HGM (410 stars at check). HF paper page 22 upvotes; githubRepo linked; no linked models/datasets. Discord posted abs. Uses public SWE-bench/Polyglot; no new corpus, so no datasets_local row. License field left blank per catalog convention.