The Illusion of Diminishing Returns: Measuring Long Horizon Execution in LLMs
- Type
- other
- Venue
- arXiv / University of Cambridge / Max Planck Institute for Intelligent Systems / University of Stuttgart / ELLIS Institute Tübingen
Summary
Gives the model the knowledge and plan, then measures retrieve-then-compose over a key→int dictionary. Horizon length Hs=ln(s)/ln(p) grows hyperbolically; short-task benches can hide compounding. Qwen3/Gemma3: near-perfect first-step accuracy but Qwen3-32B drops below 50% by ~15 turns; larger models execute more turns. Per-step accuracy degrades via self-conditioning on own errors, not just long context; scaling size does not fix it, thinking/RL does. Single-turn without CoT: even DeepSeek-V3/Kimi-K2 fail past ~6 steps; GPT-5 thinking 2176 steps vs Claude-4 Sonnet 432 / Grok 4 384 / Gemini 2.5 Pro 120. Published at ICLR 2026. Code https://github.com/long-horizon-execution/measuring-execution.
Keywords
long-horizon · execution · self-conditioning · iclr · scaling · test-time-compute · gpt-5 · qwen3 · gemma3
Topics
long-horizon execution, scaling laws, self-conditioning, test-time compute
Research notes
- Primary: arxiv abs (cs.AI). License not stated on abs/HTML at check. Comment: Published at ICLR 2026. Sinha Cambridge; Arun/Staab Stuttgart; Goel/Geiping MPI-IS / ELLIS Tübingen / Tübingen AI Center. Correspondence via author affiliations. Code https://github.com/long-horizon-execution/measuring-execution (57 stars at check). HF paper page 37 upvotes; githubRepo linked; 17 unofficial linked models not copied into hf_* fields. Companion HF arvindh75/Long-Horizon-Execution is a 100-row synthetic slice (15 likes / 256 downloads at check), not a substantial hosted corpus, so no datasets_local row. Discord posted abs. License field left blank per catalog convention.