When Agents Slow Down: Understanding LLM Agents' Test-Time Strategies via Elo-per-token Analysis
- Type
- paper
- Venue
- arXiv
- Year
- 2026
- Source
- arxiv
- Access
- public
- Language
- en
- Added
- 2026-09-30
- Verified
- 2026-09-30
Summary
LLM agents allocate test-time compute adaptively as they revise solutions, use tools, explore alternatives, and decide when to stop, making it hard to measure how agent performance scales. The paper studies open-ended tasks with continuous scores for intermediate submissions and proposes Elo-per-token analysis: track the best solution found at each token budget and use a Bradley-Terry model to aggregate within-task orderings into Elo ratings across tasks with different score scales. Applied to four general-purpose agents on four open-ended benchmarks with sessions up to 100M tokens, plus three feedback-driven LLM optimization harnesses in controlled single-task interventions. Independent sampling provides a theoretically characterized reference for which Elo grows linearly with log compute; against it, agents initially convert tokens into Elo faster but their marginal gains diminish and eventually fall below the reference. The strongest historical human contestants instead improve superlinearly over contest time on shared AtCoder Heuristic Contest tasks, showing continual learning and substantial headroom after agents slow down. The scaling inflection point is defined as the per-session budget where marginal Elo gains match the independent-sampling reference; splitting 100M tokens across parallel sessions on FrontierCS Polyomino Packing gains +264 Elo over one long session and +355 over ten short sessions.
Keywords
agents · test-time scaling · Elo · evaluation · hybrid agents · benchmark scaling
Topics
agents, test-time scaling, evaluation, Elo
Research notes
- Discovery: Wenhao Chai (@wenhaocha1, verified; PhD PrincetonCS; prev GoogleDeepMind; RSI is AI for AI R&D) on 2026-09-30: https://x.com/wenhaocha1/status/2105168366126117102
- Shared alongside METR's 2026-07-21 expenditure-horizon blog post (separately cataloged) with the comment that AI+Human hybrid beats AI or Human alone and that placing humans in the loop needs design.
- Submitted 2026-09-14, cs.CL.
- Connects to the collection's agent, test-time scaling, and evaluation entries.