Thinking Before Thinking: Scaling Agentic Inference Through Meta-Reasoning
- Type
- paper
- Venue
- arXiv:2609.38147 (cs.AI), submitted 29 Sep 2026
- Year
- 2026
- Source
- arxiv
- Access
- free
- Language
- English
- Added
- 2026-10-02
- Verified
- 2026-10-02
Summary
Introduces agentic meta-reasoning, an inference-time harness that makes agent control choices — which partial work to build on, whether to start fresh, when to stop — an explicit, structured reasoning process. Workers carry out task-level computation while a controller consolidates what the run has established, explores next options, assesses each option's worth under the remaining budget, and dispatches chosen work with context from persistent memory; between decisions the controller carries only a compact account of the run rather than replaying full history. Against production coding agents (Codex, Claude Code), research harnesses, and a Direct Control Agent baseline with the same workers and compute allowance, meta-reasoning reaches 71.5% on ProgramBench with GPT-5.5 vs 58.0% for Codex, and 67.2% with Opus 4.8 vs 65.5% for Claude Code; on other benchmarks (abstract reasoning, multi-domain long-horizon reasoning, proof generation) it gains 3.6-4.2 points over direct control averaged across three frontier models. It keeps improving over tested budget ranges where direct control plateaus, though its overhead can hurt at small budgets; artifact-graph analysis shows more reuse of earlier work and higher coverage of correct solutions.
Keywords
meta-reasoning · agentic inference · agent controller · long-horizon · ProgramBench · inference-time scaling · direct control · artifact graph
Topics
agents, meta-reasoning, inference-time scaling, long-horizon tasks, agent control, ProgramBench
Research notes
- Discovery: @arXivBangers X post 2026-10-02 (https://x.com/arXivBangers/status/2105903402400264595)
- Controller/worker design: controller consolidates run state, explores options, values them under remaining budget, dispatches with persistent-memory context; carries a compact run account between decisions instead of full history
- Baselines: production coding agents (Codex GPT-5.5 58.0%, Claude Code Opus 4.8 65.5% on ProgramBench) and a Direct Control Agent using the same workers and compute budget allowance; meta-reasoning: 71.5% (GPT-5.5) and 67.2% (Opus 4.8) on ProgramBench; +3.6 to +4.2 points over direct control on abstract reasoning, multi-domain long-horizon reasoning, and proof generation, averaged across three frontier models
- Scaling behavior: keeps improving over tested budget ranges where direct control plateaus; controller overhead can hurt at small budgets
- Artifact-graph analysis: more reuse of earlier work, higher coverage of correct solutions in most settings, nonuniform gains in final selection
- Related catalog entries: ProgramBench codebases in Datasets row 82; harness-focused Papers 920 (multi-harness RL) and 932 (Harness Learning)
- License: CC BY 4.0