Agentic Systems as Boosting Weak Reasoning Models
- Type
- other
- Venue
- arXiv / Texas A&M University / MIT
Summary
Separates proposal coverage, local identifiability, progress, and diversity. Coverage can be amplified by sampling, but critics/comparators need a local soundness signal (execution, tests, proof/type checking). Rank-based bounds for composing local selection errors; oracle best-of-k only covers task slices with nonzero useful proposal probability. On SWE-bench Verified, GPT-5.4 nano 67.0% → critic–comparator orchestration 76.4% at k=8, matching Gemini 3 Pro and Claude Opus 4.5 Thinking vs 79.0% oracle Bo8. Remaining failures are mostly shared proposal blind spots. No official code on abs.
Keywords
boosting · committee-search · swe-bench · inference-time · gpt-5.4-nano · tamu · mit · verifier
Topics
inference-time compute, multi-agent, SWE-bench, boosting
Research notes
- Primary: arxiv abs (cs.AI). CC BY 4.0 on HTML. Texas A&M / MIT; Sunkaraneni/Beneventano equal contrib (coin-flip order). Corresponding galanti@tamu.edu. No official code on abs. HF has no paper page (API 404). Discord posted abs (edited). Uses public SWE-bench Verified; no new corpus, so no datasets_local row. License field left blank per catalog convention.