← Back to explorer

Agentic Systems as Boosting Weak Reasoning Models

Type
other
Venue
arXiv / Texas A&M University / MIT

Summary

Separates proposal coverage, local identifiability, progress, and diversity. Coverage can be amplified by sampling, but critics/comparators need a local soundness signal (execution, tests, proof/type checking). Rank-based bounds for composing local selection errors; oracle best-of-k only covers task slices with nonzero useful proposal probability. On SWE-bench Verified, GPT-5.4 nano 67.0% → critic–comparator orchestration 76.4% at k=8, matching Gemini 3 Pro and Claude Opus 4.5 Thinking vs 79.0% oracle Bo8. Remaining failures are mostly shared proposal blind spots. No official code on abs.

Keywords

boosting · committee-search · swe-bench · inference-time · gpt-5.4-nano · tamu · mit · verifier

Topics

inference-time compute, multi-agent, SWE-bench, boosting

Research notes

  • Primary: arxiv abs (cs.AI). CC BY 4.0 on HTML. Texas A&M / MIT; Sunkaraneni/Beneventano equal contrib (coin-flip order). Corresponding galanti@tamu.edu. No official code on abs. HF has no paper page (API 404). Discord posted abs (edited). Uses public SWE-bench Verified; no new corpus, so no datasets_local row. License field left blank per catalog convention.