Thinking Longer, Not Larger: Enhancing Software Engineering Agents via Scaling Test-Time Compute
- Type
- paper
- Venue
- arXiv / Tongyi Lab, Alibaba
- Year
- 2026
- Source
- arxiv
- Access
- free
- Language
- en
- Added
- 2026-08-14T16:36:00Z
- Verified
- 2026-08-14T16:36:00Z
Summary
Unified TTC for SWE agents: internal TTC (development-contextualized long-CoT trajectories bootstrapped with DeepSeek R1 from GitHub issues, then rejection-sampled) plus external TTC (PRM-guided search at repository understanding, fault localization, and patch generation, with execution verification and a DPO ORM). On SWE-bench Verified, Qwen2.5-Coder-32B SWE-Reasoner reports 37.6% with internal TTC and 46% at external budget 8, matching Claude 3.5 Sonnet v2 and beating o1/DeepSeek-R1 in the paper's 2025 comparison. Code https://github.com/yingweima2022/SWE-Reasoner.
Keywords
swe-bench · test-time-compute · swe-reasoner · qwen2.5-coder · tongyi · long-cot · prm · agent
Topics
software engineering agents, test-time compute
Research notes
- Primary: arxiv abs (default nonexclusive-distrib 1.0, cs.SE/AI). Tongyi Lab Alibaba; corresponding Yongbin Li (mayingwei.myw@alibaba-inc.com). Dong and Jiang interned from Peking University. Code https://github.com/yingweima2022/SWE-Reasoner (25 stars at check). Paper claims public release of training data/models; an Apr 2025 HF issue still asked for Hub checkpoints. 46% SOTA claim is the authors' 2025 comparison. HTML still has leftover ACM template (Woodstock 2018). Discord posted abs link.