← Back to explorer

Thinking Longer, Not Larger: Enhancing Software Engineering Agents via Scaling Test-Time Compute

Type
paper
Venue
arXiv / Tongyi Lab, Alibaba
Year
2026
Source
arxiv
Access
free
Language
en
Added
2026-08-14T16:36:00Z
Verified
2026-08-14T16:36:00Z

Summary

Unified TTC for SWE agents: internal TTC (development-contextualized long-CoT trajectories bootstrapped with DeepSeek R1 from GitHub issues, then rejection-sampled) plus external TTC (PRM-guided search at repository understanding, fault localization, and patch generation, with execution verification and a DPO ORM). On SWE-bench Verified, Qwen2.5-Coder-32B SWE-Reasoner reports 37.6% with internal TTC and 46% at external budget 8, matching Claude 3.5 Sonnet v2 and beating o1/DeepSeek-R1 in the paper's 2025 comparison. Code https://github.com/yingweima2022/SWE-Reasoner.

Keywords

swe-bench · test-time-compute · swe-reasoner · qwen2.5-coder · tongyi · long-cot · prm · agent

Topics

software engineering agents, test-time compute

Research notes

  • Primary: arxiv abs (default nonexclusive-distrib 1.0, cs.SE/AI). Tongyi Lab Alibaba; corresponding Yongbin Li (mayingwei.myw@alibaba-inc.com). Dong and Jiang interned from Peking University. Code https://github.com/yingweima2022/SWE-Reasoner (25 stars at check). Paper claims public release of training data/models; an Apr 2025 HF issue still asked for Hub checkpoints. 46% SOTA claim is the authors' 2025 comparison. HTML still has leftover ACM template (Woodstock 2018). Discord posted abs link.