Reasoning Models Can Be Effective Without Thinking
- Type
- other
- Venue
- arXiv / UC Berkeley / Allen Institute for AI
Summary
NoThinking prefills an empty thinking block on DeepSeek-R1-Distill-Qwen-32B so the model writes the solution directly. At matched budget it often beats Thinking, especially low-budget (AMC 2023 51.3 vs 28.9 at ~700 tokens) and as pass@k grows. Parallel NoThinking + best-of-N (verifier or self-certainty) matches sequential Thinking at up to 7–9× lower latency and, with verifiers, ~4× fewer tokens (MiniF2F/ProofNet). LiveCodeBench is the main exception. No official code on abs.
Keywords
nothinking · reasoning · r1-distill · test-time-compute · best-of-n · latency · berkeley · ai2
Topics
reasoning models, test-time compute, inference efficiency
Research notes
- Primary: arxiv abs (cs.AI; also cs.CL). License not stated on abs/HTML at check. UC Berkeley / Ai2. Correspondence {windsey,jingxuan.he,csnell22,tgriggs,sewonm,matei}@berkeley.edu. No official code on abs. HF paper page 11 upvotes; no linked models/datasets. Discord posted PDF. Uses public AIME/AMC/OlympiadBench/LiveCodeBench/MiniF2F/ProofNet; no new corpus, so no datasets_local row. License field left blank per catalog convention.