← Back to explorer

Reasoning Models Can Be Effective Without Thinking

Type
other
Venue
arXiv / UC Berkeley / Allen Institute for AI

Summary

NoThinking prefills an empty thinking block on DeepSeek-R1-Distill-Qwen-32B so the model writes the solution directly. At matched budget it often beats Thinking, especially low-budget (AMC 2023 51.3 vs 28.9 at ~700 tokens) and as pass@k grows. Parallel NoThinking + best-of-N (verifier or self-certainty) matches sequential Thinking at up to 7–9× lower latency and, with verifiers, ~4× fewer tokens (MiniF2F/ProofNet). LiveCodeBench is the main exception. No official code on abs.

Keywords

nothinking · reasoning · r1-distill · test-time-compute · best-of-n · latency · berkeley · ai2

Topics

reasoning models, test-time compute, inference efficiency

Research notes

  • Primary: arxiv abs (cs.AI; also cs.CL). License not stated on abs/HTML at check. UC Berkeley / Ai2. Correspondence {windsey,jingxuan.he,csnell22,tgriggs,sewonm,matei}@berkeley.edu. No official code on abs. HF paper page 11 upvotes; no linked models/datasets. Discord posted PDF. Uses public AIME/AMC/OlympiadBench/LiveCodeBench/MiniF2F/ProofNet; no new corpus, so no datasets_local row. License field left blank per catalog convention.