s1: Simple test-time scaling
- Type
- paper
- Venue
- arXiv / Stanford / Ai2
- Year
- 2026
- Source
- arxiv
- Access
- free
- Language
- en
- Added
- 2026-08-14T20:26:00Z
- Verified
- 2026-08-14T20:26:00Z
Summary
Stanford/Ai2/UW (cite Muennighoff et al.; arXiv 2501.19393). Curates s1K: 1,000 hard/diverse/quality questions with Gemini Thinking traces, distilled from a 59K pool (NuminaMATH, AIME 1983-2021, OlympicArena, OmniMath, AGIEval, plus original s1-prob/s1-teasers). SFT of Qwen2.5-32B-Instruct for 26 min on 16 H100s. Budget forcing: append end-of-think to cap tokens or suppress it and append "Wait" to extend thinking. s1-32B exceeds o1-preview on MATH/AIME24 by up to 27%; AIME24 50% without BF to 56.7% with BF (extrapolates 50%→57%). Ablations: random/diverse/longest 1K all worse (~−30% AIME24); 59K-full is not substantially better. Follow-up s1.1 regenerates traces with DeepSeek r1. Open model/data/code.
Keywords
s1 · test-time-scaling · budget-forcing · s1k · reasoning · qwen2.5 · distillation · stanford
Topics
test-time scaling, reasoning, sample-efficient SFT
Research notes
- Primary: arxiv abs 2501.19393 (cs.CL; arXiv perpetual non-exclusive 1.0). Code https://github.com/simplescaling/s1 (Apache-2.0). Data https://huggingface.co/datasets/simplescaling/s1K (Apache-2.0, 1,000 rows). Model simplescaling/s1-32B. Discord posted HTML v1. s1K is a tiny SFT set, not a 27-field scorecard corpus.