Open-Reasoner-Zero
- Type
- repo
- Venue
- StepFun / Tsinghua (GitHub Open-Reasoner-Zero)
- Year
- 2026
- Source
- github
- Access
- free
- Language
- en
- Added
- 2026-08-14T20:50:00Z
- Verified
- 2026-08-14T20:50:00Z
Summary
PPO reasoning RL on Qwen2.5-0.5B/1.5B/7B/32B base (same 32B base as DeepSeek-R1-Zero-Qwen-32B). Claims superior AIME2024, MATH500, and GPQA Diamond versus that pipeline at about one-tenth the steps. Releases training scripts (including 0.5B on a single A800), critic models, and curated math data: original 57k (AIME through 2023, MATH, Numina, Tulu3 MATH) plus extended 72k cleaned from OpenR1-Math-220k equals 129k, plus 13k hard mined for 32B annealing. Later ORZ-R1-Distill-Qwen-14B (Jun 2025) beats Distill-Qwen-32B on AIME 2024/2025 and MATH500. Discord posted the data tree. Paper arXiv 2503.24290. MIT.
Keywords
open-reasoner-zero · ppo · math-rl · qwen2.5 · stepfun · orz
Topics
reinforcement learning, math reasoning, LLM post-training
Research notes
- Primary: GitHub README (MIT, 2098 stars / 120 forks at check) plus arXiv 2503.24290. Discord posted the data tree. Curated 129k math data is substantial and not in datasets_local. Paper/repo not previously in papers_local. source_url is repo root (canonical) not the data subtree.