← Back to explorer

Open-Reasoner-Zero

Type
repo
Venue
StepFun / Tsinghua (GitHub Open-Reasoner-Zero)
Year
2026
Source
github
Access
free
Language
en
Added
2026-08-14T20:50:00Z
Verified
2026-08-14T20:50:00Z

Summary

PPO reasoning RL on Qwen2.5-0.5B/1.5B/7B/32B base (same 32B base as DeepSeek-R1-Zero-Qwen-32B). Claims superior AIME2024, MATH500, and GPQA Diamond versus that pipeline at about one-tenth the steps. Releases training scripts (including 0.5B on a single A800), critic models, and curated math data: original 57k (AIME through 2023, MATH, Numina, Tulu3 MATH) plus extended 72k cleaned from OpenR1-Math-220k equals 129k, plus 13k hard mined for 32B annealing. Later ORZ-R1-Distill-Qwen-14B (Jun 2025) beats Distill-Qwen-32B on AIME 2024/2025 and MATH500. Discord posted the data tree. Paper arXiv 2503.24290. MIT.

Keywords

open-reasoner-zero · ppo · math-rl · qwen2.5 · stepfun · orz

Topics

reinforcement learning, math reasoning, LLM post-training

Research notes

  • Primary: GitHub README (MIT, 2098 stars / 120 forks at check) plus arXiv 2503.24290. Discord posted the data tree. Curated 129k math data is substantial and not in datasets_local. Paper/repo not previously in papers_local. source_url is repo root (canonical) not the data subtree.