OpenR1-Distill-7B
- Type
- other
- Venue
- HuggingFace (open-r1)
- Year
- 2026
- Source
- huggingface
- Access
- free
- Language
- English
- Added
- 2026-07-17T20:18:03.617716+00:00
- Verified
- 2026-07-17T20:18:03.617716+00:00
Summary
OpenR1-Distill-7B is a 7B-parameter language model post-trained from a variant of Qwen2.5-Math-7B (with RoPE extended to 300k for 32k context) on the Mixture-of-Thoughts dataset—350k verified reasoning traces distilled from DeepSeek-R1 spanning mathematics, coding, and science. It matches or exceeds DeepSeek-R1-Distill-Qwen-7B on benchmarks (AIME 2024: 52.7, MATH-500: 89.0, GPQA Diamond: 52.8, LiveCodeBench v5: 39.4) while being fully open and reproducible. The model is part of the Open R1 project, which aims to reproduce DeepSeek-R1's reasoning training pipeline openly.
Keywords
reasoning distillation deepseek-r1 qwen math coding science rlvr open-source llm
Topics
NLP / Reasoning
Research notes
- Training logs available at wandb.ai/huggingface/open-r1. Evaluation logs at huggingface.co/datasets/open-r1/details-open-r1_OpenR1-Distill-7B. Code at github.com/huggingface/open-r1.