← Back to explorer

OpenR1-Distill-7B

Type
other
Venue
HuggingFace (open-r1)
Year
2026
Source
huggingface
Access
free
Language
English
Added
2026-07-17T20:18:03.617716+00:00
Verified
2026-07-17T20:18:03.617716+00:00

Summary

OpenR1-Distill-7B is a 7B-parameter language model post-trained from a variant of Qwen2.5-Math-7B (with RoPE extended to 300k for 32k context) on the Mixture-of-Thoughts dataset—350k verified reasoning traces distilled from DeepSeek-R1 spanning mathematics, coding, and science. It matches or exceeds DeepSeek-R1-Distill-Qwen-7B on benchmarks (AIME 2024: 52.7, MATH-500: 89.0, GPQA Diamond: 52.8, LiveCodeBench v5: 39.4) while being fully open and reproducible. The model is part of the Open R1 project, which aims to reproduce DeepSeek-R1's reasoning training pipeline openly.

Keywords

reasoning distillation deepseek-r1 qwen math coding science rlvr open-source llm

Topics

NLP / Reasoning

Research notes

  • Training logs available at wandb.ai/huggingface/open-r1. Evaluation logs at huggingface.co/datasets/open-r1/details-open-r1_OpenR1-Distill-7B. Code at github.com/huggingface/open-r1.