DeepScaleR-1.5B-Preview
- Type
- other
- Venue
- Hugging Face (agentica-org)
- Year
- 2026
- Source
- huggingface
- Access
- free
- Language
- English
- Added
- 2026-07-17T20:18:03.590945+00:00
- Verified
- 2026-07-17T20:18:03.590945+00:00
Summary
DeepScaleR-1.5B-Preview is a 1.5B-parameter language model fine-tuned from DeepSeek-R1-Distilled-Qwen-1.5B using distributed reinforcement learning (GRPO) with progressive context length extension (8K→16K→24K). It achieves 43.1% Pass@1 accuracy on AIME 2024, surpassing OpenAI's O1-Preview with just 1.5B parameters, and was trained on approximately 40,000 problem-answer pairs from AIME, AMC, Omni-MATH, and Still datasets.
Keywords
math reasoning reinforcement-learning grpo small-lm aime distill preview
Topics
Math / Reasoning
Research notes
- Model checkpoint, not a dataset. Trained on 8-32 A100-80GB GPUs. Uses DeepSeek's Group Relative Policy Optimization (GRPO).