← Back to explorer

DeepScaleR-1.5B-Preview

Type
other
Venue
Hugging Face (agentica-org)
Year
2026
Source
huggingface
Access
free
Language
English
Added
2026-07-17T20:18:03.590945+00:00
Verified
2026-07-17T20:18:03.590945+00:00

Summary

DeepScaleR-1.5B-Preview is a 1.5B-parameter language model fine-tuned from DeepSeek-R1-Distilled-Qwen-1.5B using distributed reinforcement learning (GRPO) with progressive context length extension (8K→16K→24K). It achieves 43.1% Pass@1 accuracy on AIME 2024, surpassing OpenAI's O1-Preview with just 1.5B parameters, and was trained on approximately 40,000 problem-answer pairs from AIME, AMC, Omni-MATH, and Still datasets.

Keywords

math reasoning reinforcement-learning grpo small-lm aime distill preview

Topics

Math / Reasoning

Research notes

  • Model checkpoint, not a dataset. Trained on 8-32 A100-80GB GPUs. Uses DeepSeek's Group Relative Policy Optimization (GRPO).