← Back to explorer

rLLM / DeepScaleR: Agentic RL framework (Agentica)

Type
repo
Venue
Agentica / Berkeley Sky Computing Lab (GitHub rllm-org)
Year
2026
Source
github
Access
free
Language
en
Added
2026-08-14T20:50:00Z
Verified
2026-08-14T20:50:00Z

Summary

agentica-project/deepscaler redirected to rllm-org/rllm. Original DeepScaleR work (Feb 2025 blog) scaled RL on DeepSeek-R1-Distill-Qwen-1.5B using ~40k unique math Q-A pairs (AIME 1984-2023, AMC pre-2023, Omni-MATH, Still). Current rLLM is a general agentic RL stack: GRPO/REINFORCE/RLOO/SFT/on-policy distillation; backends verl, tinker, fireworks; 60+ evals. Later Agentica results include DeepCoder-14B and DeepSWE-32B. Discord posted the old deepscaler/verl tree.

Keywords

deepscaler · rllm · agentica · grpo · math-rl · verl · berkeley-sky

Topics

reinforcement learning, LLM post-training, math reasoning, agentic RL

Research notes

  • Primary: GitHub README + API (Apache-2.0). Discord posted https://github.com/agentica-project/deepscaler/tree/main/verl which now 301s to rllm-org/rllm. Blog https://pretty-radio-b75.notion.site/DeepScaleR-Surpassing-O1-Preview-with-a-1-5B-Model-by-Scaling-RL-19681902c1468005bed8ca303013a4e2. Preview dataset/model not in datasets_local as themselves (only uiuc-kang-lab noisy-label derivatives tied to papers_local 481). Not previously cataloged as DeepScaleR paper/repo.