rLLM / DeepScaleR: Agentic RL framework (Agentica)
- Type
- repo
- Venue
- Agentica / Berkeley Sky Computing Lab (GitHub rllm-org)
- Year
- 2026
- Source
- github
- Access
- free
- Language
- en
- Added
- 2026-08-14T20:50:00Z
- Verified
- 2026-08-14T20:50:00Z
Summary
agentica-project/deepscaler redirected to rllm-org/rllm. Original DeepScaleR work (Feb 2025 blog) scaled RL on DeepSeek-R1-Distill-Qwen-1.5B using ~40k unique math Q-A pairs (AIME 1984-2023, AMC pre-2023, Omni-MATH, Still). Current rLLM is a general agentic RL stack: GRPO/REINFORCE/RLOO/SFT/on-policy distillation; backends verl, tinker, fireworks; 60+ evals. Later Agentica results include DeepCoder-14B and DeepSWE-32B. Discord posted the old deepscaler/verl tree.
Keywords
deepscaler · rllm · agentica · grpo · math-rl · verl · berkeley-sky
Topics
reinforcement learning, LLM post-training, math reasoning, agentic RL
Research notes
- Primary: GitHub README + API (Apache-2.0). Discord posted https://github.com/agentica-project/deepscaler/tree/main/verl which now 301s to rllm-org/rllm. Blog https://pretty-radio-b75.notion.site/DeepScaleR-Surpassing-O1-Preview-with-a-1-5B-Model-by-Scaling-RL-19681902c1468005bed8ca303013a4e2. Preview dataset/model not in datasets_local as themselves (only uiuc-kang-lab noisy-label derivatives tied to papers_local 481). Not previously cataloged as DeepScaleR paper/repo.