← Back to explorer

Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning

Type
paper
Venue
arXiv / Inclusion AI
Year
2026
Source
arxiv
Access
free
Language
en
Added
2026-08-14T16:54:04Z
Verified
2026-08-14T16:54:04Z

Summary

Trains Ring-2.5-1T-Zero from Ling-2.5-1T-Base (1T MoE, 63B active) on 320xH200. Four-stage pipeline: clipped importance-sampling PG with training-inference ratio correction and token-level loss; self-distillation to compress CoT and reset the engine gap; sample-level loss; tier-based Low/Med/High depth. 1T vs 104B flash: higher ceiling and sample efficiency. Pass@1024 rises then plateaus (discovery then sharpening). Spontaneous behaviors without extra rewards: anthropomorphism, structured formatting, self-verification, parallel reasoning, context anxiety. Stage-1 AIME 2026 84.2%; stage-2 Yarn=2 AIME24 94.1 / AIME26 93.2. CoT quality (comprehensibility/reproducibility/efficiency): 6368 avg tokens vs more than 2x baselines; 100K-trace distill into Qwen2.5-32B 78.4 vs R1 72.6. Third-stage High sits slightly below the stage-2 peak (negative transfer). No official training-code repo on the abs page.

Keywords

ring-zero · zero-rl · rlvr · trillion-scale · inclusion-ai · ling-2.5 · aime

Topics

LLM reasoning, zero RL, RLVR

Research notes

  • Primary: arxiv abs (CC BY 4.0, cs.CL). Inclusion AI / Ling Team; corresponding Zhiqiang Zhang, Jun Zhou. No official training-code repo on abs (cites Areal, slime, Megatron/SGLang). HF paper page 100 upvotes. Related weights inclusionAI/Ring-2.5-1T (243 likes at check); Zero checkpoint not separately confirmed, so hf_* left blank. Training math data not released as a standalone corpus; no datasets_local row. Discord posted AlphaXiv 2607.12395v2.