Scaling Laws for Looped Mixture of Experts
- Type
- paper
- Venue
- arXiv:2609.40316 (cs.LG), submitted 30 Sep 2026
- Year
- 2026
- Source
- arxiv
- Access
- free
- Language
- English
- Added
- 2026-10-01
- Verified
- 2026-10-01
Summary
Introduces Loop Scaling Laws, the first scaling law to jointly model recurrence (looping) and MoE sparsity alongside model size and data. Its core is a bounded, sparsity-conditional recurrence mapping that characterizes the effective-parameter gain from looping and how sparsity raises that gain; the laws predict held-out loss of looped models more accurately than prior alternatives and recover standard dense and MoE scaling laws as special cases. Downstream, sparsity delivers ~3x active-parameter efficiency and recurrence ~2x total-parameter efficiency on reasoning, and at trillion-token scale a looped MoE with law-derived recurrence matches a ~2x larger non-looped MoE at matched training compute while enabling test-time scaling through recurrence.
Keywords
Loop Scaling Laws · looped transformers · mixture of experts · recurrence · sparsity · effective parameters · test-time scaling · parameter efficiency
Topics
scaling laws, looped transformers, mixture of experts, recurrence, sparsity, test-time scaling
Research notes
- Discovery: @anirudhg9119 X quote-post 2026-10-01 (https://x.com/anirudhg9119/status/2105797229214838997) of lead author @yanbei_c's thread (https://x.com/yanbei_c/status/2105787039853813989)
- Core: a bounded, sparsity-conditional recurrence mapping characterizes the effective-parameter gain from looping and how MoE sparsity raises that gain; looping trades extra compute for effective capacity but gains diminish and saturate; sparse routing sends tokens to different experts across loops, traversing more parameters
- Laws predict held-out loss of looped models more accurately than prior alternatives and recover standard dense and MoE scaling laws as special cases; designed to guide selection of recurrence and sparsity under fixed training compute and memory budgets
- Downstream: sparsity ~3x active-parameter efficiency, recurrence ~2x total-parameter efficiency on reasoning (complementary axes); at matched training compute and trillion-token scale, looped MoE (0.3B-active/1.3B-total) with law-derived recurrence reaches reasoning performance of a ~2x larger non-looped MoE (0.6B-active/2.9B-total), while unlocking test-time scaling by varying recurrence at inference
- Authors listed as Meta AI in the @fly51fly arXiv listing in the thread
- No code release mentioned
- License: CC BY 4.0
- Additional discovery: @fly51fly arXiv listing 2026-10-01 (https://x.com/fly51fly/status/2105772301858025607)
- From the paper's highlight figures: looping gains are bounded, against the unbounded-gain assumption in prior scaling-law work (e.g. Parcae, Iso-Depth); the mechanism is Expert-Path Diversity — sparse routing sends tokens through different experts across loops, traversing more parameters; IsoFLOP recipe: favor recurrence under tight weight-memory budgets and sparsity under large memory budgets; measured effective-parameter gain from looping saturates with recurrence and is highest at the highest sparsity (S=32)