← Back to explorer

Scaling Laws for Looped Mixture of Experts

Type
paper
Venue
arXiv:2609.40316 (cs.LG), submitted 30 Sep 2026
Year
2026
Source
arxiv
Access
free
Language
English
Added
2026-10-01
Verified
2026-10-01

Summary

Introduces Loop Scaling Laws, the first scaling law to jointly model recurrence (looping) and MoE sparsity alongside model size and data. Its core is a bounded, sparsity-conditional recurrence mapping that characterizes the effective-parameter gain from looping and how sparsity raises that gain; the laws predict held-out loss of looped models more accurately than prior alternatives and recover standard dense and MoE scaling laws as special cases. Downstream, sparsity delivers ~3x active-parameter efficiency and recurrence ~2x total-parameter efficiency on reasoning, and at trillion-token scale a looped MoE with law-derived recurrence matches a ~2x larger non-looped MoE at matched training compute while enabling test-time scaling through recurrence.

Keywords

Loop Scaling Laws · looped transformers · mixture of experts · recurrence · sparsity · effective parameters · test-time scaling · parameter efficiency

Topics

scaling laws, looped transformers, mixture of experts, recurrence, sparsity, test-time scaling

Research notes

  • Discovery: @anirudhg9119 X quote-post 2026-10-01 (https://x.com/anirudhg9119/status/2105797229214838997) of lead author @yanbei_c's thread (https://x.com/yanbei_c/status/2105787039853813989)
  • Core: a bounded, sparsity-conditional recurrence mapping characterizes the effective-parameter gain from looping and how MoE sparsity raises that gain; looping trades extra compute for effective capacity but gains diminish and saturate; sparse routing sends tokens to different experts across loops, traversing more parameters
  • Laws predict held-out loss of looped models more accurately than prior alternatives and recover standard dense and MoE scaling laws as special cases; designed to guide selection of recurrence and sparsity under fixed training compute and memory budgets
  • Downstream: sparsity ~3x active-parameter efficiency, recurrence ~2x total-parameter efficiency on reasoning (complementary axes); at matched training compute and trillion-token scale, looped MoE (0.3B-active/1.3B-total) with law-derived recurrence reaches reasoning performance of a ~2x larger non-looped MoE (0.6B-active/2.9B-total), while unlocking test-time scaling by varying recurrence at inference
  • Authors listed as Meta AI in the @fly51fly arXiv listing in the thread
  • No code release mentioned
  • License: CC BY 4.0
  • Additional discovery: @fly51fly arXiv listing 2026-10-01 (https://x.com/fly51fly/status/2105772301858025607)
  • From the paper's highlight figures: looping gains are bounded, against the unbounded-gain assumption in prior scaling-law work (e.g. Parcae, Iso-Depth); the mechanism is Expert-Path Diversity — sparse routing sends tokens through different experts across loops, traversing more parameters; IsoFLOP recipe: favor recurrence under tight weight-memory budgets and sparsity under large memory budgets; measured effective-parameter gain from looping saturates with recurrence and is highest at the highest sparsity (S=32)