Mixture-of-Recursions: Learning Dynamic Recursive Depths for Adaptive Token-Level Computation
- Type
- paper
- Venue
- arXiv
- Year
- 2026
- Source
- arxiv
- Access
- public
- Language
- en
- Added
- 2026-09-29
- Verified
- 2026-09-29
Summary
Shares layers across recursion steps and routes individual tokens to different recursion depths, giving adaptive per-token compute: easy tokens get shallow passes, hard tokens recurse deeper. Experiments span 135M-1.7B parameter models.
Keywords
architecture · adaptive-compute · recursion · efficiency · pretraining
Topics
architecture, adaptive-compute, recursion, efficiency, pretraining
Research notes
- Discovery: Posted in #random-papers on 2026-09-28; arXiv abstract page fetched on 2026-09-29.
- Method: Mixture-of-Recursions (MoR): layer sharing across recursion steps plus a routing mechanism that assigns tokens to different recursion depths dynamically.
- Key findings: Reported better validation perplexity, few-shot accuracy, model size, and throughput Pareto trade-offs than vanilla and recursive baselines across 135M-1.7B scales.
- Limitations: Code URL not verified from available sources. Evaluation reported up to 1.7B parameters; results at larger scales and full limitation discussion not extracted.
- Latest revision 2025-10-25 (v2). Related conceptually to adaptive-depth / early-exit work but distinct in token-level recursion-depth routing.