Mixture-of-Depths: Dynamically allocating compute in transformer-based language models
- Type
- other
- Venue
- arXiv / Google DeepMind
Summary
MoD caps each (or every other) block to k tokens for attention+MLP; others take a residual. Expert-choice top-k keeps a static graph. IsoFLOP: 12.5% capacity every other block matches or beats vanilla at fewer FLOPs/step (up to ~50–66% faster sampling). An auxiliary router predictor enables causal decode (~99% top-k match). No official code on abs. Not the Luo/Specia Mixture-of-Depths Ensemble (papers_local 576).
Keywords
mixture-of-depths · mod · conditional-compute · routing · deepmind · isoflop
Topics
conditional computation, Mixture-of-Depths, efficient transformers
Research notes
- Primary: arxiv abs (cs.LG; also cs.CL). License CC BY 4.0 on HTML at check. Google DeepMind; Richards also McGill/Mila. No official code on abs. HF paper page 107 upvotes; unofficial linked models/datasets not copied into hf_* fields and not substantial, so no datasets_local row. Discord posted abs with the paper title. Name collides with papers_local 576 (MoDE ensemble).