← Back to explorer

Mixture-of-Depths: Dynamically allocating compute in transformer-based language models

Type
other
Venue
arXiv / Google DeepMind

Summary

MoD caps each (or every other) block to k tokens for attention+MLP; others take a residual. Expert-choice top-k keeps a static graph. IsoFLOP: 12.5% capacity every other block matches or beats vanilla at fewer FLOPs/step (up to ~50–66% faster sampling). An auxiliary router predictor enables causal decode (~99% top-k match). No official code on abs. Not the Luo/Specia Mixture-of-Depths Ensemble (papers_local 576).

Keywords

mixture-of-depths · mod · conditional-compute · routing · deepmind · isoflop

Topics

conditional computation, Mixture-of-Depths, efficient transformers

Research notes

  • Primary: arxiv abs (cs.LG; also cs.CL). License CC BY 4.0 on HTML at check. Google DeepMind; Richards also McGill/Mila. No official code on abs. HF paper page 107 upvotes; unofficial linked models/datasets not copied into hf_* fields and not substantial, so no datasets_local row. Discord posted abs with the paper title. Name collides with papers_local 576 (MoDE ensemble).