Tuning Language Models by Mixture-of-Depths Ensemble
- Type
- other
- Venue
- arXiv / University of Cambridge / Imperial College London
Summary
Not the DeepMind Mixture-of-Depths (dynamic compute); Discord flagged the name collision. MoDE reads logits off late layers (logit-lens style), trains a router, and adds per-layer norm/distillation. Plug-in on LoRA/DoRA. LLaMA2-7B arithmetic avg 54.8 (LoRA+MoDE) / 55.1 (DoRA+MoDE) vs LoRA ALL 54.5 / DoRA 54.7; LLaMA3-8B 77.6 / 77.9 vs 77.2 / 77.5. Replacing late-layer LoRA with MoDE is +0.04% params vs +10.3%. Sparse top-k routing keeps most accuracy and speeds generation up to 1.6×. No official code on abs.
Keywords
mode · mixture-of-depths · peft · lora · logit-lens · late-layers · cambridge · imperial
Topics
parameter-efficient fine-tuning, late-layer logits, ensemble routing
Research notes
- Primary: arxiv abs (cs.CL; also cs.AI). License CC BY 4.0 on HTML at check. Cambridge (hl678@cam.ac.uk) / Imperial (l.specia@imperial.ac.uk). No official code on abs. HF has no paper page (API 404). Discord posted PDF with a note that this is a different Mixture of Depths than the other one. Uses public math/commonsense sets; no new corpus, so no datasets_local row.