← Back to explorer

Dr.LLM: Dynamic Layer Routing in LLMs

Type
other
Venue
arXiv / Parameter Lab / MBZUAI / NAVER AI Lab

Summary

Attaches a tiny MLP router to each frozen block; routers are supervised on 4k length-aware MCTS paths (skip/execute/repeat) from ARC and DART-Math, then run search-free at inference. On six LLaMA-3.2/Qwen-2.5 models, in-domain accuracy rises in all cases (up to +3.4%p / +4.0%p on DART) with ~3–11 fewer layers per query; OOD drop 0.85%p average. Beats LayerSkip/ShortGPT/MindSkip/FlexiDepth by up to +7.7%p avg on GSM8k/MMLU/HellaSwag/HumanEval. Code https://github.com/parameterlab/dr-llm.

Keywords

dr-llm · layer-routing · mcts · adaptive-depth · skip-repeat · iclr · mbzuai · naver

Topics

adaptive depth, layer routing, efficient inference

Research notes

  • Primary: arxiv abs (cs.CL; also cs.AI, cs.LG). License not stated on abs/HTML at check. Parameter Lab / MBZUAI / NAVER AI Lab / University of Tübingen / Tübingen AI Center. Heakl corresponding. Code https://github.com/parameterlab/dr-llm (57 stars at check). HF paper page 32 upvotes; githubRepo linked; no linked models/datasets. Discord posted abs (Substack UTM). MCTS supervision is derived from public ARC/DART-Math, not a hosted corpus, so no datasets_local row. License field left blank per catalog convention.