← Back to explorer

Dynamic Mixture of Experts: An Auto-Tuning Approach for Efficient Transformer Models

Type
other
Venue
arXiv / CUHK-Shenzhen / Westlake University / Tencent AI Lab

Summary

DynMoE replaces fixed top-k with top-any gating (cosine similarity vs a trainable per-expert threshold) so tokens activate a variable number of experts, plus an adaptive add/remove process when tokens activate none or experts go unused. A diversity+simplicity auxiliary loss encourages sparse orthogonal expert representations. Competitive with tuned GMoE on DomainBed and MoE-LLaVA on VQA-style benches while activating fewer params (StableLM-1.6B: avg k=1.25 / 1.75B vs k=2 / 2.06B). Visualizations suggest bottom layers need MoE more than top layers and a shared easy-to-activate expert per layer. Code https://github.com/LINs-lab/DynMoE.

Keywords

dynmoe · moe · top-any · adaptive-experts · moe-llava · glue · domainbed · iclr · westlake · tencent

Topics

mixture of experts, adaptive computation, transformers

Research notes

  • Primary: arxiv abs (cs.LG; also cs.AI). CC BY 4.0 on HTML. ICLR 2025. Guo/Cheng equal contrib; Tang/Lin corresponding. CUHK-Shenzhen / AIRS / Westlake / SII / Tencent AI Lab. Code Apache-2.0 https://github.com/LINs-lab/DynMoE (161 stars at check). HF paper page 3 upvotes; official models LINs-lab/DynMoE-{Phi-2-2.7B,StableLM-1.6B,Qwen-1.8B} not copied into hf_* fields. Discord posted PDF. Uses public GLUE/DomainBed/LLaVA-Finetuning; no new corpus, so no datasets_local row. License field left blank per catalog convention.