More Expressive Feedforward Layers: Part I. Token-Adaptive Mixing of Activations
- Type
- other
- Venue
- arXiv / ByteDance Seed / Peking University
Summary
MoA mixes a small activation dictionary (GELU/SiLU/ReLU²/etc.) with lightweight per-token gates on shared FFN projections; LA is the input-independent linear-combination ablation. Finite-width theory: fixed-activation FFN ⊊ LA ⊊ MoA. Pretrains dense Llama 0.12B–0.5B (AdamW, cosine, 20/100 TPP) and LlamaMoE 0.25B–2B (Muon, WSD, ~100 tokens/active param). Type-II one-MoA/qd-MoA beat well-tuned SwiGLU/Muon baselines by >0.01 terminal loss with ~1.03–1.13× wall-clock and negligible params; LlamaMoE-2B zero-shot avg 42.20→42.86/42.96. Also helps MAE ViT-B. No official code on abs.
Keywords
moa · mixture-of-activations · swiglu · ffn · bytedance-seed · muon · llama-moe
Topics
FFN activations, token-adaptive mixing, LLM pretraining
Research notes
- Primary: arxiv abs (cs.LG/AI/stat.ML). ByteDance Seed / Peking University; correspondence wangmingze.999@bytedance.com, zhongshu@bytedance.com. No official code on abs. HF has no paper page (API 404). Discord posted abs. Pretrains on an unnamed high-quality corpus; no standalone public dataset release, so no datasets_local row. ArXiv license widget not visible in converted abs HTML, so paper license left blank.