← Back to explorer

More Expressive Feedforward Layers: Part I. Token-Adaptive Mixing of Activations

Type
other
Venue
arXiv / ByteDance Seed / Peking University

Summary

MoA mixes a small activation dictionary (GELU/SiLU/ReLU²/etc.) with lightweight per-token gates on shared FFN projections; LA is the input-independent linear-combination ablation. Finite-width theory: fixed-activation FFN ⊊ LA ⊊ MoA. Pretrains dense Llama 0.12B–0.5B (AdamW, cosine, 20/100 TPP) and LlamaMoE 0.25B–2B (Muon, WSD, ~100 tokens/active param). Type-II one-MoA/qd-MoA beat well-tuned SwiGLU/Muon baselines by >0.01 terminal loss with ~1.03–1.13× wall-clock and negligible params; LlamaMoE-2B zero-shot avg 42.20→42.86/42.96. Also helps MAE ViT-B. No official code on abs.

Keywords

moa · mixture-of-activations · swiglu · ffn · bytedance-seed · muon · llama-moe

Topics

FFN activations, token-adaptive mixing, LLM pretraining

Research notes

  • Primary: arxiv abs (cs.LG/AI/stat.ML). ByteDance Seed / Peking University; correspondence wangmingze.999@bytedance.com, zhongshu@bytedance.com. No official code on abs. HF has no paper page (API 404). Discord posted abs. Pretrains on an unnamed high-quality corpus; no standalone public dataset release, so no datasets_local row. ArXiv license widget not visible in converted abs HTML, so paper license left blank.