← Back to explorer

Transformers without Normalization

Type
other
Venue
arXiv / FAIR Meta / NYU / MIT / Princeton University

Summary

Observes that LN input–output maps are tanh-like S-curves and replaces LN/RMSNorm with DyT(x)=γ·tanh(αx)+β (α learnable scalar). Matches or beats LN with original hparams: ImageNet ViT-B 82.5 vs 82.3, ViT-L 83.6 vs 83.1; MAE/DINO on par; DiT FID comparable; LLaMA 7B–70B on 200B Pile tokens match RMSNorm zero-shot (70B both 0.549 / loss 1.45). Default α0=0.5 except LLMs (attention vs other α0 split). Tanh ablation needed for stability; identity diverges. Does not replace BN in ResNet-50 (76.2→68.9). CVPR 2025. Code https://github.com/jiachenzhu/DyT; project https://jiachenzhu.github.io/DyT/.

Keywords

dyt · layernorm · rmsnorm · transformers · cvpr · fair · meta · mit · princeton · nyu

Topics

normalization, Transformer architecture, Dynamic Tanh

Research notes

  • Primary: arxiv abs (cs.LG; also cs.AI, cs.CL, cs.CV). License not stated on abs/HTML at check. Comment: CVPR 2025; project https://jiachenzhu.github.io/DyT/. Affiliations FAIR/Meta, NYU, MIT, Princeton. Correspondence jiachen.zhu@nyu.edu, zhuangl@princeton.edu. Liu project lead. Code https://github.com/jiachenzhu/DyT (1,042 stars at check). HF paper page 171 upvotes; githubRepo linked; no linked models/datasets. Discord posted PDF v1. Trains on public ImageNet/Pile/LibriSpeech; no new corpus, so no datasets_local row. License field left blank per catalog convention.