← Back to explorer

ResiDual: Transformer with Dual Residual Connections

Type
other
Venue
arXiv / Microsoft Research / Microsoft Azure Translation / Renmin University of China

Summary

PPLN keeps a Post-LN path (diverse hidden states) plus a Pre-LN dual residual (lower-bounded gradients). Theory: Post-LN grads decay ~O((1/2)^{(N-k)/2}); Pre-LN |Δh|~O(1/√k) collapse; ResiDual inherits the better of each. IWSLT-14 E6D6 BLEU 35.63 vs Post-LN 35.37 / Pre-LN 35.12; E12D12 36.09 vs Pre-LN 35.18 (Post-LN fails). WMT DE→EN E18D18 27.65 vs Pre-LN 26.57 / B2T 27.30. OPUS-100 E18D18 ALL 31.0 vs Pre-LN 30.3, matching a 100-layer DeepNet 31.1. Trains without LR warmup on IWSLT. Code https://github.com/microsoft/ResiDual.

Keywords

residual · pre-ln · post-ln · ppln · machine-translation · microsoft · fairseq

Topics

Transformer architecture, residual connections, layer normalization

Research notes

  • Primary: arxiv abs (cs.CL; also cs.AI, cs.LG, cs.NE). License not stated on abs/HTML at check. Correspondence Xu Tan xuta@microsoft.com and Rui Yan ruiyan@ruc.edu.cn. Microsoft Research / Azure Translation / Renmin University. Code https://github.com/microsoft/ResiDual (98 stars at check; MIT; archived). HF has no paper page (API 404). Discord posted PDF. Uses public IWSLT/WMT/OPUS; no new corpus, so no datasets_local row. License field left blank per catalog convention.