Training nGPT
- Type
- paper
- Venue
- arXiv / NVIDIA
- Year
- 2026
- Source
- arxiv
- Access
- free
- Language
- en
- Added
- 2026-08-14T16:24:56Z
- Verified
- 2026-08-14T16:24:56Z
Summary
Extends nGPT (unit-hypersphere parameters and activations) from dense Transformers to Nemotron-3-style hybrid Mamba-2–Transformer MoE. Recipe: Logit Gradient Preconditioning, logarithmic LR decay, GatedAdamW (a=0.5), angular step cap, optional tangent projection / exploration noise / second-moment growth clipping. On the same architecture, 14B-total nGPT reaches the GPT+AdamW validation loss with about half the tokens; across 1B-14B total (0.21B-1.74B active) validation loss is ~2.5-3.4% lower. Limited-budget study; not an exhaustive ablation.
Keywords
ngpt · hypersphere · gatedadamw · moe · mamba · nemotron · nvidia · training-recipe
Topics
LLM training, hyperspherical Transformers, MoE
Research notes
- Primary: arxiv abs (CC BY 4.0, cs.LG / cs.AI). NVIDIA (iloshchilov, bginsburg). No official code on abs; HF has no paper page. Discord posted abs link.