← Back to explorer

Training nGPT

Type
paper
Venue
arXiv / NVIDIA
Year
2026
Source
arxiv
Access
free
Language
en
Added
2026-08-14T16:24:56Z
Verified
2026-08-14T16:24:56Z

Summary

Extends nGPT (unit-hypersphere parameters and activations) from dense Transformers to Nemotron-3-style hybrid Mamba-2–Transformer MoE. Recipe: Logit Gradient Preconditioning, logarithmic LR decay, GatedAdamW (a=0.5), angular step cap, optional tangent projection / exploration noise / second-moment growth clipping. On the same architecture, 14B-total nGPT reaches the GPT+AdamW validation loss with about half the tokens; across 1B-14B total (0.21B-1.74B active) validation loss is ~2.5-3.4% lower. Limited-budget study; not an exhaustive ablation.

Keywords

ngpt · hypersphere · gatedadamw · moe · mamba · nemotron · nvidia · training-recipe

Topics

LLM training, hyperspherical Transformers, MoE

Research notes

  • Primary: arxiv abs (CC BY 4.0, cs.LG / cs.AI). NVIDIA (iloshchilov, bginsburg). No official code on abs; HF has no paper page. Discord posted abs link.