Scaling Pre-Training in Practice: A Hierarchical Approach
- Type
- blog
- Venue
- Aleph Alpha blog 2026-09-30
- Year
- 2026
- Source
- blog
- Access
- public
- Language
- en
- Added
- 2026-10-01
- Verified
- 2026-10-01
Summary
Aleph Alpha's technical write-up of a practical, hierarchical approach to finding an efficient LLM pre-training configuration and scaling it from a small cluster to a large one. Three degrees of freedom are swept hierarchically: (1) parallelism scheme (FSDP/DP degrees), (2) activation checkpointing (full-AC, selective-AC, or none), (3) local batch size. The sweep runs on 16 GPUs (1/32 of target), with profiler traces (PyTorch profiler + Perfetto) used to pick a scheme that fits within ~95% HBM; the configuration is then verified and re-profiled at 128 GPUs and again at 512 GPUs. Demonstrated on a 30B-A3B MoE (30B total params, 3B active per token) scaled 16 -> 512 NVIDIA B200 GPUs (64 nodes x 8, InfiniBand): 16 GPUs: 28.4k TPS/GPU, 37.5% MFU, 92% HBM; 128 GPUs: 28.0k TPS/GPU, 37.0% MFU, 93% HBM; 512 GPUs: 26.7k TPS/GPU, 35.3% MFU, 94% HBM -- only ~6% per-GPU efficiency drop from 16 to 512 GPUs, claimed near-linear scaling. Training stack is Aleph Alpha's own optimized fork of PyTorch's torchtitan (no link to the fork given). Stated scope: a single architecture and a single sequence length (4096 tokens); cluster sizes assumed known in advance. No model, code, weights, or dataset released in the post.
Keywords
LLM pretraining · distributed training · MoE · FSDP · activation checkpointing · MFU · B200 · torchtitan
Topics
LLM pretraining, distributed training, MoE, FSDP, activation checkpointing, MFU, B200, torchtitan, hierarchical scaling
Research notes
- Discovery: shared directly in chat (2026-10-01) via Aleph Alpha's X post
- Blog author Jordan Sassoon (credited in thread as @resiliem); acknowledgements to Steffen Hirschmann, Samuel Weinbach, Yasser Jadidi, Max Hoeth, Fabien Benureau, Alessio Serra
- No license stated; no paper, model, code, or dataset released
- References open-source torchtitan (github.com/pytorch/torchtitan) as the base of their fork
- Highly relevant to the user's own pretraining work: concrete MFU/scaling numbers and a practical recipe for 16->512 GPU scale-up.