← Back to explorer

Scaling Pre-Training in Practice: A Hierarchical Approach

Type
blog
Venue
Aleph Alpha blog 2026-09-30
Year
2026
Source
blog
Access
public
Language
en
Added
2026-10-01
Verified
2026-10-01

Summary

Aleph Alpha's technical write-up of a practical, hierarchical approach to finding an efficient LLM pre-training configuration and scaling it from a small cluster to a large one. Three degrees of freedom are swept hierarchically: (1) parallelism scheme (FSDP/DP degrees), (2) activation checkpointing (full-AC, selective-AC, or none), (3) local batch size. The sweep runs on 16 GPUs (1/32 of target), with profiler traces (PyTorch profiler + Perfetto) used to pick a scheme that fits within ~95% HBM; the configuration is then verified and re-profiled at 128 GPUs and again at 512 GPUs. Demonstrated on a 30B-A3B MoE (30B total params, 3B active per token) scaled 16 -> 512 NVIDIA B200 GPUs (64 nodes x 8, InfiniBand): 16 GPUs: 28.4k TPS/GPU, 37.5% MFU, 92% HBM; 128 GPUs: 28.0k TPS/GPU, 37.0% MFU, 93% HBM; 512 GPUs: 26.7k TPS/GPU, 35.3% MFU, 94% HBM -- only ~6% per-GPU efficiency drop from 16 to 512 GPUs, claimed near-linear scaling. Training stack is Aleph Alpha's own optimized fork of PyTorch's torchtitan (no link to the fork given). Stated scope: a single architecture and a single sequence length (4096 tokens); cluster sizes assumed known in advance. No model, code, weights, or dataset released in the post.

Keywords

LLM pretraining · distributed training · MoE · FSDP · activation checkpointing · MFU · B200 · torchtitan

Topics

LLM pretraining, distributed training, MoE, FSDP, activation checkpointing, MFU, B200, torchtitan, hierarchical scaling

Research notes

  • Discovery: shared directly in chat (2026-10-01) via Aleph Alpha's X post
  • Blog author Jordan Sassoon (credited in thread as @resiliem); acknowledgements to Steffen Hirschmann, Samuel Weinbach, Yasser Jadidi, Max Hoeth, Fabien Benureau, Alessio Serra
  • No license stated; no paper, model, code, or dataset released
  • References open-source torchtitan (github.com/pytorch/torchtitan) as the base of their fork
  • Highly relevant to the user's own pretraining work: concrete MFU/scaling numbers and a practical recipe for 16->512 GPU scale-up.