← Back to explorer

Supercharging Olmo-core for Efficient and Scalable MoE Training

Type
paper
Venue
Allen AI technical report, 2026-10
Year
2026
Source
paper
Access
public
Language
en
Added
2026-10-01
Verified
2026-10-01

Summary

Supercharging Olmo-core for Efficient and Scalable MoE Training (Allen AI technical report, October 2026, 168 pages; no arXiv listing at research time, lives only on allenai.org; no license notice visible on the report). Olmo-core 3 is a redesigned open MoE training stack, moving from Fully Sharded Data Parallel (FSDP) -- which gathered/resharded weights per microbatch -- to Distributed Data Parallel (DDP) + Expert Parallelism (EP) + Pipeline Parallelism (PP) + a distributed optimizer + activation recompute. Experts stay resident on GPUs and training data is routed to them, avoiding repeated weight gathering. Key techniques: NVSHMEM-based rowwise EP (tokens written directly into expert buffers), GPU-resident routing metadata, device-scheduled grouped GEMMs, and MXFP8 low-precision formats. Results on NVL8 B300 nodes: configurations from 12.9B to 1.2T total parameters on up to 512 GPUs; 858 useful-model TFLOP/s/GPU at 1.2T with MXFP8 and per-layer recompute; an experimental DeepEP v2 backend capacity test reached 2.38T total parameters. ~2.7x throughput vs Olmo-core 2's FSDP stack (47B-param MoE on 8x B300: ~52,000 vs 19,400 tokens/sec/GPU); expert pool grown 8->128 with only 4 experts/token and <5% throughput loss; ~21% higher throughput with MXFP8 vs BF16 (peak memory 103->95 GiB). Failure-mode studies include 'token gerrymandering' (the MoE router learns to hack the load-balancing loss so the score improves while the actual workload gets less balanced) and a finding that overlapping communication with computation can slow the overlapped kernels by more than it hides. Code released as Olmo-core v3.0.0 (Apache-2.0). Designed to scale into the trillion-parameter range; the next-generation Olmo will use an MoE architecture.

Keywords

MoE training · expert parallelism · distributed training · NVSHMEM · MXFP8 · grouped GEMM · load balancing · token gerrymandering · Olmo · open infrastructure · B300 · pipeline parallelism

Topics

MoE training, expert parallelism, distributed training, NVSHMEM, MXFP8, grouped GEMM, load balancing, token gerrymandering, Olmo, open infrastructure, B300

Research notes

  • Discovery: shared directly in chat (2026-10-01) via Allen AI release announcement
  • Lead author/primary developer Tianhua Tao; contact tianhuat@cs.washington.edu
  • Associated: blog 'Introducing Olmo-core 3' (allenai.org/blog/olmocore3, 2026-10-01), interactive explainer 'How do 512 GPUs train one AI model?' (narrative.allen.ai/scaling-up-training, 6 chapters + Build-a-run playground), GitHub allenai/olmo-core (Apache-2.0, ~1.6k stars, 320 forks, 908 commits, 44 contributors), docs olmo-core.readthedocs.io
  • Benchmarks on NVIDIA B300; Docker images tested on in-house H100 (adapt Dockerfile for other hardware)
  • Highly relevant to user's pretraining work: open MoE training stack, EP scaling techniques, honest failure-mode analysis.