← Back to explorer

Enabling Frontier-Scale Training for the GLM-5.3 Family

Type
paper
Venue
Trajectory field report, 2 Oct 2026
Year
2026
Source
web
Access
free
Language
English
Added
2026-10-02
Verified
2026-10-02

Summary

Field report on the training infrastructure that made GLM-5.3 and GLM-5.3 Flash trainable with RL. Two camps of the 'expedition': (1) silent numerical mismatches — the Megatron trainer and vLLM sampler can drift into different policies without crashing (flat reward curve hours later). A 5-minute numerical regression test compares trainer/sampler logprobs token-by-token, deliberately perturbs only the trainer's adapter to confirm the stale sampler disagrees, then resyncs and confirms agreement — ~28x faster than a ~150-minute training cycle. It caught two real bugs (shared-expert scaling, cached prefill), fixed upstream in SkyRL. (2) slow handoffs — LoRA adapter sync took >6.5 minutes on full GLM-5.3 because the shared-expert adapter was duplicated per expert and written to disk. An in-memory path dedupes tensors and uses NCCL/CUDA IPC, cutting sync from ~390s to ~13s (~29.4x). Validation trek: trained both models on DAPO-math with a tight 8,192-token response limit; over 20 steps AIME 2024 accuracy rose 57%->81% (GLM-5.3) and 54%->83% (Flash), while Flash's share of capped responses fell 46%->6% and average length 5,217->2,424 tokens.

Keywords

frontier-scale training · GLM-5.3 · training loop · numerical regression · weight sync · NCCL · SkyRL · RLHF · AIME

Topics

RL training loop, vLLM, Megatron, numerical regression, weight synchronization, NCCL, LoRA, MoE, GLM-5.3, AIME, DAPO-math

Research notes

  • Discovery: @trajectorylabs (Trajectory) 7-post X thread 2026-10-02 2:19 PM ET (https://x.com/trajectorylabs/status/2106086640032858340)
  • Field report at trajectory.ai; several OSS contributions upstreamed to SkyRL (numerical regression tests, in-memory weight sync, shared-expert fix, Flash integration, full-model and two-node recipes)
  • Setup: correctness rewards with soft length penalty, group-relative advantages, asymmetric policy clipping, token-mean loss; 128 prompts/step (16 responses/prompt GLM-5.3, 12 Flash); 4 optimizer updates/step, lr 1e-5; full GLM rank-32 LoRA + 40-update warmup, Flash rank-64 LoRA
  • Data and methods in a linked gist