Enabling Frontier-Scale Training for the GLM-5.3 Family
- Type
- paper
- Venue
- Trajectory field report, 2 Oct 2026
- Year
- 2026
- Source
- web
- Access
- free
- Language
- English
- Added
- 2026-10-02
- Verified
- 2026-10-02
Summary
Field report on the training infrastructure that made GLM-5.3 and GLM-5.3 Flash trainable with RL. Two camps of the 'expedition': (1) silent numerical mismatches — the Megatron trainer and vLLM sampler can drift into different policies without crashing (flat reward curve hours later). A 5-minute numerical regression test compares trainer/sampler logprobs token-by-token, deliberately perturbs only the trainer's adapter to confirm the stale sampler disagrees, then resyncs and confirms agreement — ~28x faster than a ~150-minute training cycle. It caught two real bugs (shared-expert scaling, cached prefill), fixed upstream in SkyRL. (2) slow handoffs — LoRA adapter sync took >6.5 minutes on full GLM-5.3 because the shared-expert adapter was duplicated per expert and written to disk. An in-memory path dedupes tensors and uses NCCL/CUDA IPC, cutting sync from ~390s to ~13s (~29.4x). Validation trek: trained both models on DAPO-math with a tight 8,192-token response limit; over 20 steps AIME 2024 accuracy rose 57%->81% (GLM-5.3) and 54%->83% (Flash), while Flash's share of capped responses fell 46%->6% and average length 5,217->2,424 tokens.
Keywords
frontier-scale training · GLM-5.3 · training loop · numerical regression · weight sync · NCCL · SkyRL · RLHF · AIME
Topics
RL training loop, vLLM, Megatron, numerical regression, weight synchronization, NCCL, LoRA, MoE, GLM-5.3, AIME, DAPO-math
Research notes
- Discovery: @trajectorylabs (Trajectory) 7-post X thread 2026-10-02 2:19 PM ET (https://x.com/trajectorylabs/status/2106086640032858340)
- Field report at trajectory.ai; several OSS contributions upstreamed to SkyRL (numerical regression tests, in-memory weight sync, shared-expert fix, Flash integration, full-model and two-node recipes)
- Setup: correctness rewards with soft length penalty, group-relative advantages, asymmetric policy clipping, token-mean loss; 128 prompts/step (16 responses/prompt GLM-5.3, 12 Flash); 4 optimizer updates/step, lr 1e-5; full GLM rank-32 LoRA + 40-update warmup, Flash rank-64 LoRA
- Data and methods in a linked gist