← Back to explorer

SkyRL v0.4.0: Large-Scale Post-Training Release

Type
paper
Venue
SkyRL blog, 2 Oct 2026
Year
2026
Source
web
Access
free
Language
English
Added
2026-10-02
Verified
2026-10-02

Summary

SkyRL v0.4.0 release, focused on large-scale post-training. New model support: GLM 5.3 Flash, GLM 5.2/5.3 (with Trajectory — links to their 'Enabling Frontier-Scale Training for the GLM-5.3 Family' field report), Kimi K2.6/2.7 (1T+ params trainable with LoRA on just 2 B300 nodes via INT4 serving + trainer fake-quantization), Qwen3.8, Nemotron 3.5 Lightning. RL stability: Top-P sampler replay (Keep Sampling Mask from DeepSeek-V3.2 — vLLM returns the surviving support set per token, the trainer normalizes logprobs over it; keeps entropy from collapsing, works on Megatron and FSDP) and optimized R3 (rollout routing replay) data packing — compact packed arrays, zero-copy ray object store views, turn-only route deltas — up to ~4x throughput for large MoEs (Nemotron 3 Super: 111 -> 483 tokens/s/GPU; routing data was nearly 900x the tokens). R3 + top-p replay keeps the logprob gap under 0.1 with reward still climbing at step 160, vs. gap >0.8 and reward collapse without them (validated with Lila Sciences). Full FP8/MXFP8 RL: blockwise FP8 and MXFP8 training/rollout with on-policy weight sync, tracking BF16 convergence while cutting end-to-end step time by up to 23% (Anyscale blog by @JinghanYao). Weight syncing reworked: delta weight sync (byte-exact XOR deltas vs. previous BF16 checkpoint, compressed, manifest published; ~1-3% of weights updated per step; disk/S3/GCS paths, receiver fetch overlapped with generation) and sharded RDT+NIXL transfer (all trainer ranks send only needed shards, skip PP gathering and expert-layer gathering — BF16 Kimi K2 on 48 8xH100 nodes transfers in 7.53s). In-memory LoRA weight syncing via NCCL/CUDA IPC (vLLM-side changes to be upstreamed). Scalability: native SFTTrainer now handles 10B-1T+ token datasets (memory-mapped pretokenized loading, custom sampling, dataset mixing); Tinker API server now handles 131,072 concurrent trajectories with 212k-token outputs (up from 1,024 requests; sampling no longer touches SQLite — thanks to Trajectory AI).

Keywords

SkyRL · RL training · MoE · FP8 RL · weight sync · RDT · NIXL · top-p sampler replay · R3

Topics

RL training frameworks, SkyRL, MoE, FP8 RL, weight sync, RDT, NIXL, sampler replay, large-scale post-training

Research notes

  • Discovery: @erictang000 (Eric Tang) X thread 2026-10-02 3:57 PM ET (https://x.com/erictang000/status/2106111313533673812)
  • Release blog: blog.skyrl.ai/blog/skyrl-v0-4-0
  • Repo: github.com/NovaSky-AI/SkyRL
  • Community Slack: join.slack.com/t/skyrl/shared_invite/zt-3f6ncn5b8-QawzK3uks6ka3KWoLwsi5Q
  • OSS collaborators: Trajectory (GLM 5.3 enablement; @hersh_godse, @j316chuck, @neilkale, @yapdianang), @casper_hansen_ (Kimi K2.6/2.7), @DyurkLila / @LilaSciences (R3 + sampler replay validation on Nemotron 3 Super 120B)
  • Companion Anyscale blog on FP8 RL by @JinghanYao: anyscale.com/blog/fp8-reinfinforcement-learning-in-skyrl
  • Sharded weight-transfer work led by @AaronHao4 (@sumanthrh post)
  • Directly relevant to the user's RL/post-training work and the 948 Trajectory field report