← Back to explorer

Seer: Online Context Learning for Fast Synchronous LLM Reinforcement Learning

Type
other
Venue
arXiv / Moonshot AI / Tsinghua University

Summary

Rollout is 63–87% of RL iteration time. Seer exploits intra-group (G=8–16) length/pattern similarity: divided chunk-level rollout with a global Mooncake KVCache, context-aware longest-first scheduling via a speculative probe request, and adaptive grouped speculative decoding with compressed suffix trees. On production GRPO workloads (Moonlight, Qwen2-VL-72B, Kimi-K2) up to 2.04× end-to-end rollout throughput vs veRL and 72–94% less long-tail latency, while keeping on-policy semantics. Context-aware scheduling alone cuts tail time ~89%; grouped SD adds up to 1.3× vs vanilla SD.

Keywords

seer · grpo · rollout · speculative-decoding · kvcache · mooncake · moonshot · tsinghua · on-policy

Topics

RL systems, LLM serving, speculative decoding

Research notes

  • Primary: arxiv abs (cs.DC; also cs.LG). License not stated on abs/HTML at check. Moonshot AI / Tsinghua. Correspondence zhang_mingxing@mail.tsinghua.edu.cn. No official code on abs. HF paper page 3 upvotes; no linked models/datasets. Discord posted abs (Substack UTM). Systems paper over production RL traces; no new public corpus, so no datasets_local row. License field left blank per catalog convention.