Breaking Entropy Bounds: Accelerating RL Training via MTP with Rejection Sampling
- Type
- paper
- Venue
- arXiv / Qwen Team, Alibaba
- Year
- 2026
- Source
- arxiv
- Access
- free
- Language
- en
- Added
- 2026-08-14T16:58:34Z
- Verified
- 2026-08-14T16:58:34Z
Summary
Bebop shows MTP accept length falls linearly with policy entropy under target-only sampling or CE/KL-trained drafts. Switching verification to rejection sampling (accept = 1 − TV(p,q)) and training drafts with an end-to-end TV loss that directly maximizes multi-step overlap cuts the entropy–accept slope ~95% and adds ~3–10 pp accept vs CE (up to 95% on agent tasks). Pre-RL TV SFT plus rejection sampling is enough; online MTP updates during RL are unnecessary under RS because draft–target mismatch from weight updates is negligible. Up to 1.8× async RL e2e on Qwen3.5/3.6/3.7; agentic rollouts up to 2.4×. SGLang RS implementation https://github.com/sgl-project/sglang/pull/26312. No official training-code repo on the abs page.
Keywords
bebop · mtp · speculative-decoding · rejection-sampling · tv-loss · qwen · rl-rollout
Topics
LLM RL systems, speculative decoding, MTP
Research notes
- Primary: arxiv abs (default nonexclusive-distrib 1.0, cs.LG/CL). Qwen Team Alibaba. Discord posted AlphaXiv 2606.12370; canonical abs recorded. HF paper page 21 upvotes, org Qwen; no linked models/datasets. RS code in SGLang PR 26312; no standalone training-code repo on abs. Not a dataset; no datasets_local row.