← Back to explorer

Breaking Entropy Bounds: Accelerating RL Training via MTP with Rejection Sampling

Type
paper
Venue
arXiv / Qwen Team, Alibaba
Year
2026
Source
arxiv
Access
free
Language
en
Added
2026-08-14T16:58:34Z
Verified
2026-08-14T16:58:34Z

Summary

Bebop shows MTP accept length falls linearly with policy entropy under target-only sampling or CE/KL-trained drafts. Switching verification to rejection sampling (accept = 1 − TV(p,q)) and training drafts with an end-to-end TV loss that directly maximizes multi-step overlap cuts the entropy–accept slope ~95% and adds ~3–10 pp accept vs CE (up to 95% on agent tasks). Pre-RL TV SFT plus rejection sampling is enough; online MTP updates during RL are unnecessary under RS because draft–target mismatch from weight updates is negligible. Up to 1.8× async RL e2e on Qwen3.5/3.6/3.7; agentic rollouts up to 2.4×. SGLang RS implementation https://github.com/sgl-project/sglang/pull/26312. No official training-code repo on the abs page.

Keywords

bebop · mtp · speculative-decoding · rejection-sampling · tv-loss · qwen · rl-rollout

Topics

LLM RL systems, speculative decoding, MTP

Research notes

  • Primary: arxiv abs (default nonexclusive-distrib 1.0, cs.LG/CL). Qwen Team Alibaba. Discord posted AlphaXiv 2606.12370; canonical abs recorded. HF paper page 21 upvotes, org Qwen; no linked models/datasets. RS code in SGLang PR 26312; no standalone training-code repo on abs. Not a dataset; no datasets_local row.