DualKV: Shared-Prompt Flash Attention for Efficient RL Training with Large Rollouts and Long Contexts
- Type
- other
- Venue
- arXiv / Amazon Web Services
Summary
Standard FA2 packs N same-prompt rollouts as N(P+R) tokens, recomputing the prompt N times. DualKV repacks to P+NR and splits attention into one prompt self-attention plus a fused kernel that reads shared KV once and accumulates N-way dK/dV with fp32 atomics; mathematically equivalent, no approximation. On Qwen3-8B GRPO (8xH100, N=32, 8K-context) policy-update is 1.63–2.09x faster, MFU 36%→76%; DAPO 2.47x / 77% MFU. At 30B MoE on 16xH100, 3.82x policy-update and 3.38x step vs FA2 that needs 4-way Ulysses SP. Also supports Gemma-4 hybrid d=512 / sliding-window at 64K. Code https://github.com/amazon-science/dualkv-flash-attn-for-rl.
Keywords
dualkv · flashattention · grpo · dapo · kv-cache · rl-training · verl · amazon · prefix-sharing
Topics
RL training, FlashAttention, KV cache, long context
Research notes
- Primary: arxiv abs (cs.LG). AWS (Gai/Zhang/Wang equal contrib); Song at Google (work done at AWS); Karypis at UMN. Correspondence jiadingg/shuaizs/yuyawang@amazon.com, xiangsx@google.com, karypis@umn.edu. Code CC-BY-NC-4.0 https://github.com/amazon-science/dualkv-flash-attn-for-rl (4 stars at check); gemma4-dev branch for hd512/SWA. HF paper page 1 upvote. Discord posted PDF. Uses LongReason/GSM8K; no new corpus, so no datasets_local row. ArXiv license widget not visible in converted abs HTML, so paper license left blank.