FlashAR: Efficient Post-Training Acceleration for Autoregressive Image Generation
- Type
- other
- Venue
- arXiv / University of Adelaide / Zhejiang University
Summary
Keeps the original AR head for row-wise prediction and branches a vertical head from an intermediate layer, fused by a learnable gate so decoding walks anti-diagonals (HW steps → H+W−1). Two-stage post-training: freeze backbone to init the vertical head, then joint finetune. On LlamaGen ImageNet 256, FlashAR-L FID 3.16 / IS 289.0 vs BlockDiffusion 4.55 / 243.5 after 25 vs 75 epochs; FlashAR-B 447 img/s. On Emu3.5-Image-34B at 512×512, 22.9× wall-clock (130.10s → 5.68s, 1024 → 63 steps) with GenEval 80.48→80.29 using ~80K pairs (0.05% of pretrain data). FlexAttention sparse 2D masks + batched KV. Code https://github.com/lxazjk/Emu3.5-FlashAR; project https://lxazjk.github.io/FlashAR/.
Keywords
flashar · autoregressive-image · diagonal-decoding · llamagen · emu3.5 · post-training · adelaide · zhejiang
Topics
autoregressive image generation, parallel decoding, post-training
Research notes
- Primary: arxiv abs (cs.CV). Adelaide / Zhejiang; Zhou/He/Chen equal contrib, Wang project lead, Zhuang corresponding. Code Apache-2.0 https://github.com/lxazjk/Emu3.5-FlashAR (32 stars at check) plus https://github.com/lxazjk/LlamaGen-FlashAR (MIT, 5 stars). Project https://lxazjk.github.io/FlashAR/. HF paper page 1 upvote. Linked models lxazjk/Emu3.5-Image-FlashAR (~37.5B, 13 dl / 3 likes at check) and lxazjk/LlamaGen-FlashAR. Discord posted abs. Uses ImageNet and existing OpenGPT-4o-Image/ShareGPT-4o-Image; no new corpus, so no datasets_local row. ArXiv license widget not visible in converted abs HTML, so paper license left blank.