← Back to explorer

Looped Diffusion Language Models

Type
other
Venue
arXiv / KAIST / KRAFTON / UC Berkeley

Summary

LoopMDM shares a small early-middle block (typically 2 layers) for S applications while keeping head/tail unshared; S is sampled uniformly in {1..S_max} at train time. Iso-parameter 170M DiT matches same-size MDM test NLL with up to 3.3x fewer training FLOPs on FineWeb-Edu/OWT/LM1B; GSM8K +8.5pp vs same-size MDM and beats a deeper 21-layer MDM at matched per-step FLOPs. Adaptive loop stopping (~5 vs 12 loops) keeps downstream acc. Attention analysis: looping raises mask-to-mask attention, using masked positions as a workspace (Sudoku under forced L2R order). No official code on abs.

Keywords

loopmdm · masked-diffusion · looped-transformer · fineweb-edu · gsm8k · krafton · kaist · test-time-compute

Topics

masked diffusion, looped transformers, language modeling

Research notes

  • Primary: arxiv abs (cs.LG). CC BY 4.0 on abs HTML. KAIST (Lee internship at KRAFTON) / KRAFTON / UC Berkeley. Correspondence lsh83210/hoarer/seungryong.kim@kaist.ac.kr, jonghyunlee/dongmin.park@krafton.com, jjhpark@berkeley.edu. No official code on abs. HF paper page 0 upvotes. Discord posted PDF. Pretrains on public FineWeb-Edu/OWT/LM1B/TinyGSM; no new corpus, so no datasets_local row.