← Back to explorer

LT2: Linear-Time Looped Transformers

Type
other
Venue
arXiv / Rice University / Apple / UC Santa Cruz / Carnegie Mellon University

Summary

LT2 loops subquadratic mixers: DPLR linear attention gets a rank-T state update; sparse windows get Tw receptive field. On FineWeb-Edu 100B tokens, 1.3B T=4: Looped Hybrid (Full+GDN) 62.89 avg zero-shot vs looped Transformer 59.27, with ~2.7–5x decode speedup; GDN+DSA matches quality with no full attention (~2.9x at 32k). Distills Ouro-1.4B into Ouro-hybrid-1.4B with ~1B tokens; competitive with 1B–4B industry models. Code https://github.com/chili-lab/LT2; checkpoint https://huggingface.co/chili-lab/Ouro-hybrid-1.4B.

Keywords

lt2 · looped-transformer · gdn · dsa · ouro · linear-attention · rice · apple

Topics

looped transformers, linear attention, efficient LMs

Research notes

  • Primary: arxiv abs. Correspondence chunyuan.deng@rice.edu, hanjie@rice.edu; Zhang at Apple, Zhu at UCSC, Liu at CMU. Code https://github.com/chili-lab/LT2 (51 stars at check); license not stated on the GitHub landing page, so left blank. HF checkpoint chili-lab/Ouro-hybrid-1.4B. HF paper page 0 upvotes. Discord posted abs. Pretrains on public FineWeb-Edu; no new corpus, so no datasets_local row. ArXiv license widget not visible in converted abs HTML, so paper license left blank.