LT2: Linear-Time Looped Transformers
- Type
- other
- Venue
- arXiv / Rice University / Apple / UC Santa Cruz / Carnegie Mellon University
Summary
LT2 loops subquadratic mixers: DPLR linear attention gets a rank-T state update; sparse windows get Tw receptive field. On FineWeb-Edu 100B tokens, 1.3B T=4: Looped Hybrid (Full+GDN) 62.89 avg zero-shot vs looped Transformer 59.27, with ~2.7–5x decode speedup; GDN+DSA matches quality with no full attention (~2.9x at 32k). Distills Ouro-1.4B into Ouro-hybrid-1.4B with ~1B tokens; competitive with 1B–4B industry models. Code https://github.com/chili-lab/LT2; checkpoint https://huggingface.co/chili-lab/Ouro-hybrid-1.4B.
Keywords
lt2 · looped-transformer · gdn · dsa · ouro · linear-attention · rice · apple
Topics
looped transformers, linear attention, efficient LMs
Research notes
- Primary: arxiv abs. Correspondence chunyuan.deng@rice.edu, hanjie@rice.edu; Zhang at Apple, Zhu at UCSC, Liu at CMU. Code https://github.com/chili-lab/LT2 (51 stars at check); license not stated on the GitHub landing page, so left blank. HF checkpoint chili-lab/Ouro-hybrid-1.4B. HF paper page 0 upvotes. Discord posted abs. Pretrains on public FineWeb-Edu; no new corpus, so no datasets_local row. ArXiv license widget not visible in converted abs HTML, so paper license left blank.