← Back to explorer

Leave No Context Behind: Efficient Infinite Context Transformers with Infini-attention

Type
other
Venue
arXiv / Google

Summary

Infini-attention keeps local softmax attention plus a compressive associative memory updated from the same QKV (linear or delta rule) and a learned gate β. 114× smaller memory than a 65k Memorizing Transformer; PG19 PPL 9.65 vs 11.37. A 1B model solves 1M passkey after 5k-length FT; 8B BookSum 500k SOTA Rouge overall 18.5. No official code on abs.

Keywords

infini-attention · long-context · compressive-memory · google · passkey · booksum

Topics

long context, compressive memory, linear attention

Research notes

  • Primary: arxiv abs (cs.CL; also cs.AI, cs.LG, cs.NE). License: arXiv.org perpetual non-exclusive on HTML at check. Comment: 9 pages, 4 figures, 4 tables (v2 adds background, implementation details, recent citations and acknowledgments). Google; correspondence tsendsuren@google.com. No official code on abs. HF paper page 111 upvotes; unofficial linked model mustafaaljadery/gemma-2B-10M not copied into hf_* fields. Discord posted HTML. Uses public PG19/Arxiv-math/BookSum rather than a new hosted corpus, so no datasets_local row. License field left blank per catalog convention.