Memory Caching: RNNs with Growing Memory
- Type
- paper
- Venue
- arXiv / Google
- Year
- 2026
- Source
- arxiv
- Access
- free
- Language
- en
- Added
- 2026-08-14T16:36:00Z
- Verified
- 2026-08-14T16:36:00Z
Summary
Memory Caching (MC) segments the sequence, caches each segment's memory state, and aggregates online plus cached memories at retrieval. Four variants: Residual Memory, Gated Residual Memory (GRM), Memory Soup, and Sparse Selective Caching (SSC). Complexity O(NL) interpolates RNN O(L) and Transformer O(L^2). Applied to Linear Attention, SWLA, DLA, and Titans. On 760M/30B-token and 1.3B/100B-token FineWeb training, MC lifts language-modeling and commonsense averages; Titans+GRM is strongest. Closes part of the recall gap vs Transformers on NIAH, in-context retrieval, LongBench, and MQAR. No official code on the abs page.
Keywords
memory-caching · titans · linear-attention · rnn · long-context · google · ssc · grm
Topics
sequence models, recurrent memory, long-context
Research notes
- Primary: arxiv abs (CC BY 4.0, cs.LG/AI). Correspondence alibehrouz@google.com. No official code on abs; unofficial community reimplementations exist and were not treated as the authors' release. Discord posted abs link.