← Back to explorer

Titans: Learning to Memorize at Test Time

Type
other
Venue
arXiv / Google

Summary

Treats attention as short-term memory and a deep MLP as long-term memory updated by surprise (gradient of associative ||M(k)-v||^2) with momentum and adaptive forget/weight-decay, plus persistent task tokens. Three hybrids: MAC, MAG, MAL. 760M MAG Wiki ppl 18.61 vs Transformer++ 25.21 and Gated DeltaNet-H2 19.88; MAC best on long-context NIAH. BABILong MAC beats GPT-4 and Llama3.1-8B+RAG at far fewer params. Claims >2M context. Code "available soon" on abs.

Keywords

titans · test-time-memory · surprise · mac · mag · mal · long-context · mamba · ttt · google

Topics

long-context memory, test-time training, hybrid architectures

Research notes

  • Primary: arxiv abs (cs.LG; also cs.AI, cs.CL). License not stated on abs/HTML at check. Google Research; correspondence {alibehrouz, peilinz, mirrokni}@google.com. No official code on abs (authors say coming soon). HF paper page 31 upvotes; 4 unofficial linked models and 2 unofficial datasets not copied into hf_* fields and not substantial, so no datasets_local row. Trains on FineWeb-Edu / Pile slices rather than releasing a new corpus. Discord posted abs.