← Back to explorer

Memory Attention

Type
paper
Venue
arXiv:2609.28399 (via alphaXiv)
Year
2026
Source
web
Access
public
Language
en
Added
2026-09-29
Verified
2026-09-29

Summary

Memory Attention (Jiale Kang) replaces the Transformer's learned value projection with V = K + Norm(E[s]) — the sum of contextual keys and layer-specific token-indexed memory vectors. Under matched training budgets it beats standard attention on WikiText perplexity and downstream average while enabling CPU offloading of the memory tables, trading parameter storage for lookup-based capacity.

Keywords

attention · memory · architecture · efficiency

Topics

attention, memory, architecture, efficiency

Research notes

  • Method: Remove the learned value projection; instead retrieve a learned vector by token ID from a per-layer embedding table, RMSNorm it per token/head, and add it to the contextual key. Keep standard attention weighting/aggregation (RoPE applied to Q/K before values are formed).
  • Key findings: V = K + Norm(E[s]): values formed by adding layer-specific token-indexed memory to contextual keys; each layer has its own table, same token → same row across positions/contexts; At matched 10B-token training budgets: better WikiText perplexity and higher downstream average than standard attention, at the cost of more parameters; At inference, normalization folds into the tables → value construction is lookup + addition; MA-Offload: tables stored on CPU/SSD with prefetching, cutting GPU parameter storage; MA-Recall: reconstruct values from retained keys + memory to cut value-cache storage
  • Limitations: Single-author paper; more total parameters than baseline; latency cost of CPU offload depends on retrieval/transfer overlap with compute — not yet benchmarked at scale.
  • Related to Value Embedding, DeepEmbed, Per-Layer Embeddings, STEM, Engram — but replaces rather than supplements the value projection.