Memory Attention
- Type
- paper
- Venue
- arXiv:2609.28399 (via alphaXiv)
- Year
- 2026
- Source
- web
- Access
- public
- Language
- en
- Added
- 2026-09-29
- Verified
- 2026-09-29
Summary
Memory Attention (Jiale Kang) replaces the Transformer's learned value projection with V = K + Norm(E[s]) — the sum of contextual keys and layer-specific token-indexed memory vectors. Under matched training budgets it beats standard attention on WikiText perplexity and downstream average while enabling CPU offloading of the memory tables, trading parameter storage for lookup-based capacity.
Keywords
attention · memory · architecture · efficiency
Topics
attention, memory, architecture, efficiency
Research notes
- Method: Remove the learned value projection; instead retrieve a learned vector by token ID from a per-layer embedding table, RMSNorm it per token/head, and add it to the contextual key. Keep standard attention weighting/aggregation (RoPE applied to Q/K before values are formed).
- Key findings: V = K + Norm(E[s]): values formed by adding layer-specific token-indexed memory to contextual keys; each layer has its own table, same token → same row across positions/contexts; At matched 10B-token training budgets: better WikiText perplexity and higher downstream average than standard attention, at the cost of more parameters; At inference, normalization folds into the tables → value construction is lookup + addition; MA-Offload: tables stored on CPU/SSD with prefetching, cutting GPU parameter storage; MA-Recall: reconstruct values from retained keys + memory to cut value-cache storage
- Limitations: Single-author paper; more total parameters than baseline; latency cost of CPU offload depends on retrieval/transfer overlap with compute — not yet benchmarked at scale.
- Related to Value Embedding, DeepEmbed, Per-Layer Embeddings, STEM, Engram — but replaces rather than supplements the value projection.