Continual Learning via Sparse Memory Finetuning
- Type
- other
- Venue
- arXiv / Meta FAIR / UC Berkeley
Summary
Replaces one mid-stack FFN with a 1M-slot memory layer (k=32, 4 heads) and finetunes only the top-t slots that are highly accessed on the current batch relative to DCLM pretraining (TF-IDF). On a 1.3B model, TriviaQA 1K fact stream: NaturalQuestions F1 drops 11% vs 89% full finetuning and 71% LoRA at matched target learning; document-stream SimpleQA shows the same Pareto pattern. Naive memory FT and TF-only ranking forget more. No official code on abs.
Keywords
sparse-memory-finetuning · continual-learning · catastrophic-forgetting · memory-layers · lora · tf-idf · fair · meta
Topics
continual learning, memory layers, catastrophic forgetting
Research notes
- Primary: arxiv abs (cs.CL; also cs.AI). License not stated on abs/HTML at check. FAIR at Meta / UC Berkeley. Correspondence jessy_lin@berkeley.edu, {vincentpierre,barlaso}@meta.com. No official code on abs. HF paper page 3 upvotes; no linked models/datasets. Discord posted abs (Substack UTM). Uses public TriviaQA/NaturalQuestions/SimpleQA/DCLM; no new corpus, so no datasets_local row. License field left blank per catalog convention.