MeMo: Memory as a Model
- Type
- other
- Venue
- arXiv / NUS / A*STAR / University of Tokyo / Liquid AI / SMART / MIT CSAIL
Summary
MeMo distills a target corpus into reflection QA (fact extract, consolidate, verify, entity-surface, cross-document synthesis) and SFT a dedicated Memory model (e.g. Qwen2.5-14B) while the Executive stays frozen and queries it via a 3-stage multi-turn protocol. Retrieval cost is independent of corpus size; plug-and-play with closed-source Executives. NarrativeQA 26.85%/53.58% (Qwen2.5-32B / Gemini-3-Flash) vs HippoRAG2 21.39%/23.21%; MuSiQue 48.30%/60.20% vs 42.17%/57.00%; BrowseComp-Plus 54.22%/66.67% (trails HippoRAG2 56.11% on Qwen). Robust to distractor docs (±1.8pp vs ~5–6pp RAG drops). TIES-merge of two NarrativeQA halves cuts compute 33% at K=2 but loses 11–19pp vs full retrain.
Keywords
memo · parametric-memory · rag · reflection-qa · hipporag · narrativeqa · musique · browsecomp · nus · mit
Topics
LLM memory, knowledge integration, RAG alternatives
Research notes
- Primary: arxiv abs (cs.CL; also cs.AI, cs.LG). CC BY 4.0 on HTML. Quek equal contrib and corresponding (ryanquekweiheng@u.nus.edu). NUS IDS/ISEP/CS, A*STAR, UTokyo, Liquid AI, SMART, MIT CSAIL, AI Singapore. Code/datasets claimed in supplementary, not a public HF/GitHub pointer on abs. HF paper page 0 upvotes; no linked models/datasets. Discord posted abs. Evaluates on public BrowseComp-Plus/NarrativeQA/MuSiQue; synthesized reflection QA is not a standalone public release, so no datasets_local row. License field left blank per catalog convention.