← Back to explorer

Can a Language Model Learn Facts Continually in Its Weights?

Type
paper
Venue
arXiv / Baseten
Year
2026
Source
arxiv
Access
free
Language
en
Added
2026-08-14T16:40:00Z
Verified
2026-08-14T16:40:00Z

Summary

Writes invented facts into Qwen3-4B (8B replication of the entailment gap) via LoRA/full FT and context distillation, then follows them through 20–100 later writes with five held-out question types against a fact-in-prompt ceiling. Training-data breadth, not the objective label, creates usable knowledge: diverse restatements cut the recitation-to-use gap from 27.4 to 5.4 points without showing conclusions. After 20 sequential writes, bare-statement facts retain 1% vs 46% for study facts (plateau ~25–28% by 100 study writes). Forgotten facts keep 57–67% of their log-prob lift, so access fails rather than storage; incoming writes cause interference. A frozen original-model teacher preserves capability; no tested intervention keeps earlier facts reachable. Context remains the reliable channel for composition and survival.

Keywords

continual-learning · knowledge-writing · qwen3 · lora · context-distillation · forgetting · baseten · cortex

Topics

continual learning, knowledge writing, LLM memory

Research notes

  • Primary: arxiv abs (CC BY 4.0, cs.CL/LG). Baseten. Code https://github.com/basetenlabs/cortex. Paper lists eval data at https://huggingface.co/datasets/baseten/cortex; HF API returned 401 at catalog time so public access was not confirmed. 247-fact eval set is small, not added to datasets_local.csv. Discord posted abs link (also a duplicate abs post 1527767116177215709).