Local Support Learning
- Type
- paper
- Venue
- arXiv:2610.02126 (cs.LG), submitted 1 Oct 2026
- Year
- 2026
- Source
- arxiv
- Access
- free
- Language
- English
- Added
- 2026-10-02
- Verified
- 2026-10-02
Summary
Local Support Learning (LSL) augments gradient-based training so LLMs of up to 7B parameters learn new tasks at full capacity while retaining prior capabilities — without access to prior data. The core idea: cast forgetting as a geometric problem and optimize for the worst case. A gradient update ΔW is a matrix that changes the output for every input not orthogonal to it, giving each update a global effect; LSL constrains the update to act only on the adapter's training distribution. Each learning phase pairs a standard weight adapter (e.g., LoRA), trained as usual, with a tailored gate enabling the adapter only on inputs from its training distribution — a GMM-based gate whose likelihood decays away from its training data, so it naturally stays closed on prior data it has never seen (conventional MLP classifiers lack this guarantee). Experiments fine-tune LLMs on chemistry, English-to-Igbo translation, and cybersecurity while measuring retention of pretraining skills (math, coding, instruction following): LSL achieves near-optimal retention with strong new-task performance across multiple phases, is robust to hyperparameter choice (no LR/batch-size/rank retention tradeoff), scales from 88% (1.5B) to 97% (3B) to 99% (7B) retention, and costs about the same as standard LoRA in memory and compute.
Keywords
Local Support Learning · LSL · catastrophic forgetting · continual learning · LoRA adapters · GMM gate · fine-tuning retention · weight-matrix geometry
Topics
catastrophic forgetting, continual learning, LoRA, adapters, gating, GMM, fine-tuning
Research notes
- Discovery: @abk_tau (Assaf Ben Kish) 7-part X thread 2026-10-02 (https://x.com/abk_tau/status/2106005515532632069)
- Website and code: https://assafbk.github.io/lsl
- Mechanism detail: a single weight-matrix gradient update changes the output for every input not orthogonal to it, including inputs from previous learning phases; LSL makes the update local to the adapter's training distribution via a GMM gate
- Theoretical motivation: at pretraining scale prior data is too large/unavailable to revisit, so optimize for the worst case — a prior sample could lie anywhere outside current data; for unrelated prior tasks forgetting happens off the training data (excess), but the gate's objective only sees current data, so the objective alone cannot solve forgetting — choose the hypothesis class smartly (density estimators, whose likelihood decays away from training data)
- Results: near-optimal retention of math/coding/instruction-following plus strong new-task performance on chemistry, English->Igbo, cybersecurity across multiple finetuning phases; retention 88% (1.5B) -> 97% (3B) -> 99% (7B); no hyperparameter retention tradeoff; memory/compute overhead on par with standard LoRA
- Collaborators: @akarshkumar0101, @RGiryes, James Glass
- License: CC BY 4.0