Sharpness-Aware Pretraining Mitigates Catastrophic Forgetting
- Type
- other
- Venue
- arXiv / Carnegie Mellon University
Summary
Treats forgetting as local Hessian curvature along the fine-tuning direction. SAM, larger peak LR, and shorter WSD annealing each improve the learning–forgetting Pareto frontier even when they do not lower base pretraining loss. OLMo-60M / 192B tokens: SAM cuts StarCoder forgetting ~80% at matched FT loss; gap widens with token budget. Late-only SAM during WSD decay (~10% of steps) recovers much of full-SAM robustness. OLMo-2-1B mid-train 50B tokens: SAM vs OLMo recipe reduces forgetting 31% after MetaMath SFT and 40% after 4-bit NF4 quantization despite a slightly weaker base (42.9 vs 43.2 avg). Hessian analysis: SAM and large peak LR lower fine-tuning-directional sharpness. No official code on abs.
Keywords
sam · catastrophic-forgetting · sharpness · pretraining · quantization · olmo · dclm · icml · cmu
Topics
pretraining, catastrophic forgetting, sharpness-aware minimization
Research notes
- Primary: arxiv abs (cs.LG; also cs.CL). License not stated on abs/HTML at check. CMU (Raghunathan group); NSF GRFP DGE2140739; support Apple/Google/Jane Street/FLAME. No official code on abs. HF paper page 0 upvotes; no linked models/datasets. Discord posted abs. Uses public DCLM/Dolmino and FT sets; no new corpus, so no datasets_local row. License field left blank per catalog convention.