← Back to explorer

The Atlas Neuron (Arth Singh: continual learning with no backprop, no replay, no task labels)

Type
blog
Venue
arthsingh.com (blog post, 2026-09-30)
Year
2026
Source
blog
Access
public
Language
en
Added
2026-09-30
Verified
2026-09-30

Summary

Can a model learn tasks one after another without forgetting, with no replay, no task labels, and no backprop? The Atlas neuron: keeps one page of memory per "world"; when several batches in a row look unfamiliar (3 consecutive batches below 85% of the world's frozen usual familiarity, or batches with only never-seen labels), it forks a new page, and only the current page learns, so old pages can't be overwritten. Each image is compressed by a fixed random projection to 256 numbers; per class, each page stores 128 landmarks (online k-means) and a chart in pixel space = the class mean plus its top 16 principal directions (incremental PCA, one streaming pass). To classify, each chart bends from its mean along its 16 directions toward the input; the leftover squared distance is what the chart cannot explain; the class score combines the smallest unexplained distance among its charts with its closest landmark similarity across ALL pages (no routing to a world first). The two vote weights are the only numbers learned by gradient descent (self-calibrated from inverse spreads). Results in the Mammoth continual-learning library (one pass, no task labels, 5 seeds): 96.7% after 20 tasks of Permuted MNIST vs 45% for a sequentially trained MLP; beats DER++ with a 5,120-image buffer on 5 of 6 benchmarks (exception: Rotated MNIST 94.2 vs 94.5), e.g. 96.7 vs 92.3 on Permuted MNIST; also beats the joint-trained MLP reference on 5 of 6. A landmark-only lookup-table ablation with 2x memory matches 3/6 tasks but collapses on Split CIFAR-10 (17.0 vs 41.5) -- the charts do the work. Costs: 9.2M stored floats on permuted tasks vs DER++'s 4.2M; a leaner shared-charts variant matches on permuted tasks at 2.4M but loses on Split CIFAR-100. On natural images it only partly works: with frozen unsupervised patch features it reaches 41.5% on Split CIFAR-10 but stays weak on Split CIFAR-100 (14.4% vs 37.5% DER++); streaming LDA beats it above 2M floats. Caveats: rotated nearby angles merge into one page; Permuted KMNIST is the one benchmark where the lookup table stays ahead; memory grows with the number of worlds. Partly old lineage: subspace classifiers back to CLAFIC (1960s), Hinton/Dayan/Revow's 1997 wake-sleep charts; new is the automatic forking rule bolted onto modern continual-learning benchmarks.

Keywords

continual learning · catastrophic forgetting · online learning · subspace classifiers · benchmarks

Topics

continual learning, catastrophic forgetting, online learning

Research notes

  • Discovery: Arth Singh (@iarthsingh, verified, AI Safety Research Scientist, U Toronto/Vector under Zhijing Jin) 2026-09-30 thread: https://x.com/iarthsingh/status/2105296609655631908
  • Full post with all tables and reproduction details: https://www.arthsingh.com/blog/atlas-neuron
  • 15-minute read; no code repository linked on the page.
  • Connects to the collection's continual-learning, catastrophic-forgetting, and online-learning entries.