Emergent Capabilities Arise Randomly from Learning Sparse Attention Patterns
- Type
- paper
- Venue
- arXiv / NYU
- Year
- 2026
- Source
- arxiv
- Access
- free
- Language
- en
- Added
- 2026-08-14T19:25:00Z
- Verified
- 2026-08-14T19:25:00Z
Summary
NYU (18 pp / 13 figs). On Pythia 14M–410M × 10 seeds, copy, in-context repetition, pattern completion, and IOI appear at random training steps; larger models find them earlier and more often. Patching the top-K post-emergence attention maps into the previous checkpoint elicits the skill, so attention-pattern search is the bottleneck rather than a metric illusion. Synthetic linear-map (sparse Ax mod 2) and cellular-automata tasks: difficulty is governed by context length and medium sparsity; biasing attention logits to the ground-truth pattern removes the loss plateau. More heads help; MLP-Mixer beats a transformer by ~10× on linear maps but loses on automata. Dyck/NCA pre-pretraining has mixed effects on emergence timing. No official code on the abs.
Keywords
emergence · sparse-attention · pythia · mlp-mixer · nyu · x
Topics
emergence, mechanistic interpretability, attention
Research notes
- Primary: arxiv abs 2606.25010 (cs.LG/cs.CL). Discord/X https://x.com/che_shr_cat/status/2084239996790120638 via fxtwitter (third-party recap; author also writes arxiviq.substack.com). Correspondence vatsalbaherwani@nyu.edu. Author page https://vatsal0.github.io/ links paper+blog; no code URL at check. Synthetic training data not released. Not a hosted corpus, so no datasets_local row.