← Back to explorer

Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activations

Type
paper
Venue
Anthropic / Transformer Circuits
Year
2026
Source
web
Access
free
Language
en
Added
2026-08-14T19:20:00Z
Verified
2026-08-14T19:20:00Z

Summary

Anthropic Transformer Circuits method (May 2026). An activation verbalizer maps a target residual to text and an activation reconstructor maps that text back; both are copies of the target LLM jointly trained with RL on reconstruction (FVE ~0.6-0.8) after a summarization warm-start. Explanations grow more informative over training on Haiku 3.5/4.5 and Opus 4.6. Used in Opus 4.6 / Mythos Preview audits: unverbalized evaluation awareness, language-switching traced to malformed SFT data, rhyme planning, and a hidden-motivation auditing game where NLA-equipped agents win 12-15% without training-data access vs <3% without NLAs. Limitations: confabulation of specifics, cost (two-model RL; hundreds of tokens per activation), single-layer reads. Open training code and NLAs for popular models; Neuronpedia demo.

Keywords

nla · interpretability · anthropic · transformer-circuits · sae · auditing · evaluation-awareness · blog

Topics

interpretability, activation explanations, model auditing

Research notes

  • Primary: Transformer Circuits paper (cite frasertaliente2026nla). Discord posted https://www.anthropic.com/research/natural-language-autoencoders. Code https://github.com/kitft/natural_language_autoencoders (Apache-2.0, 922 stars / 127 forks at check). Equal-contribution AV authors Fraser-Taliente, Kantamneni, Ong; correspondence subhash@anthropic.com. Neuronpedia interactive frontend. Method/interpretability item, not a new hosted corpus, so no datasets_local row.