Targeted Neuron Modulation via Contrastive Pair Search
- Type
- other
- Venue
- arXiv / Nous Research
Summary
Contrastive neuron attribution (CNA) ranks MLP neurons by harmful-vs-benign last-token activation difference using only forward passes; top 0.1% form the circuit. Ablating it cuts JBB-Behaviors refusal >50% on Llama/Qwen instruct 1B–72B while n-gram quality stays >0.96 and MMLU is preserved; CAA matches the refusal drop but collapses quality at high α. Matched base models have similar late-layer discrimination structure but steering only shifts content, not refusal — alignment crystallizes a sparse gate. Code https://github.com/NousResearch/neural-steering.
Keywords
cna · refusal · neuron-ablation · caa · jailbreakbench · llama · qwen · nous-research · interpretability
Topics
interpretability, refusal circuits, activation steering
Research notes
- Primary: arxiv abs (cs.LG). ArXiv HTML states perpetual non-exclusive license. Nous Research; correspondence nightwing/jake/karan@nousresearch.com. Code MIT https://github.com/NousResearch/neural-steering (33 stars / 15 forks at check). HF paper page 17 upvotes; linked models are unofficial abliterated forks (Carlosian/*), not recorded in hf_* fields. Discord posted abs (substack UTM). Uses public JBB-Behaviors; no new corpus, so no datasets_local row. License field left blank per catalog convention.