Inoculation Adapters: Improved Selective Generalization of Capabilities with Fewer Surprising Backdoors
- Type
- paper
- Venue
- arXiv / Center on Long-Term Risk
- Year
- 2026
- Source
- arxiv
- Access
- free
- Language
- en
- Added
- 2026-08-14T16:44:00Z
- Verified
- 2026-08-14T16:44:00Z
Summary
IA/GIA/CGIA: train a LoRA on the undesired trait, freeze it while training a task LoRA on mixed desired+undesired data, drop the IA at deploy. Compared with inoculation prompting, preventative steering, CAFT, and KL across 9 setups and 5 model families. The family occupies the observed suppression–retention Pareto frontier (wide CIs). Vanilla IA strongly suppresses undesired traits and adds fewer surprising backdoors than IP on negated/structure/keyword probes. Works on unelicitable traits (new cipher capability, hate speech under refusal, base-model sycophancy) where IP fails. GIA/CGIA trade some backdoor-resistance for better desired-trait retention. Authors note the magnitude vs best baselines is uncertain.
Keywords
inoculation-adapters · emergent-misalignment · lora · selective-generalization · clr · inoculation-prompting
Topics
AI alignment, selective generalization, emergent misalignment
Research notes
- Primary: arxiv abs (CC BY 4.0, cs.AI). Code https://github.com/longtermrisk/inoculation-adapters (2 stars at check). Correspondence maxime.riche@longtermrisk.org. HF paper page 0 upvotes. Discord posted abs link. Training data synthesized from public instruction sets; no standalone public dataset release confirmed, so no datasets_local row.