← Back to explorer

Inoculation Adapters: Improved Selective Generalization of Capabilities with Fewer Surprising Backdoors

Type
paper
Venue
arXiv / Center on Long-Term Risk
Year
2026
Source
arxiv
Access
free
Language
en
Added
2026-08-14T16:44:00Z
Verified
2026-08-14T16:44:00Z

Summary

IA/GIA/CGIA: train a LoRA on the undesired trait, freeze it while training a task LoRA on mixed desired+undesired data, drop the IA at deploy. Compared with inoculation prompting, preventative steering, CAFT, and KL across 9 setups and 5 model families. The family occupies the observed suppression–retention Pareto frontier (wide CIs). Vanilla IA strongly suppresses undesired traits and adds fewer surprising backdoors than IP on negated/structure/keyword probes. Works on unelicitable traits (new cipher capability, hate speech under refusal, base-model sycophancy) where IP fails. GIA/CGIA trade some backdoor-resistance for better desired-trait retention. Authors note the magnitude vs best baselines is uncertain.

Keywords

inoculation-adapters · emergent-misalignment · lora · selective-generalization · clr · inoculation-prompting

Topics

AI alignment, selective generalization, emergent misalignment

Research notes

  • Primary: arxiv abs (CC BY 4.0, cs.AI). Code https://github.com/longtermrisk/inoculation-adapters (2 stars at check). Correspondence maxime.riche@longtermrisk.org. HF paper page 0 upvotes. Discord posted abs link. Training data synthesized from public instruction sets; no standalone public dataset release confirmed, so no datasets_local row.