Minimally Invasive Steering of Language Models (MISVO)
- Type
- paper
- Venue
- NeurIPS 2026 (arXiv 2026-09-24, v1)
- Year
- 2026
- Source
- paper
- Access
- public
- Language
- en
- Added
- 2026-09-30
- Verified
- 2026-09-30
Summary
Pre-logit steering adapts a frozen LLM to a test-time reward by adding vectors to its final hidden states -- no parameter updates. MISVO (Minimally Invasive Steering Vector Optimization) penalizes interventions using the local KL geometry of the induced token distribution: a steering vector's magnitude need not reflect how much it changes the model, so MISVO uses a Fisher-quadratic penalty for distribution shift, which admits an analytic gradient via matrix-vector products with the frozen LM head. Includes an exact decomposition of the sequence-level KL gradient into an analytic Fisher term and a suffix score-function term; theory shows the suffix term is second order in steering magnitude and three Fisher surrogates agree with the full KL gradient to first order. Across preference and code-generation tasks on ~1B-14B models, MISVO achieves the highest mean reward in six of seven model-task settings, with diversity and coherence close to Best-of-N. Uses the frozen-reference surrogate to optimize position-specific interventions.
Keywords
steering vectors · inference-time steering · KL geometry · Fisher information · alignment
Topics
steering vectors, inference-time steering, KL geometry, Fisher information, alignment, reward optimization
Research notes
- Discovery: Jack Jingyu Zhang (@jackjingyuzhang, verified) 2026-09-30 post: https://x.com/jackjingyuzhang/status/2105372147699093637?s=20 quoting Mahyar Fazlyab's thread https://x.com/MFazlyab/status/2105023751091896325
- No code repo, dataset, demo, or product link in the announcement.
- License: arXiv non-exclusive distribution only (no open CC license).
- Connects to the collection's steering/intervention entries.