← Back to explorer

Minimally Invasive Steering of Language Models (MISVO)

Type
paper
Venue
NeurIPS 2026 (arXiv 2026-09-24, v1)
Year
2026
Source
paper
Access
public
Language
en
Added
2026-09-30
Verified
2026-09-30

Summary

Pre-logit steering adapts a frozen LLM to a test-time reward by adding vectors to its final hidden states -- no parameter updates. MISVO (Minimally Invasive Steering Vector Optimization) penalizes interventions using the local KL geometry of the induced token distribution: a steering vector's magnitude need not reflect how much it changes the model, so MISVO uses a Fisher-quadratic penalty for distribution shift, which admits an analytic gradient via matrix-vector products with the frozen LM head. Includes an exact decomposition of the sequence-level KL gradient into an analytic Fisher term and a suffix score-function term; theory shows the suffix term is second order in steering magnitude and three Fisher surrogates agree with the full KL gradient to first order. Across preference and code-generation tasks on ~1B-14B models, MISVO achieves the highest mean reward in six of seven model-task settings, with diversity and coherence close to Best-of-N. Uses the frozen-reference surrogate to optimize position-specific interventions.

Keywords

steering vectors · inference-time steering · KL geometry · Fisher information · alignment

Topics

steering vectors, inference-time steering, KL geometry, Fisher information, alignment, reward optimization

Research notes

  • Discovery: Jack Jingyu Zhang (@jackjingyuzhang, verified) 2026-09-30 post: https://x.com/jackjingyuzhang/status/2105372147699093637?s=20 quoting Mahyar Fazlyab's thread https://x.com/MFazlyab/status/2105023751091896325
  • No code repo, dataset, demo, or product link in the announcement.
  • License: arXiv non-exclusive distribution only (no open CC license).
  • Connects to the collection's steering/intervention entries.