Pre-logit steering adapts a frozen language model to a test-time reward by adding vectors to its final hidden states. Unregularized reward optimization can substantially alter the output distribution and degrade generation quality. We propose Minimally Invasive Steering Vector Optimization (MISVO), which penalizes interventions using the local KL geometry of the induced token distribution. The resulting Fisher quadratic measures distributional sensitivity and admits an analytic gradient computed through matrix--vector products with the frozen language-model head. We derive an exact decomposition of the sequence-level KL gradient into an analytic Fisher term and a suffix score-function term. For a fixed generation horizon, we show that the suffix term is second order in the steering magnitude and that three Fisher surrogates agree with the full KL gradient to first order. MISVO uses the frozen-reference surrogate to optimize position-specific interventions without updating model parameters. Across preference and code-generation tasks on models with approximately 1B--14B parameters, MISVO achieves the highest mean reward in six of seven model--task settings, with diversity and coherence scores close to those of Best-of-N.
Figures & tables
Figure 1: Minimally invasive test-time steering. MISVO adds a position-specific vector ut to the hidden state ht entering the frozen LM head (W,b) . The vectors are optimized for reward with a Fisher-quadratic penalty that locally approximates KL divergence from the reference policy.
Model
Method
Reward ↑
Diversity ↑
Coherence ↑
LFM2.5 1.2B
BoN (top- p )
8.70±0.11
0.688±0.004
0.671±0.004
AISP
9.09±0.03
0.726±0.015
0.663±0.009
MISVO
9.29±0.04
0.719±0.002
0.654±0.005
Gemma3 4B
BoN (top- p )
9.10±0.11
0.721±0.002
0.632±0.006
AISP
9.27±0.04
0.741±0.005
0.631±0.005
MISVO
9.56±0.04
0.728±0.005
0.632±0.003
Table 1: Results on SHP across model scales. We sample 500 prompts and score with Skywork-Reward-V2-Qwen3-0.6B. Reward (best-so-far over the optimization trace) is the primary metric, reported as mean ± std over 3 seeds ( 42,1340,2026 ); bold indicates the best per model. MISVO uses the established Fisher-quadratic in ( 9 ) Fˉt . All methods see a matched per-prompt rollout budget of K⋅N samples.
Model
Method
Pass rate ↑
Diversity ↑
Coherence ↑
LFM2.5 1.2B
BoN (top- p )
0.703±0.013
0.613±0.018
0.605±0.010
AISP ( λ=0.1 )
0.675±0.014
0.603±0.023
0.677±0.006
AISP ( λ=1 )
0.669±0.017
0.601±0.015
0.679±0.004
MISVO
0.739±0.027
0.588±0.008
0.680±0.007
Gemma3 4B
BoN (top- p )
0.739±0.010
0.596±0.031
0.594±0.005
AISP ( λ=0.1 )
0.742±0.008
0.610±0.016
0.593±0.004
Table 2: Results on MBPP+ across model scales. Metric is the pass rate (fraction of unit tests passed), reported as mean ± std over 3 seeds ( 42,1340,2026 ); bold indicates the best per model. AISP is shown at two guidance strengths λ∈{0.1,1.0} . MISVO uses the established Fisher-quadratic in ( 9 ) Fˉt .
Figure 2: Reward trajectory for LFM2.5-1.2B/SHP as a function of the cumulative number of sampled completions. Left: mean reward within each step. Middle: maximum reward within each step. Right: cumulative maximum reward.
Method / regularizer
Reward ↑
Final ∥u∥
Final KL ↓
MISVO
9.60±3.38
339.9±22.4
47.5±12.9
AISP
9.28±3.10
2631.7±86.3
303.9±61.6
Unregularized steering
9.31±3.30
474.6±24.9
162.2±38.3
Direct KL regularizer
8.49±3.07
459.3±29.1
120.1±29.7
Hidden-space ℓ2
8.39±2.86
48.3±20.4
0.92±1.1
BoN ( N=256 )
8.89±2.69
n/a
n/a
Table 3: Measured policy deviation on SHP with LFM2.5-1.2B. KL is KLref(u) from ( 17 ); ± denotes variation across prompts.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 3: Reward trajectory for Gemma 3 4B/SHP as a function of the cumulative number of sampled completions. Left: mean reward within each step. Middle: maximum reward within each step. Right: cumulative maximum reward.
Figure 4: Reward trajectory for Llama 3 8B Instruct/SHP as a function of the cumulative number of sampled completions. Left: mean reward within each step. Middle: maximum reward within each step. Right: cumulative maximum reward.
Figure 5: Reward trajectory for Phi 4/SHP as a function of the cumulative number of sampled completions. Left: mean reward within each step. Middle: maximum reward within each step. Right: cumulative maximum reward.
Regularizer
Reward ↑
Diversity ↑
Coherence ↑
off-policy rollouts, off-policy Fisher Fˉt
9.26
0.723
0.661
on-policy rollouts, off-policy Fisher Fˉtu
9.06
0.721
0.659
on-policy rollouts, on-policy Fisher F~tu
8.92
0.710
0.656
KL
8.35
0.725
0.660
Appendix
Table 4: Regularizer ablation on SHP with LFM2.5-1.2B. Settings: K=N=16 , η=0.1 , λ=1.0 , seed 42, 500 prompts. The three Fisher surrogates lie within close values of each other (Theorem 2 ).
LFM2.5-1.2B
Gemma3-4B-IT
Method
Hyperparams
Time (s)
Tok/s
Peak (GB)
Time (s)
Tok/s
Peak (GB)
BoN
N=256
24.4
5385
3.7
79.1
1658
12.2
AISP
K=16,N=16
43.3
3053
3.5
195.7
670
10.5
MISVO
K=16,N=16
49.3
2666
10.7
167.1
784
38.8
Appendix
Table 5: Per-prompt runtime, generation throughput, and peak GPU memory for each test-time alignment method on the SHP prompt set. Numbers are means over 10 timed prompts (after 5 warmup) on a single NVIDIA H200 NVL, max_new_tokens=512 , mean prompt length 131 tokens, no cross-prompt batching. AISP uses K=16 Gaussian perturbations per iteration over N=16 iterations; MISVO uses K=16 trajectories over N=16 gradient steps with frozen-Fisher regularization; BoN draws N=256 top- p samples.
Recently, steering vectors (SVs) have emerged as an effective and lightweight approach to steer behaviors of large language models (LLMs), among which fine-tuned SVs are more effective than optimization-free ones. However, current approaches to fine-tuned SVs suffer from two limitations. First, they require careful selection of steering factors on a per-SV basis to balance steering effectiveness and generation quality at inference time. Second, they operate as full-sequence SVs (FSSVs), which can sacrifice generation quality regardless of factor selection due to excessive intervention on the model generation process. To address the first limitation, we propose joint training of steering factors and directions, such that post-hoc factor selection is no longer required. Using neural network scaling theory, we find that moderately large initialization sizes and learning rates for steering factors are essential for stability and efficiency of joint training. To tackle the second limitation, we draw inspiration from representation fine-tuning and introduce Prompt-only SV (PrOSV), an SV that intervenes only on a few prompt tokens. Our empirical results show that PrOSV outperforms traditional FSSVs on AxBench when using our joint training scheme. We also find that PrOSV achieves a better tradeoff between general model utility and adversarial robustness than FSSV.
Yuntai Bao, Qinfeng Li, Xinyan Yu +6
Zhejiang University · Innovation and Management Center, School of Software Technology (Ningbo), Zhejiang University · Ant Group
Large language models can be steered at inference time through prompting or activation interventions, but activation steering methods often underperform compared to prompt-based approaches. We propose a framework that formulates prompt steering as a form of activation steering and investigates whether distilling successful prompt steering behavior into simpler, interpretable models can close this gap. Our analysis reveals that popular activation steering methods are not faithful to the mechanics of prompt steering, which applies strong interventions on some tokens while barely affecting others. Based on these insights, we introduce Prompt Steering Replacement (PSR) models that estimate token-specific steering coefficients from the activations themselves and are trained to imitate prompt-based interventions. Experiments on three steering benchmarks across multiple language models show that PSR models outperform existing activation steering methods, especially when controlling for high-coherence completions, and also compare favorably to prompting on AxBench and persona steering.
Applying steering vectors to large language models (LLMs) is an efficient and effective model alignment technique, but we lack an interpretable explanation for how it works--specifically, what internal mechanisms steering vectors affect and how this results in different model outputs. To investigate the causal mechanisms underlying the effectiveness of steering vectors, we conduct a comprehensive case study on refusal. We propose a multi-token activation patching framework and discover that different steering methodologies leverage functionally interchangeable circuits when applied at the same layer. These circuits reveal that steering vectors primarily interact with the attention mechanism through the OV circuit while largely ignoring the QK circuit. Freezing all attention scores during steering drops performance by only 8.83% across three model families. A mathematical decomposition of the steered OV circuit further reveals semantically interpretable concepts, even in cases where the steering vector itself does not. Leveraging the activation patching results, we show that steering vectors can be sparsified by up to 85-96% while retaining most performance, and that different steering methodologies agree on a subset of important dimensions.