Inference-time steering offers cheap, fine-grained control over a language model's outputs by estimating a concept's representation in activation space and shifting activations towards it. Existing methods build these representations from activation averages over contrastive datasets. These averages incorporate unrelated concepts and noise, and are dominated by a few tokens, meaning the activation transport encodes token-level rather than thematic concepts. In this work, we steer towards examples that most express a concept thematically, rather than towards an expectation over all. We identify these examples using influence functions, which estimate how much each data point contributes to a model's representation of a concept. Unlike simple model activation similarity, they incorporate the curvature of the model's loss landscape, allowing them to capture concept-relevant relationships beyond superficial token-level similarity. We then propose influence-weighted activation transport, which uses optimal transport to steer activations of non-concept text towards those of concept text, weighting concept examples by their influence scores. We evaluate on toxicity suppression (Jigsaw), object-based concept induction (OneSec) and truthfulness induction (TruthfulQA), outperforming existing activation-transport baselines. We track capability after steering using perplexity and MMLU accuracy, finding that our method improves steering while largely preserving model quality. We further show that influence functions capture concept-relevant information that activation-based methods miss with the two approaches ranking data points significantly differently. Together, these results demonstrate the value of curvature-aware influence information for activation steering.
Figures & tables
Figure 1 : Overview of IF-weighted activation transport.
PPL ↓
Method
Concept score (LLM) ↑
Flu.
Wiki
MMLU ↑
Gemma-2-2B (unsteered)
0.018 ± 0.019
6.58 ± 0.00
32.30 ± 0.00
0.549 ± 0.000
Act
0.708 ± 0.211
11.24 ± 2.78
91.18 ± 78.23
0.466 ± 0.031
IF-Act
0.721 ± 0.194
32.27 ± 20.74
1010.35 ± 927.72
0.267 ± 0.050
LinEAS
0.544 ± 0.243
8.38 ± 1.08
55.36 ± 21.98
0.499 ± 0.024
IF-LinEAS
0.584 ± 0.274
8.49 ± 1.00
61.55 ± 24.78
0.489 ± 0.025
Table 1 : OneSec concept induction: average over concepts. Steering on all layers; steered values at λ = 1, the unsteered row is λ = 0. PPL columns: Flu. = fluency perplexity (of the generated continuation), Wiki = held-out WikiText perplexity. Each value is the mean over OneSec concepts.
MC1 ↑
MC2 ↑
Method
n=32
n=1024
n=32
n=1024
Δ MC2 ↑
Gemma-2-2B (unsteered)
0.209 ± 0.407
0.209 ± 0.407
0.336 ± 0.395
0.336 ± 0.395
0.000
Act
0.245 ± 0.429
0.256 ± 0.436
0.396 ± 0.413
0.400 ± 0.406
0.003
IF-Act
0.278 ± 0.448
0.255 ± 0.436
0.459 ± 0.449
0.482 ± 0.455
0.023
LinEAS
0.212 ± 0.408
0.214 ± 0.410
0.348 ± 0.396
0.353 ± 0.398
0.005
IF-LinEAS
0.229 ± 0.420
0.203 ± 0.402
0.370 ± 0.400
0.356 ± 0.405
-0.013
Table 2 : TruthfulQA: effect of the number of contrastive pairs. Steering on all layers; steered values at λ = 1, the unsteered row is λ = 0. Fold-averaged MC1/MC2 at λ = 1. Δ = n=1024 minus n=32.
Toxicity ↓
PPL ↓
RTP ↓
TET ↓
Method
LLM
RoB.
Flu.
Wiki
MMLU ↑
LLM
RoB.
LLM
RoB.
Gemma-2-2B (unsteered)
0.344
0.233
27.11
32.30
0.550
0.107
0.033
0.219
0.133
Act
0.083
0.044
27.45
32.43
0.524
0.016
0.005
0.098
0.048
IF-Act
0.031
0.001
31.13
51.60
0.438
0.016
0.003
0.059
0.018
LinEAS
0.095
0.045
27.69
32.44
0.520
0.041
0.006
0.083
0.036
IF-LinEAS
0.018
0.009
27.58
34.57
0.520
0.034
0.008
0.073
0.033
Table 3 : Toxicity suppression. Steering on all layers; steered values at λ = 1, the unsteered row is λ = 0. PPL columns: Flu. = fluency perplexity (of the generated continuation), Wiki = held-out WikiText perplexity. RoB. = RoBERTa toxicity classifier judge.
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 2 : The mean cosine similarity between vectors in each group. ‘concept’ are the gradients of the concept completion. ‘control’ are the gradients of the non-concept completion. ‘paired’ are the gradients of the concept completion − non-concept completion. ‘Aggregate’ is the mean of the ‘paired’ gradients.
Figure 3 : ‘concept’ are the gradients of the concept completion. ‘control’ are the gradients of the non-concept completion. ‘paired’ are the gradients of the concept completion − non-concept completion. ‘Aggregate’ is the mean of the ‘paired’ gradients.
Figure 4 : Here we vary the number of contrastive prompt completion pairs whilst classifying toxicity using EK-FAC and MLP parameters
Model
Hook points
Gemma-2-2B
model.layers.{i}.post_attention_layernorm
model.layers.{i}.post_feedforward_layernorm
Llama-3-8B
model.layers.{i}.mlp.down_proj
Appendix
Table 4 : Hook points used for the Gemma and Llama continuations, applied at every transformer block.
Parameter
Value
temperature
1.0
top_p
0.3
top_k
50
repetition_penalty
1.2
max_new_tokens
50
stop_at_eos
True (no forced min_new_tokens )
Appendix
Table 5 : Generation hyperparameters used for the judge model evaluation.
Parameter
Value
temperature
1.0
top_p
0.3
top_k
50
repetition_penalty
1.2
max_new_tokens
32
stop_at_eos
True (no forced min_new_tokens )
Appendix
Table 6 : Generation hyperparameters used for both Gemma and Llama in experiments
Figure 5 : Effective sample size for our temperature scores over the layers for 1000 data points of the jigsaw toxicity dataset
Figure 6 : Overview of the IF weighted transport matrix Γ . Data points with greater influence score will have larger steps in the CDF function and therefore a larger quantile in the quantile function that results in more mass being transported to it in Γ .
PPL ↓
Layers
Method
Concept score (LLM) ↑
Flu.
Wiki
MMLU ↑
Gemma-2-2B
–
Unsteered
0.018 ± 0.019
6.58 ± 0.00
32.30 ± 0.00
0.549 ± 0.000
All layers
All
Act
0.708 ± 0.211
11.24 ± 2.78
91.18 ± 78.23
0.466 ± 0.031
All
IF-Act
0.721 ± 0.194
32.27 ± 20.74
1010.35 ± 927.72
0.267 ± 0.050
Appendix
Table 7 : OneSec concept induction: average over concepts. Steered values at λ = 1; the unsteered row is λ = 0. Layers: All = steering on all layers, Excl. = first and last layer excluded. PPL columns: Flu. = fluency perplexity (of the generated continuation), Wiki = held-out WikiText perplexity. Each value is the mean over OneSec concepts. Full results with standard deviations.
Figure 7 : IF weighted steering methods achieve the best concept indiction scores in most settings. Apart from Llama IF-Act. We again see that while IF-Act gets good concept score performance it disrupts the model indiction much more than the other methods and causes increasingly poor perplexity scores.
Figure 8 : Object based concepts achieve various levels of success. While concepts such as ’church’ , ’book’, ’flower’ and ’football’ achieve good steering performance, the remaining concepts are steered relatively poorly, with Llama ’Balloon’ completely failing any sort of concept indiction.
Figure 9 : Llama and Gemma’s performance diverges when first and last layers are not steered upon While Gemma adapts well to not steering on the first and last layers, Llama largely collapses and mostly fails to induce any concept meaningfully.
Toxicity ↓
PPL ↓
RTP ↓
TET ↓
Layers
Method
LLM
RoB.
Flu.
Wiki
MMLU ↑
LLM
RoB.
LLM
RoB.
Gemma-2-2B
–
Unsteered
0.344 ± 0.439
0.233 ± 0.407
27.11 ± 20.63
32.30 (1.25)
0.550 ± 0.498
0.107 ± 0.309
0.033 ± 0.179
0.219 ± 0.413
0.133 ± 0.340
All layers
All
Act
0.083 ± 0.252
0.044 ± 0.187
27.45 ± 20.41
32.43 (1.23)
0.524 ± 0.499
0.016 ± 0.125
0.005 ± 0.071
0.098 ± 0.298
0.048 ± 0.214
All
IF-Act
0.031 ± 0.145
0.001 ± 0.019
31.13 ± 20.09
51.60 (1.34)
0.438 ± 0.496
0.016 ± 0.125
0.003 ± 0.055
0.059 ± 0.235
0.018 ± 0.133
Appendix
Table 8 : Toxicity suppression. Steered values at λ = 1; the unsteered row is λ = 0. Layers: All = steering on all layers, Excl. = first and last layer excluded. PPL columns: Flu. = fluency perplexity (of the generated continuation), Wiki = held-out WikiText perplexity (a side-effect check on general language modelling). RoB. = RoBERTa toxicity classifier judge. IF-LinEAS = weighted-Wasserstein LinEAS Full results with standard deviations.
Figure 10 : IF weighted steering methods achieve the best concept suppression and induction scores, when compared to their non-weighted baselines. Again we also see that while IF-Act achieves a strong toxicity indiction score, it does so at the cost of a very high perplexity score in both models. overall the suppression arms maintain a largely consistent perplexity, whilst this is largely explained by the lover overall change in its objective of suppression than the more volatile indiction objective.
Figure 11 : As in Figure 10 , IF-weighted steering methods achieve the best concept suppression and induction scores in most settings, with the exception of IF-Act on Llama supression. IF-Act again reaches a strong toxicity induction score, but only at the cost of very high perplexity in both models. The perplexity trends also match the earlier results. Under the suppression objective, perplexity stays roughly constant as steering strength increases. Under the induction objective, it rises steadily.
MC1 ↑
MC2 ↑
Layers
Method
n=32
n=1024
n=32
n=1024
Δ MC2 ↑
Gemma-2-2B (unsteered)
0.209 ± 0.407
0.209 ± 0.407
0.336 ± 0.395
0.336 ± 0.395
0.000
All
Act
0.245 ± 0.429
0.256 ± 0.436
0.396 ± 0.413
0.400 ± 0.406
0.003
IF-Act
0.278 ± 0.448
0.255 ± 0.436
0.459 ± 0.449
0.482 ± 0.455
0.023
LinEAS
0.212 ± 0.408
0.214 ± 0.410
0.348 ± 0.396
0.353 ± 0.398
0.005
IF-LinEAS
0.229 ± 0.420
0.203 ± 0.402
0.370 ± 0.400
0.356 ± 0.405
-0.013
Appendix
Table 9 : TruthfulQA: effect of the number of contrastive pairs. Steered values at λ = 1; the unsteered row is λ = 0. Layers: All = steering on all layers, Excl. = first and last layer excluded. PPL columns: Flu. = fluency perplexity (of the generated continuation), Wiki = held-out WikiText perplexity (a side-effect check on general language modelling). Fold-averaged MC1/MC2 at λ = 1. Δ = n=1024 minus n=32. Full results, including layer-exclusion ablation.
Figure 12 : Influence functions localise the toxicity concept to the middle layers. We plot the mean influence weight that each data point assigns to each layer of the model, for each concept. A non-uniform distribution means that the influence score has identified layers where steering on that data point is more beneficial. For OneSec (object-based concepts) and TruthfulQA (truthfulness), the weights are close to uniform across layers. For toxicity, they are clearly concentrated in the middle layers.
Method pair
ρ * (candidate axis)
ρ * (layer axis)
Core methods
Projection proxy vs. Gradient norm
0.039
0.172
EK-FAC influence vs. Projection proxy
0.317
0.504
EK-FAC influence vs. Gradient norm
0.079
0.346
Appendix
Table 10 : Influence vs. projection proxy: method-pair agreement (Gemma-2-2B, toxic). ρ is the Spearman correlation between two methods’ scores, corrected for both methods’ split-half unreliability.
Activation steering has emerged as a promising alternative for controlling language-model behavior at inference time by modifying intermediate representations while keeping model parameters frozen. However, large-scale evaluations such as AxBench show that existing steering methods are often outperformed by simple in-context prompting and generalize poorly to unseen concepts. We hypothesize that these limitations arise from unvalidated simplifying assumptions shared across prior methods, which typically restrict steering interventions to fixed, single-step, position-invariant transforms. We propose FLAS (Flow-based Activation Steering), which learns a general, concept-conditioned velocity field vt(h,t,c) that transports unsteered activations to steered ones without relying on these assumptions. On AxBench, FLAS is the first learned method to consistently outperform prompting, reaching held-out harmonic means of 1.015 on Gemma-2-2B-IT and 1.113 on Gemma-2-9B-IT without per-concept tuning. Analysis of the learned flow shows curved, multi-step, token-varying trajectories, which suggests that previous hypotheses on activation space geometry might be incomplete.
Activation-based control steers large language models (LLMs) by intervening on their internal representations during inference, and has emerged as an effective paradigm for controlling behaviors such as persona and style. However, existing methods often rely on fixed steering directions or task-specific intervention modules, making them difficult to adapt to fine-grained concepts and compositional constraints. We propose UniSteer, a text-guided activation flow matching model that learns a conditional distribution over residual-stream activations from natural-language conditions. Instead of fitting a separate intervention for each target behavior, UniSteer learns a universal conditional velocity field in activation space. At inference time, UniSteer performs flow inversion by partially transporting a source activation toward a latent state and regenerating it under a target textual condition before injecting it back into the frozen LLM. The same conditional model supports activation-space classification by selecting the textual label with the lowest reconstruction energy. Experiments on three target LLMs show that UniSteer provides a unified interface across behavioral control, truthfulness steering, fine-grained concept steering, multi-constraint instruction following, and activation-space classification.
Large language models can be steered at inference time through prompting or activation interventions, but activation steering methods often underperform compared to prompt-based approaches. We propose a framework that formulates prompt steering as a form of activation steering and investigates whether distilling successful prompt steering behavior into simpler, interpretable models can close this gap. Our analysis reveals that popular activation steering methods are not faithful to the mechanics of prompt steering, which applies strong interventions on some tokens while barely affecting others. Based on these insights, we introduce Prompt Steering Replacement (PSR) models that estimate token-specific steering coefficients from the activations themselves and are trained to imitate prompt-based interventions. Experiments on three steering benchmarks across multiple language models show that PSR models outperform existing activation steering methods, especially when controlling for high-coherence completions, and also compare favorably to prompting on AxBench and persona steering.