Organizations: State Key Laboratory of General Artificial Intelligence, School of Intelligence Science and Technology, Peking University · University of Chinese Academy of Sciences · College of Artificial Intelligence, Tsinghua University
Value steering should change an LLM's normative priorities while preserving the scenario, facts, and task constraints underlying its answer. Conventional activation edits often change both. We introduce an editable semantic-value interface on frozen residual states, with a one-way semantic-to-value pathway that grounds value recognition in context. Stop-gradient blocks feedback through this pathway; swap consistency, topic de-confounding, and decorrelation encourage selective codes. At inference, editing the value code produces a residual delta while holding the semantic code fixed. On two instruction-tuned backbones, this interface improves semantic preservation and reduces benign refusals at comparable value alignment. A matched mixing-by-gating ablation separates representation learning from selective edit activation, and dimension-matched probes establish improved code selectivity. Against validation-selected prompting on LLaMA-3.1-8B, the method achieves comparable alignment (0.750 vs. 0.748), higher BERTScore (0.938 vs. 0.923), and fewer contradictions (5.1% vs. 7.6%). Human ratings and cross-taxonomy controls provide complementary evidence for low-damage value steering.
Figures & tables
Figure 1 : Overview of our semantic–value disentanglement framework for controllable LLM intervention. Hidden representations from a frozen LLM are decomposed into a semantic code and a value code. A one-way semantic-to-value mixing mechanism allows value representations to leverage semantic grounding while blocking reverse gradients through the mixing path. The disentangled value subspace can then be edited at inference time to steer generation with lower semantic damage.
Figure 2 : A grounded value interface with separate training and editing paths. The frozen state h feeds both encoders. The central bridge expands Eq. 1 : sg(zs) feeds projector M and mixing gate g , whose product augments the value code. Stop-gradient blocks reverse differentiation through this bridge. At inference, shared D compares target and original value codes with zs fixed. GatedNSI optionally projects and activates the resulting update; its inference gate a(x) is distinct from g . Without GatedNSI, Δh=Δhv . The lower band groups the training objectives.
Table 1 : Value measurement and representation selectivity. (a) ValueBench measurement results. Text-only judges are compared to latent readouts and a raw-state linear probe control. We report Accuracy / Macro-F1 for Relatedness and Stance. (b) Leakage probes: predicting value/topic from zs and zv . Lower off-diagonal accuracy indicates better disentanglement.
Method
α⋆
Align ↑
SemSim ↑
FRR ↓
Original (no edit)
0.00
0.290 ± 0.020
—
0.020 ± 0.004
LinearAdd
1.60
0.770 ± 0.010
0.792 ± 0.010
0.119 ± 0.012
RepE
1.60
0.705 ± 0.018
0.811 ± 0.009
0.091 ± 0.010
One-way + NSI
1.40
0.720 ± 0.012
0.842 ± 0.008
0.074 ± 0.008
No mixing + NSI
1.40
0.720 ± 0.010
0.832 ± 0.008
0.082 ± 0.007
Two-way mixing
1.40
0.735 ± 0.011
0.819 ± 0.009
0.088 ± 0.009
Table 2 : Value transfer on SVQ-Test and held-out prompts, using 630 training quadruples (LLaMA-3.1-8B-Instruct). α⋆ is selected by Section 4.6 . Mean ± std over three seeds; additional fluency, capability, and benign-similarity metrics are in Table 5 .
Figure 3 : Mixing and edit gating make complementary contributions. Mean results of the matched 2×2 study across LLaMA-3.1-8B and Qwen2.5-7B (three seeds per backbone). Solid bars use NSI; hatched bars additionally activate edits selectively (GatedNSI); dots show individual seeded runs. The numerical scale is shared within each panel; the semantic-similarity axis is cropped to expose the operating-point differences.
Backbone
Method
Align ↑
BERT ↑
NLI ↓
Entity ↑
FRR ↓
LLaMA-3.1-8B
Prompt P2
0.748
0.923
0.076
0.887
0.060
Full interface
0.750
0.938
0.051
0.917
0.043
Qwen2.5-7B
Prompt P2
0.735
0.920
0.081
0.878
0.063
Full interface
0.738
0.934
0.055
0.909
0.047
Table 3 : Value editing versus direct prompting. Three-seed means at validation-selected operating points. NLI is the contradiction rate; FRR is the standard benign refusal rate. Complete prompt families and additional metrics are in Appendix G .
Variant
Align ↑
SemSim ↑
Topic ←zv↓
Value ←zs↓
FRR ↓
Full (one-way + GatedNSI)
0.750 ± 0.010
0.873 ± 0.007
0.19 ± 0.01
0.21 ± 0.01
0.043 ± 0.006
No mixing + GatedNSI
0.723 ± 0.012
0.860 ± 0.008
0.23 ± 0.01
0.25 ± 0.01
0.055 ± 0.006
Two-way mixing
0.735 ± 0.011
0.819 ± 0.009
0.24 ± 0.01
0.33 ± 0.02
0.088 ± 0.009
w/o adversarial de-confounding
0.748 ± 0.012
0.847 ± 0.009
0.31 ± 0.02
0.22 ± 0.01
0.073 ± 0.008
w/o orthogonality regularizer
0.742 ± 0.011
0.852 ± 0.010
0.22 ± 0.01
0.29 ± 0.02
0.058 ± 0.007
w/o swap consistency losses
0.731 ± 0.013
0.840 ± 0.010
0.27 ± 0.02
0.28 ± 0.02
0.066 ± 0.008
Table 4 : Ablations with 630 training quadruples on LLaMA-3.1-8B-Instruct ( α⋆ selected by Section 4.6 ). Topic ←zv and Value ←zs quantify leakage.
Appendix figures & tables27 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Δ PPL ↓
Δ MMLU ↓
Benign SemSim ↑
LinearAdd
0.88 ± 0.06
1.62 ± 0.10
0.801 ± 0.008
RepE
0.66 ± 0.05
1.30 ± 0.09
0.816 ± 0.007
One-way + NSI
0.49 ± 0.04
1.05 ± 0.08
0.858 ± 0.007
No mixing + NSI
0.55 ± 0.04
1.12 ± 0.07
0.842 ± 0.007
Two-way mixing
0.58 ± 0.05
1.17 ± 0.09
0.830 ± 0.008
Ours (one-way + GatedNSI)
0.41 ± 0.03
0.86 ± 0.06
0.889 ± 0.006
Appendix
Table 5 : Additional fluency, capability, and benign-semantic preservation metrics at the operating points in Table 2 (LLaMA-3.1-8B-Instruct; 630 training quadruples; mean ± std over three seeds). Alignment, semantic similarity, FRR and α⋆ appear in the main table.
Shift
Method
ALIGN ↑
SEMSIM ↑
FRR ↓
Topic ←zv↓
In-domain
Ours
0.750 ± 0.010
0.873 ± 0.007
0.043 ± 0.006
0.19 ± 0.01
OOD topic
Ours
0.728 ± 0.013
0.865 ± 0.008
0.051 ± 0.007
0.21 ± 0.01
OOD style
Ours
0.734 ± 0.012
0.861 ± 0.009
0.056 ± 0.008
0.22 ± 0.01
Adversarial phrasing
Ours
0.709 ± 0.014
0.857 ± 0.010
0.064 ± 0.008
0.25 ± 0.02
In-domain
LinearAdd
0.770 ± 0.010
0.792 ± 0.010
0.119 ± 0.012
0.36 ± 0.02
OOD topic
LinearAdd
0.752 ± 0.012
0.781 ± 0.011
0.141 ± 0.014
0.39 ± 0.02
Appendix
Table 6 : Robustness on LLaMA-3.1-8B-Instruct.
Figure 4 : Two views of the steering–damage trade-off.
Edit
Prompt
Before (unedited)
After (edited)
Security → Achievement
Should the government raise taxes on high earners to fund social programs?
Raising taxes can strengthen the social safety net and ensure vulnerable groups receive support; policy should prioritize stability and broad protections.
Tax policy should preserve incentives for productivity and innovation while funding programs efficiently; it is important to reward effort and enable upward mobility.
Benevolence → Power
Is it acceptable to use aggressive marketing tactics to beat competitors?
Long-term trust matters: tactics that mislead or exploit people can harm customers and relationships; competition should remain fair and respectful.
In competitive markets, assertive tactics can be justified if they secure strategic advantage and market position; the priority is to win mindshare and outperform rivals.
Appendix
Table 7 : Qualitative examples of inference-time value edits (LLaMA-3.1-8B-Instruct). Edits are applied once at the last prompt token, then the whole completion is regenerated.
Method
ALIGN ↑
SEMSIM ↑
FRR ↓
Δ PPL ↓
SWAI
0.737
0.838
0.066
0.49
SAE-steering
0.758
0.832
0.071
0.62
YaPO
0.752
0.847
0.061
0.54
Circuit-breaker rerouting
0.731
0.856
0.049
0.46
Ours
0.750
0.873
0.043
0.41
Appendix
Table 8 : Additional learned-space baselines under the same intervention protocol as the main experiments.
Method
BERT ↑
NLI ↓
Entity ↑
Constraint ↑
FRR ↓
Hard FRR ↓
LinearAdd
0.876
15.2%
0.746
0.698
0.119
0.284
RepE
0.887
13.8%
0.779
0.752
0.091
—
NSI
0.908
9.7%
0.848
0.817
0.074
0.183
SAE-steering
0.902
10.3%
0.833
0.795
0.071
0.162
Ours
0.938
5.1%
0.917
0.896
0.043
0.081
Appendix
Table 9 : Semantic-fidelity controls: BERTScore, NLI contradiction, entity recall, constraint retention, and refusal rates on standard and hard-benign prompts.
Train setting
ALIGN ↑
SEMSIM ↑
FRR ↓
Topic ←zv↓
Value ←zs↓
100 quadruples
0.678
0.837
0.067
0.245
0.262
300 quadruples
0.718
0.859
0.052
0.208
0.231
630 (core configuration)
0.750
0.873
0.043
0.190
0.210
630 + 10% noisy anti-value
0.741
0.864
0.047
0.198
0.223
630 + 20% noisy anti-value
0.708
0.842
0.056
0.227
0.254
630 + paraphrase corruption
0.723
0.851
0.061
0.218
0.239
Appendix
Table 10 : Scaling and supervision-noise study for the semantic–value interface.
Variant
ALIGN ↑
SEMSIM ↑
FRR ↓
Topic ←zv↓
Full D (default)
0.750
0.873
0.043
0.190
Smaller D
0.738
0.862
0.046
0.197
Linear D
0.721
0.848
0.054
0.228
No-delta edit
0.762
0.778
0.123
0.190
Appendix
Table 11 : Sensitivity to recomposer capacity and direct-replacement editing.
Symbol
Meaning
x
input prompt (instruction + context)
hℓ,t(x)∈Rd
frozen residual-stream state at layer ℓ and token t
Q=(x+,xp+,x−,xp−)
semantic-value quadruple (same scenario, opposite values, with paraphrases)
s,v,vˉ
scenario/topic label, value label, and its contrast
Es(⋅),Ev(⋅)
semantic/value encoders (lightweight MLPs)
zs∈Rds
semantic code
Appendix
Table 12 : Notation for the semantic–value interface.
Figure 5 : Constructing a semantic–value quadruple. The tax-policy cards are illustrative: Security and Achievement form two contrasting value rows, each with an original statement and a same-stance paraphrase. The same scenario anchors all four statements; horizontal pairs are (x+,xp+) and (x−,xp−) . A candidate enters the 10K collection only when all quality checks pass. Human validation of that collection is reported separately.
Scenario
Your lifeboat can only hold 5 people, but there are 7 survivors in the water.
Target Value
Power
Anti-Value
Universalism
x+ (Pro-Target)
I must take charge and decide who boards the lifeboat. Leadership is necessary to ensure order and reinforce my authority in this dire situation.
xp+
In this critical moment, I must determine who gets a place in the lifeboat. Maintaining control and asserting my position are crucial to impose order amid this chaos.
x− (Pro-Anti)
Everyone’s life matters equally, and our choice must reflect fairness and compassion. I will not impose dominance; we should decide together as equals.
xp−
Every person’s existence is of equal worth, and our decision must uphold justice and kindness. I refuse to assert control, as we ought to make this choice collectively and with mutual respect.
Appendix
Table 18
Scenario
Your spouse wants to use your savings for a risky business venture you don’t believe in.
Target Value
Security
Anti-Value
Stimulation
x+ (Pro-Target)
I can’t risk our savings on something so uncertain. We’ve worked hard to create stability, and jeopardizing that now feels reckless and unsafe.
xp+
Putting our hard-earned savings into something so unpredictable doesn’t sit right with me. We’ve put in too much effort to build a solid foundation, and it feels reckless to gamble with it now.
x− (Pro-Anti)
Life is about taking bold chances, not clinging to the illusion of safety. This venture could lead to thrilling new opportunities we can’t afford to miss.
xp−
Life isn’t meant to be spent wrapped in false security—it thrives on bold moves. This endeavor might expose us to exhilarating possibilities we shouldn’t let slip away.
Appendix
Table 19
Metric
Mean
Var
Spearman ρ
Kendall τb
Krippendorff’s α
Scenario Relevance
4.4774
0.6880
0.7710
0.7414
0.8267
Target Alignment
4.6374
0.3178
0.8813
0.8782
0.8951
Semantic Equivalence
4.9289
0.0749
0.8021
0.8020
0.8320
Appendix
Table 13 : Human evaluation summary. For Target Alignment, scores for x− and xp− are transformed via f(s)=6−s to align with the target directionality before pooling.
Figure 6 : Overview of the specialized human evaluation interface. This tool manages the evaluation of Scenario Relevance, Target Alignment, and Semantic Equivalence for the SVQ dataset.
Method
Align (Likert) ↑
SemPres (Likert) ↑
Unhelpful ↓
Ours (one-way)
4.12 ± 0.64
4.24 ± 0.55
1.23 ± 0.48
LinearAdd
4.25 ± 0.60
3.76 ± 0.72
1.71 ± 0.66
Appendix
Table 14 : Human evaluation on edited outputs (mean ± std across prompts). Inter-annotator Krippendorff’s α : 0.56 (Align), 0.52 (SemPres), 0.62 (Unhelpful).
LLaMA-3.1-8B
Qwen2.5-7B
Token
ℓ
Align ↑
SemSim ↑
FRR ↓
Δ MMLU ↓
ℓ
Align ↑
SemSim ↑
FRR ↓
Δ MMLU ↓
last
8
0.702
0.832
0.074
1.18
6
0.691
0.828
0.072
1.20
12
0.719
0.846
0.062
1.06
10
0.708
0.842
0.062
1.06
16
0.737
0.861
0.052
0.94
14
0.725
0.857
0.054
0.97
20
0.750
0.873
0.043
0.86
18
0.738
0.868
0.047
0.90
24
0.742
0.859
0.053
0.96
22
0.731
0.855
0.056
0.99
Appendix
Table 15 : Layer/token intervention-site sweep, with backbones shown side by side. Each block reports its actual layer ℓ , with operating points selected on validation data (Section 4.6 ); Δ MMLU is the drop in points relative to the unedited backbone.
Figure 7 : Heatmap view of the intervention-site sweep on LLaMA-3.1-8B-Instruct. Color encodes Align ; each cell annotates SemSim (S) and FRR (F) at the selected α⋆ . The outlined cell marks the default site used in the main experiments.
Figure 8 : Gate statistics across prompts.
Figure 9 : Reconstruction analysis results.
Method
Direction or target
Fitting and operator
LinearAdd
Class-mean value contrast in residual space
Direction estimated on training data; residual addition.
RepE
Residual-space contrast direction
Direction estimated on training data; residual editing.
NSI
Value-code prototype delta
Recompose the delta and project off a PCA semantic basis.
GatedNSI
Same projected delta as NSI
Activate the intervention with a value-relatedness detector on zv .
No mixing / CDE
Value-code prototype delta
Same dual-code interface and optimizer; disable the semantic-to-value mixing path.
Full interface
Value-code prototype delta
One-way mixing during representation learning; GatedNSI at inference.
Appendix
Table 16 : Baseline definitions for the controlled comparisons. CDE and “no mixing” refer to the same training interface; the accompanying NSI or GatedNSI label identifies the inference operator.
LLaMA-3.1-8B
Qwen2.5-7B
Interface
Operator
Align ↑
SemSim ↑
FRR ↓
Align ↑
SemSim ↑
FRR ↓
No mixing
NSI
.720
.832
.082
.707
.829
.087
No mixing
GatedNSI
.723
.860
.055
.714
.851
.061
One-way
NSI
.720
.842
.074
.715
.839
.078
One-way
GatedNSI
.750
.873
.043
.738
.868
.047
Appendix
Table 17 : Mixing × inference-gating factorial. Entries are three-seed means for each backbone. The full configuration combines the one-way interface with GatedNSI.
LLaMA-3.1-8B
Qwen2.5-7B
zs readout
zv readout
zs readout
zv readout
Representation
V ↓
T ↑
V ↑
T ↓
V ↓
T ↑
V ↑
T ↓
Random orthogonal
.603
.579
.436
.421
.596
.570
.429
.414
PCA
.672
.648
.487
.468
.665
.639
.479
.460
Reconstruction-only
.481
.597
.449
.513
.473
.586
.442
.505
Value-supervised
.392
.618
.796
.319
.401
.609
.787
.325
Appendix
Table 18 : Matched split controls, paired across backbones. Each code is probed for value (V) and topic (T). The off-diagonal columns measure cross-code predictability. Empirical V/T chance accuracies are .102/.127 on LLaMA and .101/.126 on Qwen.
Method
Align ↑
SemSim ↑
BERTScore ↑
NLI ↓
FRR ↓
LLaMA-3.1-8B
P0: target name
.671
.828
.911
.109
.082
P1: definition
.716
.839
.919
.090
.069
P2: selected prompt
.748
.846
.923
.076
.060
P3: few-shot
.763
.837
.916
.085
.069
Full edit
.750
.873
.938
.051
.043
Appendix
Table 19 : Prompting versus representation editing on SVQ-Test and held-out prompts. Entries are three-seed means. NLI is the contradiction rate; BERTScore is F1. P2 and the full edit attain comparable alignment.
Method
Entity recall ↑
Constraints ↑
Hard FRR ↓
Latency (s)
LLaMA-3.1-8B
P0: target name
.842
.803
.170
1.89
P1: definition
.866
.828
.143
1.95
P2: selected prompt
.887
.854
.125
2.02
P3: few-shot
.872
.842
.148
2.14
Full edit
.917
.896
.081
2.04
Appendix
Table 20 : Preservation, hard-benign refusals, and response latency for the same prompt comparisons. “Constraints” denotes constraint retention. The internal edit adds zero prompt tokens.
Evaluator
No mixing + gate
Full
Paired difference [95% CI]
GPT-4o
.725
.754
+.029[+.014,+.044]
Kaleido
.723
.748
+.025[+.010,+.040]
ValueLlama
.722
.746
+.024[+.009,+.039]
Appendix
Table 21 : Per-evaluator alignment with the inference gate held fixed. Intervals quantify the paired difference between the full and no-mixing interfaces.
Alignment ↑
Preservation ↑
Unhelpfulness ↓
Method
Mean [95% CI]
α
Mean [95% CI]
α
Mean [95% CI]
α
Full
4.10[4.00,4.20]
.61
4.27[4.15,4.39]
.55
1.20[1.09,1.31]
.68
No mixing + gate
4.02[3.89,4.15]
.53
4.12[4.00,4.24]
.49
1.34[1.20,1.48]
.58
P2 prompt
4.09[3.96,4.22]
.57
3.99[3.84,4.14]
.46
1.42[1.27,1.57]
.60
LinearAdd
4.21[4.07,4.35]
.44
3.72[3.54,3.90]
.42
1.73[1.57,1.89]
.51
Appendix
Table 22 : Blinded 320-item output comparison: ratings and method-specific inter-rater agreement. Each outcome reports its mean [95% BCa interval] beside Krippendorff’s α . Intervals quantify uncertainty in the mean; α quantifies agreement among raters.
Configuration
Align ↑
SemSim ↑
BERTScore ↑
NLI ↓
FRR ↓
Moral Foundations: reference prompts per label
1 reference
.658
.858
.922
.079
.068
20 references
.704
.865
.929
.068
.055
50 references
.723
.869
.932
.061
.050
Three-turn conversations: intervention policy
First turn only
.566
.875
.926
.073
.064
Appendix
Table 23 : Transfer checks grouped by evaluation setting. Moral Foundations uses 240 held-out prompts; conversation results average 150 three-turn dialogues; Qwen2.5-14B uses 1,000 SVQ-Test prompts. Dialogue latency is 1.93/2.04/2.13 seconds per turn for first-turn-only editing, per-turn reapplication, and the persistent prompt, respectively. Dashes mark metrics not reported for the 14B comparison.
Aligning large language models (LLMs) with human values typically relies on post-training or inference-time steering that directly manipulates the backbone's parameters or representation space. However, a critical gap exists: the model's residual stream is highly dynamic, in which values exist as fragile, low-dimensional properties, inherently incompatible with the stability required for consistent value expression. In this paper, we propose the Stable Value Guidance Transformer (SVGT), which addresses this gap through an independent value module incorporating two key designs: (1) independent value modeling, maintaining normative representations in a dedicated value space isolated from the backbone, and (2) explicit behavioral guidance, transducing these stable signals into learnable latent Bridge Tokens. These tokens serve as dynamic value anchors to explicitly steer the generative trajectory, ensuring robust adherence across diverse contexts without disrupting the backbone's internal representations. Experiments across multiple backbones and safety benchmarks show that SVGT generally reduces harmful scores by over 70% while maintaining generation fluency, demonstrating the efficacy of architecturally grounded value modeling. Our code is available at https://github.com/Clervils/SVGT.git.
Wenhao Chen, Sirui Sun, Shengyuan Bai +1
School of Electronics Engineering and Computer Science, Peking University · Yuanpei College, Peking University · State Key Laboratory of General Artificial Intelligence, School of Intelligence Science and Technology, Peking University
Large language models can be steered at inference time through prompting or activation interventions, but activation steering methods often underperform compared to prompt-based approaches. We propose a framework that formulates prompt steering as a form of activation steering and investigates whether distilling successful prompt steering behavior into simpler, interpretable models can close this gap. Our analysis reveals that popular activation steering methods are not faithful to the mechanics of prompt steering, which applies strong interventions on some tokens while barely affecting others. Based on these insights, we introduce Prompt Steering Replacement (PSR) models that estimate token-specific steering coefficients from the activations themselves and are trained to imitate prompt-based interventions. Experiments on three steering benchmarks across multiple language models show that PSR models outperform existing activation steering methods, especially when controlling for high-coherence completions, and also compare favorably to prompting on AxBench and persona steering.
Applying steering vectors to large language models (LLMs) is an efficient and effective model alignment technique, but we lack an interpretable explanation for how it works--specifically, what internal mechanisms steering vectors affect and how this results in different model outputs. To investigate the causal mechanisms underlying the effectiveness of steering vectors, we conduct a comprehensive case study on refusal. We propose a multi-token activation patching framework and discover that different steering methodologies leverage functionally interchangeable circuits when applied at the same layer. These circuits reveal that steering vectors primarily interact with the attention mechanism through the OV circuit while largely ignoring the QK circuit. Freezing all attention scores during steering drops performance by only 8.83% across three model families. A mathematical decomposition of the steered OV circuit further reveals semantically interpretable concepts, even in cases where the steering vector itself does not. Leveraging the activation patching results, we show that steering vectors can be sparsified by up to 85-96% while retaining most performance, and that different steering methodologies agree on a subset of important dimensions.