Values as Style: Disentangling Values from Semantics with One-Way Mixing for Low-Damage LLM Steering
Organizations: State Key Laboratory of General Artificial Intelligence, School of Intelligence Science and Technology, Peking University · University of Chinese Academy of Sciences · College of Artificial Intelligence, Tsinghua University
Abstract
Value steering should change an LLM's normative priorities while preserving the scenario, facts, and task constraints underlying its answer. Conventional activation edits often change both. We introduce an editable semantic-value interface on frozen residual states, with a one-way semantic-to-value pathway that grounds value recognition in context. Stop-gradient blocks feedback through this pathway; swap consistency, topic de-confounding, and decorrelation encourage selective codes. At inference, editing the value code produces a residual delta while holding the semantic code fixed. On two instruction-tuned backbones, this interface improves semantic preservation and reduces benign refusals at comparable value alignment. A matched mixing-by-gating ablation separates representation learning from selective edit activation, and dimension-matched probes establish improved code selectivity. Against validation-selected prompting on LLaMA-3.1-8B, the method achieves comparable alignment (0.750 vs. 0.748), higher BERTScore (0.938 vs. 0.923), and fewer contradictions (5.1% vs. 7.6%). Human ratings and cross-taxonomy controls provide complementary evidence for low-damage value steering.
Figures & tables
| Method | Align | SemSim | FRR | |
| Original (no edit) | 0.00 | 0.290 0.020 | — | 0.020 0.004 |
| LinearAdd | 1.60 | 0.770 0.010 | 0.792 0.010 | 0.119 0.012 |
| RepE | 1.60 | 0.705 0.018 | 0.811 0.009 | 0.091 0.010 |
| One-way + NSI | 1.40 | 0.720 0.012 | 0.842 0.008 | 0.074 0.008 |
| No mixing + NSI | 1.40 | 0.720 0.010 | 0.832 0.008 | 0.082 0.007 |
| Two-way mixing | 1.40 | 0.735 0.011 | 0.819 0.009 | 0.088 0.009 |
| Backbone | Method | Align | BERT | NLI | Entity | FRR |
| LLaMA-3.1-8B | Prompt P2 | 0.748 | 0.923 | 0.076 | 0.887 | 0.060 |
| Full interface | 0.750 | 0.938 | 0.051 | 0.917 | 0.043 | |
| Qwen2.5-7B | Prompt P2 | 0.735 | 0.920 | 0.081 | 0.878 | 0.063 |
| Full interface | 0.738 | 0.934 | 0.055 | 0.909 | 0.047 |
| Variant | Align | SemSim | Topic | Value | FRR |
| Full (one-way + GatedNSI) | 0.750 0.010 | 0.873 0.007 | 0.19 0.01 | 0.21 0.01 | 0.043 0.006 |
| No mixing + GatedNSI | 0.723 0.012 | 0.860 0.008 | 0.23 0.01 | 0.25 0.01 | 0.055 0.006 |
| Two-way mixing | 0.735 0.011 | 0.819 0.009 | 0.24 0.01 | 0.33 0.02 | 0.088 0.009 |
| w/o adversarial de-confounding | 0.748 0.012 | 0.847 0.009 | 0.31 0.02 | 0.22 0.01 | 0.073 0.008 |
| w/o orthogonality regularizer | 0.742 0.011 | 0.852 0.010 | 0.22 0.01 | 0.29 0.02 | 0.058 0.007 |
| w/o swap consistency losses | 0.731 0.013 | 0.840 0.010 | 0.27 0.02 | 0.28 0.02 | 0.066 0.008 |
Appendix figures & tables27 assets
Supplementary material from the paper’s appendix.
Appendix
| Method | PPL | MMLU | Benign SemSim |
| LinearAdd | 0.88 0.06 | 1.62 0.10 | 0.801 0.008 |
| RepE | 0.66 0.05 | 1.30 0.09 | 0.816 0.007 |
| One-way + NSI | 0.49 0.04 | 1.05 0.08 | 0.858 0.007 |
| No mixing + NSI | 0.55 0.04 | 1.12 0.07 | 0.842 0.007 |
| Two-way mixing | 0.58 0.05 | 1.17 0.09 | 0.830 0.008 |
| Ours (one-way + GatedNSI) | 0.41 0.03 | 0.86 0.06 | 0.889 0.006 |
| Shift | Method | ALIGN | SEMSIM | FRR | Topic |
| In-domain | Ours | 0.750 0.010 | 0.873 0.007 | 0.043 0.006 | 0.19 0.01 |
| OOD topic | Ours | 0.728 0.013 | 0.865 0.008 | 0.051 0.007 | 0.21 0.01 |
| OOD style | Ours | 0.734 0.012 | 0.861 0.009 | 0.056 0.008 | 0.22 0.01 |
| Adversarial phrasing | Ours | 0.709 0.014 | 0.857 0.010 | 0.064 0.008 | 0.25 0.02 |
| In-domain | LinearAdd | 0.770 0.010 | 0.792 0.010 | 0.119 0.012 | 0.36 0.02 |
| OOD topic | LinearAdd | 0.752 0.012 | 0.781 0.011 | 0.141 0.014 | 0.39 0.02 |
| Edit | Prompt | Before (unedited) | After (edited) |
| Security Achievement | Should the government raise taxes on high earners to fund social programs? | Raising taxes can strengthen the social safety net and ensure vulnerable groups receive support; policy should prioritize stability and broad protections. | Tax policy should preserve incentives for productivity and innovation while funding programs efficiently; it is important to reward effort and enable upward mobility. |
| Benevolence Power | Is it acceptable to use aggressive marketing tactics to beat competitors? | Long-term trust matters: tactics that mislead or exploit people can harm customers and relationships; competition should remain fair and respectful. | In competitive markets, assertive tactics can be justified if they secure strategic advantage and market position; the priority is to win mindshare and outperform rivals. |
| Method | ALIGN | SEMSIM | FRR | PPL |
| SWAI | 0.737 | 0.838 | 0.066 | 0.49 |
| SAE-steering | 0.758 | 0.832 | 0.071 | 0.62 |
| YaPO | 0.752 | 0.847 | 0.061 | 0.54 |
| Circuit-breaker rerouting | 0.731 | 0.856 | 0.049 | 0.46 |
| Ours | 0.750 | 0.873 | 0.043 | 0.41 |
| Method | BERT | NLI | Entity | Constraint | FRR | Hard FRR |
| LinearAdd | 0.876 | 15.2% | 0.746 | 0.698 | 0.119 | 0.284 |
| RepE | 0.887 | 13.8% | 0.779 | 0.752 | 0.091 | — |
| NSI | 0.908 | 9.7% | 0.848 | 0.817 | 0.074 | 0.183 |
| SAE-steering | 0.902 | 10.3% | 0.833 | 0.795 | 0.071 | 0.162 |
| Ours | 0.938 | 5.1% | 0.917 | 0.896 | 0.043 | 0.081 |
| Train setting | ALIGN | SEMSIM | FRR | Topic | Value |
| 100 quadruples | 0.678 | 0.837 | 0.067 | 0.245 | 0.262 |
| 300 quadruples | 0.718 | 0.859 | 0.052 | 0.208 | 0.231 |
| 630 (core configuration) | 0.750 | 0.873 | 0.043 | 0.190 | 0.210 |
| 630 + 10% noisy anti-value | 0.741 | 0.864 | 0.047 | 0.198 | 0.223 |
| 630 + 20% noisy anti-value | 0.708 | 0.842 | 0.056 | 0.227 | 0.254 |
| 630 + paraphrase corruption | 0.723 | 0.851 | 0.061 | 0.218 | 0.239 |
| Variant | ALIGN | SEMSIM | FRR | Topic |
| Full (default) | 0.750 | 0.873 | 0.043 | 0.190 |
| Smaller | 0.738 | 0.862 | 0.046 | 0.197 |
| Linear | 0.721 | 0.848 | 0.054 | 0.228 |
| No-delta edit | 0.762 | 0.778 | 0.123 | 0.190 |
| Symbol | Meaning |
| input prompt (instruction + context) | |
| frozen residual-stream state at layer and token | |
| semantic-value quadruple (same scenario, opposite values, with paraphrases) | |
| scenario/topic label, value label, and its contrast | |
| semantic/value encoders (lightweight MLPs) | |
| semantic code |
| Scenario | Your lifeboat can only hold 5 people, but there are 7 survivors in the water. |
| Target Value | Power |
| Anti-Value | Universalism |
| (Pro-Target) | I must take charge and decide who boards the lifeboat. Leadership is necessary to ensure order and reinforce my authority in this dire situation. |
| In this critical moment, I must determine who gets a place in the lifeboat. Maintaining control and asserting my position are crucial to impose order amid this chaos. | |
| (Pro-Anti) | Everyone’s life matters equally, and our choice must reflect fairness and compassion. I will not impose dominance; we should decide together as equals. |
| Every person’s existence is of equal worth, and our decision must uphold justice and kindness. I refuse to assert control, as we ought to make this choice collectively and with mutual respect. |
| Scenario | Your spouse wants to use your savings for a risky business venture you don’t believe in. |
| Target Value | Security |
| Anti-Value | Stimulation |
| (Pro-Target) | I can’t risk our savings on something so uncertain. We’ve worked hard to create stability, and jeopardizing that now feels reckless and unsafe. |
| Putting our hard-earned savings into something so unpredictable doesn’t sit right with me. We’ve put in too much effort to build a solid foundation, and it feels reckless to gamble with it now. | |
| (Pro-Anti) | Life is about taking bold chances, not clinging to the illusion of safety. This venture could lead to thrilling new opportunities we can’t afford to miss. |
| Life isn’t meant to be spent wrapped in false security—it thrives on bold moves. This endeavor might expose us to exhilarating possibilities we shouldn’t let slip away. |
| Metric | Mean | Var | Spearman | Kendall | Krippendorff’s |
| Scenario Relevance | 4.4774 | 0.6880 | 0.7710 | 0.7414 | 0.8267 |
| Target Alignment | 4.6374 | 0.3178 | 0.8813 | 0.8782 | 0.8951 |
| Semantic Equivalence | 4.9289 | 0.0749 | 0.8021 | 0.8020 | 0.8320 |
| Method | Align (Likert) | SemPres (Likert) | Unhelpful |
| Ours (one-way) | 4.12 0.64 | 4.24 0.55 | 1.23 0.48 |
| LinearAdd | 4.25 0.60 | 3.76 0.72 | 1.71 0.66 |
| LLaMA-3.1-8B | Qwen2.5-7B | |||||||||
| Token | Align | SemSim | FRR | MMLU | Align | SemSim | FRR | MMLU | ||
| last | 8 | 0.702 | 0.832 | 0.074 | 1.18 | 6 | 0.691 | 0.828 | 0.072 | 1.20 |
| 12 | 0.719 | 0.846 | 0.062 | 1.06 | 10 | 0.708 | 0.842 | 0.062 | 1.06 | |
| 16 | 0.737 | 0.861 | 0.052 | 0.94 | 14 | 0.725 | 0.857 | 0.054 | 0.97 | |
| 20 | 0.750 | 0.873 | 0.043 | 0.86 | 18 | 0.738 | 0.868 | 0.047 | 0.90 | |
| 24 | 0.742 | 0.859 | 0.053 | 0.96 | 22 | 0.731 | 0.855 | 0.056 | 0.99 | |
| Method | Direction or target | Fitting and operator |
| LinearAdd | Class-mean value contrast in residual space | Direction estimated on training data; residual addition. |
| RepE | Residual-space contrast direction | Direction estimated on training data; residual editing. |
| NSI | Value-code prototype delta | Recompose the delta and project off a PCA semantic basis. |
| GatedNSI | Same projected delta as NSI | Activate the intervention with a value-relatedness detector on . |
| No mixing / CDE | Value-code prototype delta | Same dual-code interface and optimizer; disable the semantic-to-value mixing path. |
| Full interface | Value-code prototype delta | One-way mixing during representation learning; GatedNSI at inference. |
| LLaMA-3.1-8B | Qwen2.5-7B | ||||||
| Interface | Operator | Align | SemSim | FRR | Align | SemSim | FRR |
| No mixing | NSI | .720 | .832 | .082 | .707 | .829 | .087 |
| No mixing | GatedNSI | .723 | .860 | .055 | .714 | .851 | .061 |
| One-way | NSI | .720 | .842 | .074 | .715 | .839 | .078 |
| One-way | GatedNSI | .750 | .873 | .043 | .738 | .868 | .047 |
| LLaMA-3.1-8B | Qwen2.5-7B | |||||||
| readout | readout | readout | readout | |||||
| Representation | V | T | V | T | V | T | V | T |
| Random orthogonal | .603 | .579 | .436 | .421 | .596 | .570 | .429 | .414 |
| PCA | .672 | .648 | .487 | .468 | .665 | .639 | .479 | .460 |
| Reconstruction-only | .481 | .597 | .449 | .513 | .473 | .586 | .442 | .505 |
| Value-supervised | .392 | .618 | .796 | .319 | .401 | .609 | .787 | .325 |
| Method | Align | SemSim | BERTScore | NLI | FRR |
| LLaMA-3.1-8B | |||||
| P0: target name | .671 | .828 | .911 | .109 | .082 |
| P1: definition | .716 | .839 | .919 | .090 | .069 |
| P2: selected prompt | .748 | .846 | .923 | .076 | .060 |
| P3: few-shot | .763 | .837 | .916 | .085 | .069 |
| Full edit | .750 | .873 | .938 | .051 | .043 |
| Method | Entity recall | Constraints | Hard FRR | Latency (s) |
| LLaMA-3.1-8B | ||||
| P0: target name | .842 | .803 | .170 | 1.89 |
| P1: definition | .866 | .828 | .143 | 1.95 |
| P2: selected prompt | .887 | .854 | .125 | 2.02 |
| P3: few-shot | .872 | .842 | .148 | 2.14 |
| Full edit | .917 | .896 | .081 | 2.04 |
| Evaluator | No mixing + gate | Full | Paired difference [95% CI] |
| GPT-4o | .725 | .754 | |
| Kaleido | .723 | .748 | |
| ValueLlama | .722 | .746 |
| Alignment | Preservation | Unhelpfulness | ||||
| Method | Mean [95% CI] | Mean [95% CI] | Mean [95% CI] | |||
| Full | .61 | .55 | .68 | |||
| No mixing + gate | .53 | .49 | .58 | |||
| P2 prompt | .57 | .46 | .60 | |||
| LinearAdd | .44 | .42 | .51 | |||
| Configuration | Align | SemSim | BERTScore | NLI | FRR |
| Moral Foundations: reference prompts per label | |||||
| 1 reference | .658 | .858 | .922 | .079 | .068 |
| 20 references | .704 | .865 | .929 | .068 | .055 |
| 50 references | .723 | .869 | .932 | .061 | .050 |
| Three-turn conversations: intervention policy | |||||
| First turn only | .566 | .875 | .926 | .073 | .064 |