cs.AISep 30, 2026

Values as Style: Disentangling Values from Semantics with One-Way Mixing for Low-Damage LLM Steering

Authors: Jiale Dai, Hongcan Deng, Liuxian Ma, Xiaoke Niu, Guojie Song

Organizations: State Key Laboratory of General Artificial Intelligence, School of Intelligence Science and Technology, Peking University · University of Chinese Academy of Sciences · College of Artificial Intelligence, Tsinghua University

Abstract

Value steering should change an LLM's normative priorities while preserving the scenario, facts, and task constraints underlying its answer. Conventional activation edits often change both. We introduce an editable semantic-value interface on frozen residual states, with a one-way semantic-to-value pathway that grounds value recognition in context. Stop-gradient blocks feedback through this pathway; swap consistency, topic de-confounding, and decorrelation encourage selective codes. At inference, editing the value code produces a residual delta while holding the semantic code fixed. On two instruction-tuned backbones, this interface improves semantic preservation and reduces benign refusals at comparable value alignment. A matched mixing-by-gating ablation separates representation learning from selective edit activation, and dimension-matched probes establish improved code selectivity. Against validation-selected prompting on LLaMA-3.1-8B, the method achieves comparable alignment (0.750 vs. 0.748), higher BERTScore (0.938 vs. 0.923), and fewer contradictions (5.1% vs. 7.6%). Human ratings and cross-taxonomy controls provide complementary evidence for low-damage value steering.

Figures & tables

Appendix figures & tables27 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Toward Stable Value Alignment: Introducing Independent Modules for Consistent Value Guidance

    May 12, 2026Wenhao Chen, Sirui Sun, Shengyuan Bai +1Large Language Model AlignmentLarge Language Model Safety

  2. What Drives Representation Steering? A Mechanistic Case Study on Steering Refusal

    Apr 9, 2026Stephen Cheng, Sarah Wiegreffe, Dinesh ManochaLinear Activation SteeringSteering