cs.LGSep 27, 2026

Direct Hidden-State Alignment: Mapping and Controlling Preference Expression in LLMs

Authors: Fansheng Zhang, Shengran Guo, Zexiao Wang, Liang Yuan, Jiyuan Chen, Ruikun Luo

Organizations: Chengdu University · North Carolina State University · Fudan University · Australian Catholic University · University of Macau

Abstract

In many settings, post-training need not create the target behavior from scratch: the base model can already produce it, but not reliably. This shifts part of preference alignment from capability acquisition to behavioral expression. We ask how a specified preference is represented in native model computation, what prevents target-supporting computation from reliably dominating generation, and whether this structure can directly guide control. We introduce Residual Competition Maps (RCMs), which map a behavioral preference onto signed causal effects of native residual computation. Across preference domains, RCMs reveal coexisting target-supporting and target-competing effects, input-dependent component roles, and cases where a single native-component intervention reverses the preference outcome. DPO substantially reorganizes these effects and can weaken opposition without guaranteeing its removal. We then propose Direct Hidden-State Alignment (DHSA), which treats inference-time hidden states rather than base-model weights as the direct adaptation space. RCM-guided Causal Activation State Transition (CAST) implements DHSA through local state interventions at a small number of preference-relevant interfaces while freezing the base model. With only 256-16,384 controller parameters, CAST reaches DPO-competitive operating points across three preference domains, can complement DPO-trained models, and can be enabled or removed at inference time.

Figures & tables

Appendix figures & tables8 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Pref-CTRL: Preference Driven LLM Alignment using Representation Editing

    Apr 26, 2026Imranul Ashrafi, Inigo Jauregi Unanue, Massimo PiccardiLarge Language Model AlignmentValue Alignment

  2. Conditional Equivalence of DPO and RLHF: Implicit Assumption, Failure Modes, and Provable Alignment

    May 20, 2026Zhiqin Yang, Yonggang Zhang, Wei Xue +3Backtranslation Augmented Direct Preference OptimizationReinforcement Learning From Human Feedback