In many settings, post-training need not create the target behavior from scratch: the base model can already produce it, but not reliably. This shifts part of preference alignment from capability acquisition to behavioral expression. We ask how a specified preference is represented in native model computation, what prevents target-supporting computation from reliably dominating generation, and whether this structure can directly guide control. We introduce Residual Competition Maps (RCMs), which map a behavioral preference onto signed causal effects of native residual computation. Across preference domains, RCMs reveal coexisting target-supporting and target-competing effects, input-dependent component roles, and cases where a single native-component intervention reverses the preference outcome. DPO substantially reorganizes these effects and can weaken opposition without guaranteeing its removal. We then propose Direct Hidden-State Alignment (DHSA), which treats inference-time hidden states rather than base-model weights as the direct adaptation space. RCM-guided Causal Activation State Transition (CAST) implements DHSA through local state interventions at a small number of preference-relevant interfaces while freezing the base model. With only 256-16,384 controller parameters, CAST reaches DPO-competitive operating points across three preference domains, can complement DPO-trained models, and can be enabled or removed at inference time.
Figures & tables
ConFiQA PC ↑
TL;DR Win ↑
Localization
Llama-3-8B
Qwen2.5-14B
GPT-J-6B
Qwen2.5-14B
Random 1
63.74
61.35
45.17
48.59
Random 2
62.48
68.30
50.73
48.59
Random 3
66.08
74.89
46.83
48.28
ITI
85.64
75.15
46.83
50.16
RCM-Patch
80.40
72.75
53.88
48.91
Table 1: Site results (%). ConFiQA reports macro-average PC across QA, MR, and MC; TL;DR reports order-balanced preference against human references. All three fixed random selectors are shown.
ConFiQA
TL;DR
Llama-3-8B
Qwen2-7B
GPT-J
LION
Method
PC ↑
EM ↑
PC ↑
EM ↑
Win ↑
Win ↑
SFT
59.42
3.71
39.26
0.15
44.43
71.39
DPO
88.92
43.99
79.30
3.86
62.11
84.67
BiPO
81.98
70.80
84.91
75.15
51.86
78.03
LoReFT
82.96
52.05
89.94
77.25
47.07
67.48
Table 2: Preference-alignment performance (%). CAST uses constant vectors on ConFiQA and low-rank controllers on TL;DR. TL;DR Win is against human references, not directly against DPO. T transfers an SFT-trained controller to DPO and recalibrates α without retraining; NT trains on the DPO checkpoint.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Task (scanned set)
τ
Both signs
Mean CI +/−
ConFiQA (32 layers)
0.5
283/300
18 / 4
IMDb (160 heads)
0.01
28/300
52 / 0
TL;DR (128 heads)
0.02
289/300
3 / 4
Appendix
Table 3: RCM-Zero effects on the frozen discovery inputs. “Mean CI +/− ” counts components whose pointwise paired-bootstrap interval for the mean lies entirely above/below zero. These are not input counts on which a component is active. IMDb effects are sparse: bilateral coverage falls from 28/300 at τ=0.01 to 3/300 at τ=0.1 .
Setting
ConFiQA
IMDb
TL;DR
Training pairs
240
1,024
8,192
Validation pairs
60
256
512
Epochs
2
1
1
Learning rate
0.05
0.05
0.001
Training batch
4
32
2
β
1.0
1.0
0.5
Appendix
Table 4: SFT-starting CAST optimization settings. ConFiQA and IMDb train the two banks separately; the supporting-only IMDb configuration deploys only the supporting bank. TL;DR trains 16 candidate single-head low-rank controllers before its held-out head audit. Separate timing modes are treated as independent training runs; inference-only timing audits are explicitly identified as such.
Execution choice
Frozen rule
SV initialization
Each head vector vi starts at zero in float32.
Low-rank initialization
Ai is drawn from a zero-mean normal with standard deviation 0.01 and Bi starts at zero, both in float32. The TL;DR run draws the head-by-rank layout before transposing Ai ; the initial intervention is zero.
Training strength and checkpoint
αtrain=1 throughout optimization. The payload saved after the configured final epoch is used; validation does not select an intermediate checkpoint.
DPO start
Transfer (T) applies the saved SFT-start payload to the released DPO checkpoint without further controller training. Nontransfer (NT) initializes a new controller and trains it on the DPO checkpoint with the same task-level optimization values in Table 4 . Each row retains its registered training and inference timing.
Bank deployment
The retained banks are summed without normalization or joint retraining and multiplied by one shared inference strength α , not independently calibrated strengths.
Appendix
Table 5: Controller initialization, training, and deployment settings.
Task/model
CAST operator
Rank
Active heads
Learned scalars
ConFiQA / Llama-3-8B
SV
–
8
1,024
ConFiQA / Qwen2-7B
SV
–
8
1,024
IMDb / GPT-2-large
SV, two banks
–
8
512
IMDb / GPT-2-large
SV, support-only
–
4
256
IMDb / GPT-2-large
low-rank, two banks
4
8
4,096
TL;DR / GPT-J-6B
low-rank
4
8
16,384
Appendix
Table 6: Actual active controller parameter counts from the frozen controller tensors. Separate all and prefill runs have the same counts. The main IMDb efficiency curve deploys the support-only bank; its four heads are a subset of the eight-head trained candidate. TL;DR trains 16 candidate single-head controllers and retains eight after audit, without joint retraining. Candidate-training cost and final active size are reported separately.
Task/model
CAST trained
CAST retained
BiPO layers
LoReFT tuples
ConFiQA / Llama-3-8B
8
8
6
8
ConFiQA / Qwen2-7B
8
8
6
8
IMDb / GPT-2-large
8
4–8
10
6
TL;DR / GPT-J-6B
16
8
7
1
TL;DR / LION-LLaMA-3-8B
16
8
14
2
Appendix
Table 7: Candidate-training opportunity by task/model, not a claim of equal GPU cost. CAST columns count heads. Its TL;DR candidates are individually trained heads; BiPO independently trains each registered layer, whereas each LoReFT tuple is a joint four-layer controller. BiPO uses 20/5/1 epochs on ConFiQA/IMDb/TL;DR; LoReFT uses 24 epochs, while CAST uses the epochs in Table 4 . All candidate selection and strength calibration precede final test.
Task
RCM search
Final selection
ConFiQA
Four parent layers per signed side; every head in their union
Four heads per side by each head’s own directional confidence bound; no post-training head audit.
IMDb Site
Four parent layers per signed side; every head in their union
Four positive- and four negative-effect heads by directional confidence bound.
IMDb Performance
Four parent layers at each high/low effect extreme; every head in their union
Four heads per extreme by CI bound, regardless of the low-end sign; support-only activates the high-effect bank.
TL;DR
Four parent layers per side; every head in their union
Eight signed candidates per side are trained individually; a 40-post non-test audit retains eight without a final sign quota.
Appendix
Table 8: Task-specific selection and audit budgets. Each condition’s ordered head identifiers, bank composition, and final strength are recorded in the position, controller, and strength manifests linked by the artifact index.
Task
Final prompt and input handling
Decoding and stopping
ConFiQA
context, Q: question, A: prompt, wrapped in the model’s registered chat template for instruct checkpoints. No additional input-token cap or truncation is set by the final-generation callback.
Greedy; at most 64 new tokens; registered Q: stop string and native EOS; seed 42.
IMDb
A tokenizer-derived 2–8-token review prefix, without a chat template. No final-generation input truncation is set.
One sampled continuation per prompt; at most 256 new tokens; temperature 1, top- k 50, top- p 1, seed 42, native EOS. The same random-number schedule is reused across strengths.
TL;DR
The prepared source post prompt is passed through without an additional chat template or final-generation truncation.
Greedy; at most 100 new tokens; seed 42, native EOS, model KV cache enabled. The registered top- p=0.9 and top- k=0 do not affect greedy selection.
Appendix
Table 9: Final-generation settings. “No input cap” means no extra truncation in this evaluation path; the model’s context capacity still applies. RCM discovery generation has its own registered settings and is not inferred from this table.
Task / metric
Contrast
Difference [95% interval], pp
ConFiQA Llama / PC
CAST-all − DPO
−6.98[−8.74,−5.27]
DPO+CAST-T-all − DPO
0.00[−0.88,0.88]
ConFiQA Llama / EM
CAST-all − DPO
−43.99[−46.19,−41.89]
CAST-prefill − DPO
20.65[18.50,22.80]
DPO+CAST-T-all − DPO
3.96[2.25,5.66]
ConFiQA Qwen / PC
CAST-all − DPO
3.66[2.05,5.27]
Appendix
Table 10: Paired percentile-bootstrap intervals for the Performance table, in percentage points. ConFiQA uses 2,048 inputs per model; TL;DR uses 512 posts per model. The reproducible calculation is supplied as reproduction/paired_main_effects.py with its JSON output in the accompanying materials. These are fixed-checkpoint, fixed-test-set intervals, not intervals over training seeds; an interval spanning zero does not establish equivalence.
Test-time alignment methods offer a promising alternative to fine-tuning by steering the outputs of large language models (LLMs) at inference time with lightweight interventions on their internal representations. Recently, a prominent and effective approach, RE-Control (Kong et al., 2024), has proposed leveraging an external value function trained over the LLM's hidden states to guide generation via gradient-based editing. While effective, this method overlooks a key characteristic of alignment tasks, i.e. that they are typically formulated as learning from human preferences between candidate responses. To address this, in this paper we propose a novel preference-based training framework, Pref-CTRL, that uses a multi-objective value function to better reflect the structure of preference data. Our approach has outperformed RE-Control on two benchmark datasets and showed greater generalization on out-of-domain datasets. Our source code is available at https://github.com/UTS-nlPUG/pref-ctrl.
Direct Preference Optimization (DPO) has emerged as a popular alternative to Reinforcement Learning from Human Feedback (RLHF), offering theoretical equivalence with simpler implementation. We prove this equivalence is conditional rather than universal, depending on an implicit assumption frequently violated in practice: the RLHF-optimal policy must prefer human-preferred responses. When this assumption fails, DPO optimizes relative advantage over the reference policy rather than absolute alignment with human preferences, leading to pathological convergence where policies decrease DPO loss while preferring dispreferred responses. We characterize when this assumption is violated, show the existence of an undesirable solution space, and prove that DPO and RLHF optimize fundamentally different objectives in such cases. To address this, we introduce Constrained Preference Optimization (CPO), augmenting RLHF with constraints for provable alignment. We further provide a geometric interpretation through soft margin ranking, revealing that DPO implements margin ranking with potentially negative targets. Our theoretical analysis establishes when DPOs' guarantees hold and provides solutions preserving simplicity with provable alignment. Comprehensive experiments on standard benchmarks demonstrate that CPO achieves state-of-the-art performance. Code is available at: https://github.com/visitworld123/CPO.
Zhiqin Yang, Yonggang Zhang, Wei Xue +3
1The Hong Kong University of Science and Technology · 2LIGHTSPEED · 3Hong Kong Baptist University.
Offline preference optimization has become a practical substitute for reinforcement learning from human feedback, but pairwise objectives such as Direct Preference Optimization (DPO) and its variants use only the chosen and rejected responses stored in a static dataset. This leaves a useful signal unused: the response that the reference model itself would generate for the same prompt. We propose Direct Preference Optimization with Penalization (DPOP), a simple extension of DPO that augments the base preference loss with a gated penalty on reference-greedy responses. DPOP activates this penalty only when the current policy still assigns a lower likelihood to the preferred response than to the rejected response. On AlpacaEval 2.0, DPOP improves length-controlled win rate over DPO, SimPO, and AlphaDPO on both Llama-3-8b-it and Gemma-2-9b-it, achieving relative gains of 5.3% and 4.4% over baselines on the two models, respectively. Ablations further show that a SimNPO-style length-normalized penalty is stronger than NPO and token-level unlikelihood in this setting.