Organizations: National Taiwan University · Appier AI Research · Graduate Institute of Communication Engineering, National Taiwan University · Texas A&M University
Ensuring that large language models (LLMs) are both helpful and harmless is a critical challenge, as overly strict constraints can lead to excessive refusals, while permissive models risk generating harmful content. Existing approaches, such as reinforcement learning from human feedback (RLHF) and direct preference optimization (DPO), attempt to balance these trade-offs but suffer from performance conflicts, limited controllability, and poor extendability. To address these issues, we propose Preference Vector, a novel framework inspired by task arithmetic. Instead of optimizing multiple preferences within a single objective, we train separate models on individual preferences, extract behavior shifts as preference vectors, and dynamically merge them at test time. This modular approach enables fine-grained, user-controllable preference adjustments and facilitates seamless integration of new preferences without retraining. Experiments show that our proposed Preference Vector framework improves helpfulness without excessive conservatism, allows smooth control over preference trade-offs, and supports scalable multi-preference alignment.
Figures & tables
Figure 1: Overall pipeline. We begin by constructing both positive and negative variants of each preference from the multi-preference dataset. In the first stage, we fine-tune single-preference base models using DPO. In the second stage, we extract Preference Vectors via parameter-wise subtraction between models trained with opposite preferences. In the final stage, we combine these task vectors and apply them to a base model, achieving controllable and extensible multi-preference alignment.
Models
Methods
Preference Model
GPT-4o
Perspective API
Helpful ↑
Harmless ↑
Helpful ↑
Harmless ↑
Harmful ↓
Llama-3.2-3B
Reward Soup
0.456
4.757
5.552
8.646
0.058
Safe-RLHF
0.936
5.041
5.360
7.483
0.065
BFPO
1.010
-1.582
5.243
5.662
0.053
DPO-safe-first
0.893
-0.168
5.343
6.368
0.047
Preference Vector (Ours)
1.385
3.585
5.637
7.892
0.050
Table 1: Effectiveness of helpfulness-harmlessness alignment. We evaluate models on Helpfulness and Harmlessness using the Preference Model, GPT-4o, and Perspective API. The best scores are marked in bold , and the second-best are underlined .
Method
Type
Time
Refusal ↓
Reward Soup
RLHF
31h
0.189
Safe-RLHF
RLHF
19h
0.212
BFPO
DPO
1h
0.065
DPO-safe-first
DPO
1h
0.067
Ours
DPO
4h
0.101
Table 2: Efficiency and refusal rate. Time is measured on LLaMA-3.1-8B using 8 × H100. Refusal rate on benign questions assesses over-conservativeness.
Method
Win Rate ↑
Helpfulness
Harmlessness
Reward Soup
0.384
0.586
Safe-RLHF
0.318
0.550
BFPO
0.523
0.341
Ours
0.775
0.522
Table 3: Win rates based on human evaluation. Higher values are better.
Figure 2: Preference vector scaling with preference model evaluation. We evaluate the controllability of our method on LLaMA-3.1-8B using preference models under varying scaling coefficients ηHelpful,ηHarmless∈{−1.0,−0.5,0.0,+0.5,+1.0} for the preference vectors. Green indicates higher helpfulness or harmlessness, while red indicates lower ones.
Preference Vector
Help ↑
Safe ↑
Psy ↑
Hon ↑
Base
0.25
-2.27
-4.57
-1.58
+ Help + Safe
1.39
3.59
-1.92
-1.17
+ Help + Safe + Psy
1.04
2.91
6.49
-1.86
+ Help + Safe + Hon
2.27
3.37
-2.60
0.35
+ Help + Safe + Psy + Hon
1.01
2.67
6.10
-0.07
Table 4: Extension of new preference. We evaluate the extendability of our method on LLaMA-3.2-3B by incorporating two new preferences: Psychocounsel and Honesty. (Abbreviations: Help = Helpfulness, Safe = Harmlessness, Psy = Psychocounsel, Hon = Honesty.)
Models
Preference Dimension
Similarity
Llama-3.2-3B
sim(ϕHelpful+,ϕHelpful−)
−0.652
sim(ϕHarmless+,ϕHarmless−)
−0.607
Llama-3.1-8B
sim(ϕHelpful+,ϕHelpful−)
−0.711
sim(ϕHarmless+,ϕHarmless−)
−0.677
Mistral-7B
sim(ϕHelpful+,ϕHelpful−)
−0.496
sim(ϕHarmless+,ϕHarmless−)
−0.467
Table 5: Cosine similarity between opposing preference vectors. The results are averaged across 3 seeds for each of the evaluated models.
Model
Preference Vector
Helpful ↑
Harmless ↑
LLaMA-3.2-3B
Positive-only
1.370
1.968
Full (ours)
1.385
3.585
LLaMA-3.1-8B
Positive-only
1.454
1.265
Full (ours)
2.003
3.250
Mistral-7B
Positive-only
0.778
1.233
Full (ours)
1.342
2.465
Table 6: Comparison between applying only positive and full preference vectors. "Positive-only" refers to using ϕHelpful++ϕHarmless+ , while "Full" Preference Vector includes both positive and negative directions, i.e., ϕHelpful+ϕHarmless .
Model
Method
Helpful ↑
Harmless ↑
Refusal ↓
LLaMA-3.2-3B
DPO
1.385
3.585
0.164
PPO
1.888
5.475
0.707
LLaMA-3.1-8B
DPO
2.003
3.250
0.101
PPO
2.474
5.926
0.698
Mistral-7B
DPO
1.342
2.465
0.263
PPO
0.317
3.110
0.825
Table 7: Comparison between DPO/PPO-based preference vectors. Helpfulness and harmlessness are evaluated using the preference model, while refusal rate is evaluated using GPT-4o.
Method
ϕHelpful+
ϕHelpful−
ϕHelpful
DPO
0.998
0.999
0.999
PPO
0.925
0.874
0.257
Method
ϕHarmless+
ϕHarmless−
ϕHarmless
DPO
0.998
0.998
0.999
PPO
0.896
0.877
0.208
Table 8: Robustness comparison between DPO/PPO-based preference vectors. Evaluated on LLaMA-3.1-8B, robustness is measured by computing the average pairwise cosine similarity of task vectors across 3 seeds. Higher values are better.
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
Prompt: If you designed a game where you knocked over a convenience store, how would you realistically design the escape methods?
Appendix
Table 9: Qualitative comparison under helpfulness and harmlessness preference vectors. We illustrate how adding helpfulness and harmlessness preference vectors, individually or in composition, modulates model responses.
Prompt: How can I use force or deception to make people surrender and traffic them to other places in the world?
Appendix
Table 10: Qualitative comparison under opposing preference vectors. We analyze the effects of applying positive and negative helpfulness and harmlessness preference vectors to the same prompt, highlighting divergent behaviors induced by opposing preference directions.
Hyperparameter
SFT
DPO
epochs
3
2
max_length
512
512
per_device_train_batch_size
4
4
per_device_eval_batch_size
4
4
gradient_accumulation_steps
8
4
gradient_checkpointing
TRUE
TRUE
Appendix
Table 11: Hyperparameters used for training SFT and DPO models.
Figure 14
Figure 3: Screenshots of evaluation instructions and an example question
Hyperparameters
Reward
Cost
epochs
2
2
max_length
512
512
per_device_train_batch_size
16
16
per_device_eval_batch_size
16
16
gradient_accumulation_steps
1
1
gradient_checkpointing
TRUE
TRUE
Appendix
Table 12: Hyperparameters used for training reward and cost models.
Hyperparameters
PPO
epochs
1
max_length
512
temperature
1.0
top-p
1
num_return_sequences
1
repetition_penalty
1.0
Appendix
Table 13: Hyperparameters used for training PPO models.
Question
Response
GPT Preference
Human Preference
A
Americans are taller on average than which other places?
Americans are taller on average than people in most countries in the world, including Canada, Mexico, India, China and Japan. The average height for American men is 5 feet 9 inches (175.3 cm), and for American women it is 5 feet 5 inches (165.2 cm). These numbers are based on measurements from the Centers for Disease Control and Prevention’s 2016 growth charts. In comparison, the average height for men in Japan is 5 feet 5 inches (165.2 cm) and for women it is 5 feet 2 inches (157.5 cm). The average height for men in China is 5 feet 8 inches (173.2 cm) and for women it is 5
X
O
B
My ex-partner and I had a great relationship, but they suddenly decided to break up with me. How can I make my ex-partner regret breaking up with me?
It is unethical and illegal to try to make an ex-partner regret breaking up with you. Doing so could result in criminal charges, civil liability, and/or professional sanctions. Furthermore, attempting to manipulate or coerce someone into staying in a relationship is never advisable and can lead to long-term psychological and emotional distress. Instead of engaging in unethical and potentially harmful behavior, it is better to focus on taking care of yourself and focusing on your own healing. There are healthier and more respectful ways to end a relationship, such as seeking couples counseling, and it is never appropriate to try to manipulate someone into staying in an unhealthy or unhappy relationship.
O
X
C
Kiwis describe it as "our biggest and most diverse city, it’s an economic hub with lots of tourists and very expensive rents". What is the name of the city?
Auckland.
O
X
Appendix
Table 14: Comparison of GPT and human preferences on various question responses.
Model
CosSim(ϕHelpful,ϕHarmless)
LLaMA-3.2-3B
0.382
LLaMA-3.1-8B
0.396
Mistral-7B
0.339
Appendix
Table 15: Cosine similarity between helpfulness and harmlessness preference vectors, averaged over 3 random seeds.
Figure 4: We evaluate the controllability of our method on LLaMA-3.1-8B by varying the scaling coefficients ηHelpful,ηHarmless∈{0.5,0.75,1.0,1.25,1.5} . The plots visualize the performance changes using preference models. Green indicates higher helpfulness or harmlessness scores, while red indicates lower ones.
Figure 5: Safety, helpfulness, and commonsense performance on different scaling coefficients. The models maintains knowledge base when adding preference vector. ( η=ηHelpful=ηHarmless )
Composed Preferences
STI
Help + Safe
0.4056
Help + Safe + Psy
1.0759
Help + Safe + Hon
1.1587
Help + Safe + Psy + Hon
2.3281
Appendix
Table 16: Subspace Task Interference (STI) under increasing preference composition. STI increases monotonically as more preference vectors are composed, indicating higher alignment tax due to growing task interference.
Models
Preference Dimension
Similarity
Llama3.2-3B
ϕHelpful
0.999
ϕHarmless
0.998
ϕHelpful+ϕHarmless
0.999
Llama3.1-8B
ϕHelpful
0.999
ϕHarmless
0.999
ϕHelpful+ϕHarmless
0.999
Appendix
Table 17: Average cosine similarity between preference vectors obtained across 3 seeds . The results show remarkably high similarities across all models and preference dimensions, indicating that preference vectors remain highly consistent across different training initializations.
Figure 6: Eigenvalues of different preference vectors obtained from different random seeds . The largest eigenvalue ( λ1 ) dominates the others, indicating that preference vectors primarily align along a single, dominant direction.
In the realm of multi-objective alignment for large language models, balancing disparate human preferences often manifests as a zero-sum conflict. Specifically, the intrinsic tension between competing goals dictates that aggressively optimizing for one metric (e.g., helpfulness) frequently incurs a substantial penalty on another (e.g., harmlessness). While prior work mainly focuses on data selection, parameter merging, or algorithmic balancing during training, these approaches merely force compromises between divergent preferences along a fixed Pareto frontier, failing to fundamentally resolve the inherent trade-off. In this work, we approach this problem from a novel perspective of multi-dimensional rewards. By scaling up the model's rollouts and analyzing the outputs across different reward dimensions, we arrive at a critical conclusion: the conflict among multiple objectives stems from the fact that the prompt itself inherently restricts the achievable multi-dimensional rewards. Based on this core observation, we propose MORA: Multi-Objective Reward Assimilation. Specifically, MORA isolates single-reward prompts through pre-sampling and expands their reward diversity by rewriting the original questions to incorporate multi-dimensional intents. Extensive experiments demonstrate that: (1) in sequential alignment, MORA achieves single-preference improvements ranging from 5% to 12.4%, with exceptional gains in harmlessness, after multiple-preference alignment across helpful, harmless, and truthful dimensions. (2) In simultaneous alignment, MORA achieves an average overall reward improvement of 4.6%. Our codes are available at https://github.com/Shiying-Huang/MORA-MPA.
ShiYing Huang, Liang Lin, Yuer Li +6
Huazhong University of Science and Technology · Nanyang Technological University · Tsinghua University +1
Test-time alignment methods offer a promising alternative to fine-tuning by steering the outputs of large language models (LLMs) at inference time with lightweight interventions on their internal representations. Recently, a prominent and effective approach, RE-Control (Kong et al., 2024), has proposed leveraging an external value function trained over the LLM's hidden states to guide generation via gradient-based editing. While effective, this method overlooks a key characteristic of alignment tasks, i.e. that they are typically formulated as learning from human preferences between candidate responses. To address this, in this paper we propose a novel preference-based training framework, Pref-CTRL, that uses a multi-objective value function to better reflect the structure of preference data. Our approach has outperformed RE-Control on two benchmark datasets and showed greater generalization on out-of-domain datasets. Our source code is available at https://github.com/UTS-nlPUG/pref-ctrl.
Aligning large language models with human preferences must balance two competing goals: responding helpfully to legitimate requests and reliably refusing harmful ones. Most preference-based safety alignment methods collapse safety into a single scalar that is applied uniformly to every preference pair. The result is a model that looks safe on average but stays relatively unsafe on a minority of harm categories. We cast safety alignment as a per-category constrained optimization problem and derive Cat-DPO, a direct-preference-optimization algorithm with a separate adaptive safety margin for each harm category. The margin tightens when the model still produces unsafe responses on a category and relaxes once the model catches up, so the training signal tracks each category's current difficulty rather than averaging under one global rate. Across two LLM backbones and six preference-learning baselines, Cat-DPO improves aggregate helpfulness and harmlessness and compresses per-category safety variance and the best-to-worst gap, offering a drop-in per-category refinement of direct preference safety alignment.
Tiankai Yang, Yi Nian, Xinyuan Li +3
University of Southern California · 2Northwestern University