cs.LGApr 27, 2025

Adaptive Helpfulness-Harmlessness Alignment with Preference Vectors

Authors: Ren-Wei Liang, Chin-Ting Hsu, Chan-Hung Yu, Saransh Agrawal, Shih-Cheng Huang, Chieh-Yen Lin, Shang-Tse Chen, Kuan-Hao Huang, +1 more

Organizations: National Taiwan University · Appier AI Research · Graduate Institute of Communication Engineering, National Taiwan University · Texas A&M University

Abstract

Ensuring that large language models (LLMs) are both helpful and harmless is a critical challenge, as overly strict constraints can lead to excessive refusals, while permissive models risk generating harmful content. Existing approaches, such as reinforcement learning from human feedback (RLHF) and direct preference optimization (DPO), attempt to balance these trade-offs but suffer from performance conflicts, limited controllability, and poor extendability. To address these issues, we propose Preference Vector, a novel framework inspired by task arithmetic. Instead of optimizing multiple preferences within a single objective, we train separate models on individual preferences, extract behavior shifts as preference vectors, and dynamically merge them at test time. This modular approach enables fine-grained, user-controllable preference adjustments and facilitates seamless integration of new preferences without retraining. Experiments show that our proposed Preference Vector framework improves helpfulness without excessive conservatism, allows smooth control over preference trade-offs, and supports scalable multi-preference alignment.

Figures & tables

Appendix figures & tables14 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion

    May 12, 2026ShiYing Huang, Liang Lin, Yuer Li +6Preference AlignmentMulti-Objective Reinforcement Learning

  2. Pref-CTRL: Preference Driven LLM Alignment using Representation Editing

    Apr 26, 2026Imranul Ashrafi, Inigo Jauregi Unanue, Massimo PiccardiLarge Language Model AlignmentValue Alignment

  3. Cat-DPO: Category-Adaptive Safety Alignment

    Apr 19, 2026Tiankai Yang, Yi Nian, Xinyuan Li +3Safety AlignmentLarge Language Model Alignment