cs.LGOct 5, 2026

Collaborative Personalized Preference Alignment for LLMs under Data Deficiency

Authors: Liyan Yang, Yige Yuan, Zhiqin Yang

Organizations: City University of Hong Kong · University of Washington · The Hong Kong University of Science and Technology

Abstract

Real-world users often exhibit highly heterogeneous preferences over multiple objectives for LLM responses. A lightweight aligner can tailor these responses to individual preferences, but scarce user-specific feedback makes personalized training difficult. Learning shared initializations across users can support few-shot adaptation. However, heterogeneous preferences and competing objectives cause gradient conflicts across users and within each user, hindering effective initialization learning. This raises a central question: \textbf{how can we collaboratively learn aligner initializations that support few-shot adaptation to diverse user preferences?} To answer this question, we propose \textbf{A}pproximate \textbf{P}areto \textbf{O}ptimality (APO). We first group users whose updates are compatible, so that their information can be combined with less interference. Within each group, we combine gradient descent with controlled ascent to coordinate competing objectives and move towards preference-specific points on the Pareto front. This produces an initialization that is close to the optima of the users in the group. We then iteratively refine it using updates from few-shot local adaptation, making it more effective for personalization. Furthermore, we establish conditional suboptimality bounds for a one-local-step collaborative update and characterize how initialization error affects subsequent stochastic adaptation. Experiments on Fed-ChatbotPA and UltraFeedback show consistent improvements over existing methods using only 20 local examples.

Figures & tables

Appendix figures & tables1 asset

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. FaST: Feature-aware Sampling and Tuning for Personalized Preference Alignment with Limited Data

    Aug 6, 2025Thibaut Thonet, Germán Kruszewski, Jos Rozen +2Large Language Model PersonalizationPreference Alignment

  2. Learning Preference Adaptation for Large Language Model Personalization via Verbal Reinforcement Learning

    Aug 10, 2026Yuting Liu, Wei Wu, Jianzhe Zhao +1Large Language Model PersonalizationPreference Alignment

  3. MATO: Multi-objective Personalized Alignment with Test-time Optimization for Large Language Models

    May 25, 2026Linhao Luo, Thuy-Trang Vu, Van-Anh Nguyen +3Large Language Model PersonalizationPreference Alignment Learning