cs.CLFeb 6, 2026

Uncovering Cross-Objective Interference in Multi-Objective Alignment

Authors: Yining Lu, Meng Jiang

Organizations: University of Notre Dame

Abstract

We study a persistent failure mode in multi-objective alignment for large language models (LLMs), in which scalarized training improves only some objectives while the others degrade. We formalize this phenomenon as cross-objective interference and, to our knowledge, conduct the first systematic study of scalarization algorithms for multi-objective LLM alignment. The study shows that interference is pervasive across algorithms yet strongly model-dependent. To understand how interference arises, we derive a local covariance law stating that an objective improves or degrades at first order according to the sign of the covariance between its reward and the scalarized score. We extend this law to the clipped surrogate objectives of modern reinforcement fine-tuning and show that it still holds under mild conditions. Building on this law, we propose COVariance-floor Enforced Reweighting (COVER), a one-sided controller that raises an objective's weight only when the covariance between its reward and the clipped advantage weight falls below a target. Through extensive experiments, we find that COVER can mitigate cross-objective interference while matching linear scalarization when objectives already co-improve. Finally, to explain why interference is model-dependent, we complement the local covariance law with a global convergence analysis. This analysis gives sufficient conditions for the non-convex scalarized objective to satisfy the Polyak--Łojasiewicz condition and relates interference to model geometry.

Figures & tables

Appendix figures & tables13 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. MATO: Multi-objective Personalized Alignment with Test-time Optimization for Large Language Models

    May 25, 2026Linhao Luo, Thuy-Trang Vu, Van-Anh Nguyen +3Large Language Model PersonalizationPreference Alignment Learning

  2. EvoPref: Multi-Objective Evolutionary Optimization Discovers Diverse LLM Alignments Beyond Gradient Descent

    May 10, 2026Dongxin Guo, Jikun Wu, Siu Ming YiuLarge Language Model AlignmentLarge Language Model Adaptation