Uncovering Cross-Objective Interference in Multi-Objective Alignment
Organizations: University of Notre Dame
Abstract
We study a persistent failure mode in multi-objective alignment for large language models (LLMs), in which scalarized training improves only some objectives while the others degrade. We formalize this phenomenon as cross-objective interference and, to our knowledge, conduct the first systematic study of scalarization algorithms for multi-objective LLM alignment. The study shows that interference is pervasive across algorithms yet strongly model-dependent. To understand how interference arises, we derive a local covariance law stating that an objective improves or degrades at first order according to the sign of the covariance between its reward and the scalarized score. We extend this law to the clipped surrogate objectives of modern reinforcement fine-tuning and show that it still holds under mild conditions. Building on this law, we propose COVariance-floor Enforced Reweighting (COVER), a one-sided controller that raises an objective's weight only when the covariance between its reward and the clipped advantage weight falls below a target. Through extensive experiments, we find that COVER can mitigate cross-objective interference while matching linear scalarization when objectives already co-improve. Finally, to explain why interference is model-dependent, we complement the local covariance law with a global convergence analysis. This analysis gives sufficient conditions for the non-convex scalarized objective to satisfy the Polyak--Łojasiewicz condition and relates interference to model geometry.
Figures & tables
| Mean covariance ( ) | Negative fraction | Test performance | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Dataset | Objective | Linear | Dynamic | COVER | Linear | Dynamic | COVER | Linear | Dynamic | COVER |
| MATH-500 | Accuracy | |||||||||
| Clarity | ||||||||||
| Conciseness † | ||||||||||
| IFEval | Constraint following | |||||||||
| Helpfulness | ||||||||||
| Objectives | Final accuracy |
|---|---|
| Accuracy only | |
| Clarity only | |
| Conciseness only | |
| All three (Linear) |
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
| Accuracy | Clarity | Conciseness | Length | |||
|---|---|---|---|---|---|---|
| Objectives | Early | Final | Drop | Final | Final | Final |
| Initialization | ||||||
| Accuracy only | ||||||
| Clarity only | ||||||
| Conciseness only | ||||||
| All three (Linear) | ||||||
| Accuracy weight | Accuracy | Conciseness | Length | |||
|---|---|---|---|---|---|---|
| Target | Final (%) | Peak | Final | Drop | Final | Final |
| (Linear) | ||||||
| Method | Helpfulness | Constraint following | Conciseness |
|---|---|---|---|
| Linear | |||
| Dynamic | |||
| Lagrangian | |||
| Tchebycheff | |||
| PAMA | |||
| MGDA |
| Weights | Helpfulness | Constraint following | Conciseness |
|---|---|---|---|
| COVER (Adaptive) |
| Mean covariance | Negative fraction | Test performance | COVER | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| Linear | Linear | Linear | COVER | Engaged | Final | ||||||
| Dataset | Objective | Start | Peak | Final | Start | Peak | Final | (%) | |||
| GSM8K | Accuracy | ||||||||||
| Clarity | |||||||||||
| Conciseness † | |||||||||||
| MBPP | Pass rate | ||||||||||
| Steps 0-210 | Full run (315 steps) | |||
|---|---|---|---|---|
| Target | Peak | Accuracy | Drop | Final |
| (Linear) | ||||
| Test accuracy | ||||||
| Objectives | Penalty | Gradient norm | Peak (step) | Step 25 | Final | Final entropy |
| Accuracy only | (165) | |||||
| Accuracy only | (5) | |||||
| Accuracy only | (5) | |||||
| Accuracy only | (10) | |||||
| Accuracy only | (10) | |||||
| Method | Hyperparameters |
|---|---|
| Lagrangian primal-dual | On MATH-500 the primary objective is accuracy, with constraints on conciseness and clarity (targets ); on IFEval the primary objective is constraint following, with constraints on helpfulness ( ) and conciseness ( ); dual learning rate ; KL coefficient . |
| Smooth Tchebycheff | Preference weights ; OMD learning rate . |
| MGDA | Same shared settings, except KL coefficient . |
| GradNorm | Exponent ; weight learning rate . |
| Nash-MTL | Nash optimization iterations ; weight learning rate ; numerical constant ; maximum weight . |
| FAMO | Weight learning rate ; weight decay ; numerical constant ; loss margin . |