cs.LGSep 28, 2026

ORPG: Reconciling Multiple Reward Objectives through Objective-wise Policy Gradients

Authors: Shicheng Fang, Yiwen Zhao, Wenbo Tian, Jiahao Lu, Yining Zheng, Yuxin Wang, Xipeng Qiu

Organizations: Fudan University · Shanghai Innovation Institute

Abstract

Multi-reward policy optimization requires a joint update that reflects both the learning signals and the intended relationships among objectives. We introduce Objective-wise Reconciled Policy Gradient (ORPG), which constructs a separate clipped policy objective for each reward and reconciles the resulting gradients into one policy update. For compatible gradients, a cosine-dependent interpolation coordinates their contributions through a partially normalized reference while preserving the norm of their sum. We characterize this update as the unique solution of a spherical directional compromise. For conflicting gradients, projection follows the task's priorities. We evaluate the same compatible rule in helpfulness--safety alignment and correctness--cost optimization for mathematical reasoning. ORPG substantially improves average Useful and Harmless scores over the strongest external baseline on each axis. In mathematics, it achieves the highest average full-budget accuracy and three-budget hypervolume among the compared methods, with more accurate and shorter responses than the initial policy. Component comparisons and training dynamics show the larger contribution of compatible coordination and a complementary benefit from conflict handling. These results support gradient reconciliation for objectives with equal standing and for objectives with an explicit priority.

Figures & tables

Appendix figures & tables10 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Learn What's Left, Not What's Mastered: Saturation Aware Advantage Reweighting for Multi-Reward Policy Optimization

    Aug 17, 2026Yixuan Wang, Yifei Chen, Haichao Zhang +6

  2. Not All Preferences Deserve Gradients: Understanding Gradient Utility in Offline Reasoning Alignment

    Feb 1, 2026Hui Wu, Hengyi Cai, Jinman Zhao +6GradientPreference Alignment Learning

  3. Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL

    Jul 31, 2026Ruiming Liang, Yi Zhong, Yizhen Yuan +6Prism