ORPG: Reconciling Multiple Reward Objectives through Objective-wise Policy Gradients
Organizations: Fudan University · Shanghai Innovation Institute
Abstract
Multi-reward policy optimization requires a joint update that reflects both the learning signals and the intended relationships among objectives. We introduce Objective-wise Reconciled Policy Gradient (ORPG), which constructs a separate clipped policy objective for each reward and reconciles the resulting gradients into one policy update. For compatible gradients, a cosine-dependent interpolation coordinates their contributions through a partially normalized reference while preserving the norm of their sum. We characterize this update as the unique solution of a spherical directional compromise. For conflicting gradients, projection follows the task's priorities. We evaluate the same compatible rule in helpfulness--safety alignment and correctness--cost optimization for mathematical reasoning. ORPG substantially improves average Useful and Harmless scores over the strongest external baseline on each axis. In mathematics, it achieves the highest average full-budget accuracy and three-budget hypervolume among the compared methods, with more accurate and shorter responses than the initial policy. Component comparisons and training dynamics show the larger contribution of compatible coordination and a complementary benefit from conflict handling. These results support gradient reconciliation for objectives with equal standing and for objectives with an explicit priority.
Figures & tables
| Method | Alpaca | HH-RLHF | PKU-SafeRLHF | Avg | ||||
|---|---|---|---|---|---|---|---|---|
| U | H | U | H | U | H | U | H | |
| Base | \text{2.536}_{\text{\tiny\pm0.025}} | \text{2.941}_{\text{\tiny\pm0.028}} | \text{2.855}_{\text{\tiny\pm0.002}} | \text{3.977}_{\text{\tiny\pm0.002}} | \text{4.644}_{\text{\tiny\pm0.002}} | \text{6.404}_{\text{\tiny\pm0.004}} | \text{3.345}_{\text{\tiny\pm0.009}} | \text{4.441}_{\text{\tiny\pm0.010}} |
| GRPO | \text{5.232}_{\text{\tiny\pm0.036}} | \text{6.212}_{\text{\tiny\pm0.080}} | \text{4.470}_{\text{\tiny\pm0.037}} | \text{6.082}_{\text{\tiny\pm0.044}} | \text{5.620}_{\text{\tiny\pm0.011}} | \text{6.938}_{\text{\tiny\pm0.002}} | \text{5.107}_{\text{\tiny\pm0.028}} | \text{6.411}_{\text{\tiny\pm0.038}} |
| GDPO | \text{5.335}_{\text{\tiny\pm0.010}} | \text{6.250}_{\text{\tiny\pm0.121}} | \text{4.546}_{\text{\tiny\pm0.004}} | \text{6.116}_{\text{\tiny\pm0.019}} | \text{5.640}_{\text{\tiny\pm0.008}} | \text{6.939}_{\text{\tiny\pm0.005}} | \text{5.174}_{\text{\tiny\pm0.005}} | \text{6.435}_{\text{\tiny\pm0.047}} |
| GD 2 PO | \text{5.332}_{\text{\tiny\pm0.010}} | \text{6.297}_{\text{\tiny\pm0.033}} | \text{4.541}_{\text{\tiny\pm0.007}} | \text{6.138}_{\text{\tiny\pm0.010}} | \text{5.631}_{\text{\tiny\pm0.001}} | \text{6.939}_{\text{\tiny\pm0.002}} | \text{5.168}_{\text{\tiny\pm0.005}} | \text{6.458}_{\text{\tiny\pm0.013}} |
| ORPG | \text{{5.943}}_{\text{\tiny\pm0.002}} | \text{{7.061}}_{\text{\tiny\pm0.002}} | \text{{5.044}}_{\text{\tiny\pm0.002}} | \text{{6.679}}_{\text{\tiny\pm0.003}} | \text{{5.781}}_{\text{\tiny\pm0.0001}} | \text{{6.972}}_{\text{\tiny\pm0.0002}} | \text{{5.589}}_{\text{\tiny\pm0.0004}} | \text{{6.904}}_{\text{\tiny\pm0.001}} |
| Method | A24 | AMC | MATH | Min. | Oly. | Avg | ||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Acc (%) | Len | Acc (%) | Len | Acc (%) | Len | Acc (%) | Len | Acc (%) | Len | Acc (%) | Len | |
| Base | \text{53.9}_{\text{\tiny\pm3.4}} | \text{5698}_{\text{\tiny\pm60}} | \text{{82.4}}_{\text{\tiny\pm0.7}} | \text{3351}_{\text{\tiny\pm86}} | \text{87.6}_{\text{\tiny\pm0.1}} | \text{1462}_{\text{\tiny\pm6}} | \text{37.3}_{\text{\tiny\pm0.3}} | \text{1495}_{\text{\tiny\pm22}} | \text{66.4}_{\text{\tiny\pm0.2}} | \text{3821}_{\text{\tiny\pm9}} | \text{65.5}_{\text{\tiny\pm0.7}} | \text{3165}_{\text{\tiny\pm12}} |
| GRPO | \text{42.2}_{\text{\tiny\pm1.0}} | \text{2541}_{\text{\tiny\pm135}} | \text{75.5}_{\text{\tiny\pm2.3}} | \text{1708}_{\text{\tiny\pm58}} | \text{86.0}_{\text{\tiny\pm0.2}} | \text{939}_{\text{\tiny\pm40}} | \text{37.9}_{\text{\tiny\pm0.6}} | \text{964}_{\text{\tiny\pm91}} | \text{61.8}_{\text{\tiny\pm0.3}} | \text{1688}_{\text{\tiny\pm71}} | \text{60.7}_{\text{\tiny\pm0.5}} | \text{1568}_{\text{\tiny\pm66}} |
| GDPO | \text{38.6}_{\text{\tiny\pm1.3}} | \text{{2255}}_{\text{\tiny\pm105}} | \text{74.7}_{\text{\tiny\pm0.6}} | \text{{1540}}_{\text{\tiny\pm26}} | \text{85.5}_{\text{\tiny\pm0.1}} | \text{{833}}_{\text{\tiny\pm12}} | \text{37.3}_{\text{\tiny\pm0.6}} | \text{833}_{\text{\tiny\pm25}} | \text{60.6}_{\text{\tiny\pm0.6}} | \text{{1501}}_{\text{\tiny\pm46}} | \text{59.3}_{\text{\tiny\pm0.1}} | \text{{1392}}_{\text{\tiny\pm38}} |
| GD 2 PO | \text{41.9}_{\text{\tiny\pm5.4}} | \text{2470}_{\text{\tiny\pm90}} | \text{75.1}_{\text{\tiny\pm1.5}} | \text{1583}_{\text{\tiny\pm51}} | \text{85.4}_{\text{\tiny\pm0.1}} | \text{846}_{\text{\tiny\pm12}} | \text{37.9}_{\text{\tiny\pm1.0}} | \text{{826}}_{\text{\tiny\pm34}} | \text{60.9}_{\text{\tiny\pm0.3}} | \text{1553}_{\text{\tiny\pm29}} | \text{60.3}_{\text{\tiny\pm0.8}} | \text{1456}_{\text{\tiny\pm29}} |
| ORPG | \text{{56.4}}_{\text{\tiny\pm1.9}} | \text{4613}_{\text{\tiny\pm94}} | \text{{82.4}}_{\text{\tiny\pm0.6}} | \text{2748}_{\text{\tiny\pm39}} | \text{{88.0}}_{\text{\tiny\pm0.1}} | \text{1229}_{\text{\tiny\pm3}} | \text{{39.3}}_{\text{\tiny\pm0.5}} | \text{1287}_{\text{\tiny\pm27}} | \text{{67.2}}_{\text{\tiny\pm0.6}} | \text{2893}_{\text{\tiny\pm23}} | \text{{66.7}}_{\text{\tiny\pm0.3}} | \text{2554}_{\text{\tiny\pm22}} |
| Method | A24 | AMC | MATH | Min. | Oly. | Avg |
|---|---|---|---|---|---|---|
| Base | \text{0.280}_{\text{\tiny\pm0.015}} | \text{0.608}_{\text{\tiny\pm0.006}} | \text{0.768}_{\text{\tiny\pm0.001}} | \text{0.325}_{\text{\tiny\pm0.002}} | \text{0.485}_{\text{\tiny\pm0.0005}} | \text{0.493}_{\text{\tiny\pm0.004}} |
| GRPO | \text{0.311}_{\text{\tiny\pm0.011}} | \text{0.620}_{\text{\tiny\pm0.014}} | \text{0.769}_{\text{\tiny\pm0.003}} | \text{0.336}_{\text{\tiny\pm0.003}} | \text{0.507}_{\text{\tiny\pm0.007}} | \text{0.509}_{\text{\tiny\pm0.003}} |
| GDPO | \text{0.295}_{\text{\tiny\pm0.011}} | \text{0.625}_{\text{\tiny\pm0.005}} | \text{0.774}_{\text{\tiny\pm0.001}} | \text{0.337}_{\text{\tiny\pm0.006}} | \text{0.507}_{\text{\tiny\pm0.006}} | \text{0.507}_{\text{\tiny\pm0.002}} |
| GD 2 PO | \text{0.315}_{\text{\tiny\pm0.040}} | \text{0.627}_{\text{\tiny\pm0.014}} | \text{0.773}_{\text{\tiny\pm0.001}} | \text{0.342}_{\text{\tiny\pm0.009}} | \text{0.508}_{\text{\tiny\pm0.002}} | \text{0.513}_{\text{\tiny\pm0.005}} |
| ORPG | \text{{0.339}}_{\text{\tiny\pm0.010}} | \text{{0.637}}_{\text{\tiny\pm0.004}} | \text{{0.777}}_{\text{\tiny\pm0.0005}} | \text{{0.344}}_{\text{\tiny\pm0.005}} | \text{{0.516}}_{\text{\tiny\pm0.004}} | \text{{0.523}}_{\text{\tiny\pm0.002}} |
| Update | Useful | Harmless |
|---|---|---|
| ORPG | \text{{5.589}}_{\text{\tiny\pm0.0004}} | \text{{6.904}}_{\text{\tiny\pm0.001}} |
| Without compatible coordination | \text{5.212}_{\text{\tiny\pm0.062}} | \text{6.710}_{\text{\tiny\pm0.032}} |
| Without conflict resolution | \text{5.563}_{\text{\tiny\pm0.0006}} | \text{6.859}_{\text{\tiny\pm0.001}} |
| Without either component | \text{5.193}_{\text{\tiny\pm0.073}} | \text{6.695}_{\text{\tiny\pm0.085}} |
| CAGrad | \text{5.183}_{\text{\tiny\pm0.002}} | \text{6.626}_{\text{\tiny\pm0.002}} |
| Aligned-MTL | \text{5.065}_{\text{\tiny\pm0.002}} | \text{6.501}_{\text{\tiny\pm0.005}} |
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
| Parameter | Helpfulness–safety | Correctness–cost |
|---|---|---|
| Training steps | 100 | 100 |
| Batch / minibatch / rollout group | 512 / 128 / 4 | 512 / 64 / 8 |
| PPO epochs | 1 | 1 |
| Learning rate | ||
| Optimizer | AdamW | AdamW |
| Schedule / warmup | Constant / none | Constant / none |
| Recorded operation | Seconds per step | Total hours |
|---|---|---|
| Total training step | 788.9 | 21.91 |
| Rollout generation | 211.1 | 5.86 |
| Actor update | 474.4 | 13.18 |
| Old-policy log probabilities | 48.6 | 1.35 |
| Reference-policy log probabilities | 46.4 | 1.29 |
| Advantage computation | 2.4 | 0.07 |
| Dataset | Use | Count |
|---|---|---|
| Alpaca | Training | 50,978 |
| Alpaca | Reward calibration | 512 |
| Alpaca | Evaluation | 512 |
| HH-RLHF | Evaluation | 8,520 |
| PKU-SafeRLHF | Evaluation | 8,211 |
| AIME-24 | Evaluation | 30 |
| Method | 2048 | 4096 | 8192 | |||
|---|---|---|---|---|---|---|
| Acc | Len | Acc | Len | Acc | Len | |
| Base | \text{43.76}_{\text{\tiny\pm0.34}} | \text{1434}_{\text{\tiny\pm2}} | \text{53.66}_{\text{\tiny\pm0.48}} | \text{2232}_{\text{\tiny\pm1}} | \text{65.52}_{\text{\tiny\pm0.70}} | \text{3165}_{\text{\tiny\pm12}} |
| GRPO | \text{51.09}_{\text{\tiny\pm0.66}} | \text{1299}_{\text{\tiny\pm54}} | \text{{60.40}}_{\text{\tiny\pm0.48}} | \text{1556}_{\text{\tiny\pm70}} | \text{60.69}_{\text{\tiny\pm0.52}} | \text{1568}_{\text{\tiny\pm66}} |
| GDPO | \text{51.94}_{\text{\tiny\pm0.19}} | \text{{1192}}_{\text{\tiny\pm19}} | \text{59.00}_{\text{\tiny\pm0.41}} | \text{{1385}}_{\text{\tiny\pm36}} | \text{59.35}_{\text{\tiny\pm0.12}} | \text{{1392}}_{\text{\tiny\pm38}} |
| GD 2 PO | \text{{52.31}}_{\text{\tiny\pm1.27}} | \text{1204}_{\text{\tiny\pm19}} | \text{59.84}_{\text{\tiny\pm1.03}} | \text{1439}_{\text{\tiny\pm33}} | \text{60.26}_{\text{\tiny\pm0.85}} | \text{1456}_{\text{\tiny\pm29}} |
| ORPG | \text{47.04}_{\text{\tiny\pm0.32}} | \text{1406}_{\text{\tiny\pm3}} | \text{59.00}_{\text{\tiny\pm0.27}} | \text{2076}_{\text{\tiny\pm8}} | \text{{66.66}}_{\text{\tiny\pm0.26}} | \text{2554}_{\text{\tiny\pm22}} |
| Update | Accuracy (%) | Length | HV |
|---|---|---|---|
| ORPG | \text{{66.66}}_{\text{\tiny\pm0.26}} | \text{2554.14}_{\text{\tiny\pm21.83}} | \text{0.5226}_{\text{\tiny\pm0.0017}} |
| Without compatible coordination | \text{66.15}_{\text{\tiny\pm0.19}} | \text{2447.87}_{\text{\tiny\pm30.61}} | \text{{0.5270}}_{\text{\tiny\pm0.0025}} |
| Without conflict resolution | \text{65.46}_{\text{\tiny\pm0.55}} | \text{2456.62}_{\text{\tiny\pm18.64}} | \text{0.5203}_{\text{\tiny\pm0.0026}} |
| Without either component | \text{66.04}_{\text{\tiny\pm0.14}} | \text{{2399.56}}_{\text{\tiny\pm5.39}} | \text{0.5236}_{\text{\tiny\pm0.0021}} |
| Useful | Harmless | ||
|---|---|---|---|
| 0.25 | 0.25 | \text{5.5197}_{\text{\tiny\pm0.0005}} | \text{6.8409}_{\text{\tiny\pm0.0012}} |
| 0.5 | 0.25 | \text{{5.5891}}_{\text{\tiny\pm0.0004}} | \text{{6.9038}}_{\text{\tiny\pm0.0014}} |
| 0.75 | 0.25 | \text{5.5124}_{\text{\tiny\pm0.0004}} | \text{6.8248}_{\text{\tiny\pm0.0011}} |
| 0.5 | 0.125 | \text{5.4816}_{\text{\tiny\pm0.0017}} | \text{6.8064}_{\text{\tiny\pm0.0015}} |
| 0.5 | 0.5 | \text{5.5542}_{\text{\tiny\pm0.0012}} | \text{6.8620}_{\text{\tiny\pm0.0009}} |
| Accuracy (%) | Length | HV | ||
|---|---|---|---|---|
| 0.25 | 0.25 | \text{65.58}_{\text{\tiny\pm0.33}} | \text{2425.42}_{\text{\tiny\pm38.69}} | \text{0.5227}_{\text{\tiny\pm0.0023}} |
| 0.5 | 0.25 | \text{{66.66}}_{\text{\tiny\pm0.26}} | \text{2554.14}_{\text{\tiny\pm21.83}} | \text{0.5226}_{\text{\tiny\pm0.0017}} |
| 0.75 | 0.25 | \text{66.60}_{\text{\tiny\pm0.51}} | \text{2613.33}_{\text{\tiny\pm12.82}} | \text{{0.5243}}_{\text{\tiny\pm0.0043}} |
| 0.5 | 0.125 | \text{65.43}_{\text{\tiny\pm0.29}} | \text{{2413.83}}_{\text{\tiny\pm14.08}} | \text{0.5205}_{\text{\tiny\pm0.0028}} |
| 0.5 | 0.5 | \text{65.52}_{\text{\tiny\pm0.49}} | \text{2434.32}_{\text{\tiny\pm11.65}} | \text{0.5218}_{\text{\tiny\pm0.0036}} |
| 2048 | 4096 | 8192 | |||||
|---|---|---|---|---|---|---|---|
| Acc. | Length | Acc. | Length | Acc. | Length | ||
| 0.25 | 0.25 | \text{{47.77}}_{\text{\tiny\pm0.40}} | \text{{1374.0}}_{\text{\tiny\pm1.7}} | \text{59.31}_{\text{\tiny\pm0.44}} | \text{{2015.2}}_{\text{\tiny\pm11.0}} | \text{65.58}_{\text{\tiny\pm0.33}} | \text{2425.4}_{\text{\tiny\pm38.7}} |
| 0.5 | 0.25 | \text{47.04}_{\text{\tiny\pm0.32}} | \text{1406.3}_{\text{\tiny\pm3.2}} | \text{59.00}_{\text{\tiny\pm0.27}} | \text{2075.9}_{\text{\tiny\pm7.9}} | \text{{66.66}}_{\text{\tiny\pm0.26}} | \text{2554.1}_{\text{\tiny\pm21.8}} |
| 0.75 | 0.25 | \text{47.26}_{\text{\tiny\pm0.61}} | \text{1390.6}_{\text{\tiny\pm3.2}} | \text{59.33}_{\text{\tiny\pm0.36}} | \text{2079.0}_{\text{\tiny\pm19.7}} | \text{66.60}_{\text{\tiny\pm0.51}} | \text{2613.3}_{\text{\tiny\pm12.8}} |
| 0.5 | 0.125 | \text{47.34}_{\text{\tiny\pm0.49}} | \text{1408.5}_{\text{\tiny\pm3.1}} | \text{{60.55}}_{\text{\tiny\pm0.77}} | \text{2035.7}_{\text{\tiny\pm14.5}} | \text{65.43}_{\text{\tiny\pm0.29}} | \text{{2413.8}}_{\text{\tiny\pm14.1}} |
| 0.5 | 0.5 | \text{47.27}_{\text{\tiny\pm0.44}} | \text{1380.7}_{\text{\tiny\pm2.0}} | \text{59.44}_{\text{\tiny\pm0.21}} | \text{2031.8}_{\text{\tiny\pm2.9}} | \text{65.52}_{\text{\tiny\pm0.49}} | \text{2434.3}_{\text{\tiny\pm11.6}} |
| Update | Steps | Cosine | Conflict | Projection | Useful RMS |
|---|---|---|---|---|---|
| ORPG | 1–33 | 0.661 | 0.000 | 0.000 | 0.490 |
| 34–66 | 0.584 | 0.000 | 0.000 | 0.500 | |
| 67–100 | 0.265 | 0.191 | 0.191 | 0.857 | |
| Without compatible coordination | 1–33 | 0.690 | 0.000 | 0.000 | 0.492 |
| 34–66 | 0.616 | 0.000 | 0.000 | 0.488 | |
| 67–100 | 0.522 | 0.000 | 0.000 | 0.507 |
| Method | Useful | Harmless | Useful RMS |
|---|---|---|---|
| GDPO | 1.1568 | 1.4582 | – |
| GD 2 PO | 1.1577 | 1.4620 | – |
| Without either component | 1.1532 | 1.7368 | 0.5207 |
| Without compatible coordination | 1.1738 | 1.7676 | 0.5278 |
| Without conflict resolution | 1.4924 | 1.9990 | 0.8600 |
| ORPG | 1.5217 | 2.0051 | 0.8868 |