Uncertainty-Normalized Margins for Direct Preference Optimization
Organizations: School of Computer and Communication Sciences, EPFL, Lausanne, Switzerland · University of Helsinki, Helsinki, Finland
Abstract
Direct preference optimization (DPO) models binary preferences through a Bradley-Terry model with a common noise scale, without explicitly accounting for preference strength or prompt-dependent uncertainty from human feedback. We introduce uncertainty-normalized margin DPO (UNM-DPO), which combines strength-dependent margins with a learned prompt scale. Motivated by a heteroskedastic Bradley-Terry model, we develop two training objectives. Both compare the implicit rewards of preferred and rejected responses, derived from response log-probability ratios to a reference policy. Advantage-only (AO) divides this reward difference by the prompt scale before subtracting the margin; whole-residual (WR) subtracts the margin before dividing by the scale. For the WR comparison model, we establish a necessary and sufficient condition under which known margins make the prompt scale identifiable. We introduce a practical procedure for learning the scale. Building on WR, we introduce ULNM-DPO-WR, which normalizes each response's implicit reward by its length. We evaluate our methods against DPO and related baselines on HelpSteer2 and HelpSteer3, using the Skywork reward model as a judge. With Llama-3.1-8B-Instruct, ULNM-DPO-WR achieves tie-adjusted win rates against matched DPO of 68.00% and 65.31% on evaluation panels, with higher mean rewards and shorter responses on average. On AlpacaEval with a GPT-4.1 judge and GPT-4-Turbo reference answers, the same 8B policy achieves a length-controlled win rate of 21.62%, compared with 16.39% for DPO and 15.30% for SimPO. These results demonstrate the potential of combining preference-strength margins, learned prompt scales, and length normalization for policy optimization.
Figures & tables
| 400 prompts | 800 prompts | |||||
| Method | Win (%) | Win (%) | ||||
| DPO (ref.) | 50.00 | 0.000 | 1.000 | 50.00 | 0.000 | 1.000 |
| Fixed-margin DPO | 50.50 [45.63, 55.38] | +0.341 [-0.284, +0.988] | 1.074 | 50.62 [47.19, 53.94] | +0.216 [-0.281, +0.716] | 1.094 |
| ODPO | 53.88 [49.00, 58.63] | +0.529 [-0.092, +1.140] | 1.056 | 51.38 [47.94, 54.81] | +0.272 [-0.219, +0.750] | 1.069 |
| MMPO | 44.50 [39.63, 49.25] | -0.271 [-0.819, +0.297] | 0.985 | 49.13 [45.75, 52.44] | -0.289 [-0.703, +0.119] | 0.981 |
| -DPO | 25.13 [21.00, 29.50] | -9.018 [-10.216, -7.825] | 1.056 | 24.50 [21.56, 27.56] | -8.606 [-9.410, -7.777] | 1.050 |
| 400 prompts | 800 prompts | |||||
| Method | Win (%) | Win (%) | ||||
| DPO (ref.) | 50.00 | 0.000 | 1.000 | 50.00 | 0.000 | 1.000 |
| Fixed-margin DPO | 58.13 [53.25, 62.88] | +1.120 [+0.453, +1.799] | 1.024 | 54.75 [51.38, 58.19] | +0.322 [-0.174, +0.816] | 1.004 |
| ODPO | 55.50 [50.63, 60.25] | +1.042 [+0.412, +1.692] | 1.012 | 52.81 [49.38, 56.19] | +0.073 [-0.407, +0.546] | 1.016 |
| MMPO | 48.88 [44.13, 53.63] | -0.287 [-0.876, +0.306] | 1.036 | 45.75 [42.44, 49.19] | -0.691 [-1.111, -0.266] | 1.023 |
| -DPO | 8.63 [6.00, 11.38] | -25.450 [-26.980, -23.917] | 2.266 | 8.69 [6.81, 10.69] | -25.533 [-26.604, -24.434] | 2.246 |
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
| Evidence bin | Mean evidence | Held-out votes | Prompts | Agreement (95% CI) |
|---|---|---|---|---|
| Low | ||||
| Middle | ||||
| High |
| Model | Frozen revision |
|---|---|
| Llama-3.2-1B-Instruct | 9213176726f574b556790deb65791e0c5aa438b6 |
| Llama-3.1-8B-Instruct | 0e9e39f249a16976918f6564b8830bc894c89659 |
| Method | Implemented per-pair loss |
|---|---|
| DPO | , with . |
| ODPO | . |
| MMPO | . |
| SimPO | . |
| SPO-basic | . |
| -DPO | on the selected batch subset. |
| Method and parameter | Candidate values | Selected / fixed |
|---|---|---|
| Earlier AO: margin | ||
| ODPO: strength multiplier | ||
| MMPO: target slope | ||
| -DPO: adaptation coefficient | ||
| -PO (DPO): base margin | ||
| SimPO | , , LR | , , |
| Setting | Median tokens | |
|---|---|---|
| HS3 1B | 364.5 | 18.225 |
| HS2 1B | 272 | 13.6 |
| HS3 8B | 364.5 | 18.225 |
| Scale estimator | Supervision and fitting schedule |
|---|---|
| Pooled OOF, WR-fit | One prompt head fitted to all pooled OOF values using the WR loss, then frozen; this is the main formulation. |
| Pooled OOF, AO-fit | One prompt head fitted to all pooled OOF values using the AO loss, then frozen. |
| Cross-fitted, AO-fit | Five prompt or pair heads fitted on four folds of OOF values each; held-out predictions are pooled and frozen. |
| Joint, AO-fit | A prompt or pair head fitted alongside the final policy using detached current-policy values . |
| Entropy-CF | Five prompt or pair heads regress annotation dispersion on four folds; pooled held-out predictions are transformed and frozen. |
| EMA-joint | A prompt or pair head uses current-policy values and moving-average centering during policy training. |
| 400 prompts | 800 prompts | |||||
| Method | Win (%) | Win (%) | ||||
| Fixed-margin DPO | 50.50 [45.63, 55.38] | +0.341 [-0.284, +0.988] | 1.074 | 50.62 [47.19, 53.94] | +0.216 [-0.281, +0.716] | 1.094 |
| Pooled prompt scale: AO-fitted | ||||||
| Pooled prompt AO (AO-fit) | 57.00 [52.13, 61.75] | +1.173 [+0.536, +1.825] | 1.072 | 55.06 [51.56, 58.44] | +0.666 [+0.159, +1.170] | 1.082 |
| Pooled prompt WR (AO-fit) | 54.50 [49.63, 59.38] | +0.818 [+0.155, +1.462] | 1.085 | 55.38 [51.88, 58.81] | +0.633 [+0.115, +1.150] | 1.089 |
| Pooled prompt scale: WR-fitted | ||||||
| Method | Win (%) | ||
| 400 prompts, seed 42 | |||
| EMA-joint pair AO | 52.75 [47.88, 57.63] | — | |
| Entropy-CF pair AO | 51.88 [47.13, 56.75] | — | |
| Entropy-CF prompt AO | 52.00 [47.12, 56.88] | — | |
| Entropy-CF prompt WR ∗ | 56.00 [51.13, 60.88] | +0.996 [+0.309, +1.683] | 1.070 |
| Entropy-CF pair WR | 48.25 [43.38, 53.00] | — | |
| 400 prompts | 800 prompts | |||
| Scale fitting | Win (%) | Win (%) | ||
| Llama-3.2-1B-Instruct | ||||
| Centered, regularized | 62.63 [57.88, 67.25] | 1.007 | 61.50 [58.13, 64.81] | 0.998 |
| Uncentered, regularized | 64.13 [59.50, 68.88] | 1.024 | 62.81 [59.50, 66.13] | 1.007 |
| Uncentered, unregularized | 68.75 [64.25, 73.25] | 1.039 | 61.00 [57.63, 64.38] | 1.024 |
| Llama-3.1-8B-Instruct | ||||
| 400 prompts | 800 prompts | |||||
|---|---|---|---|---|---|---|
| Prompt scale | Win (%) | Win (%) | ||||
| Constant | 62.00 [57.25, 66.63] | +1.641 [+0.890, +2.386] | 0.991 | 56.56 [53.13, 60.00] | +0.527 [-0.071, +1.139] | 0.992 |
| Learned | 68.00 [63.38, 72.38] | +2.372 [+1.661, +3.080] | 0.985 | 65.31 [62.00, 68.56] | +1.796 [+1.241, +2.356] | 0.964 |
| Panel | Scale fitting | Win (%) | ||
|---|---|---|---|---|
| 400 | Centered, regularized | 58.25 [53.50, 63.00] | +1.040 | 1.064 |
| 400 | Uncentered, regularized | 57.13 [52.37, 62.00] | +1.233 | 1.103 |
| 400 | Uncentered, unregularized | 54.75 [49.88, 59.63] | +0.787 | 1.107 |
| 800 | Centered, regularized | 56.94 [53.56, 60.31] | +0.929 | 1.085 |
| 800 | Uncentered, regularized | 56.56 [53.19, 59.94] | +0.598 | 1.117 |
| 800 | Uncentered, unregularized | 55.88 [52.50, 59.25] | +0.725 | 1.112 |
| Method | Win (%) | Length ratio | |
|---|---|---|---|
| ULNM-DPO-WR | 70.13 | +2.949 | 1.034 |
| [66.94, 73.31] | [2.388, 3.511] | ||
| DPO | 59.44 | +1.152 | 1.072 |
| [56.19, 62.88] | [0.630, 1.679] | ||
| SimPO | 62.94 | +2.377 | 1.042 |
| [59.62, 66.19] | [1.707, 3.034] |
| Method | Win (%) | (point) | |
|---|---|---|---|
| DPO (reference) | 50.00 | 0.000 | 1.000 |
| Fixed-margin DPO | 52.23 [45.76, 58.71] | +0.825 | 1.023 |
| ODPO | 49.33 [42.63, 55.80] | +0.759 | 1.018 |
| MMPO | 51.56 [44.87, 58.04] | 0.965 | |
| -DPO | 57.59 [51.34, 63.84] | +1.173 | 1.016 |
| -PO (DPO) | 55.80 [49.33, 62.28] | +0.786 | 1.020 |
| Policy | Learning rate | Updates | Win (%) | ||
|---|---|---|---|---|---|
| AO | 52 | 54.69 [48.21, 61.16] | +0.241 | 0.965 | |
| AO | 52 | 56.92 [50.45, 63.17] | +1.041 | 0.967 | |
| AO | 52 | 55.36 [48.66, 61.61] | +0.203 | 0.958 | |
| AO | 104 | 56.03 [49.55, 62.50] | +0.784 | 0.984 | |
| AO | 104 | 55.80 [49.55, 62.05] | +0.883 | 0.974 | |
| AO | 104 | 54.69 [48.21, 61.17] | +0.873 | 0.996 |
| Win (%) | |||
|---|---|---|---|
| 0.5 | 54.46 [48.21, 60.71] | +0.639 | 0.979 |
| 1.0 | 60.27 [54.01, 66.52] | +0.824 | 0.975 |
| 1.5 | 54.24 [47.54, 60.49] | +0.717 | 0.989 |
| Method | LC win (%) | Mean characters |
|---|---|---|
| DPO | 16.39 | 2148 |
| Fixed-margin DPO | 15.72 | 2348 |
| ODPO | 17.29 | 2284 |
| MMPO | 13.71 | 2492 |
| -DPO | 0.49 | 8445 |
| -PO (DPO) | 14.80 | 2198 |