cs.LGSep 29, 2026

Uncertainty-Normalized Margins for Direct Preference Optimization

Authors: Sadegh Khorasani, Petrus Mikkola, Matthias Grossglauser

Organizations: School of Computer and Communication Sciences, EPFL, Lausanne, Switzerland · University of Helsinki, Helsinki, Finland

Abstract

Direct preference optimization (DPO) models binary preferences through a Bradley-Terry model with a common noise scale, without explicitly accounting for preference strength or prompt-dependent uncertainty from human feedback. We introduce uncertainty-normalized margin DPO (UNM-DPO), which combines strength-dependent margins with a learned prompt scale. Motivated by a heteroskedastic Bradley-Terry model, we develop two training objectives. Both compare the implicit rewards of preferred and rejected responses, derived from response log-probability ratios to a reference policy. Advantage-only (AO) divides this reward difference by the prompt scale before subtracting the margin; whole-residual (WR) subtracts the margin before dividing by the scale. For the WR comparison model, we establish a necessary and sufficient condition under which known margins make the prompt scale identifiable. We introduce a practical procedure for learning the scale. Building on WR, we introduce ULNM-DPO-WR, which normalizes each response's implicit reward by its length. We evaluate our methods against DPO and related baselines on HelpSteer2 and HelpSteer3, using the Skywork reward model as a judge. With Llama-3.1-8B-Instruct, ULNM-DPO-WR achieves tie-adjusted win rates against matched DPO of 68.00% and 65.31% on evaluation panels, with higher mean rewards and shorter responses on average. On AlpacaEval with a GPT-4.1 judge and GPT-4-Turbo reference answers, the same 8B policy achieves a length-controlled win rate of 21.62%, compared with 16.39% for DPO and 15.30% for SimPO. These results demonstrate the potential of combining preference-strength margins, learned prompt scales, and length normalization for policy optimization.

Figures & tables

Appendix figures & tables16 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. ξξ-DPO: Direct Preference Optimization via Ratio Reward Margin

    May 9, 2026Zhengyuan Fan, Zhonghua Wu, Yuxuan Du +1Backtranslation Augmented Direct Preference Optimization

  2. TUR-DPO: Topology- and Uncertainty-Aware Direct Preference Optimization

    Apr 30, 2026Abdulhady Abas Abdullah, Fatemeh Daneshfar, Seyedali Mirjalili +1Reasoning SkillsTopology