cs.LGAug 27, 2026

Disentangling Optimization Scale from Preference Scale in DPO

Authors: Ivan Kruzhilov

Abstract

Direct Preference Optimization (DPO) is a widely used objective for aligning language models from preference data, with the coefficient ββ commonly interpreted as controlling the KL constraint to a reference policy. We show that ββ entangles two distinct roles: it governs the effective inverse preference-noise scale and simultaneously rescales the optimization dynamics, coupling this scale with the effective step size. As a consequence, at a fixed learning rate the achieved policy deviation is non-monotone in ββ: it vanishes in a dead zone at small ββ, reaches a peak at an intermediate value, and decreases again for larger ββ. Moreover, standard DPO loss values are not comparable across ββ: runs with nearly identical loss curves can differ several-fold in KL divergence from the reference model. This entanglement obscures the role of ββ, increases sensitivity to hyperparameter choices, and complicates learning-rate scheduling. We propose a centered-softplus reformulation that is argmin-equivalent to DPO for β>0β>0, while making the inverse preference-noise-scale and learning-rate effects explicit and independently tunable. The normalized centered-softplus objective also admits a continuous β→0β\to0 endpoint that reduces to a linear preference-margin objective.

Explore similar work

CardsList
  1. LSC-DPO: Learning-Signal-Controlled Direct Preference Optimization

    Oct 6, 2026Yang Qu, Yusheng Han, Chengjia Feng +1

  2. Uncertainty-Normalized Margins for Direct Preference Optimization

    Sep 29, 2026Sadegh Khorasani, Petrus Mikkola, Matthias GrossglauserPairwise Preference LearningDirect Preference Optimization