stat.MLMay 28, 2026

Reward Learning from Best-of-N Preference Data: Targets, Tradeoffs, and Design Principles

Authors: Rattana PukdeeMaria-Florina BalcanPradeep Ravikumar

Organizations: Machine Learning Department Carnegie Mellon University

Abstract

Best-of-NN sampling is widely used to construct pairwise preference data: NN candidates are drawn from a base distribution, and the best is paired with a rejected response. Despite its widespread use, what Bradley--Terry (BT) reward learning extracts from such data, and how to choose NN and the base distribution, remain unclear. We specialize a recent analysis of preference data via its induced conditional distribution to Best-of-NN. For independent-reference variants, we derive closed-form reward targets as explicit functions of NN and the base distribution, and show that they preserve the latent reward ranking. For the practical Best-vs-Random and Best-vs-Worst variants, chosen and rejected responses are coupled through the same candidate set, so exact BT representability generally fails; nevertheless, bounded-class minimizers approach the reference targets as NN grows. Although margin and connectivity are known to govern sample efficiency in pairwise preference learning, Best-of-NN couples them through NN in opposing directions: larger NN widens pairwise margins but reduces connectivity. This trade-off yields two design principles: use larger NN when preference labels are the bottleneck, smaller NN when generation is the bottleneck; and shape the base distribution to place mass between the responses whose comparison matters most at test time. Experiments on synthetic and real preference data support the predicted dependence on sample size and base-distribution shape.

Explore similar work

CardsList