cs.AISep 30, 2026

On the Complexity of Preference-Based Bandits

Authors: Ahmed Ben Yahmed, Marc Abeille, Clément Calauzènes

Organizations: CREST, ENSAE Paris, FAIRPLAY · FAIRPLAY

Abstract

We study preference-based bandits with general reward function classes, where a learner sequentially selects pairs of arms and observes binary preference feedback governed by the Bradley--Terry model. This setting naturally arises in applications such as recommender systems, tournament ranking, and learning from human feedback, where relative preferences are easier to elicit than absolute rewards. The observation model inherits the logistic bandit challenge of handling the problem-dependent constant κκ, which accounts for the non-linearity of the link function and can grow arbitrarily large. Moreover, prior work has predominantly focused on linear or kernelized reward models, precluding the use of richer function classes. To address these limitations, we consider general reward function classes and introduce the \emph{locally sensitive eluder dimension}, a novel complexity measure tailored to the logistic structure of preference feedback that yields fine-grained regret guarantees without unfavorable dependence on κκ. Building on this notion, we propose \textbf{GINOP} (Generic INformative OPtimism), an algorithm that constructs log-loss confidence sets and jointly selects arm pairs to balance optimism and informative exploration. We establish a first-order regret bound that, in contrast with what previous results suggest, demonstrates that learning with preference feedback is as statistically efficient as learning from direct reward observation. Finally, we corroborate our theoretical findings with empirical evaluations against competitive baselines.

Explore similar work

CardsList
  1. When Greedy Sampling Explores: KL-Regularized Contextual Bandits without Eluder-Dimension Dependence

    Sep 11, 2026Zichen Wang, Haoyang Hong, Huazheng WangContextual Bandit FrameworkKullback-Leibler Regularization

  2. Reward Learning from Best-of-NN Preference Data: Targets, Tradeoffs, and Design Principles

    May 28, 2026Rattana Pukdee, Maria-Florina Balcan, Pradeep RavikumarBest-Of-NPreference Learning

  3. Learning the Preferences of a Learning Agent

    May 9, 2026Karim Abdel Sadek, Mark Bedaywi, Rhys Gould +1Preference LearningReinforcement Learning From Human Feedback