cs.LGAug 15, 2026

Online Convex Optimization with Dueling Feedback

Authors: Yiyang Lu, Hareshkumar Jadav, Mohammad Pedramfar, Ranveer Singh, Vaneet Aggarwal

Organizations: Purdue University · IIT Indore · Mila - Quebec AI Institute/McGill University

Abstract

Noisy binary comparison between two candidates is a common interface between human and learning systems, especially in modern large language model (LLM) post-training alignment. We study online convex optimization with dueling (pairwise comparison) feedback, where the learner observes only a binary preference between two queried points. We consider adversarial sequences of convex losses and measure regret with the loss at both queried points, under a comparison link with a known nonzero slope at the origin. We propose a simple reduction that converts dueling feedback into approximate gradients, enabling the use of standard first-order methods. We show that regret guarantees transfer under this reduction, yielding O(T3/4)\mathcal O(T^{3/4}) static and adaptive regret, and O(T3/41+PT/D)\mathcal O(T^{3/4}\sqrt{1+P_T/D}) dynamic regret with unknown comparator path length PTP_T. For strongly convex losses, the static and adaptive bounds improve to O~(T2/3)\widetilde{\mathcal O}(T^{2/3}). For smooth losses, we presents unified dueling ellipsoidal FTRL, and proves O~(T2/3)\widetilde{\mathcal O}(T^{2/3}) static regret, which improves to O~(T)\widetilde{\mathcal O}(\sqrt T) under additional strong convexity.

Figures & tables

Appendix figures & tables1 asset

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Online Learning on Hidden-Convex Losses via Algorithmic Equivalence: Optimal Regret, Geometric Barrier, and Bandit Feedback

    May 25, 2026Anas Barakat, Andreas Kontogiannis, Vasilis Pollatos +2Convex Loss\Widetilde{\Mathcal{O}}(\Sqrt{T})$ Regret

  2. Online Convex Optimization with Sublinear Noisy Probes

    Jun 12, 2026Simone Di Gregorio, Anupam Gupta, Stefano Leonardi +1\Widetilde{\Mathcal{O}}(\Sqrt{T})$ RegretConvex Optimization

  3. Small Gradient Norm Regret for Online Convex Optimization

    Jan 20, 2026Wenzhi Gao, Chang He, Madeleine UdellInterval RegretRegret