stat.MLOct 5, 2026

A Query Is Not a Commitment: Learning to Correct Expert Answers in Online Deferral

Authors: Yannis Montreuil, Axel Carlier, Lai Xing Ng, Wei Tsang Ooi

Organizations: School of Computing, National University of Singapore · IRIT, Toulouse INP, France · Institute for Infocomm Research, A*STAR, Singapore

Abstract

An inaccurate expert can still provide useful information after correction. We study online learning to defer in which the learner chooses an expert and fixes a correction function before purchasing its answer, then applies that function to the answer received. The difficulty is that observed losses reflect both expert quality and an unfinished correction: early errors can discourage queries that would be valuable after learning. We propose ORUCB, which pools shared and expert-specific polynomial responses. A bound on cumulative response-learning error calibrates confidence-weighted risk regression and exploration, allowing the router to account for this error when deciding which answers to buy. Under bounded residuals and disagreements, a fixed feasible model of optimal responses, and linear models of free and optimal queried risk, the calibrated algorithm achieves high-probability pseudo-regret O(Tlog⁡(T+1))O(\sqrt T\log(T+1)) over TT rounds for fixed problem parameters. The guarantee permits singular answer distributions and misspecified shared responses; optimality is relative to the bounded response class. On four test streams, the selected cubic policy has lower fee-inclusive cost than seven baselines that deploy answers unchanged. Comparisons with a common correction learner examine routing, while six-price comparisons measure cost and query rates.

Figures & tables

Appendix figures & tables7 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Online Learning-to-Defer with Varying Experts

    May 12, 2026Dang Hoang Duy, Yannis Montreuil, Maxime Meyer +3Learning-Augmented AlgorithmsExperts

  2. Prediction with Expert Advice: Anytime Regret with Many Experts Matches the Fixed-Time Constant

    Sep 23, 2026Yang Cai, Vineet Gupta, Yanchen Jiang +4\Widetilde{\Mathcal{O}}(\Sqrt{T})$ RegretRegret

  3. Learning-to-Defer in Non-Stationary Time Series via Switching State-Space Models

    Jan 30, 2026Yannis Montreuil, Letian Yu, Axel Carlier +2Non-Stationarity