Regret Minimization

Momentum

12 papers in the last four weeks, up 71% on the four weeks before. 0.1% of all new papers.

Jul 13Week of Sep 28

Latest papers 81

May 10, 2026stat.ML

Optimal Regret for Single Index Bandits

We study the single-index bandit\textit{single-index bandit} problem, where rewards depend on an unknown one-dimensional projection of high-dimensional contexts through an unknown reward function. This model extends linear and generalized linear bandits to a nonparametric setting, and is particularly relevant when the reward function is not known in advance. While optimal regret guarantees are known for monotone reward functions, the general non-monotone case remains poorly understood, with the best known bound being O~(T3/4)\tilde{\mathcal{O}}(T^{3/4}) (under standard boundedness and Lipschitz assumptions on the reward function [Kang et al., 2025]). We close this gap by establishing the optimal regret for general single-index bandits. We propose a simple two-phase algorithm, namely, Zoomed Single Index Bandit with Upper Confidence Bound (ZoomSIB-UCB\texttt{ZoomSIB-UCB}), that first estimates the projection direction via a normalized Stein estimator, and then reduces the problem to a one-dimensional bandit using discretization and finally use UCB. This approach achieves a regret of O~(T2/3)\tilde{\mathcal{O}}(T^{2/3}), and improves significantly upon prior work without any additional assumptions. We also prove a matching minimax lower bound of Ω~(T2/3)\tildeΩ(T^{2/3}), showing that the upper bound is essentially tight. Our upper and lower bounds together provide a sharp characterization of the regret in single-index bandits. Moreover, the empirical results further demonstrate the effectiveness and robustness of our approach.
May 9, 2026stat.ML

Tight Generalization Bounds for Noiseless Inverse Optimization

Inverse optimization (IO) seeks to infer the parameters of a decision-maker's objective from observed context--action data. We study noiseless IO, where demonstrations are generated by a ground-truth objective. We provide a high-probability O(dT){O}(\frac{d}{T}) generalization bound for the induced action set, where dd is the number of unknown parameters and TT is the size of the training dataset. We strengthen these guarantees under additional conditions that ensure uniqueness of the chosen action, bringing our IO guarantees in line with best-arm identification results in the bandit literature. We further show that the O(dT){O}(\frac{d}{T}) rate is tight over all consistent estimators considered here, and extend the result to both instantaneous and cumulative regret. Notably, the resulting regret lower bound matches the corresponding upper bounds in the adversarial setting, indicating that the stochastic IO setting is effectively adversarial for the class of estimators studied here. Finally, we propose a parameter-free algorithm with lower per-iteration complexity than generic solvers. Experiments validate the predicted rates and illustrate the tightness of our bounds.
May 6, 2026cs.LG

Online Nonstochastic Prediction: Logarithmic Regret via Predictive Online Least Squares

We study online prediction for marginally stable, partially observed linear dynamical systems under nonstochastic disturbances. Our objective is to minimize the cumulative squared prediction loss and compete with the best-in-hindsight Luenberger predictor. Standard online learning methods typically rely on bounded domains/gradients, and thus their guarantees may fail to deal with potentially unbounded trajectories in marginally stable systems. In this paper, we introduce an unconstrained online least squares method that stabilizes the learning process via tailored predictive hints. With model knowledge, we prove that hints constructed from any stabilizing Luenberger predictor render the hint residuals uniformly bounded, achieving logarithmic regret despite unbounded trajectory growth. We also discuss model-free prediction and introduce a simple universal hint for symmetric systems, under which logarithmic regret is maintained without model knowledge. Our results provide an adaptive, instance-wise optimal online predictor compared to classical fixed-gain observers under nonstochastic disturbances.
May 1, 2026cs.LG

Trading off rewards and errors in multi-armed bandits

In multi-armed bandits, the most-explored arms are the most informative, while reward maximization typically pulls only the best arm. We study the tradeoff between identifying arm means accurately and accumulating reward, and present an algorithm with regret guarantees that interpolates between the two objectives. We provide both upper and lower bounds and validate empirically.
Apr 29, 2026cs.LG

A Note on How to Remove the ln⁡ln⁡T\ln\ln T Term from the Squint Bound

In Orabona and Pál [2016], we introduced the shifted KT potentials, to remove the ln⁡ln⁡T\ln \ln T factor in the parameter-free learning with expert bound. In this short technical note, I show that this is equivalent to changing the prior in the Krichevsky--Trofimov algorithm. Then, I show how to use the same idea to remove the ln⁡ln⁡T\ln \ln T factor in the data-independent bound for the Squint algorithm.
Apr 27, 2026cs.LG

Prior-Agnostic Robust Forecast Aggregation

Robust forecast aggregation combines the predictions of multiple information sources to perform well in the worst case across all possible information structures. Previous work largely focuses on settings with a known binary state space, where the state is either 0 or 1. We study prior-agnostic robust forecast aggregation in which the aggregator observes only experts' reports, yet is ignorant of both the underlying joint information structure and the full prior, including the underlying state space. Unlike the standard model that fixes the binary state space {0, 1}, we allow the (binary) unknown state values to be arbitrary numbers in [0, 1], so the same reported probability may correspond to very different realized outcome frequencies across environments. Our main contribution is a simple, explicit, closed-form log-odds aggregator that linearly pools forecasts in logit space, together with (nearly-)tight minimax-regret guarantees across three knowledge regimes. We first show that under conditionally independent (CI) signals, robust aggregation with an unknown state space is strictly harder than in the known-state setting by establishing a larger lower bound, and our aggregation rule can achieve a worst-case regret of 0.0255. Along the way, we also characterize tight regret bounds for Blackwell-ordered structures and for general information structures. In the classical setting with known state space {0,1}, our aggregator achieves regret strictly below 0.0226 for CI structures. To the best of our knowledge, this is the first explicit closed-form aggregator that achieves a regret upper bound strictly less than 0.0226. Finally, we extend the model where the aggregator additionally knows each expert's marginal forecast distribution; in this setting, with the CI structures, we show that a generalized log-odds rule achieves regret of 0.0228, complementing with a lower bound of 0.0225.
Apr 22, 2026cs.LG

Cover meets Robbins while Betting on Bounded Data: ln⁡n\ln n Regret and Almost Sure ln⁡ln⁡n\ln\ln n Regret

Consider betting against a sequence of data in [0,1][0,1], where one is allowed to make any bet that is fair if the data have a conditional mean m0∈(0,1)m_0 \in (0,1). Cover's universal portfolio algorithm delivers a worst-case regret of O(ln⁡n)O(\ln n) compared to the best constant bet in hindsight, and this bound is unimprovable against adversarially generated data. In this work, we present a novel mixture betting strategy that combines insights from Robbins and Cover, and exhibits a different behavior: it eventually produces a regret of O(ln⁡ln⁡n)O(\ln \ln n) on almost all paths (a measure-one set of paths if each conditional mean equals m0m_0 and intrinsic variance increases to ∞\infty), but has an O(log⁡n)O(\log n) regret on the complement (a measure zero set of paths). Our paper appears to be the first to point out the value in hedging two very different strategies to achieve a best-of-both-worlds adaptivity to stochastic data and protection against adversarial data. We contrast our results to those in Agrawal and Ramdas [2026] for a sub-Gaussian mixture on unbounded data: their worst-case regret has to be unbounded, but a similar hedging delivers both an optimal betting growth-rate and an almost sure ln⁡ln⁡n\ln\ln n regret on stochastic data. Finally, our strategy witnesses a sharp game-theoretic upper law of the iterated logarithm, analogous to Shafer and Vovk [2005].
Apr 21, 2026cs.LG

An Efficient Black-Box Reduction from Online Learning to Multicalibration, and a New Route to ΦΦ-Regret Minimization

We give a Gordon-Greenwald-Marks (GGM) style black-box reduction from online learning to online multicalibration. Concretely, we show that to achieve high-dimensional multicalibration with respect to a class of functions H, it suffices to combine any no-regret learner over H with an expected variational inequality (EVI) solver. We also prove a converse statement showing that efficient multicalibration implies efficient EVI solving, highlighting how EVIs in multicalibration mirror the role of fixed points in the GGM result for ΦΦ-regret. This first set of results resolves the main open question in Garg, Jung, Reingold, and Roth (SODA '24), showing that oracle-efficient online multicalibration with T\sqrt{T}-type guarantees is possible in full generality. Furthermore, our GGM-style reduction unifies the analyses of existing online multicalibration algorithms, enables new algorithms for challenging environments with delayed observations or censored outcomes, and yields the first efficient black-box reduction between online learning and multiclass omniprediction. Our second main result is a fine-grained reduction from high-dimensional online multicalibration to (contextual) ΦΦ-regret minimization. Together with our first result, this establishes a new route from external regret to Phi-regret that bypasses sophisticated fixed-point or semi-separation machinery, dramatically simplifies a result of Daskalakis, Farina, Fishelson, Pipis, and Schneider (STOC '25) while improving rates, and yields new algorithms that are robust to richer deviation classes, such as those belonging to any reproducing kernel Hilbert space.
Apr 19, 2026cs.LG

Demystifying the unreasonable effectiveness of online alignment methods

Iterative alignment methods based on purely greedy updates are remarkably effective in practice, yet existing theoretical guarantees of O(log⁡T)O(\log T) KL-regularized regret can seem pessimistic relative to their empirical performance. In this paper, we argue that this mismatch arises from the regret criterion itself: KL-regularized regret conflates the statistical cost of learning with the exploratory randomization induced by the softened training policy. To separate these effects, we study the traditional temperature-zero regret criterion, which evaluates only the top-ranked response at inference time. Under this decision-centric notion of performance, we prove that standard greedy online alignment methods, including online RLHF and online DPO, achieve constant (O(1))(O(1)) cumulative regret. By isolating the cost of identifying the best response from the stochasticity induced by regularization, our results provide a sharper theoretical explanation for the practical superb efficiency of greedy alignment.
Apr 17, 2026stat.ML

Adaptive multi-fidelity optimization with fast learning rates

In multi-fidelity optimization, biased approximations of varying costs of the target function are available. This paper studies the problem of optimizing a locally smooth function with a limited budget, where the learner has to make a tradeoff between the cost and the bias of these approximations. We first prove lower bounds for the simple regret under different assumptions on the fidelities, based on a cost-to-bias function. We then present the Kometo algorithm which achieves, with additional logarithmic factors, the same rates without any knowledge of the function smoothness and fidelity assumptions, and improves previously proven guarantees. We finally empirically show that our algorithm outperforms previous multi-fidelity optimization methods without the knowledge of problem-dependent parameters.
Mar 6, 2026stat.ML

Bilateral Trade Under Heavy-Tailed Valuations: Minimax Regret without a Variance Bound

In contextual bilateral trade under full feedback, the posted price does not affect which valuations are observed. We show that in this model such action-independent feedback removes the polynomial adaptation penalty familiar from heavy-tailed bandits: fully parameter-free algorithms attain the oracle minimax TT-exponents up to logarithmic factors, with no knowledge of the moment order p∈(1,2)p \in (1,2) or its scale σpσ_p, and -- in the nonparametric case -- none of the effective Hölder smoothness β∈(0,1]β\in (0,1]. The statistic that makes model selection possible is a paired squared-loss difference, whose noise-square term cancels exactly, leaving noise damped by the candidate gap. The resulting bilateral-trade regret rates are new. Trader valuations have bounded conditional densities and heavy tails -- finite pp-th moments for some p∈(1,2)p \in (1,2), with possibly infinite variance. An epoch-based algorithm with truncated means achieves regret O~(T(2−p)/p)\widetilde{O}(T^{(2-p)/p}) in the parametric model and O~(T1−2β(p−1)/(βp+d(p−1)))\widetilde{O}(T^{1-2β(p-1)/(βp + d(p-1))}) when the market value function is ββ-Hölder, with matching Ω(⋅)Ω(\cdot) lower bounds -- under a mild nondegeneracy condition -- via Assouad's method and a fixed-support mixture construction -- characterizing the minimax rate in TT up to logarithmic factors over the effective smoothness range β∈(0,1]β\in (0,1], interpolating between the classical nonparametric rate at p=2p{=}2 and the trivial linear rate as p→1+p \to 1^+. The enabling structural step extends the self-bounding property of Bachoc et al. (ICML 2025) from bounded to real-valued valuations: within our conditionally independent, conditionally centered noise model, bounded conditional densities and finite first moments suffice for the expected regret of any price ππ to satisfy E[g(m,V,W)−g(π,V,W)]≤L∣m−π∣2\mathbb{E}[g(m,V,W) - g(π,V,W)] \le L|m-π|^2 -- no second moment is needed.
Jan 20, 2026stat.ML

Small Gradient Norm Regret for Online Convex Optimization

This paper introduces a new problem-dependent regret measure for online convex optimization with smooth losses. The notion, which we call the G⋆G^\star regret, depends on the cumulative squared gradient norm evaluated at the decision in hindsight. We show that the G⋆G^\star regret strictly refines the existing L⋆L^\star (small loss) regret, and that it can be arbitrarily sharper when the losses have vanishing curvature around the hindsight decision. We establish upper and lower bounds on the G⋆G^\star regret and extend our results to dynamic regret and bandit settings. As a byproduct, we refine the existing convergence analysis of stochastic optimization algorithms in the interpolation regime. Some experiments validate our theoretical findings.
Jan 5, 2026cs.LG

Prior Diffusiveness and Regret in the Linear-Gaussian Bandit

We prove that Thompson sampling exhibits O~(σdT+drTr(Σ0))\tilde{O}(σd \sqrt{T} + d r \sqrt{\mathrm{Tr}(Σ_0)}) Bayesian regret in the linear-Gaussian bandit with a N(μ0,Σ0)\mathcal{N}(μ_0, Σ_0) prior distribution on the coefficients, where dd is the dimension, TT is the time horizon, rr is the maximum ℓ2\ell_2 norm of the actions, and σ2σ^2 is the noise variance. In contrast to existing regret bounds, this shows that to within logarithmic factors, the prior-dependent burn-in'' term $d r \sqrt{\mathrm{Tr}(Σ_0)}$ decouples additively from the minimax (long run) regret $σd \sqrt{T}$. Previous regret bounds exhibit a multiplicative dependence on these terms. We establish these results via a new elliptical potential'' lemma, and also provide a lower bound indicating that the burn-in term is unavoidable.
Nov 29, 2025stat.ML

No-Regret Gaussian Process Optimization of Time-Varying Functions

Sequential optimization of black-box functions from noisy evaluations has been widely studied, with Gaussian Process bandit algorithms such as GP-UCB guaranteeing no-regret in stationary settings. However, for time-varying objectives, no-regret is unattainable under pure bandit feedback unless strong and often unrealistic assumptions are imposed. We propose a novel method for optimizing time-varying rewards in the frequentist setting, where the objective has bounded RKHS norm almost surely. Time variations are captured through uncertainty injection, enabling heteroscedastic Gaussian process regression that adapts past observations to the current time step. As no-regret is unattainable in general in the strict bandit setting, we relax the latter allowing additional queries on previously observed points. Building on sparse inference and the effect of uncertainty injection on regret, we propose W-SparQ-GP-UCB, an online algorithm that achieves no-regret with a vanishing number of additional queries per iteration. To assess the theoretical limits of this approach, we establish a lower bound on the number of additional queries required for no-regret, proving the efficiency of our method. Finally, we provide a comprehensive analysis linking the temporal regime of the function to achievable regret rates, together with upper and lower bounds on the number of additional queries needed in each regime.
Jul 9, 2025cs.LG

Direct Regret Optimization in Bayesian Optimization

Bayesian optimization (BO) is a powerful paradigm for optimizing expensive black-box functions. Traditional BO methods typically rely on separate hand-crafted acquisition functions and surrogate models for the underlying function, and often operate in a myopic manner. In this paper, we propose a novel direct regret optimization approach that jointly learns the optimal model and non-myopic acquisition by distilling from a set of candidate models and acquisitions, and explicitly targets minimizing the multi-step regret. Our framework leverages an ensemble of Gaussian Processes (GPs) with varying hyperparameters to generate simulated BO trajectories, each guided by an acquisition function drawn from a pool of conventional choices and terminated by a Bayesian early stop criterion. These trajectories train an end-to-end Decision Transformer that selects the next query so as to improve the ultimate objective, following a dense training sparse learning paradigm: the transformer is trained on abundant simulated data, while a limited number of real evaluations refine the GPs online. On synthetic and real-world benchmarks, our method attains the best or near-best final simple regret against standard, lookahead, trust-region and amortized BO baselines, with the largest gains in high-dimensional settings. Ablations attribute the gains jointly to region-of-interest filtering and the learned policy, and matched-budget comparisons against explicit two-step lookahead acquisitions show that the advantage is not shared by lookahead alone.
Aug 1, 2024cs.DS

Infrequent Resolving Algorithm for Online Linear Programming

Online linear programming (OLP) has gained significant attention from both researchers and practitioners due to its extensive applications such as online auctions, network revenue management, order fulfillment and advertising. Existing OLP algorithms fall into two categories: LP-based algorithms and LP-free algorithms. The former typically guarantees better performance but requires solving a large number of LPs, which could be computationally expensive. In contrast, LP-free algorithms only require first-order computations but induce a worse performance. In this work, we bridge the gap between these two extremes by proposing a well-performing algorithm that solves LPs at a few selected time points and conducts first-order computations at other time points. Specifically, for the case where the inputs are drawn from an unknown finite-support distribution, the proposed algorithm achieves a constant regret (even for the hard "degenerate" case) while solving LPs only O(log⁡log⁡T)O(\log\log T) times over the time horizon TT. Moreover, when we are allowed to solve LPs only MM times, we design the corresponding schedule such that the proposed algorithm can guarantee a nearly O(T(1/2)M−1)O\left(T^{(1/2)^{M-1}}\right) regret. Our work highlights the value of resolving both at the beginning and the end of the selling horizon, and provides a novel framework to prove the performance guarantee of the proposed policy under different infrequent resolving schedules. Numerical experiments are conducted to demonstrate the efficiency of the proposed algorithms.
Aug 1, 2024cs.LG

Online Linear Programming with Batching

We study Online Linear Programming (OLP) with batching. The planning horizon is cut into KK batches, and decisions on orders can be delayed to the end of their associated batch. The ability to delay decisions improves operational performance, as measured by regret. We study two questions: (1) What is a lower bound on the regret as a function of KK and the length of the planning horizon? (2) Which algorithms can achieve this regret lower bound? This paper analyzes these questions when the distribution of the reward has a continuous support. We provide an Ω(log⁡K)Ω(\log K) regret lower bound in the single-resource case, and we provide pricing algorithms having an O(log⁡K)O(\log K) regret in the setting with Poisson arrivals and multiple types of resources. All the algorithms update the prices at most KK times and only delay orders of the first and the last batches. All the regret bounds are independent of the length of the planning horizon. Finally, we study a more realistic large support setting where the number of distinct order types is finite but scales with the total number of orders. We prove that our Θ(log⁡K)Θ(\log K) bounds still hold for the batching operation of a multisecretary problem with discretized uniform rewards in this large support setting, provided the support grows fast enough. This suggests that the continuous support setting that is the paper's focus serves as a useful theoretical surrogate for the more realistic large finite support setting.
Jun 5, 2023cs.GT

Calibrated Stackelberg Games: Learning Optimal Commitments Against Calibrated Agents

We introduce \emph{Calibrated Stackelberg Games (CSGs)}, a generalization of the standard Stackelberg Games (SGs) framework. In CSGs, a principal repeatedly interacts with an agent who (contrary to standard SGs) does not have direct access to the principal's action but instead best-responds to calibrated forecasts about it. This framework provides a powerful and realistic modeling tool that goes beyond assuming that agents use ad hoc and highly specified algorithms for interacting in strategic settings and instead builds on statistical foundations of forecasts and calibration. We show that in CSGs, despite both the principal and the agent having less information than in standard SGs, the principal's optimal utility remains upper and lower bounded by the Stackelberg value of the one-shot game, in both finite and continuous settings. Alongside CSGs, we develop stronger notions of calibration and corresponding algorithms that address two central challenges for calibration in game-theoretic environments. First, achieving point-wise calibration typically incurs an error that scales exponentially with the dimension of the strategy space. Second, the principal's convergence rate in CSGs depends critically on the adaptivity of the agent's calibration algorithm. To address these challenges, we establish a meaningful, efficiently achievable relaxation of calibration based on conditioning on best-response regions. This yields the first notion of calibration in games with a statistical rate that only depends on the number of agents' actions rather than the dimension of the principal's strategy space and that leads to no-swap regret for the agent. We further develop adaptive calibration algorithms for the agents that provide fine-grained, any-time calibration guarantees against adversarial sequences, enabling the principal to achieve faster convergence in CSGs.
May 27, 2023cs.LG

Hierarchical Deep Counterfactual Regret Minimization

Imperfect Information Games (IIGs) are used to model games under uncertainty or lack complete information. Counterfactual Regret Minimization (CFR) is one of the most successful families of algorithms for IIGs. The integration of skill-based strategy learning with CFR could potentially mirror more human-like decision-making and improve learning on complex IIGs. It enables the learning of a hierarchical strategy, wherein low-level components represent skills for solving subgames and the high-level component manages the transition between skills. In this paper, we introduce the first hierarchical version of Deep CFR (HDCFR), an innovative method that boosts learning efficiency in tasks involving extensively large state spaces and deep game trees. Notably, HDCFR enables learning with predefined (human) expertise and extracting skills transferable to similar tasks. We first present the algorithm and establish its theory in a tabular setting, including hierarchical CFR update rules and a variance-reduced Monte Carlo sampling extension for the model-free setting, where backtracking is infeasible. We then extend HDCFR to large-scale tasks via deep learning objectives that match the tabular targets under exact function fitting. Code: https://anonymous.4open.science/r/HDCFR_RUN-677B.
Feb 3, 2023cs.LG

Robust Budget Pacing with a Single Sample

Major Internet advertising platforms offer budget pacing tools as a standard service for advertisers to manage their ad campaigns. Given the inherent non-stationarity in an advertiser's value and also competing advertisers' values over time, a commonly used approach is to learn a target expenditure plan that specifies a target spend as a function of time, and then run a controller that tracks this plan. This raises the question: how many historical samples are required to learn a good expenditure plan? We study this question by considering an advertiser repeatedly participating in TT second-price auctions, where the tuple of her value and the highest competing bid is drawn from an unknown time-varying distribution. The advertiser seeks to maximize her total utility subject to her budget constraint. Prior work has shown the sufficiency of Tlog⁡TT\log T samples per distribution to achieve the optimal O(T)O(\sqrt{T})-regret. We dramatically improve this state-of-the-art and show that just one sample per distribution is enough to achieve the near-optimal O~(T)\tilde O(\sqrt{T})-regret, while still being robust to noise in the sampling distributions.
Date pendingstat.ML

Satisficing Regret Minimization in Bandits: Constant Rate and Light-Tailed Distribution

Motivated by the concept of satisficing in decision-making, we consider the problem of satisficing regret minimization in bandit optimization. In this setting, the learner aims at selecting satisficing arms (arms with mean reward exceeding a certain threshold value) as frequently as possible. The performance is measured by satisficing regret, which is the cumulative deficit of the chosen arm's mean reward compared to the threshold. We propose SELECT, a general algorithmic template for Satisficing REgret Minimization via SampLing and LowEr Confidence bound Testing, that attains constant expected satisficing regret for a wide variety of bandit optimization problems in the realizable case (i.e., a satisficing arm exists). As a complement, SELECT also enjoys the same (standard) regret guarantee as the oracle in the non-realizable case. To further ensure stability of the algorithm, we introduce SELECT-LITE that achieves a light-tailed satisficing regret distribution plus a constant expected satisficing regret in the realizable case and a sub-linear expected (standard) regret in the non-realizable case. Notably, SELECT-LITE can operate on learning oracles with heavy-tailed (standard) regret distribution. More importantly, our results reveal the surprising compatibility between constant expected satisficing regret and light-tailed satisficing regret distribution, which is in sharp contrast to the case of (standard) regret. Finally, we conduct numerical experiments to validate the performance of SELECT and SELECT-LITE on both synthetic datasets and a real-world dynamic pricing case study.