cs.LGMay 22, 2026

Prudent-Banker: No Extra Fees for Baseline Safety in Adversarial Bandits With and Without Delays

Authors: Ting Hu, Luanda Cai, Emmanouil-Vasileios Vlatakis-Gkaragkounis

Organizations: Department of Finance University of Wisconsin–Madison · Department of Computer Sciences University of Wisconsin–Madison

Abstract

We study adversarial multi-armed bandits with and without delayed feedback under a safety-aware goal: achieving minimax-optimal worst-case regret while keeping nearly constant regret relative to a designated "safe" baseline policy. Existing approaches can balance this trade-off with immediate feedback for smooth comparators, but arbitrary delays can mistime transitions between conservatism and exploration, endangering the safety guarantee. To bridge this gap, we propose Prudent-Banker, a novel algorithm that combines a delay-adapted variant of Online Mirror Descent with a modified phased-aggression mechanism. Its key technical contribution is a delay-calibrated restart threshold that rigorously accounts for the worst-case distortion induced by unobserved feedback and reliably detects comparator suboptimality. We also establish new lower bounds for safety-constrained adversarial delayed bandits, showing that the regret guarantees of Prudent-Banker are unimprovable, up to logarithmic factors, under the baseline-safety requirement. To the best of our knowledge, Prudent-Banker is the first algorithm to achieve the optimal safety--robustness trade-off: pseudo-regret O~(T+D)\widetilde{O}(\sqrt{T}+\sqrt{D}) together with O~(1)\widetilde{O}(1) regret against the safe comparator, both with and without delays. Experiments across diverse delay distributions show that, unlike standard delay-robust baselines, Prudent-Banker effectively balances safety and learning.

Explore similar work

CardsList
  1. Near-Optimal Stochastic Linear Bandits with Delay

    Jun 15, 2026Ofir Schlisselberg, Mengxiao Zhang, Yishay MansourLinear BanditsNear-Optimal Regret Guarantees

  2. Trading off rewards and errors in multi-armed bandits

    May 1, 2026Akram Erraqabi, Alessandro Lazaric, Michal Valko +2Multi-Armed BanditsRegret