cs.LGOct 26, 2025

Managing Self-Learning Experts under Per-Round Budget Constraints

Authors: Ilgam Latypov, Alexandra Suvorikova, Alexey Kroshnin, Alexander Gasnikov, Yuriy Dorn

Organizations: AI Center, Lomonosov Moscow State University MSU Institute for Artificial Intelligence Moscow, Russia · Weierstrass Institute for Applied Analysis and Stochastics Berlin, Germany · IITP RAS Moscow, Russia · Steklov Mathematical Institute of RAS Moscow, Russia

Abstract

This paper addresses the problem of sequential decision-making under learning budget constraints. Such settings naturally arise in applications like managing a portfolio of bandit or reinforcement learning (RL) algorithms. We propose a novel UCB-type algorithm, M-LCB, designed to manage a pool of KK self-learning experts in a stochastic environment while accounting for a limited per-round learning budget MM. At each round, M-LCB selects one expert to make a decision and at most M≤KM \le K experts to learn. For selection, M-LCB uses confidence bounds constructed from limited prior knowledge about the experts (i.e., mild assumptions) and their observed training losses. We derive anytime regret bounds for M-LCB that scale with the individual regrets of the experts. In particular, if each expert has regret O~(Tα)\tilde O(T^α) by round TT, then M-LCB guarantees an overall regret of O~(KT/M+(K/M)1−αTα)\tilde O\left(\sqrt{KT/M} + (K/M)^{1-α}T^α\right) relative to the best expert in hindsight. Finally, we demonstrate the applicability of M-LCB using self-learning experts instantiated as (i) parametric models and (ii) bandit algorithms.

Figures & tables

Appendix figures & tables7 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Prediction with Expert Advice: Anytime Regret with Many Experts Matches the Fixed-Time Constant

    Sep 23, 2026Yang Cai, Vineet Gupta, Yanchen Jiang +4

  2. Self-Concordant Perturbations for Linear Bandits

    Oct 28, 2025Lucas Lévy, Jean-Lou Valeau, Arya Akhavan +1Linear BanditsEfficient Exploration

  3. Learning in Markovian bandits with non-observable states and constrained decision epochs

    Jun 25, 2026Thomas Hira, Victor Boone, Urtzi Ayesta +1O(T^Β)$ Simultaneous RegretMarkov Decision Processes