cs.LGFeb 11, 2026

Rising Multi-Armed Bandits with Known Horizons

Authors: Seockbean Song, Chenyu Gan, Youngsik Yoon, Siwei Wang, Wei Chen, Jungseul Ok

Organizations: Graduate School of AI, POSTECH, Pohang, Republic of Korea · Qiuzhen College, Tsinghua University, Beijing, China · Department of CSE, POSTECH, Pohang, Republic of Korea · Microsoft Research, Beijing, China

Abstract

Rising Multi-Armed Bandits (RMABs) model sequential decision problems where each arm's expected reward improves with repeated pulls. In such problems, the value of investing in an arm depends on how much time remains, making knowledge of the horizon useful side information, yet its benefit remains underexplored. We investigate this benefit through CURE-UCB, a horizon-aware algorithm that estimates each arm's cumulative reward over the remaining horizon. Theoretically, under structured assumptions, we prove that CURE-UCB uniformly dominates a representative horizon-agnostic algorithm and show that the advantage of horizon awareness can be substantial: on some instances, CURE-UCB incurs only O(1)O(1) regret whereas the horizon-agnostic algorithm suffers Ω(T)Ω(T). Furthermore, we establish a regret upper bound for the general concave rising bandit setting whose growth-dependent term matches the known lower bound in its dependence on TT. Empirically, across synthetic benchmarks and real-world model selection tasks, CURE-UCB achieves lower regret than both rising and non-stationary baselines over a wide range of horizons.

Figures & tables

Appendix figures & tables7 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Multi-Armed Bandits with Arriving Arms: Sequential Screening, Dynamic Regret, and Sublinear Guarantees

    Jun 8, 2026Deqi Zheng, Xiaoyang Xu, Yuhong YangStochastic Multi-Armed BanditsScreening

  2. The Greedy Advantage in Finite-Horizon Bandits

    Jul 31, 2026Kai Zhou, Michael Lingzhi Li, Kai WangStochastic Multi-Armed BanditsInterval Regret

  3. Trading off rewards and errors in multi-armed bandits

    May 1, 2026Akram Erraqabi, Alessandro Lazaric, Michal Valko +2Stochastic Multi-Armed BanditsRegret