cs.LGSep 27, 2026

Future Information-Directed Sampling for Bayesian Nonstationary Bandits

Authors: Yichen Song, Alessio Russo, Aldo Pacchiano

Organizations: Boston University · Broad Institute of MIT and Harvard

Abstract

Exploration--exploitation is a central trade-off in bandit learning. While classical algorithms such as upper confidence bound methods and Thompson Sampling effectively balance this trade-off in stationary environments, their exploration strategies mainly reduce uncertainty about the current optimal arm, which can be insufficient in nonstationary settings where future optimal arms may differ substantially from current ones. In this paper, we propose Future Information-Directed Sampling (FIDS), a new algorithm for Bayesian nonstationary bandits that explicitly explores to gather information about future optimal arms. We show that FIDS achieves regret comparable to Thompson Sampling up to a small constant factor, while being able to exploit predictive information structures that conventional exploration objectives fail to capture. To address the practical difficulty of posterior inference, we further propose a supervised-learning-based approximation framework that learns the FIDS policy from offline data, and demonstrate its effectiveness on synthetic benchmarks.

Figures & tables

Explore similar work

CardsList
  1. Flow-Corrected Thompson Sampling for Non-Stationary Contextual Bandits

    Jun 22, 2026AmirHossein Naghdi, Ali BaheriThompson Sampling

  2. Sharp Characterization of Bias in Post-Bandit Inference

    Aug 2, 2026Lisu Wang, Yilun Chen, Jiaqi LuStochastic Multi-Armed BanditsStochastic Exploration

  3. PFN-TS: Thompson Sampling for Contextual Bandits via Prior-Data Fitted Networks

    May 11, 2026Yan Shuo Tan, Kenyon Ng, Ruizhe Deng +3Thompson SamplingTabular Prior-Data Fitted Network