cs.LGOct 2, 2026

When May a Bandit Leave Its Anchor? E-Process-Authorized Thompson Sampling under Non-stationarity

Authors: Mayand Gulati, Kerong Wang, WeiChen Au

Organizations: UC Santa Barbara · Purdue University

Abstract

Stationarity rewards memory, but after a change the same history can mislead. We ask when forgetting should be permitted. E-process-authorized Thompson sampling (e-ATS) gives each arm full-history and discounted Beta states. An anytime-valid e-process first authorizes the discounted state, then a reversible relevance score controls its influence. Before authorization, e-ATS exactly follows optimistic Thompson sampling (OTS). Under a Beta-Bernoulli prior-predictive stationary model, e-ATS's probability of ever departing from OTS is at most the chosen αEα_E, without fitted thresholds. Relative to e-ATS, removing authorization increased mean normalized dynamic pseudo-regret by 38.4%38.4\% on the registered suite but reduced it by 7.5%7.5\% on the literature-derived replay suite. Therefore, evidence controls when adaptation begins, not whether it always helps.

Explore similar work

CardsList
  1. Flow-Corrected Thompson Sampling for Non-Stationary Contextual Bandits

    Jun 22, 2026AmirHossein Naghdi, Ali BaheriThompson SamplingNon-Stationary Bandits

  2. Future Information-Directed Sampling for Bayesian Nonstationary Bandits

    Sep 27, 2026Yichen Song, Alessio Russo, Aldo PacchianoMulti-Armed BanditsNon-Stationary Bandits