stat.MLOct 1, 2026

Block Optimism for Nonstationary Bandits with Latent Linear Dynamics

Authors: Taehyun Hwang, Hyunjun Choi, Heesang Ann, Min-hwan Oh

Organizations: Seoul National University Soeul, South Korea

Abstract

We study an endogenous nonstationary stochastic bandit problem with latent linear dynamics, where actions affect both immediate rewards and the future evolution of an unobserved latent state. Rewards are bilinear in the current action and latent state, inducing history-dependent rewards and a nontrivial long-horizon planning problem. The existing explore-then-commit approach achieves O~(T2/3)\tilde{O}(T^{2/3}) regret by uniformly exploring to estimate the latent dynamics and then committing to an optimized open-loop action sequence. We show that this rate can be improved via adaptive block-level optimism. Our key step is a cyclic approximation: under stable dynamics, the infinite-memory reward process can be truncated, and the open-loop benchmark can be approximated by optimizing a finite-memory block-level proxy. Building on this reduction, we propose a UCB-based block algorithm that maintains confidence sets for the truncated dynamics parameters and selects blocks optimistically. We prove a regret bound of order O~(T)\tilde{O}(\sqrt T), significantly improving over the previous O~(T2/3)\tilde{O}(T^{2/3}) guarantee for the same model. To the best of our knowledge, this is the first O~(T)\tilde{O}(\sqrt T) regret guarantee for latent linear-dynamics bandits with bilinear reward observations and an open-loop action-sequence benchmark.

Figures & tables

Appendix figures & tables1 asset

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Dynamic Regret for Non-Stationary Linear Bandits via Misspecification Reductions

    Jul 3, 2026Zihao Hu, Yuan Yao, Jiheng Zhang +1Contextual Bandit FrameworkRegret

  2. Bellman-Centric Learning: Near-Optimal Regret for Linear Bandits with Memory

    Oct 5, 2026Jingyuan Liu, Huiwen Jia

  3. Offline-to-Online Learning in Linear Bandits

    Jun 3, 2026Kushagra Chandak, Toshinori Kitamura, Xiaoqi TanContextual Bandit Framework