cs.LGMar 5, 2026

Autocorrelation effects in a stochastic-process model for solving two-armed bandit problems

Authors: Tomoki YamagamiMikio HasegawaTakatomo MihanaRyoichi HorisakiAtsushi Uchida

Organizations: Department of Information and Computer Sciences, Saitama University, 255 Shimo-Okubo, Sakura-ku, Saitama City, Saitama 338–8570, Japan. · Department of Electrical Engineering, Tokyo University of Science, 6–3–1, Niijuku, Katsushika-ku, Tokyo 125–8585, Japan. · Department of Information Physics and Computing, The University of Tokyo, 7–3–1 Hongo, Bunkyo-ku, Tokyo 113–8656, Japan.

Abstract

Decision makers exploiting photonic chaotic dynamics obtained by semiconductor lasers provide an ultrafast approach to solving multi-armed bandit problems by using a temporal optical signal as the driving source for sequential decisions. In such systems, the sampling interval of the chaotic waveform shapes the temporal correlation of the resulting time series, and experiments have reported that decision accuracy depends strongly on this autocorrelation property. However, it remains unclear whether the benefit of autocorrelation can be explained by a minimal mathematical model. Here, we analyze a stochastic-process model for solving the two-armed bandit problem based on time series, where the threshold and a two-valued Markov signal evolve jointly. Numerical results reveal an environment-dependent structure: negative (positive) autocorrelation is optimal in reward-rich (reward-poor) environments. These findings show that negative autocorrelation of the time series is advantageous when the sum of the winning probabilities is more than one, whereas positive autocorrelation is useful when the sum of the winning probabilities is less than one. Moreover, the performance is independent of autocorrelation if the sum of the winning probabilities equals one, which is mathematically clarified. This study paves the way for solving the two-armed bandit problems for reinforcement learning applications in wireless communications and robotics.

Explore similar work

CardsList
  1. Best of both worlds: Stochastic & adversarial best-arm identification

    Apr 16, 2026Yasin Abbasi-Yadkori, Peter L. Bartlett, Victor Gabillon +2BanditsStochastic

  2. Trading off rewards and errors in multi-armed bandits

    May 1, 2026Akram Erraqabi, Alessandro Lazaric, Michal Valko +2Multi-Armed BanditsRegret

  3. Near-Optimal Stochastic Linear Bandits with Delay

    Jun 15, 2026Ofir Schlisselberg, Mengxiao Zhang, Yishay MansourLinear BanditsNear-Optimal Regret Guarantees