Jan 5, 2026 · cs.LGJ/K move · Enter open · S save
Yifan Zhu, John C. Duchi, Benjamin Van Roy
Department of Electrical Engineering, Stanford University, California, United States.
We prove that Thompson sampling exhibits
O~(σdT+drTr(Σ0)) Bayesian regret in the linear-Gaussian bandit with a
N(μ0,Σ0) prior distribution on the coefficients, where
d is the dimension,
T is the time horizon,
r is the maximum
ℓ2 norm of the actions, and
σ2 is the noise variance. In contrast to existing regret bounds, this shows that to within logarithmic factors, the prior-dependent
burn-in'' term $d r \sqrt{\mathrm{Tr}(Σ_0)}$ decouples additively from the minimax (long run) regret $σd \sqrt{T}$. Previous regret bounds exhibit a multiplicative dependence on these terms. We establish these results via a new elliptical potential'' lemma, and also provide a lower bound indicating that the burn-in term is unavoidable.