cs.LGAug 2, 2026

Sharp Characterization of Bias in Post-Bandit Inference

Authors: Lisu WangYilun ChenJiaqi Lu

Abstract

Bandit algorithms generate data for downstream inference, but adaptive sampling biases post-bandit sample means. We analyze this bias for stable index algorithms, including UCB1 and its generalizations, and derive sharp leading-order expressions for the sample-mean bias and expected ZZ-statistic, in bandit experiments of fixed horizon TT. Our characterization reveals the algorithmic origin of bias through a key index-function-dependent quantity, which we term effective exploration rate. For example, under UCB1, the effective exploration rate is of order logT\sqrt{\log T}, and the standardized bias of any arm (that is not uniquely optimal) decays at the extremely slow rate 1/logT1/\sqrt{\log T}. We also show how the choice of the index function affects both regret and bias, which reveals a regret-bias trade-off: more exploratory algorithm reduces bias but increases regret. We further show how bias most severely distorts confidence intervals and hypothesis tests when the tested arm is one of the tied-optimal arms. Our sharp characterization for bias uses a novel empirical fluid approximation of the algorithm's sampling dynamics, which may be of independent interest.

Explore similar work

CardsList
  1. Optimal Regret for Single Index Bandits

    May 10, 2026Devdan Dey, Sujoy Bhore, Avishek GhoshMulti-Armed BanditsRegret

  2. Bandit Simulation for Average Reward Inference

    May 30, 2026Samya Praharaj, Chih-Yu Chang, Koulik Khamaru +1Confidence IntervalsBandits

  3. Price of Fairness in Bandits: A Tight Minimax Characterization

    Jul 15, 2026Dhruv Sarkar, Soumyadeep Dutta, Sayak Ray ChowdhuryLinear RegretBandits