stat.MLDec 12, 2024

Allocation Stability and Wald Inference under Variance-Aware UCB

Authors: Yingying Fan, Yuxuan Han, Jinchi Lv, Xiaocong Xu, Zhengyuan Zhou

Abstract

Allocation stability is often used to justify Gaussian inference from bandit data, but when is it necessary? In this paper, we address this question for a two-armed, fixed-horizon variance-aware UCB policy with bounded reward distributions that may vary with the horizon. We find a sharp criterion in terms of the reward gap and variances that determines whether the optimal-arm count admits a deterministic approximation with vanishing relative error, while the suboptimal-arm count is always stable. Despite the possible instability of the optimal-arm count, we show that the ordinary Wald statistic for a linear combination of the arm means has a standard normal limit for every fixed nonzero coefficient vector, provided the product of the pull count and reward variance diverges in probability for each arm. Under the same condition, however, this Gaussian approximation holds uniformly over deterministic nonzero coefficient vectors if and only if the optimal-arm count is stable. The analysis relies on two main ingredients: (i) a pathwise comparison with an auxiliary policy whose final optimal-arm count is asymptotically equivalent to the original count and independent of the optimal-arm reward sequence; and (ii) joint limits for the rescaled optimal-arm count and the two studentized sample-mean errors under the original policy, which yield nonstandard Wald limits for certain linear combinations of the arm means with coefficients that vary with the horizon.

Explore similar work

CardsList
  1. Sharp Characterization of Bias in Post-Bandit Inference

    Aug 2, 2026Lisu Wang, Yilun Chen, Jiaqi LuStochastic Multi-Armed BanditsStochastic Exploration

  2. Statistical Inference for Misspecified Contextual Bandits

    Jun 21, 2026Yongyi Guo, Ziping XuContextual Bandit FrameworkStatistical Inference

  3. Sampling Allocation of LinUCB: Optimal Design Limits in the Small-Gap Regime

    Oct 5, 2026Yujie Liu, Vincent Y. F. Tan, Yunbei Xu