Sampling Allocation of LinUCB: Optimal Design Limits in the Small-Gap Regime
Organizations: National University of Singapore
Abstract
We study the sampling allocation of LinUCB in the small-gap regime, where the reward gaps are of order at most over the decision horizon . This scaling captures the hard instances underlying worst-case regret lower bounds, for which LinUCB is known to be near optimal up to logarithmic factors in . Using a mean-field perspective, we characterize this allocation through the empirical sampling distribution, a macroscopic object that averages the effect of adaptive decisions over the horizon, and identify its limit as . We establish that in this regime, the empirical sampling distribution induced by LinUCB converges to the set of D-optimal designs. This central result reveals that, in the small-gap regime, LinUCB not only achieves near optimal minimax regret but also allocates samples in a way that is asymptotically efficient for learning the reward parameter, thereby connecting regret-driven online learning with information-efficient experimental design. Building on the optimal design limit, we obtain two useful consequences. First, we refine the asymptotic regret analysis of LinUCB in the small-gap regime by characterizing its leading-order constant in the limit. Second, we show that, despite LinUCB's adaptive sampling strategy, the regularized least-squares estimator satisfies a central-limit-type theorem in the small-gap regime, thereby enabling valid statistical inference for the reward parameter.
Figures & tables
| Mean-field statistical mechanics | LinUCB in small-gap regime | |
| Mean-field viewpoint | Average many interacting microscopic degrees of freedom into a macroscopic object | Average adaptive sampling decisions over the decision horizon into the empirical sampling distribution |
| Microscopic variables | Microscopic component states, e.g., | Adaptive sampling decisions |
| Macroscopic object | An order parameter or empirical observable, e.g., | The empirical sampling distribution |
| System size | Number of components | Decision horizon |
| Fixed scaling quantity | Keep suitable intensive parameters fixed, such as density or temperature | Small-gap regime: reward gaps scale with the horizon according to |
| Potential function | A variational objective characterizing the macroscopic state | The D-optimality criterion , where is given in ( 4 ) |