We study contextual bandits in which a surrogate is observed after the action but before the learner decides whether to acquire the primary outcome that defines action value and regret. The value of acquiring the primary outcome depends on both decision relevance (how much the current outcome matters for comparing policies) and the residual uncertainty after observing the surrogate. The Audited Surrogate Bandit (ASB) learns a contextual policy while allocating a budget of B primary-outcome acquisitions over T rounds. ASB sets a pre-surrogate acquisition level from current decision relevance and, after observing the surrogate, redistributes that level using an estimate of that residual uncertainty. For a finite class of N policies over K actions, ASB incurs O[KTlogN{1+T/B}] regret relative to the best policy in the class. In a two-action family where the surrogate does not reveal the better action, a learner that observes the surrogate before deciding whether to acquire can achieve bounded regret, whereas any learner that must decide before seeing the surrogate incurs Ω(T/B) worst-case regret under the same budget. Synthetic experiments show that both acquisition factors matter: ASB has lower regret than variants using only decision relevance or only residual uncertainty. On a KuaiRec benchmark of user-video interactions, the regret gap relative to relevance-only acquisition widens and then narrows as the budget grows.
Figures & tables
Figure 1: Synthetic tests of acquisition structure. Left: cumulative regret for ASB , Relevance only, Residual only, Uniform, and Fixed-N under the same exact primary-outcome budget. Middle: final cumulative regret for matched PRE and POST oracles as λ increases the surrogate’s information about residual magnitude, from none at λ=0 to exact identification of the high-residual state at λ=1 . Right: at each fixed average residual-information level, information is shifted from policy-agreement contexts (Away), through equal information in the two regions (Neutral), to policy-disagreement contexts (Toward); positive values mean lower regret than neutral placement. Bands show mean ±1 SE over 64 matched environments.
Figure 2: Post-surrogate refinement on KuaiRec. Left: paired cumulative-regret difference, Relevance only minus ASB . Right: paired exact terminal policy-value difference, ASB minus Relevance only, evaluated after the online run using the complete benchmark table. Positive values favor ASB in both panels. Error bars show mean ±1 SE over 32 paired order/action randomizations and 8 acquisition randomizations per order/action stream.