cs.LGOct 7, 2026

Budgeted Multi-Source Counterfactual Annotation for Off-Policy Evaluation

Authors: Biao Xiang, Ali Eshragh, Yuexing Li, Kai Wang

Organizations: Pennsylvania State University, State College, PA · Johns Hopkins Carey Business School, Washington, DC · International Computer Science Institute, Berkeley, CA · Georgia Institute of Technology, Atlanta, GA

Abstract

Off-policy evaluation (OPE) estimates the value of a target policy from logged data, but limited behavior-policy coverage can force high-variance reweighting or reward-model extrapolation. Counterfactual annotations can add evidence about unobserved actions, yet practical sources, including domain experts and large language models (LLMs), may be costly, biased, or noisy. We study budgeted acquisition of such annotations for contextual-bandit OPE. Given source-specific costs and error profiles, we formulate an integer allocation problem over context-action pairs and annotation sources to minimize the component of estimator variance that depends on the annotation plan. We characterize when annotations are valuable through a first-annotation threshold and local annotation-value regimes. For the coupled multi-source problem, we develop a majorization-minimization algorithm with dynamic-programming subroutines that monotonically improves the objective. Experiments in synthetic clinical and LLM-annotated education bandits show that our allocation method reduces fixed-profile mean squared error (MSE) by 20.58% and 10.77%, respectively, relative to no annotation.

Figures & tables

Appendix figures & tables11 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Optimal Sequential Annotations for Off-Policy Evaluation

    Sep 22, 2026Woojin Chae, Ezinne Nwankwo, Haitong Qin +1Off-Policy EvaluationLearning with Missing Data

  2. Cross-Domain Off-Policy Evaluation and Learning for Contextual Bandits

    Jul 24, 2026Yuta Natsubori, Masataka Ushiku, Yuta SaitoOff-Policy EvaluationContextual Bandits

  3. Variance-Optimal Off-Policy Evaluation with Conjunct Effect Modeling

    Oct 6, 2026Nicolò Felicioni, Michael Benigni, Maurizio Ferrari Dacrema +1Causal Effect EstimationOff-Policy Evaluation