cs.IRSep 28, 2026

Recommendation Ranking Off-Policy Evaluation under Ranking-Dependent Examination via Examination-Relevance Decomposition

Authors: Riki Okamura, Toshiharu Sugawara

Organizations: Department of Computer Science and Communications Engineering, Waseda University, Tokyo, Japan

Abstract

Off-policy evaluation, which estimates evaluation policy performance from logged data, is key for recommender ranking policies. However, logged clicks cannot distinguish unexamined items from examined non-clicks, causing bias in existing estimators when the assumed examination structures fail. We propose two estimators based on the decomposition of clicks into examination and relevance. First, the latent-examination independent inverse propensity score (LE-IIPS) estimator corrects the IIPS bias using policy examination probability ratios. Second, the examination-decomposed doubly robust (ED-DR) estimator extends LE-IIPS to a doubly robust framework. ED-DR is unbiased if the examination probabilities are correct regardless of relevance accuracy, or under ranking-independent examination, even if both model estimates are inaccurate. Experiments show that ED-DR achieves a lower MSE than existing methods with large sample sizes, especially when the examination depends on ranking. We also highlight its limitations under small samples or cascade user behavior conditions.

Explore similar work

CardsList
  1. Quotient DAGs for Off-Policy Evaluation:Forward-Flow Importance Sampling and Exact Slate Propensities

    May 28, 2026Ziwen Xie, Shaowen Xiang, Hongyu He +1Policy EvaluationImportance

  2. Logging Policy Design for Off-Policy Evaluation

    May 14, 2026Connor Douglas, Joel Persson, Foster ProvostOff-Policy LearningTreatment Allocation