cs.LGOct 7, 2026

Amortized Off-Policy Evaluation for LLMs

Authors: Younwoo Choi, Leo Feng, Vincent Liu, Haanvid Lee

Organizations: RBC Borealis · University of Toronto · Vector Institute

Abstract

Accurate evaluation is central to selecting which LLM to deploy, yet testing a candidate on live traffic exposes real users to an unvetted model. Teams therefore evaluate candidates offline, on data produced by already-deployed models. This is off-policy evaluation (OPE), and it faces two distribution shifts: as a model is updated in post-training, its responses diverge from the logged ones (policy shift), and the reward definition under which it is judged changes with business requirements (reward shift). Classical OPE methods are ill-suited to this continual-deployment setting because they are defined per task and require fitting from scratch on every new logged dataset or reward definition. To address this, we propose PFN-OPE, a prior-data fitted network that amortizes OPE across a distribution of contextual-bandit tasks. We pretrain it once on tasks constructed from a pool of LLM responses scored by several reward functions, in which both shifts occur. At test-time it maps a logged dataset and one sampled target response per prompt to a value estimate in a single forward pass, with no per-task fitting. On HelpSteer2 and UltraFeedback with Qwen, Llama, and Gemma policies, PFN-OPE achieves 2.0 to 9.3 times lower error than the best baselines across all tested configurations in the reward-shifted settings.

Figures & tables

Explore similar work

CardsList
  1. OTROPE: Optimal Transport-based Robust Off-policy Evaluation for Large Language Models

    Sep 28, 2026Liner Xiang, Wenbo Zhang, Hengrui CaiLLM EvaluationOff-Policy Evaluation

  2. Measuring Distribution Shift in User Prompts and Its Effects on LLM Performance

    Apr 19, 2026Parker Seegmiller, Sarah Masud PreumLLM EvaluationInstruction Following