Amortized Off-Policy Evaluation for LLMs
Organizations: RBC Borealis · University of Toronto · Vector Institute
Abstract
Accurate evaluation is central to selecting which LLM to deploy, yet testing a candidate on live traffic exposes real users to an unvetted model. Teams therefore evaluate candidates offline, on data produced by already-deployed models. This is off-policy evaluation (OPE), and it faces two distribution shifts: as a model is updated in post-training, its responses diverge from the logged ones (policy shift), and the reward definition under which it is judged changes with business requirements (reward shift). Classical OPE methods are ill-suited to this continual-deployment setting because they are defined per task and require fitting from scratch on every new logged dataset or reward definition. To address this, we propose PFN-OPE, a prior-data fitted network that amortizes OPE across a distribution of contextual-bandit tasks. We pretrain it once on tasks constructed from a pool of LLM responses scored by several reward functions, in which both shifts occur. At test-time it maps a logged dataset and one sampled target response per prompt to a value estimate in a single forward pass, with no per-task fitting. On HelpSteer2 and UltraFeedback with Qwen, Llama, and Gemma policies, PFN-OPE achieves 2.0 to 9.3 times lower error than the best baselines across all tested configurations in the reward-shifted settings.
Figures & tables
| Target: Llama-3.2-3B | Target: Gemma-3-1B | |||||
| Behavior LLM | Qwen2.5 7B | Qwen2.5 1.5B | Llama-3.2 1B | Qwen2.5 7B | Qwen2.5 1.5B | Llama-3.2 1B |
| HelpSteer2 | ||||||
| PFN-OPE (ours) | 0.041 0.006 | 0.87 0.16 | 0.97 0.23 | 1.54 0.30 | 2.01 0.21 | 2.75 0.42 |
| DM-ridge | 0.130 0.047 | 17.78 1.10 | 8.45 0.66 | 5.94 0.56 | 30.35 1.90 | 19.37 1.80 |
| DM-kernel | 0.050 0.055 | 16.57 0.91 | 7.72 0.87 | 4.91 0.73 | 30.32 1.70 | 18.98 1.90 |
| DM-MLP | 1.26 1.60 | 26.40 6.60 | 7.99 2.00 | 7.44 3.20 | 39.57 8.10 | 20.52 5.40 |