Reliable evaluation of large language models (LLMs) is essential for their development and deployment, yet is often costly, risky, and difficult to perform safely online. We study off-policy evaluation for LLMs, where limited human-labeled data from a behavior model are used to evaluate a newer target LLM. This setting is challenging because labels are scarce, behavior--target distribution shift is common, and response likelihoods are often unavailable for black-box LLMs. We propose the Optimal Transport-based Robust Off-Policy Evaluation (OTROPE), a likelihood-free evaluation that performs distributional correction in a semantic space via optimal transport to align labeled behavior-policy samples with unlabeled target-policy samples. OTROPE combines corrected human-labeled residuals with proxy predictors, yielding a doubly robust-style evaluation without behavior-policy modeling or density-ratio estimation. We theoretically characterize why baseline evaluators fail under LLM distribution shift, and establish consistency and convergence rates for OTROPE when either the reweighted behavior distribution or the proxy predictor converges. Experiments on synthetic and real LLM evaluation tasks show that OTROPE consistently outperforms baselines while enabling ensembles of weaker LLM evaluators to approach and sometimes surpass stronger evaluators. Code is available at https://github.com/LinerXiang/OTROPE.
Figures & tables
Figure 1 : (a) Density plots for behavior–target distribution shift, where the reweighted distribution by OTROPE closely matches the target one. (b) Density curve of behavior–target density ratios on labeled samples, showing a distribution concentrated near zero. (c) Mean absolute error (MAE) of different estimators with distributional shift level τ . OTROPE consistently achieves the lowest MAE.
Figure 2 : Log-scale mean absolute error under a fixed predictor g for distribution shift levels τ∈{1,2,3} . PPI++ is omitted since it nearly overlaps with PPI.
Method
Benchmark
RewardBench 2
Chatbot Arena
Target Policy ( π )
Mixed
Claude 3.5 Sonnet
GPT-4
GPT-3.5-Turbo
IS
0.263 (0.006)
0.273 (0.008)
0.414 (0.002)
0.272 (0.001)
DR
0.129 (0.008)
0.125 (0.003)
0.088 (0.007)
0.065 (0.004)
PPI
0.105 (0.006)
0.139 (0.001)
0.086 (0.007)
0.075 (0.005)
PPI++
0.116 (0.003)
0.138 (0.001)
0.301 (0.004)
0.196 (0.003)
g= Mixture-of-Experts
OTROPE
0.071 (0.046)
0.103 (0.077)
0.049 (0.018)
0.010 (0.007)
Table 1 : Mean absolute error of the estimated target policy value on RewardBench 2 and Chatbot Arena . Standard deviations are reported in parentheses. The best result in each column is highlighted in bold . The Mixed policy is the mixture of Claude 3.5 Sonnet and human responses.
Estimator
Mixture-of-Experts
Majority Voting
n=800
n=1200
n=1600
n=800
n=1200
n=1600
DM
0.252 (0.009)
0.252 (0.007)
0.256 (0.006)
0.296 (0.000)
0.296 (0.000)
0.296 (0.000)
OT
0.272 (0.073)
0.236 (0.085)
0.258 (0.050)
0.272 (0.073)
0.236 (0.085)
0.258 (0.050)
OTROPE
0.144 (0.048)
0.114 (0.036)
0.071 (0.046)
0.099 (0.083)
0.117 (0.067)
0.066 (0.059)
Table 2 : Ablation study under the Mixed policy setting on RewardBench 2 , showing the effects of the correction term and proxy predictors across different labeled sample sizes n . DM with majority voting is invariant to n as as it does not use labeled samples; OT is identical across aggregation strategies as it does not use proxy predictions.
λ
n
800
1200
1600
0.05
0.080 (0.025)
0.058 (0.028)
0.054 (0.015)
0.08
0.116 (0.073)
0.053 (0.061)
0.042 (0.024)
0.10
0.086 (0.060)
0.062 (0.025)
0.049 (0.018)
Table 3 : Stable performance with different λ and sample size n .
Table 6
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
Policy Name
Claude 3.5 Sonnet
Mixed
Target Policy
claude-3-5-sonnet-20241022
claude-3-5-sonnet-20241022 , human
Behavior Policy
Llama-3.1-70B-Instruct , Qwen2.5-72B-Instruct ,
Llama-3.1-70B-Instruct , Qwen2.5-72B-Instruct ,
Qwen2.5-7B-Instruct
Qwen2.5-7B-Instruct
Opponent Pool
Llama-3.1-70B-Instruct , Llama-3.1-8B-Instruct ,
Llama-3.1-70B-Instruct , Llama-3.1-8B-Instruct ,
Llama-3.1-Tulu-3-8B , Mistral-7B-Instruct-v0.3 ,
Llama-3.1-Tulu-3-8B , Mistral-7B-Instruct-v0.3 ,
Qwen2.5-72B-Instruct , Qwen2.5-7B-Instruct ,
Qwen2.5-72B-Instruct , Qwen2.5-7B-Instruct ,
Appendix
Table D.1 : Two policies evaluated in RewardBench 2 .
Policy Name
GPT-4
GPT-3.5-Turbo
Target Policy
gpt-4
gpt-3.5-turbo
Behavior Policy
koala-13b
koala-13b
Opponent Pool
RWKV-4-Raven-14B , alpaca-13b ,
RWKV-4-Raven-14B , alpaca-13b ,
stablelm-tuned-alpha-7b , vicuna-13b
stablelm-tuned-alpha-7b , vicuna-13b
Size D1
1,615
1,615
Size D2
998
1,080
Appendix
Table D.2 : Two policies evaluated in Chatbot Arena .
Hyperparameter
RewardBench 2
Chatbot Arena
Hidden dimensions
(64,32)
(64,32)
Optimizer
Adam
Adam
Learning rate
3×10−4
1×10−3
Batch size
256
256
Number of epochs
50
50
Temperature
{0.3,0.5,0.7,1.0}
{0.3,0.5,0.7,1.0}
Appendix
Table D.3 : Hyperparameters for MoE training on RewardBench 2 and Chatbot Arena .
Hyperparameter
Value / Description
Entropic regularization
0.1
Number of Sinkhorn iterations
1000
Outer optimizer
Adam
Number of outer iterations
500
Learning rate
0.05
Appendix
Table D.4 : Hyperparameters used for OT weight optimization in the controlled experiments, under both fixed and trained g settings.
Hyperparameter
Values
PCA dimension
128
Number of Sinkhorn iterations
2000
Outer optimizer
Adam
Number of outer iterations
1000
Learning rate
0.01
Appendix
Table D.5 : Hyperparameters for OT weight optimization on RewardBench 2 and Chatbot Arena .
τ
IS
DR
PPI
PPI++
OTROPE
1
0.908 (1.077)
0.959 (1.186)
2.026 (0.126)
2.013 (0.063)
0.236 (0.187)
2
6.906 (18.068)
6.420 (16.300)
4.005 (0.131)
3.998 (0.069)
1.037 (0.703)
3
8.680 (8.832)
8.232 (5.112)
5.992 (0.138)
5.989 (0.070)
2.880 (1.229)
4
9.865 (0.227)
9.878 (0.023)
7.886 (0.138)
7.894 (0.072)
4.635 (1.740)
5
11.450 (0.000)
11.450 (0.022)
9.459 (0.118)
9.452 (0.073)
6.205 (1.969)
Appendix
Table E.1 : Mean absolute error (MAE) over 200 replications under varying distribution shift τ . Values are reported as mean with standard deviation in parentheses.
Estimator
n
200
400
600
800
1000
τ=1
IS
1.944 (5.290)
1.463 (2.548)
1.183 (1.724)
1.085 (1.395)
0.924 (1.208)
DR
2.464 (6.989)
1.776 (3.272)
1.446 (2.233)
1.300 (1.890)
1.139 (1.656)
PPI
2.048 (0.426)
2.079 (0.316)
2.055 (0.270)
2.057 (0.240)
2.057 (0.191)
PPI++
2.009 (0.120)
2.018 (0.087)
2.016 (0.076)
2.021 (0.068)
2.018 (0.061)
Appendix
Table E.2 : Mean absolute error and standard deviation of target policy value estimation with a fixed predictor. The labeled sample size n varies.
Figure E.1 : Log-scale mean absolute error under the predictor g trained on D1 for distribution shift levels τ∈{1,2,3} . PPI++ is omitted since it nearly overlaps with PPI.
Estimator
n
200
400
600
800
1000
τ=1
IS
1.940 (5.291)
1.460 (2.550)
1.179 (1.726)
1.082 (1.398)
0.922 (1.210)
DR
0.984 (2.234)
0.788 (1.230)
0.674 (0.839)
0.624 (0.648)
0.549 (0.524)
PPI
1.998 (0.121)
2.003 (0.093)
2.001 (0.083)
2.005 (0.078)
2.001 (0.075)
PPI++
2.000 (0.120)
2.009 (0.087)
2.007 (0.076)
2.012 (0.068)
2.009 (0.061)
Appendix
Table E.3 : Mean absolute error and standard deviation of target policy value estimation with a predictor trained on D1 . The labeled sample size n varies.
τ
an
bn
bn−Cran
min∣VDR−V∗∣
Holds
1
1.373
3.965
−18.70
0.008
100%
2
2.719
5.959
−38.94
0.071
100%
3
0.177
7.948
5.03
6.119
100%
4
0.0006
9.847
9.838
9.847
100%
5
0.0000
11.415
11.415
11.415
100%
Appendix
Table E.4 : Numerical verification of the bias bound in Proposition 2 . The constant Cr=16.51 is a pooled upper bound on ∣r(z)∣ across all τ .
τ
κ
Var(VDR)
Predicted lower bound
Holds
1
0.000
30.5
−1.0×103
Yes
2
0.000
1.62×104
5.57×103
Yes
3
0.000
7.97×103
250
Yes
Appendix
Table E.5 : Numerical verification of the variance lower bound in Proposition 2 under a conservative predictor satisfying 0<cr≤r(z) .
σtgt/σref
IS
DR
PPI
OTROPE
0.50
0.226 (0.066)
0.287 (0.101)
0.083 (0.005)
0.104 (0.007)
2.00
2.163 (30.56)
2.138 (24.19)
0.120 (0.007)
0.125 (0.013)
Appendix
Table E.6 : Mean absolute error under variance-shift simulations. Standard deviations are reported in parentheses.
Dataset
Target → Behavior NN
Target → Target NN
Ratio
Fraction with ratio >2
Chatbot Arena (GPT-4)
0.855
0.871
0.98
0.0%
Chatbot Arena (GPT-3.5-Turbo)
0.865
0.883
0.98
0.0%
RewardBench 2 (Claude 3.5 Sonnet)
0.955
0.163
5.85
94.2%
RewardBench 2 (Mixed)
0.936
0.164
5.71
94.1%
Appendix
Table E.7 : Nearest-neighbor support coverage diagnostics in the embedding space. The ratio compares the target-to-behavior nearest-neighbor distance with the target-to-target nearest-neighbor distance. A ratio close to one indicates stronger local support coverage.
Dataset
d0
R2
Chatbot Arena (GPT-4)
23.98
1.0000
Chatbot Arena (GPT-3.5-Turbo)
24.01
1.0000
RewardBench 2 (Mixed)
29.23
0.9996
RewardBench 2 (Claude 3.5 Sonnet)
29.17
0.9997
Appendix
Table E.8 : Estimated intrinsic dimension of the raw pre-PCA bge-m3 embeddings. The slope of the log–log regression gives d0 .
Mixture-of-Experts
Majority Voting
Method
GPT-4
GPT-3.5
GPT-4
GPT-3.5
IS
0.408 (0.002)
0.261 (0.002)
0.408 (0.002)
0.261 (0.002)
DR
0.088 (0.009)
0.081 (0.007)
0.127 (0.002)
0.110 (0.002)
PPI
0.086 (0.009)
0.080 (0.007)
0.124 (0.002)
0.108 (0.002)
PPI++
0.302 (0.006)
0.197 (0.004)
0.365 (0.001)
0.234 (0.001)
OTROPE
0.049 (0.013)
0.076 (0.007)
0.088 (0.012)
0.106 (0.006)
Appendix
Table E.9 : Mean absolute error of the estimated target policy value on Chatbot Arena using 64-dimensional PCA-truncated bge-m3 embeddings. Standard errors are reported in parentheses.
Method
Mixture-of-Experts
Majority Voting
IS
0.272 (0.001)
0.272 (0.001)
DR
0.069 (0.008)
0.094 (0.002)
PPI
0.079 (0.008)
0.108 (0.002)
PPI++
0.197 (0.004)
0.234 (0.001)
OTROPE
0.025 (0.011)
0.034 (0.016)
Appendix
Table E.10 : Mean absolute error on Chatbot Arena under GPT-3.5-Turbo using the ℓ1 transport cost. Standard deviations are reported in parentheses.
Dept. of Computer Science University of Warwick, UK · Department of Computing Science University of Alberta, Edmonton, Canada · Max-Planck Institute for Software Systems, Germany