Reliable evaluation of large language models (LLMs) is essential for their development and deployment, yet is often costly, risky, and difficult to perform safely online. We study off-policy evaluation for LLMs, where limited human-labeled data from a behavior model are used to evaluate a newer target LLM. This setting is challenging because labels are scarce, behavior--target distribution shift is common, and response likelihoods are often unavailable for black-box LLMs. We propose the Optimal Transport-based Robust Off-Policy Evaluation (OTROPE), a likelihood-free evaluation that performs distributional correction in a semantic space via optimal transport to align labeled behavior-policy samples with unlabeled target-policy samples. OTROPE combines corrected human-labeled residuals with proxy predictors, yielding a doubly robust-style evaluation without behavior-policy modeling or density-ratio estimation. We theoretically characterize why baseline evaluators fail under LLM distribution shift, and establish consistency and convergence rates for OTROPE when either the reweighted behavior distribution or the proxy predictor converges. Experiments on synthetic and real LLM evaluation tasks show that OTROPE consistently outperforms baselines while enabling ensembles of weaker LLM evaluators to approach and sometimes surpass stronger evaluators. Code is available at https://github.com/LinerXiang/OTROPE.
Figures & tables
Figure 1 : (a) Density plots for behavior–target distribution shift, where the reweighted distribution by OTROPE closely matches the target one. (b) Density curve of behavior–target density ratios on labeled samples, showing a distribution concentrated near zero. (c) Mean absolute error (MAE) of different estimators with distributional shift level τ . OTROPE consistently achieves the lowest MAE.
Figure 2 : Log-scale mean absolute error under a fixed predictor g for distribution shift levels τ∈{1,2,3} . PPI++ is omitted since it nearly overlaps with PPI.
Method
Benchmark
RewardBench 2
Chatbot Arena
Target Policy ( π )
Mixed
Claude 3.5 Sonnet
GPT-4
GPT-3.5-Turbo
IS
0.263 (0.006)
0.273 (0.008)
0.414 (0.002)
0.272 (0.001)
DR
0.129 (0.008)
0.125 (0.003)
0.088 (0.007)
0.065 (0.004)
PPI
0.105 (0.006)
0.139 (0.001)
0.086 (0.007)
0.075 (0.005)
PPI++
0.116 (0.003)
0.138 (0.001)
0.301 (0.004)
0.196 (0.003)
g= Mixture-of-Experts
OTROPE
0.071 (0.046)
0.103 (0.077)
0.049 (0.018)
0.010 (0.007)
Table 1 : Mean absolute error of the estimated target policy value on RewardBench 2 and Chatbot Arena . Standard deviations are reported in parentheses. The best result in each column is highlighted in bold . The Mixed policy is the mixture of Claude 3.5 Sonnet and human responses.
Estimator
Mixture-of-Experts
Majority Voting
n=800
n=1200
n=1600
n=800
n=1200
n=1600
DM
0.252 (0.009)
0.252 (0.007)
0.256 (0.006)
0.296 (0.000)
0.296 (0.000)
0.296 (0.000)
OT
0.272 (0.073)
0.236 (0.085)
0.258 (0.050)
0.272 (0.073)
0.236 (0.085)
0.258 (0.050)
OTROPE
0.144 (0.048)
0.114 (0.036)
0.071 (0.046)
0.099 (0.083)
0.117 (0.067)
0.066 (0.059)
Table 2 : Ablation study under the Mixed policy setting on RewardBench 2 , showing the effects of the correction term and proxy predictors across different labeled sample sizes n . DM with majority voting is invariant to n as as it does not use labeled samples; OT is identical across aggregation strategies as it does not use proxy predictions.
λ
n
800
1200
1600
0.05
0.080 (0.025)
0.058 (0.028)
0.054 (0.015)
0.08
0.116 (0.073)
0.053 (0.061)
0.042 (0.024)
0.10
0.086 (0.060)
0.062 (0.025)
0.049 (0.018)
Table 3 : Stable performance with different λ and sample size n .
Table 6
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
Policy Name
Claude 3.5 Sonnet
Mixed
Target Policy
claude-3-5-sonnet-20241022
claude-3-5-sonnet-20241022 , human
Behavior Policy
Llama-3.1-70B-Instruct , Qwen2.5-72B-Instruct ,
Llama-3.1-70B-Instruct , Qwen2.5-72B-Instruct ,
Qwen2.5-7B-Instruct
Qwen2.5-7B-Instruct
Opponent Pool
Llama-3.1-70B-Instruct , Llama-3.1-8B-Instruct ,
Llama-3.1-70B-Instruct , Llama-3.1-8B-Instruct ,
Llama-3.1-Tulu-3-8B , Mistral-7B-Instruct-v0.3 ,
Llama-3.1-Tulu-3-8B , Mistral-7B-Instruct-v0.3 ,
Qwen2.5-72B-Instruct , Qwen2.5-7B-Instruct ,
Qwen2.5-72B-Instruct , Qwen2.5-7B-Instruct ,
Appendix
Table D.1 : Two policies evaluated in RewardBench 2 .
Policy Name
GPT-4
GPT-3.5-Turbo
Target Policy
gpt-4
gpt-3.5-turbo
Behavior Policy
koala-13b
koala-13b
Opponent Pool
RWKV-4-Raven-14B , alpaca-13b ,
RWKV-4-Raven-14B , alpaca-13b ,
stablelm-tuned-alpha-7b , vicuna-13b
stablelm-tuned-alpha-7b , vicuna-13b
Size D1
1,615
1,615
Size D2
998
1,080
Appendix
Table D.2 : Two policies evaluated in Chatbot Arena .
Hyperparameter
RewardBench 2
Chatbot Arena
Hidden dimensions
(64,32)
(64,32)
Optimizer
Adam
Adam
Learning rate
3×10−4
1×10−3
Batch size
256
256
Number of epochs
50
50
Temperature
{0.3,0.5,0.7,1.0}
{0.3,0.5,0.7,1.0}
Appendix
Table D.3 : Hyperparameters for MoE training on RewardBench 2 and Chatbot Arena .
Hyperparameter
Value / Description
Entropic regularization
0.1
Number of Sinkhorn iterations
1000
Outer optimizer
Adam
Number of outer iterations
500
Learning rate
0.05
Appendix
Table D.4 : Hyperparameters used for OT weight optimization in the controlled experiments, under both fixed and trained g settings.
Hyperparameter
Values
PCA dimension
128
Number of Sinkhorn iterations
2000
Outer optimizer
Adam
Number of outer iterations
1000
Learning rate
0.01
Appendix
Table D.5 : Hyperparameters for OT weight optimization on RewardBench 2 and Chatbot Arena .
τ
IS
DR
PPI
PPI++
OTROPE
1
0.908 (1.077)
0.959 (1.186)
2.026 (0.126)
2.013 (0.063)
0.236 (0.187)
2
6.906 (18.068)
6.420 (16.300)
4.005 (0.131)
3.998 (0.069)
1.037 (0.703)
3
8.680 (8.832)
8.232 (5.112)
5.992 (0.138)
5.989 (0.070)
2.880 (1.229)
4
9.865 (0.227)
9.878 (0.023)
7.886 (0.138)
7.894 (0.072)
4.635 (1.740)
5
11.450 (0.000)
11.450 (0.022)
9.459 (0.118)
9.452 (0.073)
6.205 (1.969)
Appendix
Table E.1 : Mean absolute error (MAE) over 200 replications under varying distribution shift τ . Values are reported as mean with standard deviation in parentheses.
Estimator
n
200
400
600
800
1000
τ=1
IS
1.944 (5.290)
1.463 (2.548)
1.183 (1.724)
1.085 (1.395)
0.924 (1.208)
DR
2.464 (6.989)
1.776 (3.272)
1.446 (2.233)
1.300 (1.890)
1.139 (1.656)
PPI
2.048 (0.426)
2.079 (0.316)
2.055 (0.270)
2.057 (0.240)
2.057 (0.191)
PPI++
2.009 (0.120)
2.018 (0.087)
2.016 (0.076)
2.021 (0.068)
2.018 (0.061)
Appendix
Table E.2 : Mean absolute error and standard deviation of target policy value estimation with a fixed predictor. The labeled sample size n varies.
Figure E.1 : Log-scale mean absolute error under the predictor g trained on D1 for distribution shift levels τ∈{1,2,3} . PPI++ is omitted since it nearly overlaps with PPI.
Estimator
n
200
400
600
800
1000
τ=1
IS
1.940 (5.291)
1.460 (2.550)
1.179 (1.726)
1.082 (1.398)
0.922 (1.210)
DR
0.984 (2.234)
0.788 (1.230)
0.674 (0.839)
0.624 (0.648)
0.549 (0.524)
PPI
1.998 (0.121)
2.003 (0.093)
2.001 (0.083)
2.005 (0.078)
2.001 (0.075)
PPI++
2.000 (0.120)
2.009 (0.087)
2.007 (0.076)
2.012 (0.068)
2.009 (0.061)
Appendix
Table E.3 : Mean absolute error and standard deviation of target policy value estimation with a predictor trained on D1 . The labeled sample size n varies.
τ
an
bn
bn−Cran
min∣VDR−V∗∣
Holds
1
1.373
3.965
−18.70
0.008
100%
2
2.719
5.959
−38.94
0.071
100%
3
0.177
7.948
5.03
6.119
100%
4
0.0006
9.847
9.838
9.847
100%
5
0.0000
11.415
11.415
11.415
100%
Appendix
Table E.4 : Numerical verification of the bias bound in Proposition 2 . The constant Cr=16.51 is a pooled upper bound on ∣r(z)∣ across all τ .
τ
κ
Var(VDR)
Predicted lower bound
Holds
1
0.000
30.5
−1.0×103
Yes
2
0.000
1.62×104
5.57×103
Yes
3
0.000
7.97×103
250
Yes
Appendix
Table E.5 : Numerical verification of the variance lower bound in Proposition 2 under a conservative predictor satisfying 0<cr≤r(z) .
σtgt/σref
IS
DR
PPI
OTROPE
0.50
0.226 (0.066)
0.287 (0.101)
0.083 (0.005)
0.104 (0.007)
2.00
2.163 (30.56)
2.138 (24.19)
0.120 (0.007)
0.125 (0.013)
Appendix
Table E.6 : Mean absolute error under variance-shift simulations. Standard deviations are reported in parentheses.
Dataset
Target → Behavior NN
Target → Target NN
Ratio
Fraction with ratio >2
Chatbot Arena (GPT-4)
0.855
0.871
0.98
0.0%
Chatbot Arena (GPT-3.5-Turbo)
0.865
0.883
0.98
0.0%
RewardBench 2 (Claude 3.5 Sonnet)
0.955
0.163
5.85
94.2%
RewardBench 2 (Mixed)
0.936
0.164
5.71
94.1%
Appendix
Table E.7 : Nearest-neighbor support coverage diagnostics in the embedding space. The ratio compares the target-to-behavior nearest-neighbor distance with the target-to-target nearest-neighbor distance. A ratio close to one indicates stronger local support coverage.
Dataset
d0
R2
Chatbot Arena (GPT-4)
23.98
1.0000
Chatbot Arena (GPT-3.5-Turbo)
24.01
1.0000
RewardBench 2 (Mixed)
29.23
0.9996
RewardBench 2 (Claude 3.5 Sonnet)
29.17
0.9997
Appendix
Table E.8 : Estimated intrinsic dimension of the raw pre-PCA bge-m3 embeddings. The slope of the log–log regression gives d0 .
Mixture-of-Experts
Majority Voting
Method
GPT-4
GPT-3.5
GPT-4
GPT-3.5
IS
0.408 (0.002)
0.261 (0.002)
0.408 (0.002)
0.261 (0.002)
DR
0.088 (0.009)
0.081 (0.007)
0.127 (0.002)
0.110 (0.002)
PPI
0.086 (0.009)
0.080 (0.007)
0.124 (0.002)
0.108 (0.002)
PPI++
0.302 (0.006)
0.197 (0.004)
0.365 (0.001)
0.234 (0.001)
OTROPE
0.049 (0.013)
0.076 (0.007)
0.088 (0.012)
0.106 (0.006)
Appendix
Table E.9 : Mean absolute error of the estimated target policy value on Chatbot Arena using 64-dimensional PCA-truncated bge-m3 embeddings. Standard errors are reported in parentheses.
Method
Mixture-of-Experts
Majority Voting
IS
0.272 (0.001)
0.272 (0.001)
DR
0.069 (0.008)
0.094 (0.002)
PPI
0.079 (0.008)
0.108 (0.002)
PPI++
0.197 (0.004)
0.234 (0.001)
OTROPE
0.025 (0.011)
0.034 (0.016)
Appendix
Table E.10 : Mean absolute error on Chatbot Arena under GPT-3.5-Turbo using the ℓ1 transport cost. Standard deviations are reported in parentheses.
The alignment of Large Language Models (LLMs) utilizes Reinforcement Learning from AI Feedback (RLAIF) for non-verifiable domains such as long-form question answering and open-ended instruction following. These domains often rely on LLM based auto-raters to provide granular, multi-tier discrete rewards (e.g., 1-10 rubrics) that are inherently stochastic due to prompt sensitivity and sampling randomness. We empirically verify the stochasticity of auto-raters that can propagate and corrupt standard advantage estimators like GRPO and MaxRL, as a noisy reward samples can skew normalization statistics and degrade the global learning signal. Empirically, sampling more rewards and taking majority voting may reduce the noise and improve performance, but this approach is computationally expensive. To address this bottleneck, we introduce Ordinal Decomposition for Robust Policy Optimization (ODRPO), a framework that structurally isolates evaluation noise by decomposing discrete rewards into a sequence of ordinal binary indicators. By independently computing and accumulating advantages across these progressively challenging success thresholds, ODRPO prevents outlier evaluations from corrupting the global update while establishing an implicit, variance-aware learning curriculum. Empirically, ODRPO achieves robust performance on Qwen2.5-7B and Qwen3-4B models, outperforming baselines with relative improvements of upto 14.8% on FACTS-grounding-v2 and 7.5% on Alpaca-Evals. Critically, these gains are achieved with negligible training-time overhead, as ODRPO requires no additional compute per step compared to standard estimators. Supported by theoretical analysis confirming its optimization stability, ODRPO provides a scalable and robust framework for aligning models within the noisy, discrete evaluation landscape of modern RLAIF.
Current Large Language Model (LLM) evaluation frameworks utilize the same static prompt template across all models under evaluation. This differs from the common industry practice of using prompt optimization (PO) techniques to optimize the prompt for each model to maximize application performance. In this paper, we investigate the effect of PO towards LLM evaluations. Our results on public academic and internal industry benchmarks show that PO greatly affects the final ranking of models. This highlights the importance of practitioners performing PO per model when conducting evaluations to choose the best model for a given task.
Reinforcement learning from human feedback (RLHF) has evolved to be one of the main methods for fine-tuning large language models (LLMs). However, existing RLHF methods are non-robust, and their performance deteriorates if the downstream task differs significantly from the preference dataset used in fine-tuning. In order to mitigate this problem, we introduce a distributionally robust RLHF for fine-tuning LLMs. In particular, our goal is to ensure that a fine-tuned model retains its performance even when the distribution of prompts significantly differs from the distribution encountered during fine-tuning. We formulate distributionally robust optimization (DRO) version of two popular fine-tuning methods -- (1) reward-based RLHF and (2) reward-free DPO (direct preference optimization). We propose a minibatch gradient descent based algorithms for both of them, and theoretically prove convergence guarantees for the algorithms. Subsequently, we evaluate our algorithms on an out-of-distribution (OOD) task by first training the model on the Unified-Feedback dataset and evaluating its performance on two different datasets. The experimental results show that our robust training improves the accuracy of the learned reward models on average, and markedly on some tasks, such as reasoning. Furthermore, we show that the robust versions of policy optimization methods, similarly improve performance on OOD tasks.
Debmalya Mandal, Paulius Sasnauskas, Goran Radanovic
Dept. of Computer Science University of Warwick, UK · Department of Computing Science University of Alberta, Edmonton, Canada · Max-Planck Institute for Software Systems, Germany