Optimal Design for Active Preference Learning with Biased LLM Judges
Organizations: Beihang University · Zhongguancun Academy · Hangzhou Innovation Institute of Beihang University
Abstract
Learning from human preferences is central to large language model (LLM) alignment, but human preference annotation is costly. Active preference learning reduces this cost by selecting informative comparisons, and LLM judges can provide additional scalable feedback. However, the preferences of the judges may deviate from those of the target human population. Even after calibration on trusted reference data, active acquisition can shift the comparison distribution and expose residual judge bias. We therefore incorporate judge deviations into the acquisition design rather than relying on a separate calibration stage. Under joint estimation, comparisons that appear highly informative about the reward may also reflect judge bias and therefore provide less information about human preferences. To address this issue, we propose Nuisance-Adjusted Optimal Design (NAOD), a comparison-selection strategy that prioritizes policy-relevant target information after nuisance adjustment and uses the Frank-Wolfe algorithm for optimization. Theoretically, we establish a sharp conditional local asymptotic minimax lower bound on policy risk and construct an estimator that attains it. We further characterize the finite-sample cost of learning the nuisance representation and show that representation error can reverse an oracle design advantage. Finally, we validate these predictions experimentally and evaluate NAOD on Chatbot Arena data across 17 judges, 15 budget configurations, and 15 random cluster-level splits. NAOD reduces the mean regret of proxy policy by 29.1% relative to a matched target-information design, outperforms existing methods, and improves human-preference prediction on held-out data.
Figures & tables
| Method | Proxy regret | Human CE | Acc. | Drop (%) | Wins |
|---|---|---|---|---|---|
| Initial | 8.86 | 0.6752 | 57.82 | 77.53 | 13/15 |
| Human-only | 18.25 | 0.6829 | 56.84 | 89.09 | 15/15 |
| Random | 4.88 | 0.6727 | 58.22 | 59.18 | 15/15 |
| Entropy | 13.42 | 0.6784 | 57.10 | 85.17 | 15/15 |
| D-opt | 2.45 | 0.6703 | 58.19 | 18.87 | 12/15 |
| PA D-opt | 2.46 | 0.6703 | 58.19 | 18.93 | 12/15 |
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
| Component | Convention |
|---|---|
| Current outcomes | Independent Bernoulli trusted and judge labels |
| Final estimator | Trusted preliminary + offset nuisance preliminary + one joint update |
| Target / nuisance radii | and , respectively, unless stated otherwise |
| Preliminary gap tolerance | divided by fit sample size |
| Nuisance-Gram threshold | |
| Joint-information guard |
| Design | Mean | MC SE | ||
|---|---|---|---|---|
| 240 | NAOD | 0.426 | 0.430 | 0.008 |
| 240 | Target-info | 1.139 | 1.164 | 0.021 |
| 240 | Random | 0.530 | 0.527 | 0.010 |
| 960 | NAOD | 0.426 | 0.441 | 0.008 |
| 960 | Target-info | 1.139 | 1.161 | 0.022 |
| 960 | Random | 0.530 | 0.518 | 0.010 |
| Prediction | Mean | MC SE | Clip (%) | ||
|---|---|---|---|---|---|
| 960 | 2.041 | 1.891 | 0.034 | 3.1 | |
| 960 | 1.276 | 1.252 | 0.023 | 0.0 | |
| 960 | 1.084 | 1.085 | 0.020 | 0.0 | |
| 3840 | 2.041 | 1.993 | 0.037 | 0.0 | |
| 3840 | 1.276 | 1.281 | 0.024 | 0.0 | |
| 3840 | 1.084 | 1.116 | 0.020 | 0.0 |
| NAOD–Target-info | Outer SE | 95% interval | |
|---|---|---|---|
| 0.237 | |||
| 0.200 | |||
| 0.100 | |||
| 0.049 |
| Fitted model | Nominal | Mean | MC SE | |
|---|---|---|---|---|
| Underfit | 0 | 0.391 | 27.562 | 0.064 |
| Correct | 1 | 0.686 | 0.699 | 0.010 |
| Overfit | 2 | 0.772 | 0.764 | 0.011 |
| Saturated | 6 | 4.000 | 4.015 | 0.059 |
| Design | Nominal risk | Mean | Pool SE | |
|---|---|---|---|---|
| 60 | NAOD | 6.108 | 6.146 | 0.081 |
| 60 | Target-info | 6.145 | 6.181 | 0.081 |
| 60 | Random | 8.945 | 8.851 | 0.117 |
| 180 | NAOD | 3.420 | 3.361 | 0.053 |
| 180 | Target-info | 3.430 | 3.373 | 0.051 |
| Role | Clusters | Use |
|---|---|---|
| Upstream | Target representation, nuisance representation, and human reference | |
| Historical initialization | Initial target and nuisance estimates | |
| Policy support | Policy weighting covariates with no labels | |
| Current human reservoir | Nested trusted budgets | |
| Candidate | Judge acquisition with cluster capacity one | |
| Test | Held-out human evaluation |
| Component | Configuration |
|---|---|
| Embedding | Qwen3-Embedding-0.6B, last-token pooling, normalization |
| Maximum input | tokens with prompt cap |
| Base reduction | Upstream-fitted whitened PCA, dimensions |
| Target representation | Judge-specific SVD span, dimension |
| Target-head penalties | |
| Nuisance representation | Intercept and standardized learned residual, dimension |
| Baseline | Proxy-regret gain | Human-CE gain | Accuracy gain |
|---|---|---|---|
| Initial | |||
| Human-only | |||
| Random | |||
| Entropy | |||
| D-opt | |||
| PA D-opt |
| Target-info–NAOD | D-opt–NAOD | ||
|---|---|---|---|
| 32 | 256 | ||
| 32 | 512 | ||
| 32 | 1024 | ||
| 64 | 256 | ||
| 64 | 512 | ||
| 64 | 1024 |
| Judge | Policy gain | |
|---|---|---|
| Athene-70B | 0.570 | |
| dolphin-2.1-mistral-7b | 0.956 | |
| dolphin-2.5-mixtral-8x7b | 0.960 | |
| Hermes-3-Llama-3.1-70B | 0.610 | |
| Meta-Llama-3-70B-Instruct | 0.459 | |
| Meta-Llama-3-8B-Instruct | 0.716 |