cs.LGSep 30, 2026

Optimal Design for Active Preference Learning with Biased LLM Judges

Authors: Zhongman Du, Huiming Zhang, Haodong Zhu, Baochang Zhang

Organizations: Beihang University · Zhongguancun Academy · Hangzhou Innovation Institute of Beihang University

Abstract

Learning from human preferences is central to large language model (LLM) alignment, but human preference annotation is costly. Active preference learning reduces this cost by selecting informative comparisons, and LLM judges can provide additional scalable feedback. However, the preferences of the judges may deviate from those of the target human population. Even after calibration on trusted reference data, active acquisition can shift the comparison distribution and expose residual judge bias. We therefore incorporate judge deviations into the acquisition design rather than relying on a separate calibration stage. Under joint estimation, comparisons that appear highly informative about the reward may also reflect judge bias and therefore provide less information about human preferences. To address this issue, we propose Nuisance-Adjusted Optimal Design (NAOD), a comparison-selection strategy that prioritizes policy-relevant target information after nuisance adjustment and uses the Frank-Wolfe algorithm for optimization. Theoretically, we establish a sharp conditional local asymptotic minimax lower bound on policy risk and construct an estimator that attains it. We further characterize the finite-sample cost of learning the nuisance representation and show that representation error can reverse an oracle design advantage. Finally, we validate these predictions experimentally and evaluate NAOD on Chatbot Arena data across 17 judges, 15 budget configurations, and 15 random cluster-level splits. NAOD reduces the mean regret of proxy policy by 29.1% relative to a matched target-information design, outperforms existing methods, and improves human-preference prediction on held-out data.

Figures & tables

Appendix figures & tables14 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. CUPID in the Model Zoo: Online Matchmaking for Selecting Your Dream LLM

    May 30, 2026Son Nguyen, Xinyuan Liu, Ransalu SenanayakeLarge Language Model AlignmentBandits

  2. Quantifying and Auditing LLM Evaluation via Positive--Unlabeled Learning

    Jun 17, 2026Zilong Zhang, Yi-Ting Hung, Lei Ding +1Large Language Model EvaluationPositive--Unlabeled Learning