cs.CLSep 18, 2026

How Many Humans Are 32 LLM Judges Worth?

Authors: Chao Li, Yingying Yu, Yunfeng Li

Organizations: Tsinghua University · University College London · Hebei University of Economics and Business

Abstract

A panel's human-equivalent size is target-specific. Matching a fixed 32-judge panel to empirical human label distributions on three ChaosNLI tasks yields two distinct effective sizes: distributional-error matching gives νMSE=2.304ν_{\mathrm{MSE}}=2.304, 3.7503.750, and 3.4453.445, whereas spectral matching gives νH=4.242ν_H=4.242, 6.4596.459, and 6.4996.499, a gap of 1.721.72--1.89×1.89\times; a binary-error diagnostic credits the same panels with only 1.9711.971--2.2272.227 effective votes. Extrapolating the distributional-error curve at fixed squared mean residual, mean member variance, and normalized mean covariance gives asymptotes of 2.3922.392, 3.9903.990, and 3.6553.655, with 32 judges already reaching 94.094.0--96.3%96.3\%. An exact spectral identity explains the gap: error depends on member energy and on the orientation of residual variation relative to averaging, information that the participation ratio (PR) discards. A realizable hard-label construction confirms that higher spectral diversity can coexist with worse distribution recovery even under equal member energies and nonnegative correlations, and the consensus direction retains γco=43.8%γ_{\mathrm{co}}=43.8\%, 33.7%33.7\%, and 35.9%35.9\% of centered residual variance. An external check on CC-1000, a 1,000-item Civil Comments subset with a different panel, gives νH=2.84ν_H=2.84. For panel choice, we establish an existence result and one feasible path: exhaustive enumeration at k∈{5,7}k\in\{5,7\} shows that panels beating the accuracy-top-kk baseline on both accuracy and νHν_H always exist, and greedily swapping at most two members reaches 24.824.8--56.0%56.0\% higher νHν_H at 0.100.10--1.101.10 percentage points higher accuracy. Our dataset and code are available at https://github.com/Chao1208/32judges-votes.

Figures & tables

Appendix figures & tables9 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. A Finite-Calibration Regime Map for LLM Judge Panels

    May 31, 2026Bin Zhu, Yanghui RaoJudgesLarge Language Model Judges