How Many Humans Are 32 LLM Judges Worth?
Organizations: Tsinghua University · University College London · Hebei University of Economics and Business
Abstract
A panel's human-equivalent size is target-specific. Matching a fixed 32-judge panel to empirical human label distributions on three ChaosNLI tasks yields two distinct effective sizes: distributional-error matching gives , , and , whereas spectral matching gives , , and , a gap of --; a binary-error diagnostic credits the same panels with only -- effective votes. Extrapolating the distributional-error curve at fixed squared mean residual, mean member variance, and normalized mean covariance gives asymptotes of , , and , with 32 judges already reaching --. An exact spectral identity explains the gap: error depends on member energy and on the orientation of residual variation relative to averaging, information that the participation ratio (PR) discards. A realizable hard-label construction confirms that higher spectral diversity can coexist with worse distribution recovery even under equal member energies and nonnegative correlations, and the consensus direction retains , , and of centered residual variance. An external check on CC-1000, a 1,000-item Civil Comments subset with a different panel, gives . For panel choice, we establish an existence result and one feasible path: exhaustive enumeration at shows that panels beating the accuracy-top- baseline on both accuracy and always exist, and greedily swapping at most two members reaches -- higher at -- percentage points higher accuracy. Our dataset and code are available at https://github.com/Chao1208/32judges-votes.
Figures & tables
| Quantity | Measurement target | Answers the question |
|---|---|---|
| PR of the normalized residual Gram matrix, matched to independent draws from . | How many independent draws have comparable spectral diversity? | |
| Squared error of the panel label distribution relative to , matched to the same draws. | How many independent draws have comparable distributional error? | |
| Share of centered residual variance along the equal-weight judge direction. | How much centered variation survives equal-weight averaging? | |
| Signed-correlation summary of binary errors relative to the supplied gold label. | How many effective votes does the binary-error diagnostic report? |
| Component | Scope and treatment |
|---|---|
| Baseline | records; 95,994 parseable. The 6 failures (0.00625%) retain flagged placeholders: 1/0/5 by task. |
| Human reference | 100 labels per item; is their empirical frequency vector. |
| Panel audit | 54,139 unique member sets per task; 150,372 fixed member additions. Shared judges and items make these overlapping comparisons. |
| Item stability | Fixed 500/500 halves; compares stability within observed data. Failed-item sensitivity is recorded separately. |
| Presentation | 122,000 records, including cache reuse; its own original-response snapshot. Complete-case sets contain 489/492/492 items and 31/31/29 judges. |
| CC-1000 check | 1,000 items, 32 judges, and 176,251 human toxicity votes; score-only public release. |
| Measurement | MNLI-m | SNLI | NLI |
|---|---|---|---|
| 1.97 | 2.23 | 2.00 | |
| 4.24 | 6.46 | 6.50 | |
| 2.30 | 3.75 | 3.44 | |
| 1.84 | 1.72 | 1.89 | |
| 0.19691 | 0.09068 | 0.04825 | |
| 43.8% | 33.7% | 35.9% |
| Anchor | MNLI-m | SNLI | NLI |
|---|---|---|---|
| Mean squared inner product | |||
| Uniform | 0.412 | 0.531 | 0.676 |
| Mode | 0.425 | 0.284 | 0.287 |
| Distribution | 0.212 | 0.128 | 0.128 |
| Effective size | |||
| Uniform | 2.75 | 2.57 | 2.33 |
| Addition | Examined | MNLI-m | SNLI | NLI |
|---|---|---|---|---|
| 54,348 | 2,116 | 309 | 1,127 | |
| 48,000 | 332 | 441 | 1,085 | |
| 32,000 | 0 | 7 | 447 | |
| 15,992 | 0 | 0 | 85 | |
| 32 | 0 | 0 | 0 |
| Measurement | CC-1000 |
|---|---|
| 2.838 | |
| (empirical) | 1.490 |
| (finite annotation) | 1.549 |
| 1.677 | |
| 57.1% |
| Baseline top- | Rule A | Rule D | |||||
|---|---|---|---|---|---|---|---|
| Dataset | acc | acc | acc | ||||
| MNLI-m | 5 | 0.7325 | 2.732 | ||||
| MNLI-m | 7 | 0.7215 | 3.085 | ||||
| SNLI | 5 | 0.8685 | 2.495 | ||||
| SNLI | 7 | 0.8590 | 2.759 | ||||
| NLI | 5 | 0.9380 | 2.651 | ||||
| Baseline top- | (%) | ||||||
|---|---|---|---|---|---|---|---|
| Dataset | A | D | E | ||||
| MNLI-m | 5 | 0.2135 | 2.125 | ||||
| MNLI-m | 7 | 0.2066 | 2.196 | ||||
| SNLI | 5 | 0.1367 | 2.487 | ||||
| SNLI | 7 | 0.1317 | 2.581 | ||||
| NLI | 5 | 0.0600 | 2.768 | ||||
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
| Dataset | Baseline and selected panels (parenthesized values are single-judge accuracy ranks) | |
|---|---|---|
| MNLI-m | 5 | : KimiK3, KimiK2.5, o4-mini, KimiK2.6, Qwen3.7-plus; A : KimiK2.5(2), Qwen3.7-plus(5), GPT5.6-terra(18), Claude-haiku4.5(24); D : o4-mini(3), Qwen3.7-plus(5), Claude-opus5(13), GPT5.6-terra(18); E : same as A. |
| MNLI-m | 7 | : KimiK3, KimiK2.5, o4-mini, KimiK2.6, Qwen3.7-plus, Qwen3-max, GLM5.3; A : KimiK2.5(2), Qwen3.7-plus(5), DeepSeekV3.2(15), Claude-haiku4.5(24); D : Qwen3-max(6), GLM5.3(7), GPT5.6-terra(18), Claude-haiku4.5(24); E : KimiK2.5(2), Qwen3-max(6), DeepSeekV3.2(15), Claude-haiku4.5(24). |
| SNLI | 5 | : Gemini3.1-pro, Claude-opus5, Qwen3.7-plus, Qwen3.8-max, GLM5.1; A : Gemini3.1-pro(1), Qwen3.8-max(4), Claude-haiku4.5(30), DeepSeekV3.2(32); D : Claude-opus5(2), Qwen3.8-max(4), GPT5.4(21), DeepSeekV3.2(32); E : Qwen3.7-plus(3), Qwen3.8-max(4), KimiK2.6(23), GPT5.6-terra(29). |
| SNLI | 7 | : Gemini3.1-pro, Claude-opus5, Qwen3.7-plus, Qwen3.8-max, GLM5.1, Gemini3.6-flash, Qwen3.5-plus; A : Gemini3.1-pro(1), Qwen3.8-max(4), Claude-haiku4.5(30), DeepSeekV3.2(32); D : Qwen3.8-max(4), Gemini3.6-flash(6), GLM5.2(11), GPT5.6-terra(29); E : Qwen3.8-max(4), Gemini3.6-flash(6), Claude-haiku4.5(30), DeepSeekV3.2(32). |
| NLI | 5 | : Grok4.6, Grok4.5, Qwen3.8-max, Claude-opus5, Gemini3.1-pro; A : Grok4.5(2), Gemini3.1-pro(5), GPT5.6-terra(24), DeepSeekV3.2(32); D : Grok4.5(2), Qwen3.8-max(3), DeepSeekV4-flash(12), MiniMaxM3(21); E : Grok4.5(2), Qwen3.8-max(3), Gemini2.5-pro(18), o4-mini(25). |
| NLI | 7 | : Grok4.6, Grok4.5, Qwen3.8-max, Claude-opus5, Gemini3.1-pro, Gemini3.6-flash, Gemini3.7-flash; A : Grok4.5(2), Gemini3.6-flash(6), GPT5.6-terra(24), DeepSeekV3.2(32); D : Gemini3.6-flash(6), Gemini3.7-flash(7), GPT5.6-terra(24), DeepSeekV3.2(32); E : Gemini3.6-flash(6), Gemini3.7-flash(7), o4-mini(25), DeepSeekV3.2(32). |
| Dataset | Signed summary | PR |
|---|---|---|
| Binary-error Pearson matrix | ||
| MNLI-m | 1.971 | 3.536 |
| SNLI | 2.227 | 4.383 |
| NLI | 1.999 | 3.792 |
| Human-residual Gram matrix | ||
| MNLI-m | 2.201 | 4.228 |
| Dataset | Random | Diverse | Difference |
|---|---|---|---|
| MNLI-m | 3.384 | 3.382 | -0.003 |
| SNLI | 4.509 | 4.615 | +0.105 |
| NLI | 4.468 | 4.577 | +0.108 |
| Dataset | Items | Ratio | Share | ||
|---|---|---|---|---|---|
| MNLI-m | 999 | 4.220 | 2.269 | 1.86 | 96.4% |
| 1,596 | 4.242 | 2.278 | 1.86 | 96.4% | |
| SNLI | 1,000 | 6.650 | 3.739 | 1.78 | 94.0% |
| 1,513 | 6.737 | 3.804 | 1.77 | 93.9% | |
| NLI | 995 | 6.283 | 3.335 | 1.88 | 94.5% |
| 1,525 | 6.287 | 3.149 | 2.00 | 94.3% |