Validating an LLM-as-a-judge requires estimating its agreement with humans, yet annotation budgets rarely allow every item to be multiply labeled. We prove that this overlap sparsity is the first-order determinant of wrong deployment decisions: at 5% pairwise overlap, wrong-decision rates reach 25% and the probability of selecting the wrong best judge among ten candidates is 65%. The two actionable levers are overlap quantity and allocation. For quantity, we derive a minimum-overlap formula showing ρ≥0.25 suffices for non-borderline judges while borderline cases remain fundamentally hard. For allocation, a zero-cost stratified scheme halves false-rejection rates relative to random sampling when strata are informative. We validate on 10 LLM judges across four evaluation matrices spanning visual assessment, causal reasoning, and summarization.
Figures & tables
Figure 1: The same annotation budget can fail in two different ways. On a real visual-assessment task, depth-first labeling biases the human-agreement ceiling (a) and inflates wrong deployment decisions (b); random spread gives each annotator independent samples, leaving too little shared overlap for reliable estimation; Stratified shared overlap draws a proportional stratum sample that all annotators label, simultaneously fixing fragmentation (maximal pairwise co-coverage) and representativeness (stratum-proportional distribution).
Figure 2: Wrong-decision rate across overlap, judge quality, and design. Each cell shows the percentage of trials whose deploy/reject outcome differs from the ground-truth reference. Row pairs vary prevalence and annotation design (Random vs. Stratified); columns vary the agreement coefficient F .
∣bias∣ of F^m
std(F^m)
Rel. ( δ=0.05 )
Dataset ( πmax )
ρ
R
S
SR
SC
R
S
SR
SC
R
S
SR
SC
cebab_stars (.26)
.05
.087
.001
.087
.105
.126
.074
.124
.117
.22
.49
.24
.22
.10
.069
.004
.070
.123
.080
.051
.079
.079
.33
.64
.32
.13
.25
.039
.001
.044
.049
.040
.028
.040
.042
.58
.93
.55
.51
summeval (.56)
.05
.020
.034
.016
.016
.030
.026
.026
.027
.82
.70
.88
.90
.10
.014
.031
.013
.012
.019
.018
.018
.017
.95
.84
.97
.98
Table 1: Test-independent design-quality metrics on real benchmarks. R= Random , S= Strat , SR= Seq-Refined , SC= Seq-Coverage . Lower is better for bias/std; higher for reliability; bold marks the best design per row.
WAX
CeBaB-asp.
CeBaB-stars
SummEval
Mean
ρ
Rank
Top-1
Rank
Top-1
Rank
Top-1
Rank
Top-1
Rank
Top-1
0.05
.297
.670
.349
.887
.342
.875
.122
.177
.278
.652
0.10
.234
.600
.279
.777
.255
.787
.083
.080
.213
.561
0.25
.146
.427
.196
.507
.156
.517
.049
.003
.137
.363
0.50
.071
.270
.138
.213
.085
.350
.031
.000
.082
.208
0.75
.031
.103
.091
.033
.053
.110
.018
.000
.048
.062
Table 2: Ranking stability under sparsity. “Rank err.” = fraction of pairwise rankings scrambled; “Top-1” = probability of selecting the wrong best judge. 10 judges per benchmark.
ρ
Design
Mean FR
Mean FA
Mean WDR
0.05
Random
.290
.112
.250
Strat
.142
.172
.184
Seq-Refined
.308
.098
.250
Seq-Coverage
.435
.066
.311
0.10
Random
.238
.082
.212
Strat
.087
.136
.137
Table 3: Wrong-decision rates by design and overlap (aggregated across the four real evaluation matrices). FR = mean false-rejection rate on 20 pass-judges ( ω≥0.60 ); FA = mean false-approval rate on 16 reject judges ( ω<0.50 ); WDR = mean over all 40 judges. Bold = best design per ρ .
Design
Category
ρ=0.05
0.10
0.25
0.50
Random
Strong pass (20)
.290
.238
.122
.021
Borderline pass (4)
.597
.597
.599
.464
Reject (16)
.112
.082
.071
.079
Strat
Strong pass (20)
.142
.087
.027
.007
Borderline pass (4)
.440
.389
.323
.295
Reject (16)
.172
.136
.114
.110
Table 4: Borderline vs. strong judges across designs: mean wrong-decision rate at increasing ρ (aggregated over the four real evaluation matrices). Strong pass ( ω≥0.60 , 20 judges), borderline pass ( ω=0.50 , 4 judges), reject ( ω<0.50 , 16 judges). Bold = best design per category and ρ .
Min. ρ for Pr[∣F^m−F∗∣>δ]≤0.05
L
Prevalence
δ=0.02
δ=0.05
δ=0.10
2
Uniform
≥ 75%
25%
5–10%
Skewed (.8/.2)
≥ 75%
25%
5–10%
5
Uniform
≥ 75%
15–25%
5–10%
Skewed (.7/.1/.1/.05/.05)
≥ 75%
15–50%
5–10%
Table 5: Minimum overlap for a precision target on the human-pool score F^m using p^o under per-rater sparse overlap ( n=500 , K=4 , F∗∈[0.6,0.9] , 300 trials, αsig=0.05 ). Ranges span the target F∗ values.
Figure 3: Reliability vs. annotator count at five overlap rates ( n=500 , ph=0.80 ). Reliability = % of trials within ±0.05 of ground truth. Grey line marks 50%.
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
Strategy
Items with ≥2 human labels
Human agreement ( α^ )
Bias vs. full
Wrong decisions
Depth-first
50
0.535
+ 0.184
75.0%
Random spread
≈14
0.340±0.183
− 0.011
39.9%
Stratified shared overlap
50
0.353±0.062
+ 0.001
26.5%
Appendix
Table 6: Exact numbers for the motivating depth-vs.-breadth example on ISIC Lesion. Agreement is measured by Krippendorff’s α ; the dense reference is α^dense=0.351 .
Parameter
Main profile
Sensitivity profile
Corpus size n
500
100–2 000
Human annotators K
4
2–10
Label categories L
2
2 or 5
Human accuracy ph
0.85
0.80
Tolerance ε
0.05
0.05
Monte Carlo trials
300
50–500
Appendix
Table 7: Default synthetic-experiment parameter profiles. Every caption states the exact settings used; this table serves as a quick reference.
F (pass rate, %)
Prevalence
pj (gt)
ρ
p^o
α^
κ^
AC1
Uniform
0.90 (pass)
0.05
94.7
87.7
87.5
87.8
0.10
98.7
94.8
94.8
94.7
0.25
100.0
98.7
98.7
99.0
0.95 (pass)
0.05
99.8
98.7
98.7
98.7
0.10
99.8
99.3
99.3
99.3
Appendix
Table 8: Pass rates (%) for ω≥0.5 when the base coefficient F is swapped. Synthetic L=2 , n=500 , K=4 , ph=0.85 , ε=0.05 , averaged over random and stratified designs, 300 trials per condition.
Table 10: Unmasked real-benchmark family pivot (mean wrong-decision rate across reject-judges, lower is better). “ − ” marks entries where Seq-Coverage could not be evaluated on SummEval because its reduced-coverage trials degenerated under SummEval’s three-annotator setting. Bold marks the per-row winner.
Dataset
πmax
ρ
Random
Strat
Seq-Ref.
Seq-Cov.
summeval@ t=3
0.55
.05
.000
.000
.000
.000
.10
.000
.000
.000
.000
.25
.000
.000
.000
.000
lesion@ t=0
0.65
.05
.110
.075
.045
.075
.10
.010
.005
.005
.010
.25
.000
.000
.000
.000
Appendix
Table 11: Skew-masked real-benchmark family pivot (300 trials per condition, mean wrong-decision rate across reject judges). Bold marks the per-row winner.
Judge
WAX
CeBaB-asp.
CeBaB-stars
SummEval
GPT-5.4
.293
.899
.635
.386
Gemini Pro
.374
.877
.585
.201
GPT-5.2
.305
.860
.675
.378
Cl. Sonnet 4.5
.325
.868
.585
.342
Cl. Opus 4.5
.350
.865
.531
.349
GPT-4o
.264
.860
.655
.263
Appendix
Table 12: Ground-truth p^o scores (dense reference, ρ=1 ) for each judge–benchmark pair. Judges sorted by mean score across the four benchmarks.
Judge (decision)
ρ
Rand.
Strat
Seq-R
Seq-C
Type
GPT-5.2 (pass, ω=1.00 )
.05
.130
.037
.103
.217
FR
.10
.037
.017
.037
.030
FR
.25
.000
.000
.000
.000
FR
.50
.000
.000
.000
.000
FR
GPT-4o (pass, ω=1.00 )
.05
.113
.020
.180
.223
FR
.10
.040
.007
.027
.060
FR
Appendix
Table 13: Wrong-decision rates on CeBaB star ratings ( F=p^o , ε=0.05 , 300 trials). Bold = best design per row.
Judge (decision)
ρ
Rand.
Strat
Seq-R
Seq-C
Type
Gemini Pro (pass, ω=.75 )
.05
.370
.033
.360
.433
FR
.10
.303
.023
.343
.493
FR
.25
.033
.000
.037
.087
FR
.50
.000
.000
.000
.000
FR
Gemini Flash (pass, ω=.75 )
.05
.553
.180
.643
.623
FR
.10
.510
.073
.503
.650
FR
Appendix
Table 14: Wrong-decision rates on WAX ( F=p^o , ε=0.05 , 300 trials). Bold = best design per row.
Judge (decision)
ρ
Rand.
Strat
Seq-R
Seq-C
Type
GPT-5.4 (pass, ω=.90 )
.05
.097
.090
.147
.293
FR
.10
.113
.040
.073
.160
FR
.25
.017
.010
.033
.017
FR
.50
.000
.000
.000
.003
FR
Cl. Sonnet 4.5 (pass, ω=.90 )
.05
.163
.090
.123
.380
FR
.10
.107
.047
.143
.317
FR
Appendix
Table 15: Wrong-decision rates on CeBaB aspects ( F=p^o , ε=0.05 , 300 trials). Bold = best design per row.
Figure 4: Reliability vs. annotator count at five overlap rates ( n=500 , ph=0.80 ). Reliability = % of trials within ±0.05 of ground truth. Grey line marks 50%.
Table 19: Comparison with related methods. “Timing” indicates whether the method operates before (prospective) or after (post-hoc) data collection. “Estimand” is what is being estimated or optimised.
Correlation
ρ=0.05
0.10
0.25
0.50
i.i.d. (500 clusters)
99.0
97.0
97.7
100.0
Mild (50 clusters)
98.0
96.7
99.3
99.7
Moderate (20 clusters)
99.7
98.3
99.0
100.0
Strong (10 clusters)
10.0
31.3
72.3
99.0
Appendix
Table 20: Sensitivity to correlated items. Reliability (% within ±0.05 of dense estimate) under varying item correlation ( n=500 , K=5 , L=2 , ph=0.85 , 300 trials).