LLM Judge Validation Under Sparse Overlap: From Inference to Design
Organizations: Adobe · Adobe Research
Abstract
Validating an LLM-as-a-judge requires estimating its agreement with humans, yet annotation budgets rarely allow every item to be multiply labeled. We prove that this overlap sparsity is the first-order determinant of wrong deployment decisions: at 5% pairwise overlap, wrong-decision rates reach 25% and the probability of selecting the wrong best judge among ten candidates is 65%. The two actionable levers are overlap quantity and allocation. For quantity, we derive a minimum-overlap formula showing suffices for non-borderline judges while borderline cases remain fundamentally hard. For allocation, a zero-cost stratified scheme halves false-rejection rates relative to random sampling when strata are informative. We validate on 10 LLM judges across four evaluation matrices spanning visual assessment, causal reasoning, and summarization.
Figures & tables
| of | Rel. ( ) | ||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Dataset ( ) | R | S | SR | SC | R | S | SR | SC | R | S | SR | SC | |
| cebab_stars (.26) | .05 | .087 | .001 | .087 | .105 | .126 | .074 | .124 | .117 | .22 | .49 | .24 | .22 |
| .10 | .069 | .004 | .070 | .123 | .080 | .051 | .079 | .079 | .33 | .64 | .32 | .13 | |
| .25 | .039 | .001 | .044 | .049 | .040 | .028 | .040 | .042 | .58 | .93 | .55 | .51 | |
| summeval (.56) | .05 | .020 | .034 | .016 | .016 | .030 | .026 | .026 | .027 | .82 | .70 | .88 | .90 |
| .10 | .014 | .031 | .013 | .012 | .019 | .018 | .018 | .017 | .95 | .84 | .97 | .98 | |
| WAX | CeBaB-asp. | CeBaB-stars | SummEval | Mean | ||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Rank | Top-1 | Rank | Top-1 | Rank | Top-1 | Rank | Top-1 | Rank | Top-1 | |
| 0.05 | .297 | .670 | .349 | .887 | .342 | .875 | .122 | .177 | .278 | .652 |
| 0.10 | .234 | .600 | .279 | .777 | .255 | .787 | .083 | .080 | .213 | .561 |
| 0.25 | .146 | .427 | .196 | .507 | .156 | .517 | .049 | .003 | .137 | .363 |
| 0.50 | .071 | .270 | .138 | .213 | .085 | .350 | .031 | .000 | .082 | .208 |
| 0.75 | .031 | .103 | .091 | .033 | .053 | .110 | .018 | .000 | .048 | .062 |
| Design | Mean FR | Mean FA | Mean WDR | |
|---|---|---|---|---|
| 0.05 | Random | .290 | .112 | .250 |
| Strat | .142 | .172 | .184 | |
| Seq-Refined | .308 | .098 | .250 | |
| Seq-Coverage | .435 | .066 | .311 | |
| 0.10 | Random | .238 | .082 | .212 |
| Strat | .087 | .136 | .137 |
| Design | Category | ||||
|---|---|---|---|---|---|
| Random | Strong pass (20) | .290 | .238 | .122 | .021 |
| Borderline pass (4) | .597 | .597 | .599 | .464 | |
| Reject (16) | .112 | .082 | .071 | .079 | |
| Strat | Strong pass (20) | .142 | .087 | .027 | .007 |
| Borderline pass (4) | .440 | .389 | .323 | .295 | |
| Reject (16) | .172 | .136 | .114 | .110 |
| Min. for | ||||
|---|---|---|---|---|
| Prevalence | ||||
| 2 | Uniform | 75% | 25% | 5–10% |
| Skewed (.8/.2) | 75% | 25% | 5–10% | |
| 5 | Uniform | 75% | 15–25% | 5–10% |
| Skewed (.7/.1/.1/.05/.05) | 75% | 15–50% | 5–10% | |
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
| Strategy | Items with human labels | Human agreement ( ) | Bias vs. full | Wrong decisions |
|---|---|---|---|---|
| Depth-first | 50 | 0.535 | 0.184 | 75.0% |
| Random spread | 0.011 | 39.9% | ||
| Stratified shared overlap | 50 | 0.001 | 26.5% |
| Parameter | Main profile | Sensitivity profile |
|---|---|---|
| Corpus size | 500 | 100–2 000 |
| Human annotators | 4 | 2–10 |
| Label categories | 2 | 2 or 5 |
| Human accuracy | 0.85 | 0.80 |
| Tolerance | 0.05 | 0.05 |
| Monte Carlo trials | 300 | 50–500 |
| (pass rate, %) | ||||||
|---|---|---|---|---|---|---|
| Prevalence | (gt) | |||||
| Uniform | (pass) | 94.7 | 87.7 | 87.5 | 87.8 | |
| 98.7 | 94.8 | 94.8 | 94.7 | |||
| 100.0 | 98.7 | 98.7 | 99.0 | |||
| (pass) | 99.8 | 98.7 | 98.7 | 98.7 | ||
| 99.8 | 99.3 | 99.3 | 99.3 | |||
| Metric | |||||
|---|---|---|---|---|---|
| (nominal) | 0.70 | 37.3 | 29.2 | 19.9 | 11.6 |
| 0.76 | 43.9 | 41.0 | 34.0 | 27.5 | |
| 0.80 | 33.9 | 29.3 | 19.6 | 9.0 | |
| 0.85 | 20.5 | 10.6 | 2.8 | 0.5 | |
| 0.90 | 8.8 | 2.8 | 0.0 | 0.0 | |
| Weighted | 0.70 | 38.7 | 31.0 | 22.7 | 13.1 |
| Dataset | Random | Strat | Seq-Ref. | Seq-Cov. | ||
|---|---|---|---|---|---|---|
| cebab_stars | 0.26 | .05 | .727 | .552 | .680 | .648 |
| .10 | .513 | .442 | .557 | .470 | ||
| .25 | .240 | .258 | .270 | .275 | ||
| summeval | 0.55 | .05 | .163 | .000 | .168 | |
| .10 | .024 | .000 | .028 | |||
| .25 | .000 | .000 | .000 |
| Dataset | Random | Strat | Seq-Ref. | Seq-Cov. | ||
|---|---|---|---|---|---|---|
| summeval@ | 0.55 | .05 | .000 | .000 | .000 | .000 |
| .10 | .000 | .000 | .000 | .000 | ||
| .25 | .000 | .000 | .000 | .000 | ||
| lesion@ | 0.65 | .05 | .110 | .075 | .045 | .075 |
| .10 | .010 | .005 | .005 | .010 | ||
| .25 | .000 | .000 | .000 | .000 |
| Judge | WAX | CeBaB-asp. | CeBaB-stars | SummEval |
|---|---|---|---|---|
| GPT-5.4 | .293 | .899 | .635 | .386 |
| Gemini Pro | .374 | .877 | .585 | .201 |
| GPT-5.2 | .305 | .860 | .675 | .378 |
| Cl. Sonnet 4.5 | .325 | .868 | .585 | .342 |
| Cl. Opus 4.5 | .350 | .865 | .531 | .349 |
| GPT-4o | .264 | .860 | .655 | .263 |
| Judge (decision) | Rand. | Strat | Seq-R | Seq-C | Type | |
|---|---|---|---|---|---|---|
| GPT-5.2 (pass, ) | .05 | .130 | .037 | .103 | .217 | FR |
| .10 | .037 | .017 | .037 | .030 | FR | |
| .25 | .000 | .000 | .000 | .000 | FR | |
| .50 | .000 | .000 | .000 | .000 | FR | |
| GPT-4o (pass, ) | .05 | .113 | .020 | .180 | .223 | FR |
| .10 | .040 | .007 | .027 | .060 | FR |
| Judge (decision) | Rand. | Strat | Seq-R | Seq-C | Type | |
|---|---|---|---|---|---|---|
| Gemini Pro (pass, ) | .05 | .370 | .033 | .360 | .433 | FR |
| .10 | .303 | .023 | .343 | .493 | FR | |
| .25 | .033 | .000 | .037 | .087 | FR | |
| .50 | .000 | .000 | .000 | .000 | FR | |
| Gemini Flash (pass, ) | .05 | .553 | .180 | .643 | .623 | FR |
| .10 | .510 | .073 | .503 | .650 | FR |
| Judge (decision) | Rand. | Strat | Seq-R | Seq-C | Type | |
|---|---|---|---|---|---|---|
| GPT-5.4 (pass, ) | .05 | .097 | .090 | .147 | .293 | FR |
| .10 | .113 | .040 | .073 | .160 | FR | |
| .25 | .017 | .010 | .033 | .017 | FR | |
| .50 | .000 | .000 | .000 | .003 | FR | |
| Cl. Sonnet 4.5 (pass, ) | .05 | .163 | .090 | .123 | .380 | FR |
| .10 | .107 | .047 | .143 | .317 | FR |
| Prevalence | ||||
|---|---|---|---|---|
| 2 | (.95, .05) | 50% | 16–50% | 5–22% |
| 5 | (.85, .05, .04, .03, .03) | 49% | 13–42% | 4–16% |
| Prevalence | |||||
|---|---|---|---|---|---|
| 100 | Uniform | 20.6 | 28.8 | 47.8 | 73.4 |
| Skewed (.90) | 10.6 | 16.0 | 39.2 | 55.0 | |
| 200 | Uniform | 25.4 | 37.8 | 61.4 | 86.6 |
| Skewed (.90) | 26.2 | 34.8 | 52.6 | 84.6 | |
| 500 | Uniform | 36.4 | 55.4 | 81.2 | 98.2 |
| Skewed (.90) | 30.4 | 50.6 | 71.4 | 97.0 |
| Overlap rate | |||||
|---|---|---|---|---|---|
| Budget | 0.05 | 0.10 | 0.25 | 0.40 | 0.50 |
| 500 | 44.5 | 38.7 | 32.0 | 30.3 | 25.9 |
| 1000 | 44.0 | 40.9 | 32.5 | 28.5 | 23.3 |
| 2000 | 43.2 | 35.6 | 30.3 | 28.3 | 20.7 |
| 5000 | 38.9 | 41.6 | 32.7 | 27.9 | 22.6 |
| Method | Estimand | Timing | Data structure | Overlap concept |
| StratPPI [ Fisch et al., 2024 ] | Population mean quality | Post-hoc | 1 human + 1 model pred. per item | Absent |
| PPI [ Angelopoulos et al., 2023 ] | Population statistic | Post-hoc | Same as StratPPI | Absent |
| SPA [ Nørregaard and Derczynski, 2022 ] | Agreement coeff. from sparse matrix | Post-hoc | Already-collected sparse matrix | Central but fixed |
| BIBD [ Fleiss et al., 2003 ] | Label allocation | Prospective | Symmetric rater pool | Symmetric blocks |
| Active allocation [ Sheng et al., 2008 ] | Redundancy for label quality | Post-hoc / online | Any; targets aggregation quality | Implicit |
| Our STRAT | Reliable deployment decision | Prospective | Asymmetric: judge + primary complete; | Central — we design it |
| Correlation | ||||
|---|---|---|---|---|
| i.i.d. (500 clusters) | 99.0 | 97.0 | 97.7 | 100.0 |
| Mild (50 clusters) | 98.0 | 96.7 | 99.3 | 99.7 |
| Moderate (20 clusters) | 99.7 | 98.3 | 99.0 | 100.0 |
| Strong (10 clusters) | 10.0 | 31.3 | 72.3 | 99.0 |