cs.AISep 25, 2026

LLM Judge Validation Under Sparse Overlap: From Inference to Design

Authors: Junxuan Li, Arko Mukherjee, Soumyabrata Pal

Organizations: Adobe · Adobe Research

Abstract

Validating an LLM-as-a-judge requires estimating its agreement with humans, yet annotation budgets rarely allow every item to be multiply labeled. We prove that this overlap sparsity is the first-order determinant of wrong deployment decisions: at 5% pairwise overlap, wrong-decision rates reach 25% and the probability of selecting the wrong best judge among ten candidates is 65%. The two actionable levers are overlap quantity and allocation. For quantity, we derive a minimum-overlap formula showing ρ≥0.25ρ\geq 0.25 suffices for non-borderline judges while borderline cases remain fundamentally hard. For allocation, a zero-cost stratified scheme halves false-rejection rates relative to random sampling when strata are informative. We validate on 10 LLM judges across four evaluation matrices spanning visual assessment, causal reasoning, and summarization.

Figures & tables

Appendix figures & tables16 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Agreement Metrics for LLM-as-Judge Evaluation: What to Report and Why

    May 25, 2026Delip Rao, Chris Callison-BurchLarge Language Model JudgesJudgement

  2. How Much Do LLM-as-a-Judge Design Choices Matter? A Systematic Comparison of Prompt Designs, Rating Scales, and Models

    Oct 4, 2026Laurène Vaugrante, Thilo Hagendorff