All Verdicts are Not Equal: Rethinking LLM Judge Reliability
Organizations: PayPal Artificial Intelligence PayPal, Bengaluru, India
Abstract
LLM-as-a-Judge is the standard paradigm for NLP evaluation, yet its systemic reliability remains poorly understood despite being widely treated as a deterministic ground truth. We present a comprehensive reliability audit, stresstesting six frontier models across four benchmarks, five prompt formats, two presentation orders, three sampling temperatures, and ten repetitions per condition. Our empirical analysis reveals severe vulnerabilities: verdicts change across identical replications at temperature zero, position-order swaps flip the majority of verdicts on challenging tasks, and the most deterministic judge achieves perfect consistency by trivially repeating incorrect verdicts, agreeing with ground truth only 51% of the time. To formalize these multi-faceted failure modes, we introduce the trustworthy verdict rate (T ), a unified metric capturing the joint probability that an evaluation is reproducible, order-invariant, and accurate. UsingT , we derive a theoretical upper bound on accuracy imposed by position bias and show that reliability is item-specific rather than modellevel. Finally, we demonstrate that shifting from pairwise win-rate to holistic rubric scoring improves trustworthiness more than any single-format prompting intervention, offering an actionable framework for robust NLP evaluation.
Figures & tables
| Judge | Format | MT-Bench | AlpacaFarm | JudgeBench | LLMBar | ||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| GPT-5.2 | Direct | 89.3 | 5.1 | 64.8 | 55.6 | 82.7 | 9.2 | 49.5 | 37.1 | 86.0 | 11.0 | 64.5 | 50.7 | 89.0 | 9.0 | 75.5 | 63.2 |
| VF – JSON | 86.2 | 7.1 | 64.3 | 52.4 | 82.1 | 6.1 | 52.0 | 40.2 | 81.0 | 10.0 | 65.0 | 48.6 | 89.0 | 9.0 | 75.5 | 63.2 | |
| VF – XML | 89.8 | 11.2 | 65.3 | 53.6 | 86.1 | 9.3 | 54.1 | 42.6 | 82.5 | 12.0 | 64.5 | 48.3 | 92.0 | 6.0 | 76.0 | 67.2 | |
| CoT – JSON | 82.1 | 6.1 | 66.8 | 52.3 | 80.4 | 15.5 | 57.7 | 40.2 | 79.0 | 13.0 | 70.5 | 50.6 | 92.5 | 13.0 | 81.5 | 69.4 | |
| CoT – XML | 84.2 | 8.2 | 65.8 | 52.0 | 81.4 | 12.4 | 56.7 | 41.1 | 75.5 | 12.0 | 68.0 | 46.8 | 90.5 | 11.1 | 82.4 | 69.5 | |
| Judge | JudgeBench – Direct | JudgeBench – FLASK | JudgeBench – G-Eval | LLMBar – Direct | LLMBar – FLASK | LLMBar – G-Eval | ||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| GPT-5.2 | 86.0 | 11.0 | 64.5 | 50.7 | 34.8 | 9.1 | 63.6 | 20.5 | 68.5 | 16.0 | 68.0 | 41.1 | 89.0 | 9.0 | 75.5 | 63.2 | 55.0 | 21.0 | 79.5 | 37.9 | 88.0 | 18.0 | 79.0 | 61.6 |
| GPT-4o | 94.5 | 24.0 | 51.0 | 36.9 | 44.4 | 47.5 | 52.5 | 12.8 | 89.0 | 61.0 | 62.0 | 28.0 | 98.5 | 8.0 | 71.5 | 66.5 | 53.0 | 31.0 | 75.0 | 31.5 | 95.5 | 28.0 | 78.0 | 61.1 |
| Gemini-2.5-Pro | 81.5 | 56.0 | 63.5 | 28.9 | 60.6 | 10.5 | 67.4 | 37.7 | 87.7 | 10.0 | 93.8 | 77.9 | 87.0 | 20.0 | 80.0 | 60.9 | 63.0 | 14.0 | 78.0 | 44.7 | 86.0 | 27.0 | 79.0 | 56.3 |
| Gemini-2.0-Flash | 99.0 | 46.0 | 55.5 | 32.2 | 90.4 | 30.3 | 57.1 | 37.9 | 94.5 | 40.4 | 59.0 | 36.7 | 98.0 | 23.0 | 61.5 | 49.0 | 66.0 | 32.0 | 65.5 | 32.7 | 99.0 | 39.0 | 74.0 | 54.0 |
| Claude-Sonnet-4.6 | 98.8 | 14.9 | 57.3 | 49.3 | 69.0 | 2.0 | 75.1 | 51.1 | 78.7 | 14.4 | 82.0 | 58.8 | 99.5 | 12.2 | 79.7 | 73.2 | 58.0 | 11.0 | 81.0 | 43.8 | 91.0 | 9.1 | 85.0 | 73.2 |
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
| Judge | Format | MT-Bench | AlpacaFarm | JudgeBench | LLMBar | |||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| C | A | C | A | C | A | C | A | |||||||||||
| GPT-4o | Direct | 98.0 | 11.2 | 64.8 | 58.0 | 95.5 | 9.1 | 58.6 | 51.6 | 94.5 | 24.0 | 51.0 | 36.9 | 98.5 | 8.0 | 71.5 | 66.5 | |
| Verdict-first | 97.4 | 12.8 | 63.0 | 55.1 | 94.2 | 13.1 | 60.6 | 50.9 | 93.8 | 33.0 | 53.5 | 34.7 | 97.0 | 6.5 | 75.2 | 69.8 | ||
| CoT | 90.3 | 17.9 | 61.2 | 47.2 | 92.6 | 21.8 | 59.6 | 45.1 | 88.0 | 46.5 | 54.0 | 27.1 | 92.8 | 22.0 | 72.0 | 56.6 | ||
| Gemini-2.0-Flash | Direct | 99.5 | 16.3 | 60.7 | 52.3 | 99.5 | 22.2 | 52.3 | 41.0 | 99.0 | 46.0 | 55.5 | 32.2 | 98.0 | 23.0 | 61.5 | 49.0 | |
| Verdict-first | 99.5 | 38.2 | 57.4 | 38.1 | 98.5 | 41.8 | 60.4 | 38.9 | 99.2 | 90.0 | 51.5 | 0 6.4 | 98.2 | 45.2 | 64.0 | 40.7 | ||
| Judge | MT Bench | AlpacaFarm | JudgeBench | LLMBar | Mean |
| GPT-5.2 | 4.8 | ||||
| GPT-4o | 11.7 | ||||
| Gemini-2.5-Pro | 24.9 | ||||
| Gemini-2.0-Flash | 19.8 | ||||
| Claude-Sonnet-4.6 | 10.8 | ||||
| Llama-3.3-70B | 7.1 |
| Source | (%) |
|---|---|
| Judge | 27.0 |
| Dataset | 26.2 |
| Judge Dataset | 28.1 |
| Judge Format | 8.3 |
| Format | 0.7 |
| Format Dataset | 1.3 |
| Main effects | Interactions + residual | ||||||
|---|---|---|---|---|---|---|---|
| Dataset | Res. | ||||||
| MT-Bench | 0.0 | 0.0 | 2.1 | 10.3 | 1.1 | 5.0 | 4.2 |
| AlpacaFarm | 2.1 | 0.0 | 1.8 | 17.5 | 1.4 | 3.8 | 6.0 |
| JudgeBench | 0.9 | 0.1 | 6.2 | 34.8 | 4.4 | 6.2 | 7.9 |
| LLMBar | 0.3 | 0.0 | 2.6 | 22.7 | 1.8 | 4.2 | 4.3 |
| Framework | Format | Rubric | Scores | Cost | Output |
|---|---|---|---|---|---|
| Pairwise | Direct | No | No | 1 | A / B |
| Pairwise | JSON VF | No | No | 1 | JSON |
| Pairwise | JSON CoT | No | No | 1 | JSON |
| Pairwise | XML VF | No | No | 1 | XML |
| Pairwise | XML CoT | No | No | 1 | XML |
| GEval | JSON CoT | Yes | 1–5 | 1 | JSON |
| Evaluation Paradigm | Cost (USD) |
|---|---|
| Pairwise Experiments | $1,450 |
| FLASK Rubric Experiments | $390 |
| G-Eval Experiments | $260 |
| Total | $2,100 |
| Dataset | Truth | Pairs | Domains | Sampling |
|---|---|---|---|---|
| MT-Bench | Subjective | 100 | 1 | Random (seed=42) |
| AlpacaFarm | Subjective | 100 | 1 | Random (seed=42) |
| JudgeBench | Objective | 100 | 17 | Stratified by domain |
| LLMBar | Objective | 100 | 5 | Stratified by split |
| Benchmark | Judge | Gap | Gap 95% CI | ||
|---|---|---|---|---|---|
| MT-Bench | GPT-5.2 | 0.54 | 0.58 | -0.04 | [-0.068, -0.012] |
| GPT-4o | 0.55 | 0.57 | -0.02 | [-0.041, -0.007] | |
| Gemini-2.5-Pro | 0.50 | 0.55 | -0.05 | [-0.072, -0.024] | |
| Gemini-2.0-Flash | 0.51 | 0.52 | -0.01 | [-0.013, -0.001] | |
| Claude-Sonnet-4.6 | 0.62 | 0.64 | -0.02 | [-0.043, -0.003] | |
| Llama-3.3-70B | 0.58 | 0.58 | 0.00 | [-0.006, +0.000] |