Conditional Accuracy Profiles: Diagnosing LLM Judges across Deployment Conditions
Organizations: University at Albany, State University of New York · West Virginia University
Abstract
LLM-as-judge is now a standard tool for scalable evaluation, but judge performance is still often summarized by a single accuracy number. This aggregate view hides the deployment conditions under which a judge succeeds or fails. We introduce \textbf{Conditional Accuracy Profiling} (CAP), a post-hoc diagnostic framework that decomposes pairwise LLM-judge accuracy into eight conditions organized into content sensitivity, robustness, and rationale quality. CAP is benchmark-agnostic: it can be applied directly when a benchmark provides the required annotations, approximately through task-subset proxies, or through controlled augmentation when perturbation pairs can be generated. We instantiate CAP on seven LLM judges across six pairwise judging benchmarks, including \textsc{judgerEva-Standard}, a controlled testbed we created to support all eight conditions. CAP exposes profile differences hidden by aggregate accuracy: on \textsc{judgerEva}'s judge-independent Hard-Constructed subset, the two judges most sensitive to omitted qualifications rank in the bottom three of seven by overall accuracy, so omission sensitivity is not predicted by aggregate accuracy. Across benchmarks, Position Robustness shows the strongest rank stability (mean Spearman ) but is itself fragile under JudgeBench-Pro adversarial stress, showing the largest mean accuracy drop among the shared conditions, though the dominant degradation channel varies by judge. Condition-level profiles provide a more actionable basis than aggregate accuracy for selecting LLM judges.
Figures & tables
| # | Category | Condition | Subset | Measurement |
|---|---|---|---|---|
| 1 | Content | Factual Sensitivity | factual-fabrication items | |
| 2 | Content | Logical Sensitivity | logical-reversal items | |
| 3 | Content | Scope-Overclaim | overclaim items | |
| 4 | Content | Scope-Omission | omission items | |
| 5 | Robustness | Style Robustness | plain/polished variants | |
| 6 | Robustness | Length Robustness | plain/padded variants |
| # | Condition | judgerEva | JudgeBench | JB-Pro | RewardBench | MT-Bench | LLMBar |
|---|---|---|---|---|---|---|---|
| 1 | Factual Sens. | Direct | Approx. | Approx. | Approx. | Approx. | — |
| 2 | Logical Sens. | Direct | Approx. | Approx. | Approx. | Approx. | — |
| 3 | Scope-Overclaim | Direct | — | — | — | — | — |
| 4 | Scope-Omission | Direct | — | — | — | — | — |
| 5 | Style Rob. | Direct | Aug. † | — | — | — | Direct |
| 6 | Length Rob. | Direct | Aug. † | — | — | — | — |
| Judge | C1 | C2 | C3 | C4 | C5 | C6 | C7 | C8 † |
|---|---|---|---|---|---|---|---|---|
| GPT-4o | 40.5 (34,46) | 53.1 (38,69) | 63.2 (51,75) | 64.1 (53,74) | 36.7 (31,42) | 20.4 (16,26) | 25.2 (22,28) | 66.7 (33,100) |
| GPT-4o-mini | 12.5 (8,17) | 53.1 (34,72) | 48.7 (41,57) | 62.8 (53,74) | 15.9 (11,21) | 10.6 (7,15) | 14.8 (12,17) | 75.0 (25,100) |
| Claude Sonnet 4.5 | 92.4 (88,96) | 86.7 (70,100) | 89.5 (80,96) | 59.2 (46,72) | 84.7 (80,89) | 84.7 (80,89) | 86.0 (83,88) | 85.7 (57,100) |
| Claude Haiku 4.5 | 82.2 (77,87) | 65.6 (44,84) | 96.1 (89,100) | 70.5 (56,83) | 71.8 (66,77) | 69.0 (63,75) | 73.4 (70,77) | 91.7 (75,100) |
| Gemini 2.5 Flash | 76.3 (70,82) | 78.1 (59,94) | 72.4 (59,83) | 56.4 (44,68) | 71.4 (65,77) | 66.4 (61,72) | 63.4 (60,67) | 44.4 (11,78) |
| Gemini 2.5 Pro | 73.2 (67,79) | 78.1 (63,91) | 90.8 (83,97) | 61.5 (49,73) | 77.5 (72,83) | 70.6 (65,76) | 62.3 (59,66) | 83.3 (50,100) |
| Judge | Overall | C1 | C4 | C7 |
|---|---|---|---|---|
| Claude Sonnet 4.5 | 76.3 (70,82) | 80.8 (69,91) | 36.7 (23,52) | 71.2 (63,78) |
| Gemini 2.5 Flash | 72.3 (65,79) | 71.8 (59,86) | 38.3 (23,53) | 66.2 (58,75) |
| Gemini 2.5 Pro | 71.2 (64,78) | 66.7 (53,78) | 38.3 (22,53) | 65.7 (58,74) |
| Claude Haiku 4.5 | 70.1 (63,77) | 76.9 (65,88) | 33.3 (18,48) | 62.6 (55,71) |
| Llama-3.3-70B | 68.3 (62,75) | 56.4 (44,69) | 56.7 (42,72) | 58.3 (51,66) |
| GPT-4o | 59.7 (53,67) | 42.3 (28,56) | 36.7 (23,52) | 51.1 (42,60) |
| Judge | C1 ≈ | C2 ≈ | C3 | C4 | C5 | C6 | C7 | C8 |
|---|---|---|---|---|---|---|---|---|
| GPT-4o | 60.7 (54,68) | 62.8 (54,70) | — | — | n.r. | n.r. | 51.1 (46,57) | n.r. |
| GPT-4o-mini | 60.7 (55,67) | 52.0 (47,57) | — | — | n.r. | n.r. | 35.7 (31,41) | n.r. |
| Claude Sonnet 4.5 | 78.8 (74,84) | 74.5 (67,82) | — | — | n.r. | n.r. | 68.3 (63,73) | n.r. |
| Claude Haiku 4.5 | 73.1 (67,79) | 70.4 (63,78) | — | — | n.r. | n.r. | 66.4 (61,71) | n.r. |
| Gemini 2.5 Flash | 77.6 (71,84) | 92.0 (86,97) | — | — | n.r. | n.r. | 81.6 (76,86) | n.r. |
| Gemini 2.5 Pro | 84.8 (79,90) | 93.5 (89,98) | — | — | n.r. | n.r. | 86.8 (83,90) | n.r. |
| Judge | C1 ≈ | C2 ≈ | C7 | C8 | Mean |
|---|---|---|---|---|---|
| GPT-4o | +37.8 | +40.4 | +40.1 | n.r. | +39.4 |
| GPT-4o-mini | +34.4 | +23.3 | +28.2 | n.r. | +28.6 |
| Claude Sonnet 4.5 | +22.0 | +35.4 | +33.8 | n.r. | +30.4 |
| Claude Haiku 4.5 | +43.4 | +46.8 | +50.8 | n.r. | +47.0 |
| Gemini 2.5 Flash | +14.4 | +14.0 | +18.5 | n.r. | +15.6 |
| Gemini 2.5 Pro | +17.7 | + -0.4 | +15.4 | n.r. | +10.9 |
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
| Domain | Source | Items |
| Medical | MedQA-USMLE ( Jin et al., 2021 ) | 300 |
| Legal | MMLU Professional Law ( Hendrycks et al., 2021 ) | 296 |
| Math | GSM8K ( Cobbe et al., 2021 ) | 356 |
| Science | MMLU College Biology/Chem./Physics | 300 |
| Logic | MMLU Logical Fallacies | 300 |
| General | TruthfulQA ( Lin et al., 2022 ) | 540 |
| Domain | Fact. | Log. | Over. | Omis. | Total |
|---|---|---|---|---|---|
| Medical | 75 | 75 | 75 | 75 | 300 |
| Legal | 74 | 74 | 74 | 74 | 296 |
| Science | 75 | 75 | 75 | 75 | 300 |
| Logic | 75 | 75 | 75 | 75 | 300 |
| General | 135 | 135 | 135 | 135 | 540 |
| Math | 89 | 89 | 89 | 89 | 356 |
| Error type | Hard-Emp. | Hard-Cons. |
|---|---|---|
| Factual Fabrication | 152 (62.0%) | 39 (28.1%) |
| Logical Reversal | 16 (6.5%) | 35 (25.2%) |
| Overclaim | 38 (15.5%) | 35 (25.2%) |
| Critical Omission | 39 (15.9%) | 30 (21.6%) |
| Total plain items | 245 | 139 |
| Judge | C1 | C2 | C3 | C4 | C5 | C6 | C7 |
|---|---|---|---|---|---|---|---|
| GPT-4o | |||||||
| GPT-4o-mini | |||||||
| Claude Sonnet 4.5 | |||||||
| Claude Haiku 4.5 | |||||||
| Gemini 2.5 Pro | |||||||
| Gemini 2.5 Flash |
| Breakdown | n | Agreement |
|---|---|---|
| By error type | ||
| factual_fabrication | 17 | 1.000 |
| overclaim | 37 | 0.946 |
| logical_reversal | 24 | 0.917 |
| critical_omission | 42 | 0.857 |
| By judge | ||
| Benchmark | Cond. | Status | Operator |
| judgerEva | C1–C4 | Direct | planted-error subsets ( plain items each); PA over A/B and B/A |
| C5, C6 | Direct | constructed polished / length-padded ( ) variants; BC vs. plain | |
| C7 | Direct | A/B and B/A orderings of every item; BC | |
| C8 | Direct | LLM evaluator on correct-verdict rationales (Appendix C ) | |
| JudgeBench | C1 | Approx. | Knowledge subset ( ); PA |
| C2 | Approx. | Reasoning subset ( ); PA |
| jE | JB | JBP | RB | MT | LB | |
|---|---|---|---|---|---|---|
| jE | — | 0.71 | 0.71 | 0.89 | 0.71 | 0.86 |
| JB | — | 1.00 | 0.89 | 0.93 | 0.86 | |
| JBP | — | 0.89 | 0.93 | 0.86 | ||
| RB | — | 0.93 | 0.96 | |||
| MT | — | 0.86 | ||||
| LB | — |
| C1 | C2 | C3 | C4 | C5 | C6 | C7 | |
|---|---|---|---|---|---|---|---|
| C2 | |||||||
| C3 | |||||||
| C4 | |||||||
| C5 | |||||||
| C6 | |||||||
| C7 |
| C5 Style | C6 Length | C7 Position | |||||||
|---|---|---|---|---|---|---|---|---|---|
| Judge | strict | cons. | strict | cons. | strict | cons. | |||
| GPT-4o | 90.3 | 95.5 | 5.2 | 62.0 | 66.7 | 4.7 | 72.9 | 87.3 | 14.4 |
| GPT-4o-mini | 82.8 | 90.6 | 7.8 | 51.0 | 58.6 | 7.6 | 61.0 | 80.9 | 19.9 |
| Claude Sonnet 4.5 | 93.7 | 95.1 | 1.4 | 92.9 | 94.1 | 1.2 | 95.3 | 96.9 | 1.6 |
| Claude Haiku 4.5 | 93.0 | 95.3 | 2.3 | 85.5 | 87.8 | 2.3 | 90.6 | 93.7 | 3.1 |
| Gemini 2.5 Flash | 93.2 | 95.4 | 2.2 | 82.0 | 84.1 | 2.1 | 86.5 | 92.5 | 6.0 |
| Judge | C1 | C2 | C3 | C4 | C7 |
|---|---|---|---|---|---|
| GPT-4o | |||||
| GPT-4o-mini | |||||
| Claude Sonnet 4.5 | |||||
| Claude Haiku 4.5 | |||||
| Gemini 2.5 Flash | |||||
| Gemini 2.5 Pro |
| Pair | jE | JB | JBP | RB | MT | LB |
|---|---|---|---|---|---|---|
| OpenAI (WF) | 93.1 | 26.2 | 13.2 | n.r. | n.r. | n.r. |
| Anthropic (WF) | 96.4 | 11.7 | 61.5 | n.r. | n.r. | n.r. |
| Cross-Family avg. | 208.1 | 63.3 | 93.4 | n.r. | n.r. | n.r. |
| Judge | Overall | C1 | C2 | C3 | C4 |
|---|---|---|---|---|---|
| GPT-4o | † | † | † | † | † |
| GPT-4o-mini | † | † | † | † | † |
| Claude Sonnet 4.5 | † | † | † | ||
| Claude Haiku 4.5 | † | † | |||
| Gemini 2.5 Flash | † | † | † | † | † |
| Gemini 2.5 Pro | † | † | † | † | † |
| C1 | C2 | C3 | C4 | C5 | C6 | C7 | C8 | |
|---|---|---|---|---|---|---|---|---|
| GPT-4o | 79.6 | 94.9 | 96.6 | 93.7 | 90.3 | 62.0 | 72.5 | 80.0 |
| Peer mean | 90.9 | 97.2 | 98.0 | 91.9 | 94.2 | 82.0 | 86.3 | 91.2 |
| Judge | ||||||||
|---|---|---|---|---|---|---|---|---|
| GPT-4o | 523 | 523 | 523 | 523 | 2092 | 2092 | 6276 | 100 |
| GPT-4o-mini | 523 | 523 | 523 | 523 | 2092 | 2092 | 6276 | 100 |
| Claude Haiku 4.5 | 523 | 523 | 523 | 523 | 2091 | 2092 | 6274 | 100 |
| Claude Sonnet 4.5 | 523 | 523 | 523 | 523 | 2092 | 2092 | 6276 | 100 |
| Gemini 2.5 Flash | 521 | 521 | 522 | 521 | 2053 | 2047 | 6139 | 100 |
| Gemini 2.5 Pro | 522 | 523 | 521 | 523 | 2062 | 2057 | 6160 | 100 |