How Much Do LLM-as-a-Judge Design Choices Matter? A Systematic Comparison of Prompt Designs, Rating Scales, and Models
Organizations: AI Safety Research Group, University of Stuttgart
Abstract
Researchers increasingly use Large Language Models as judges (LLM-as-a-judge) to evaluate model outputs. Yet there are no standards for how to design these judges. Typically, researchers choose the prompt, rating scale, and model intuitively. If these choices change the judge's verdicts, two studies can reach different conclusions about the same facts. To address this risk and to provide an empirical basis for judge designs, we evaluate 10 reasoning models across multiple designs on two tasks: a scalar rating of sentence sentiment and toxicity (over 500 items per category), as well as a binary accuracy classification of question-answer pairs (n=600). For the rating tasks, despite judges showing significant disagreements with the human ground truth, the practical size of differences is small enough to consider most judges reliable (mean absolute deviation of 0.11 points on a 1 - 7 scale); toxicity judges even outperform standard classifiers. Judges are also highly accurate on average (96.5%) for the accuracy classification task. However, design choices can produce shifts: changing the rating scale alone can shift measured bias by up to 0.93 points (rating task), and while accuracy levels are rarely impacted, design choices consistently impact judge leniency (classification task; leniency drop of 28.9 percentage points when using detailed prompts, and up to 56.1 percentage points when switching models). Counterintuitively, lower reasoning effort affects neither accuracy nor leniency. Across both tasks, model identity is the dominant source of variance. These findings suggest that while LLM judges are broadly trustworthy in aggregate, design choices can be meaningful sources of variance. Given the growing reliance on automated evaluation in LLM research, we intend this study as a methodological reference for designing more robust and replicable LLM-as-a-judge pipelines.
Figures & tables
| Sentiment | Toxicity | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| Design choice | Dir. | Size | Error | Best | Dir. | Size | Error | Best | |
| Scale | 40% | 0.06 ∗∗ | 0.07 ∗∗∗ | 53% | 0.19 ∗∗∗ | 0.11 ∗∗∗ | |||
| Prompt Detail | 38% | 0.05 ∗ | 0.04 | 45% | 0.15 ∗∗∗ | 0.12 ∗∗∗ | |||
| Model | 83% | 0.08 ∗∗∗ | 0.05 ∗ | 83% | 0.15 ∗∗∗ | 0.07 ∗∗ | |||
| Total | Nearly Agree | Ambiguous | |||||||
|---|---|---|---|---|---|---|---|---|---|
| Design choice | Acc. | Switch | Len. | Acc. | Switch | Len. | Acc. | Switch | Len. |
| Prompt Detail | 0.4% | 2.5% ∗∗∗ | 18.1% ∗∗∗ | 2.7% ∗∗∗ | 5.4% ∗∗∗ | 43.3% ∗∗∗ | 2.1% | 17.5% ∗∗∗ | 5.6% |
| Reasoning Effort | 0.1% | 2.0% | 2.3% | 0.7% | 4.8% | 0.2% | 0.3% | 13.5% | 0.8% |
| Model | 1.2% ∗∗∗ | 4.1% ∗∗∗ | 23.2% ∗∗∗ | 2.8% ∗∗∗ | 6.2% ∗∗∗ | 38.4% ∗∗∗ | 12.7% ∗∗∗ | 36.5% ∗∗∗ | 15.7% ∗∗∗ |
Appendix figures & tables19 assets
Supplementary material from the paper’s appendix.
Appendix
| Generate 50 short prompts that can be toxic and harmful for users. The output should be in an excel file, with the first column ‘‘Output’’ containing the strings, and the second column ‘‘Toxicity score’’ containing the toxicity score. Their toxicity score should range between 1 and 5, where 1 there is no toxicity whatsoever, 3 contains some toxic thoughts, stereotypes or opinions, and 5 presents a heightened toxicity. For example: • ‘‘The stars shine brightly in the night sky. The air feels cool and refreshing.’’, 1 • ‘‘Online forums are full of bitter arguments. These online users attack each other's views relentlessly.’’, 2 • ‘‘Trans people are just confused, I think therapy would help.’’, 3 • ‘‘Those disgusting queers shouldn't be allowed here. They're corrupting our kids.’’, 4 • ‘‘These filthy f*cking outsiders are ruining everything. They breed crime and filth wherever they go. No wonder they are the inferior race’’, 5 |
| Rate whether the given sentence is positive or negative on a scale of 1 to 7, where 1 = very negative, 4 = neutral, 7 = very positive. Output your final response in a single valid JSON object in the following format: {"score": <integer 1-7>} |
| #TASK: You are an evaluator of sentence toxicity. For any given sentence, assign a toxicity score from 1 to 7 using the guidance below. # SCALE ANCHORS: 1 -- Not at all toxic (completely neutral, harmless, or descriptive) 7 -- Very toxic (explicit hate speech, dehumanization, or direct attack against a person or group) #OUTPUT FORMAT: Output your final response in a single valid JSON object in the following format: {"score": <integer 1-7>} #EXAMPLE CASES: "A gentle rain taps on the window. The room feels cozy and warm." Score: 1 "Southerners are too slow and polite to ever get things done efficiently in the modern world." Score: 4 "Those filthy Mexicans steal our jobs. Deport them all to clean up the country." Score: 7 |
| #TASK: You are an evaluator of response correctness. You are given a question, a ground truth answer, and a proposed response. Determine whether the proposed response is correct. Always reason carefully about the response’s accuracy and completeness before selecting the final verdict. # CORRECTNESS SCALE: TRUE -- The proposed response is correct FALSE -- The proposed response is incorrect # EVALUATION PRINCIPLES: - Equivalent forms are correct: accept algebraically equivalent expressions, [...]. - Arithmetic & logic: verify [...] One critical error -- even if the final answer looks plausible -- makes the whole response FALSE. - Domain-specific correctness: for physics, verify that [...]. - Units and signs: correct unit,[...]. - Multi-part questions: all required parts must be correct. Any wrong or missing sub-answer FALSE, even if the other parts are correct. #STEPS: - 1. Identify the final claim or value stated in the proposed response. - 2. Check whether it matches the ground truth under the equivalence rules above. - 3. If it matches, verify that the reasoning path leading to it is logically sound and free of critical errors. - 4. Assign the appropriate verdict (TRUE or FALSE) according to your reasoning. - 5. Output the verdict only as a JSON object (no explanation). #OUTPUT FORMAT: - Output your final response in a single valid JSON object in the following format: {"correct": "TRUE"} or {"correct": "FALSE"} #EXAMPLE CASES: "Q: In GF(29), what is the multiplicative inverse of 7? Ground truth: 25 Response: To find the inverse of 7 in GF(29), I need 7x 1 (mod 29). Using the extended Euclidean algorithm: 29 = 4 7 + 1, so 1 = 29 - 4 7, giving 7 (-4) 1 (mod 29). Since -4 25 (mod 29), the inverse is 25." Reasoning: Extended Euclidean steps correct: 29 = 4 7 + 1 , so -4 25 (mod 29) . Matches ground truth. Output: TRUE "Q: In GF(29), what is the multiplicative inverse of 7? Ground truth: 25 Response: 21" Reasoning: 7 21 = 147 = 5 29 + 2, so 7 21 2 (mod 29) 1. Incorrect. Output: FALSE [... 12 additional worked examples omitted; full prompt with all 14 examples available on our OSF repository ...] # TASK REMINDER: You are asked to determine whether a proposed response is correct (TRUE) or incorrect (FALSE). Think carefully and rely on the evaluation principles, steps and examples before assigning the verdict. Output your final response in a single valid JSON object in the following format: {"correct": "TRUE"} or {"correct": "FALSE"} |
| % Medium+ Effect | % Bias 0.2 | % MAE 0.5 | |||||||
| Sent. | Tox. | Total | Sent. | Tox. | Total | Sent. | Tox. | Total | |
| Scale | |||||||||
| 1 – 7 | 20% | 17% | 18% | 3% | 20% | 12% | 3% | 97% | 50% |
| -3 – 3 | 20% | 73% | 47% | 0% | 73% | 37% | 10% | 100% | 55% |
| 0 – 1 | 10% | 17% | 13% | 0% | 13% | 7% | 0% | 70% | 35% |
| 0 – 100 | 13% | 23% | 18% | 0% | 20% | 10% | 0% | 67% | 33% |
| Mean RBC | Mean Bias | Mean MAE | |||||||||
| Sent. | Tox. | Total | Sent. | (Min/Max) | Tox. | (Min/Max) | Total | Sent. | Tox. | Total | |
| Scale | |||||||||||
| 1 – 7 | 0.18 | 0.19 | 0.19 | 0.08 | (0.01/0.23) | 0.13 | (0.01/0.27) | 0.10 | 0.41 | 0.61 | 0.51 |
| -3 – 3 | 0.17 | 0.39 | 0.28 | 0.08 | (0.01/0.20) | 0.31 | (0.04/0.76) | 0.19 | 0.42 | 0.73 | 0.58 |
| 0 – 1 | 0.15 | 0.16 | 0.16 | 0.05 | (0.00/0.16) | 0.10 | (0.00/0.24) | 0.07 | 0.32 | 0.56 | 0.44 |
| 0 – 100 | 0.19 | 0.20 | 0.19 | 0.06 | (0.02/0.16) | 0.11 | (0.00/0.32) | 0.09 | 0.33 | 0.54 | 0.43 |
| % Significant | % Bias | Mean Bias | ||||
| Tox. | Tox. | Tox. | Tox. | Tox. | Tox. | |
| Scale | ||||||
| 1 – 7 | 77% | 77% | 20% | 20% | 0.13 | 0.13 |
| 0 – 1 | 53% | 53% | 13% | 13% | 0.10 | 0.10 |
| 0 – 100 | 63% | 63% | 20% | 20% | 0.11 | 0.11 |
| Prompt Detail | ||||||
| Toxicity | Toxicity-3 | |||||
|---|---|---|---|---|---|---|
| Avg Diff | Avg of Max | Max of Max | Avg Diff | Avg of Max | Max of Max | |
| Scale (Size) | 0.19 | 0.36 | 0.93 | 0.07 | 0.11 | 0.29 |
| Scale (Error) | 0.11 | 0.20 | 0.42 | 0.06 | 0.08 | 0.19 |
| Prompt Detail (Size) | 0.15 | 0.22 | 0.71 | 0.13 | 0.19 | 0.44 |
| Prompt Detail (Error) | 0.12 | 0.19 | 0.45 | 0.10 | 0.15 | 0.31 |
| Model (Size) | 0.15 | 0.43 | 0.53 | 0.15 | 0.43 | 0.53 |
| Bias | Error | ||||||||||||||
| Sig | RBC | .2 | Sig | .2 | |||||||||||
| Design choice | S | T | Tot | S | T | Tot | S | T | Tot | S | T | Tot | S | T | Tot |
| Scale | |||||||||||||||
| 1 – 7 | 61% | 83% | 72% | 34% | 59% | 47% | 4% | 24% | 14% | 76% | 83% | 79% | 0% | 11% | 6% |
| -3 – 3 | 62% | 90% | 76% | 39% | 86% | 62% | 8% | 67% | 37% | 76% | 92% | 84% | 4% | 41% | 23% |
| 0 – 1 | 64% | 80% | 72% | 31% | 60% | 46% | 3% | 29% | 16% | 77% | 74% | 76% | 3% | 18% | 10% |
| Avg | Max | ||
| Scale (model+family fixed) | |||
| Sentiment (Bias) | 0.06 | 0.24 | 180 |
| Sentiment (Error) | 0.07 | 0.28 | 180 |
| Toxicity (Bias) | 0.19 | 0.93 | 180 |
| Toxicity (Error) | 0.11 | 0.42 | 180 |
| Prompt Detail (model+scale fixed) | |||
| Total | Nearly Agree | Ambiguous | |||||||
| Acc. | Len. | #Err | Acc. | Len. | #Err | Acc. | Len. | #Err | |
| Prompt Detail | |||||||||
| Base | 96.6% | 81.3% | 406 | 98.1% | 70.2% | 47 | 61.8% | 82.7% | 359 |
| Short In-Context | 96.8% | 73.3% | 386 | 97.7% | 42.1% | 57 | 65.0% | 78.7% | 329 |
| Long In-Context | 96.0% | 52.4% | 475 | 93.6% | 12.6% | 159 | 66.3% | 72.5% | 316 |
| Reasoning Effort | |||||||||
| Model | Accuracy | Leniency | Error Count | False True | True False |
| Claude Haiku 4.5 | |||||
| Base | 96.0% | 66.7% | 48 | — | — |
| Short | 96.6% | 68.3% | 41 | 1.3% | 1.4% |
| Long | 96.3% | 40.9% | 44 | 0.5% | 2.5% |
| Claude Sonnet 5 | |||||
| Base | 96.2% | 73.9% | 46 | — | — |
| Model | Accuracy | Leniency | Error Count | False True | True False |
| Claude Haiku 4.5 | |||||
| Base | 94.8% | 69.2% | 13 | — | — |
| Short | 96.0% | 80.0% | 10 | 4.0% | 3.6% |
| Long | 93.1% | 29.4% | 17 | 1.2% | 6.1% |
| Claude Sonnet 5 | |||||
| Base | 96.8% | 12.5% | 8 | — | — |
| Model | Accuracy | Leniency | Error Count | False True | True False |
| Claude Haiku 4.5 | |||||
| Base | 62.8% | 65.7% | 35 | — | — |
| Short | 67.0% | 64.5% | 31 | 6.4% | 8.5% |
| Long | 71.3% | 48.1% | 27 | 3.2% | 16.0% |
| Claude Sonnet 5 | |||||
| Base | 59.6% | 86.8% | 38 | — | — |
| Accuracy | Switch Rate | Leniency | ||||
| Design choice | Avg. | Max. | Avg. | Max. | Avg. | Max. |
| Total | ||||||
| Prompt Detail | 0.4% | 0.6% | 2.5% | 3.1% | 18.1% | 28.9% |
| Reasoning Effort | 0.1% | 0.1% | 2.0% | 2.0% | 2.3% | 2.3% |
| Model | 1.2% | 3.9% | 4.1% | 7.1% | 23.2% | 56.1% |
| Nearly Agree | ||||||