JudgeProfile: Understanding and Steering Subjectivity in LLM Judges
Organizations: Meta · University of California San Diego · The University of Hong Kong · The Hong Kong University of Science and Technology
Abstract
LLM judges are inherently subjective, often favoring different responses in pairwise comparison when neither option is objectively wrong. To study this subjectivity, we introduce JudgeProfile, a framework that dissects LLM evaluation into perception (how a judge compares two responses across specific attributes like clarity, correctness, and detail) and prioritization (how much each attribute influences the final choice). We curate SubjectiveSet, a dataset of 50,013 response pairs from 17 public data sources, evaluated by 21 LLM judges across 87 attributes. We find a hidden consensus in perception: judges frequently agree on attribute judgments even when their overall choices diverge. Building on this separation, we first characterize each judge's prioritization using attribute weights estimated from its own overall choices. These weights differ across judges even when estimated from the same attribute judgments. We then learn new weights from reference labels to adapt their decisions to a target evaluation standard. Reweighting perceived attributes improves average held-out agreement with reference labels from 66.48% to 71.97%, outperforming fine-tuning and rubric prompting. Our findings show that understanding and steering the subjectivity of LLM judges requires attention not only to what they perceive, but also to how they prioritize it.
Figures & tables
| Qwen3.5-4B | Qwen3.5-9B | Qwen3.5-27B | Llama-4-Scout | |||||
| Method | Acc. | Acc. | Acc. | Acc. | ||||
| Unadapted | 63.79 | – | 66.74 | – | 70.79 | – | 66.91 | – |
| SFT-256 | 65.70 | +1.91 | 66.43 | -0.31 | 70.96 | +0.17 | 64.81 | -2.10 |
| Reweighting-32 | 66.02 | +2.23 | 67.21 | +0.47 | 69.96 | -0.83 | 66.67 | -0.24 |
| Reweighting-64 | 65.98 | +2.18 | 67.83 | +1.09 | 70.55 | -0.25 | 68.35 | +1.44 |
| Reweighting-128 | 66.72 | +2.93 | 68.24 | +1.50 | 71.87 | +1.08 | 68.98 | +2.07 |
| Judge | Method | Accuracy | Weight alignment | Linear explainability | |
|---|---|---|---|---|---|
| Qwen3.5-4B | Unadapted | 63.79 | – | 58.89 | 77.10 |
| Rubric prompting (6) | 66.83 | +3.03 | 54.20 | 74.20 | |
| Reweighting (6) | 68.58 | +4.79 | 57.30 | – | |
| Reweighting (12) | 69.15 | +5.36 | 67.12 | – | |
| Reweighting (87) | 69.49 | +5.69 | 100.00 | – | |
| Qwen3.5-9B | Unadapted | 66.74 | – | 60.46 | 80.18 |
Appendix figures & tables46 assets
Supplementary material from the paper’s appendix.
Appendix
| Analysis | Population | Comparison or prediction target |
|---|---|---|
| Heatmap consensus | 19 LLMs | Full SubjectiveSet; 171 judge pairs in Figure 3 |
| Collection and subset consensus | 21 LLMs | Full collection, SubjectiveSet–LLM, and its complement; 210 judge pairs |
| Human cases | 7 raters | SubjectiveSet–Human: 21 disagreement cases and 2 controls |
| Priority profiles | 19 LLMs | Each judge’s own verdicts, using own or shared attributes |
| Attribute importance and consensus | 21 LLMs | Own-verdict weights and attribute agreement; Figure 5 |
| No-conflict comparison | 21 LLMs | Overall disagreement after requiring at least 20 shared attribute directions; Figure 7 |
| Source distribution | Label type | Pairs | Train | Test |
|---|---|---|---|---|
| ScalerLab/JudgeBench ( Tan et al., 2024 ) | Benchmark | 617 | 463 | 154 |
| IF-RewardBench ( Wen et al., 2026 ) | Benchmark | 807 | 605 | 202 |
| THU-KEG/RM-Bench ( Liu et al., 2024 ) | Benchmark | 1,146 | 860 | 286 |
| allenai/reward-bench-2 ( Malik et al., 2026 ) | Benchmark | 1,657 | 1,243 | 414 |
| lmarena/arena-expert-5k ( LMArena, n.d.a ) | Human | 1,797 | 1,348 | 449 |
| berkeley-nest/Nectar ( Zhu et al., 2023 ) | AI | 2,500 | 1,875 | 625 |
| Summarization | 1,223 | |
|---|---|---|
| WebGPT | 1,117 | |
| HH-RLHF | 1,114 | |
| HelpSteer2 | 1,035 | |
| PPE–Human | 992 | |
| Arena–140K | 915 | |
| MT-Bench | 904 |
| Attribute | |||||||
|---|---|---|---|---|---|---|---|
| Judges | Items | Overall | All | Same winner | Opp. winner | ||
| LLMs (21) | All pairs | 50,013 | 74.0 | 88.2 | 89.5 | 83.6 | |
| SubjectiveSet–LLM | 11,524 | 49.8 | 84.3 | 85.8 | 82.9 | ||
| Remaining pairs | 38,489 | 81.1 | 89.0 | 89.9 | 84.0 | ||
| Humans (7) | SubjectiveSet–Human | 23 | 62.1 | 92.3 | 92.2 | 89.9 | |
| Disagreement cases | 21 | 57.6 | 92.8 | 93.3 | 89.9 | ||
| Judge | Family | Parameters | Accuracy (%) |
|---|---|---|---|
| Qwen3.5-0.8B † ( Qwen Team, 2026a ) | Qwen | 0.8B | 47.90 |
| Qwen3.5-2B † ( Qwen Team, 2026a ) | Qwen | 2B | 59.30 |
| Qwen3.5-4B ( Qwen Team, 2026a ) | Qwen | 4B | 63.79 |
| DeepSeek-R1-Distill-7B ( DeepSeek-AI, 2025 ) | DeepSeek | 7B | 58.74 |
| Qwen3.5-9B ( Qwen Team, 2026a ) | Qwen | 9B | 66.74 |
| DeepSeek-R1-Distill-14B ( DeepSeek-AI, 2025 ) | DeepSeek | 14B | 64.54 |
| Family | Count | Attributes |
|---|---|---|
| Alignment | 6 | harmfulness, moralizing tendency, neutrality, refusal tendency, safety conservatism, sycophancy |
| Communication | 8 | accessibility, clarity, coherence, fluency, notation clarity, organization, readability, terminological precision |
| Content | 1 | elaboration |
| Correctness | 11 | assumption validity, constraint satisfaction, factual accuracy, final answer correctness, groundedness, hallucination rate, instruction adherence, internal consistency, justification quality, logical validity, step correctness |
| Creative | 9 | analogy use, creativity, humor, interestingness, novelty, originality, persuasiveness, storytelling quality, vividness |
| Epistemic | 7 | acknowledgment of limitations, assertiveness, calibration, confidence, hedging, qualification, uncertainty expression |
| Judges | Read | Pairs | Same agr. | Same cov. | Opp. agr. | Opp. cov. |
|---|---|---|---|---|---|---|
| 19 | stable | 171 | 0.919 | 0.304 | 0.859 | 0.220 |
| 19 | single | 171 | 0.785 | 0.529 | 0.695 | 0.480 |
| 21 | stable | 210 | 0.895 | 0.260 | 0.836 | 0.190 |
| 21 | single | 210 | 0.732 | 0.552 | 0.685 | 0.511 |
| Judge | Same agr. | Same cov. | Opp. agr. | Opp. cov. |
|---|---|---|---|---|
| Qwen3.5-27B | 0.945 | 0.298 | 0.897 | 0.216 |
| V4-Flash | 0.945 | 0.264 | 0.897 | 0.177 |
| Qwen3.5-122B | 0.935 | 0.322 | 0.889 | 0.245 |
| Llama-4-Scout | 0.937 | 0.295 | 0.885 | 0.185 |
| Qwen3.6-27B | 0.935 | 0.342 | 0.882 | 0.257 |
| Gemma-4-26B | 0.926 | 0.306 | 0.879 | 0.236 |
| Distribution | Pairs | Same agr. | Same cov. | Opp. agr. | Opp. cov. |
|---|---|---|---|---|---|
| MT-Bench | 4,000 | 0.937 | 0.311 | 0.897 | 0.229 |
| SHP | 4,000 | 0.928 | 0.411 | 0.874 | 0.293 |
| RewardBench | 4,481 | 0.930 | 0.351 | 0.873 | 0.279 |
| Arena-Human | 4,000 | 0.918 | 0.325 | 0.869 | 0.248 |
| Nectar | 2,500 | 0.923 | 0.307 | 0.867 | 0.234 |
| HH-RLHF | 4,000 | 0.907 | 0.329 | 0.867 | 0.276 |
| Judge | Items | Preference | All 87 | Core 20 | Cov. 87 |
|---|---|---|---|---|---|
| R1-7B | 44,590 | 0.665 | 0.833 | 0.872 | 0.211 |
| R1-14B | 45,796 | 0.744 | 0.899 | 0.928 | 0.299 |
| R1-32B | 45,220 | 0.767 | 0.906 | 0.934 | 0.314 |
| V4-Flash | 44,358 | 0.784 | 0.936 | 0.957 | 0.260 |
| V4-Pro | 46,048 | 0.795 | 0.907 | 0.945 | 0.302 |
| Gemma-3-27B | 46,970 | 0.755 | 0.913 | 0.943 | 0.341 |
| Population / protocol | Attributes | Preference | Attribute | |
|---|---|---|---|---|
| LLMs / stable | All 87 | 77.1 | 90.7 | +13.6 |
| Core 20 | 77.1 | 93.7 | +16.6 | |
| LLMs / single read | All 87 | 76.7 | 76.5 | -0.2 |
| Core 20 | 76.7 | 81.0 | +4.3 | |
| Humans / selected cases | 2–3 per item | 62.1 | 92.3 | +30.2 |
| Stratum | Stable agr. | Stable cov. | Single agr. | Observations |
|---|---|---|---|---|
| Both match reference | 0.922 | 0.322 | 0.793 | 4,436,700 |
| Both depart from reference | 0.909 | 0.266 | 0.762 | 1,436,019 |
| Opposite preferences | 0.859 | 0.220 | 0.695 | 1,888,438 |
| Subset | Items | Preference agreement | Attribute agreement |
|---|---|---|---|
| Disagreement witnesses | 21 | 0.576 (204/354) | 0.928 (942/1,015) |
| Controls | 2 | 1.000 (42/42) | 0.847 (61/72) |
| Pooled | 23 | 0.621 (246/396) | 0.923 (1,003/1,087) |
| Attribute | Resolve | Self | All agr. | All cov. | Opp. agr. | Opp. cov. |
|---|---|---|---|---|---|---|
| verbosity | 0.701 | 0.762 | 0.965 | 0.558 | 0.962 | 0.517 |
| moralizing tendency | 0.225 | 0.816 | 0.959 | 0.128 | 0.959 | 0.121 |
| technicality | 0.466 | 0.827 | 0.965 | 0.323 | 0.956 | 0.266 |
| evidence use | 0.234 | 0.825 | 0.960 | 0.127 | 0.954 | 0.109 |
| self correction | 0.089 | 0.828 | 0.928 | 0.021 | 0.950 | 0.018 |
| conceptual complexity ∗ | 0.541 | 0.819 | 0.956 | 0.381 | 0.948 | 0.314 |
| Attribute | Resolve | Self | All agr. | All cov. | Opp. agr. | Opp. cov. |
|---|---|---|---|---|---|---|
| sycophancy | 0.146 | 0.689 | 0.857 | 0.052 | 0.866 | 0.050 |
| confidence | 0.504 | 0.728 | 0.886 | 0.320 | 0.864 | 0.257 |
| personalization | 0.228 | 0.767 | 0.918 | 0.106 | 0.864 | 0.084 |
| insightfulness | 0.476 | 0.777 | 0.924 | 0.301 | 0.861 | 0.225 |
| persuasiveness | 0.518 | 0.776 | 0.922 | 0.337 | 0.861 | 0.253 |
| difficulty | 0.483 | 0.735 | 0.861 | 0.307 | 0.860 | 0.261 |
| Profile / metric | Observed | Null mean | Corrected |
|---|---|---|---|
| Perception / Pearson | 0.765 | 1.000 | 0.765 |
| Prioritization / Pearson | 0.609 | 0.944 | 0.645 |
| Perception / Spearman | 0.757 | 0.999 | 0.757 |
| Prioritization / Spearman | 0.462 | 0.847 | 0.545 |
| Judge | Raw half r | SB | Shared SB | Own CV | Shared CV |
|---|---|---|---|---|---|
| R1-7B | 0.868 | 0.929 | 0.923 | 0.718 | 0.706 |
| R1-14B | 0.905 | 0.950 | 0.871 | 0.805 | 0.798 |
| R1-32B | 0.920 | 0.958 | 0.900 | 0.824 | 0.815 |
| V4-Flash | 0.923 | 0.960 | 0.897 | 0.796 | 0.823 |
| V4-Pro | 0.924 | 0.960 | 0.872 | 0.836 | 0.845 |
| Gemma-3-27B | 0.837 | 0.911 | 0.930 | 0.846 | 0.836 |
| Attribute | Observed SD | Null mean SD | Null q95 | Ratio |
|---|---|---|---|---|
| strategy quality | 1.874 | 0.316 | 0.400 | 5.93 |
| harmfulness | 2.466 | 0.479 | 0.628 | 5.15 |
| calibration | 1.769 | 0.353 | 0.451 | 5.01 |
| relevance | 1.501 | 0.339 | 0.419 | 4.42 |
| final answer correctness | 1.600 | 0.373 | 0.490 | 4.28 |
| assumption validity | 1.243 | 0.299 | 0.414 | 4.15 |
| Distribution | P r | P ceiling | P ratio | W r | W ceiling | W ratio |
|---|---|---|---|---|---|---|
| Nectar | 0.772 | 0.998 | 0.773 | 0.305 | 0.525 | 0.581 |
| HelpSteer2 | 0.770 | 0.999 | 0.771 | 0.320 | 0.599 | 0.534 |
| RewardBench | 0.766 | 0.999 | 0.767 | 0.393 | 0.702 | 0.560 |
| MT-Bench | 0.763 | 0.999 | 0.764 | 0.175 | 0.780 | 0.224 |
| UltraFeedback | 0.762 | 0.998 | 0.763 | 0.238 | 0.410 | 0.582 |
| WebGPT | 0.762 | 0.999 | 0.763 | 0.170 | 0.384 | 0.443 |
| Attribute | Coverage | Accuracy | Strength | Sign |
|---|---|---|---|---|
| Final answer correctness | 34.9 | 78.8 | 0.288 | |
| Constraint satisfaction | 16.0 | 78.4 | 0.284 | |
| Step correctness | 27.0 | 76.0 | 0.260 | |
| Strategy quality | 54.1 | 75.3 | 0.253 | |
| Calibration | 37.7 | 75.2 | 0.252 | |
| Logical validity | 41.9 | 74.9 | 0.249 |
| Attribute | Coverage | Accuracy | Strength | Sign |
|---|---|---|---|---|
| Assertiveness | 48.7 | 64.9 | 0.149 | |
| Creativity | 38.2 | 64.8 | 0.148 | |
| Depth | 59.2 | 64.7 | 0.147 | |
| Breadth | 56.2 | 64.7 | 0.147 | |
| Elaboration | 64.9 | 64.7 | 0.147 | |
| Qualification | 48.2 | 64.6 | 0.146 |
| Attribute | Own perception | Shared perception |
|---|---|---|
| Final answer correctness | 4.29 | 3.26 |
| Strategy quality | 3.20 | 3.88 |
| Instruction adherence | 2.45 | 2.35 |
| Constraint satisfaction | 2.32 | 2.53 |
| Calibration | 1.92 | 3.26 |
| Insightfulness | 1.39 | 1.82 |
| Judge | First | Second | Third |
|---|---|---|---|
| R1-7B | Constraints (0.21) | Correctness (0.18) | step correctness (0.14) |
| R1-14B | Correctness (0.26) | Empathy (0.23) | Strategy (0.19) |
| R1-32B | Correctness (0.29) | Strategy (0.26) | Empathy (0.24) |
| V4-Flash | Strategy (0.36) | Constraints (0.20) | Empathy (0.20) |
| V4-Pro | Strategy (0.28) | Calibration (0.26) | robustness (0.23) |
| Gemma-3-27B | Constraints (0.31) | Strategy (0.29) | Empathy (0.29) |
| Attribute group | Attributes | Agreement | Coverage |
|---|---|---|---|
| Q1 (lowest fitted importance) | 22 | 87.1 | 23.6 |
| Q2 | 22 | 82.4 | 20.8 |
| Q3 | 21 | 82.4 | 16.9 |
| Q4 (highest fitted importance) | 22 | 76.8 | 14.7 |
| Final-answer correctness | 1 | 66.0 | 9.9 |
| Strategy quality | 1 | 68.4 | 21.6 |
| Family | Attributes | Importance | Agreement | Coverage |
|---|---|---|---|---|
| correctness | 11 | 1.611 | 0.722 | 0.133 |
| interpersonal | 1 | 0.419 | 0.748 | 0.225 |
| communication | 8 | 0.882 | 0.777 | 0.204 |
| relevance | 5 | 0.979 | 0.779 | 0.231 |
| reasoning | 9 | 1.214 | 0.810 | 0.173 |
| stance | 1 | 0.428 | 0.819 | 0.170 |
| Attribute | Importance | All | Same | Opposite | Coverage |
|---|---|---|---|---|---|
| harmfulness | 5.373 | 0.921 | 0.922 | 0.915 | 0.050 |
| final answer correctness | 3.944 | 0.881 | 0.919 | 0.660 | 0.099 |
| strategy quality | 2.958 | 0.873 | 0.910 | 0.684 | 0.216 |
| constraint satisfaction | 2.490 | 0.873 | 0.905 | 0.705 | 0.029 |
| hallucination rate | 2.304 | 0.890 | 0.888 | 0.892 | 0.136 |
| instruction adherence | 2.281 | 0.883 | 0.909 | 0.746 | 0.194 |
| Attribute | Importance | All | Same | Opposite | Coverage |
|---|---|---|---|---|---|
| internal consistency | 0.967 | 0.760 | 0.791 | 0.596 | 0.057 |
| respectfulness | 0.949 | 0.865 | 0.893 | 0.722 | 0.104 |
| persuasiveness | 0.943 | 0.903 | 0.918 | 0.845 | 0.213 |
| interestingness | 0.905 | 0.901 | 0.909 | 0.875 | 0.323 |
| neutrality | 0.834 | 0.843 | 0.863 | 0.748 | 0.045 |
| confidence | 0.830 | 0.870 | 0.878 | 0.844 | 0.218 |
| Attribute | Importance | All | Same | Opposite | Coverage |
|---|---|---|---|---|---|
| decomposition quality | 0.649 | 0.914 | 0.924 | 0.877 | 0.216 |
| detail level | 0.648 | 0.939 | 0.944 | 0.925 | 0.441 |
| qualification | 0.644 | 0.909 | 0.919 | 0.876 | 0.249 |
| assertiveness | 0.642 | 0.864 | 0.869 | 0.850 | 0.238 |
| conversationality | 0.623 | 0.854 | 0.857 | 0.847 | 0.242 |
| novelty | 0.603 | 0.879 | 0.883 | 0.867 | 0.097 |
| Judges | Condition | Items/pair | Pair mean | Observation-weighted |
|---|---|---|---|---|
| 21 | No conflict | 5,245 | 14.5 | 11.5 |
| 21 | At least one conflict | 17,329 | 21.8 | 19.7 |
| 19 | No conflict | — |
| Eligible | No conflict | Conflict | Pair mean | Obs.-weighted | |
|---|---|---|---|---|---|
| 5 | 36,403 | 10,948 | 25,454 | 20.8 | 18.2 |
| 10 | 30,953 | 8,309 | 22,644 | 18.4 | 15.5 |
| 20 | 22,575 | 5,245 | 17,329 | 14.5 | 11.5 |
| 30 | 15,305 | 3,204 | 12,101 | 10.9 | 8.4 |
| 40 | 8,728 | 1,748 | 6,979 | 7.5 | 5.7 |
| Eligible | No conflict | Conflict | Pair mean | Obs.-weighted | |
| 1 | 28,814 | 22,684 | 6,131 | 18.4 | 16.0 |
| 3 | 13,936 | 10,732 | 3,204 | 11.6 | 8.1 |
| 5 | 5,402 | 4,116 | 1,287 | 8.3 | 4.5 |
| 7 | 873 | 581 | 292 | 6.1 | 3.3 |
| Judge | No-conflict items/pair | Winner disagreement |
|---|---|---|
| Qwen3.5-0.8B | 762 | 35.8 |
| DeepSeek-R1-Distill-14B | 3,794 | 15.7 |
| DeepSeek-R1-Distill-32B | 3,853 | 15.2 |
| DeepSeek-R1-Distill-7B | 1,965 | 15.1 |
| Qwen3.5-2B | 1,004 | 14.9 |
| Muse-Glimmer-30B | 2,257 | 14.8 |
| Source | No-conflict items/pair | Winner disagreement |
|---|---|---|
| IF-RewardBench | 44 | 27.8 |
| hh-rlhf | 428 | 19.0 |
| arena-140k | 459 | 18.7 |
| webgpt | 442 | 17.6 |
| RM-Bench | 100 | 17.1 |
| PPE-Human | 443 | 16.6 |
| Judge | Base | Reweighted | Gain | Preference gain | Gain excl. RM |
|---|---|---|---|---|---|
| Kimi-K3 | 74.45 | 76.96 | +2.51 | +4.06 | +1.66 |
| Gemma-4-31B | 73.25 | 77.18 | +3.93 | +4.83 | +3.04 |
| Qwen3.6-27B | 73.08 | 76.64 | +3.55 | +4.05 | +2.58 |
| Qwen3.5-27B | 72.16 | 75.18 | +3.01 | +4.15 | +1.86 |
| DeepSeek-V4-Pro | 69.12 | 73.05 | +3.92 | +2.35 | +2.11 |
| Gemma-4-26B-A4B | 70.46 | 75.39 | +4.92 | +5.88 | +3.90 |
| Dataset | Base | Reweighted | Gain | Min / max gain |
|---|---|---|---|---|
| RM-Bench | 65.05 | 94.44 | +29.37 | +15.4 / +64.3 |
| MT-Human | 62.06 | 78.54 | +16.50 | +3.1 / +29.4 |
| Nectar | 64.47 | 75.09 | +10.62 | +2.1 / +16.6 |
| SHP | 58.06 | 66.85 | +8.79 | +2.1 / +13.5 |
| PPE-BoK | 66.64 | 74.49 | +7.84 | +0.2 / +19.5 |
| Summarization | 60.33 | 66.61 | +6.29 | +0.6 / +27.4 |
| Judge | JudgeBench | IF-RB | RM-Bench | RB2 | Arena-Expert | UltraFeedback |
|---|---|---|---|---|---|---|
| Kimi-K3 | ||||||
| Gemma-4-31B | ||||||
| Qwen3.6-27B | ||||||
| Qwen3.5-27B | ||||||
| DeepSeek-V4-Pro | ||||||
| Gemma-4-26B-A4B |
| Judge | Nectar | PPE-BoK | HelpSteer2 | WebGPT | MT-Human | SHP |
|---|---|---|---|---|---|---|
| Kimi-K3 | ||||||
| Gemma-4-31B | ||||||
| Qwen3.6-27B | ||||||
| Qwen3.5-27B | ||||||
| DeepSeek-V4-Pro | ||||||
| Gemma-4-26B-A4B |
| Judge | HH-RLHF | Arena-Human | Summarization | PPE-Human | RewardBench |
|---|---|---|---|---|---|
| Kimi-K3 | |||||
| Gemma-4-31B | |||||
| Qwen3.6-27B | |||||
| Qwen3.5-27B | |||||
| DeepSeek-V4-Pro | |||||
| Gemma-4-26B-A4B |
| Judge | Smallest budget beating SFT-256 | RW256 minus SFT256 | RW Full minus SFT Full | RW6 minus rubric6 |
|---|---|---|---|---|
| Qwen3.5-4B | 32 | +2.08 | -3.45 | +1.75 |
| Qwen3.5-9B | 32 | +4.37 | -1.92 | +2.80 |
| Qwen3.5-27B | 128 | +1.50 | -1.36 | -0.07 |
| Llama-4-Scout | 32 | +5.35 | +3.56 | +2.62 |
| Judge | alignment | explainability |
|---|---|---|
| Qwen3.5-4B | -4.69 | -2.90 |
| Qwen3.5-9B | +13.80 | -3.45 |
| Qwen3.5-27B | +20.32 | -2.59 |
| Llama-4-Scout | +1.98 | -3.21 |