cs.AISep 29, 2026

JudgeProfile: Understanding and Steering Subjectivity in LLM Judges

Authors: Qi Cao, Kangning Liu, Xuan Kan, Shunwen Tan, Yang Pei, Dake Chen, Yatai Ji, Zixuan Ye, +5 more

Organizations: Meta · University of California San Diego · The University of Hong Kong · The Hong Kong University of Science and Technology

Abstract

LLM judges are inherently subjective, often favoring different responses in pairwise comparison when neither option is objectively wrong. To study this subjectivity, we introduce JudgeProfile, a framework that dissects LLM evaluation into perception (how a judge compares two responses across specific attributes like clarity, correctness, and detail) and prioritization (how much each attribute influences the final choice). We curate SubjectiveSet, a dataset of 50,013 response pairs from 17 public data sources, evaluated by 21 LLM judges across 87 attributes. We find a hidden consensus in perception: judges frequently agree on attribute judgments even when their overall choices diverge. Building on this separation, we first characterize each judge's prioritization using attribute weights estimated from its own overall choices. These weights differ across judges even when estimated from the same attribute judgments. We then learn new weights from reference labels to adapt their decisions to a target evaluation standard. Reweighting perceived attributes improves average held-out agreement with reference labels from 66.48% to 71.97%, outperforming fine-tuning and rubric prompting. Our findings show that understanding and steering the subjectivity of LLM judges requires attention not only to what they perceive, but also to how they prioritize it.

Figures & tables

Appendix figures & tables46 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. SenseJudge: Human-Centric Preference-Driven Judgment Framework

    Jun 2, 2026Rui Li, Junfeng Liu, Xiangwen Kong +2Llm-As-A-JudgeLarge Language Model Decisions

  2. FairJudge: An Adaptive, Debiased, and Consistent LLM-as-a-Judge

    Feb 6, 2026Bo Yang, Lanfei Feng, Yunkui Chen +3Llm-As-A-JudgeRaw Judge Outputs

  3. Can We Trust LLM Judges: A Study of Capability-Dependent Biases and Multi-Judge Ensemble for Bias Calibration

    Sep 14, 2026Gemma Zhang, Prachi Badarayani, Asmi Kumar +2Llm-As-A-JudgeLarge Language Model Judges