We evaluate whether clinician pairwise preferences provide a reliable signal of clinical safety in large language model (LLM) evaluation using expert feedback from MOOVE (Massive Open Online Validation and Evaluation), a clinician-led platform collecting blinded pairwise preferences alongside multi-criterion rubric ratings. Clinicians assign scores on a discrete [−2,+2] scale, where negative values indicate clinically unsafe or misleading content. Using 26{,}804 pairwise judgments across outputs from 13 LLMs, contributed by more than 736 clinicians across 28+ countries, we find that clinician preference is a poor proxy for safety-critical performance. Models ranking highly under pairwise preference can still exhibit substantial rates of clinically meaningful failures (≤−1) on dimensions such as \emph{Harmlessness} and \emph{Accuracy}. These failures are unevenly distributed across specialties, creating domain-specific ``no-go zones'' not visible in aggregate rankings or single-number leaderboards. We further analyze contributing factors including prompt length, refusal and escalation behavior, and the relative contributions of safety-critical versus surface-level features. A substantial fraction of preference votes carry no positive safety signal, while feature decomposition shows that surface-level characteristics explain slightly more preference variation than safety-critical rubric differences. Finally, we introduce a clinically adjusted preference ranking combining pairwise preference with rubric-derived feedback, producing a more safety-aware ordering than raw Bradley--Terry strength alone. Our findings support evaluation practices that separate preference from safety, report safety-critical failure rates directly, and incorporate clinically grounded adjustments when ranking LLMs for clinical decision making.
Figures & tables
Figure 1 : Overview of the MOOVE evaluation setting and analysis pipeline. Clinical vignettes are presented to 13 large language models, and their responses are evaluated through MOOVE using randomized, blinded pairwise comparisons and multi-criterion rubric assessments. The snapshot analyzed in this study comprises 736+ physicians across 76 specialties and 28 countries, contributing 26,804 pairwise comparisons and 376K+ rubric evaluations. These data support the analyses in this paper, including preference–safety dissociation, co-failure structure, decomposition of preference signals, and the role of response length.
Statistic
Value
Total unique prompts
9,236
Total specialties represented
76
Median prompt length (words)
46.0
IQR prompt length (words)
63.0
Prompts >57 words
41.5%
Vignette-like prompts
41.8%
Table 1 : Prompt distribution and structural composition in the MOOVE snapshot. The evaluated prompt set spans multiple clinical specialties and heterogeneous prompt formats, including vignette-like case descriptions and shorter direct clinical questions. Reported statistics summarize the overall topical and structural variation of the prompt set used in this study.
Harmlessness
Accuracy
MM (ECG)
Preference (BT)
Model
Mean
95% CI
Fail %
Mean
95% CI
Fail %
Fail %
Score
95% CI
GPT-OSS[120B]
0.82
[0.77, 0.86]
18.0
0.92
[0.86, 0.99]
18.4
—
0.82
[0.73, 0.91]
Gemini2.5-Flash
0.92
[0.88, 0.96]
7.2
1.09
[1.02, 1.15]
5.1
—
0.73
[0.65, 0.82]
DeepSeek-Chat
1.16
[1.09, 1.22]
5.0
1.33
[1.19, 1.47]
3.3
—
0.41
[0.28, 0.53]
GPT-5
0.87
[0.82, 0.92]
14.5
1.03
[0.96, 1.09]
12.3
—
0.33
[0.22, 0.44]
Gemini2.0-Flash
0.87
[0.83, 0.92]
14.9
0.56
[0.50, 0.62]
26.3
37.2
0.23
[0.12, 0.33]
Table 2: Model performance across safety-critical rubric criteria and pairwise preference. Harmlessness and Accuracy are scored on a −2 to +2 scale. Fail (%) denotes the proportion of ratings with score ≤−1 . MM Fail (%) reports the multimodal-only (ECG) failure rate; “—” indicates models not assessed on image-based tasks. Bold indicates best metric per column; underline indicates failure rates >10% . Green shading indicates stronger rubric performance; red shading indicates higher failure risk. Color intensity reflects ordinal bins rather than continuous values. Models are ordered by Bradley–Terry preference score. *Note: Models were not evaluated on identical task distributions. In particular, multimodal-capable models were more frequently assessed on complex image-based cases. These results should not be interpreted as a direct inter-model comparison, but rather as an analysis of how preference signals behave under heterogeneous evaluation conditions.
Figure 2 : Preference–safety dissociation in failure-rate space. Each point is a model, ranked by Bradley–Terry preference strength (#1 = most preferred). The x-axis shows Harmlessness failure rate, defined as the percentage of clinician ratings with score ≤−1 . More preferred models are not consistently those with the lowest safety failure rates; in this snapshot, preference rank and safety failure remain misaligned in a deployment-relevant way. Correlation statistics are reported in-panel (Pearson r=−0.61 , Spearman ρ=−0.64 ).
Specialty
N
Harmlessness Fail %
Accuracy Fail %
Dangerous %
Highest-risk domains
Cardiology ECG
1,226
24.4%
86.9%
89.9%
Pathology
1,360
31.2%
36.2%
38.2%
Diagnostics
52
25.0%
32.7%
34.6%
Lowest-risk domains
General Surgery
41
0.0%
0.0%
0.0%
Table 3: Specialty extremes. Impact of domain on failure rates (score ≤−1 ). “Dangerous %” is the aggregate percentage of responses where either Accuracy or Harmlessness failed. N denotes the total unique model responses evaluated in each specialty.
Figure 3 : Joint co-failure structure across clinical criteria. The figure shows mean-centered scores for each criterion, conditioned on safety-critical failure in (A) Accuracy, (B) Harmlessness, and (C) joint failure of both. This conditioning reveals that degradation in these core dimensions is accompanied by broad declines across all other rubric criteria.
Stratification
Category
Metric
Short
Long
Shift (%)
Vignette length
Safety-critical failures
Accuracy Fail (%)
23.70 [22.98, 24.44]
38.74 [37.79, 39.62]
+15.04
Harmlessness Fail (%)
22.02 [21.30, 22.70]
36.93 [35.97, 37.81]
+14.91
Behavioral responses
Escalation (%)
1.11 [0.92, 1.30]
1.40 [1.18, 1.60]
+0.29
Refusal (%)
1.64 [1.42, 1.85]
1.46 [1.24, 1.70]
-0.18
Abstention (%)
0.16 [0.09, 0.23]
0.00 [0.00, 0.00]
-0.16
Response length
Safety-critical failures
Accuracy Fail (%)
24.12 [23.49, 24.75]
49.55 [48.80, 50.29]
+25.43
Table 4 : Length-based stratification of safety failure, behavior, and preference. Top block: vignette length and bottom block: response length, used to assess whether length is associated with different safety and preference patterns. Red shading highlights larger adverse shifts in safety-critical failure rates or preference. Green shows favorable preference or safety scores.
Figure 4 : Length effects on failure and preference. (a) Vignette length and (b) Response length in blinded pairwise comparisons. Together, the panels show that increasing prompt and response verbosity capture distinct but complementary contributors to the observed preference–safety dissociation.
Safety Criterion
n
SCR
SNIR
Harmlessness / Safety
5,162
4.0%
46.0%
Accuracy / Guideline Alignment
5,164
2.1%
39.0%
Uncertainty Management
3,807
4.0%
45.3%
Clinical Reasoning
3,812
2.4%
41.5%
Any criterion
5,168
8.2% strictly worse on at least one criterion
All available criteria
5,168
19.8% no safety advantage on any criterion
Table 5 : Safety contradiction and non-informativeness rates. For each safety-critical criterion, we report the proportion of pairwise votes in which the clinician’s chosen answer is rated strictly worse (SCR) or no better (SNIR) than the rejected answer, as rated by the same clinician. N=5,168 clinical comparisons.
Component
Features
R2
Share
True Safety Signal
Accuracy, Harmlessness
0.551
48.3%
Behavioral Signal
Refusal, Escalation
0.0002
0.1%
Surface Signal
Length, Clarity, Completeness
0.588
51.6%
Full Model
All features
0.731
—
Table 6 : Preference decomposition of clinical pairwise votes. Each row reports McFadden’s pseudo- R2 for a logistic model using only the features in that signal group. The Share column normalizes across the three individual-group pseudo- R2 values and should be interpreted descriptively. N=5,168 clinical pairwise comparisons.
Figure 5 : Safety-adjusted ranking shifts model order. Comparing raw Bradley–Terry ranks against a safety-adjusted rank that overrides preference when a response is explicitly flagged as unsafe or inaccurate by clinician rubrics. Red lines denote downward shifts in rank (worse), while green arrows denote upward shifts (better) under the safety-adjusted ranking. Note: This ranking should be interpreted as context-dependent rather than definitive: models were evaluated on heterogeneous task sets, with different specialties selecting models relevant to their workflows and some (e.g., multimodal models) more frequently assessed on more complex cases. As a result, changes in rank reflect both underlying model behavior and variation in task composition.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Specialty
N Prompts
Mean Words
Median Words
IQR
% >57 words
ICU / Critical Care
10
97.8
92.0
23.2
100.0%
Geriatrics
110
103.0
99.5
67.2
79.1%
Haematology
3
80.0
85.0
13.5
100.0%
Clinical Chemistry
3
64.7
73.0
14.5
66.7%
Paediatrics
11
67.6
58.0
26.5
54.5%
Neurology
94
67.3
67.5
49.5
60.6%
Appendix
Table 7 : Prompt length by specialty. The threshold of 57 words corresponds to the empirical global mean prompt length across the MOOVE snapshot. The final column reports the proportion of prompts in each specialty exceeding this threshold.
Specialty
Clinical Vignette
Diagnostic Interp.
Treatment/Mgmt
Knowledge/Guideline
Triage/Escalation
Other
N
Cardiology
36.4
47.3
0.9
5.9
9.5
0.0
220
Clinical Chemistry
0.0
100.0
0.0
0.0
0.0
0.0
35
Haematology
31.6
31.6
0.0
0.0
36.8
0.0
38
Radiology
40.4
0.0
19.1
30.9
9.6
0.0
188
ICU / Critical Care
30.1
30.1
0.0
20.4
19.4
0.0
186
Surgery
76.0
0.0
0.3
21.0
2.3
0.3
300
Appendix
Table 8 : Prompt type distribution by specialty (%). Percentages indicate the distribution of coarse task types within each specialty.
Specialty
Short ( ≤57 w)
Medium (58–114w)
Long ( >114 w)
Overall
Haematology
—
-0.395
—
-0.395
ICU / Critical Care
—
-0.250
0.342
-0.129
Oncology
0.833
0.500
—
0.819
Cardiology
0.944
0.947
0.640
0.841
Radiology
0.904
0.989
—
0.947
Surgery
1.048
0.891
1.000
0.963
Appendix
Table 9 : Mean harmlessness by specialty and prompt-length stratum. Prompt strata are defined using the empirical dataset mean: short ≤57 words, medium 58–114 words, and long >114 words.
Specialty
Multi-Constraint (%)
Missing Context (%)
Multi-Step (%)
Interp. Required (%)
Ambiguity Index
ICU / Critical Care
80.6
0.0
0.0
40.9
30.4
Cardiology
56.8
0.0
0.0
42.3
24.8
Neurology
64.9
1.2
3.8
28.8
24.7
Geriatrics
65.1
5.6
7.2
18.6
24.1
Surgery
58.0
2.0
0.0
29.7
22.4
Paediatrics
48.1
0.0
18.2
19.9
21.5
Appendix
Table 10 : Ambiguity and underspecification proxies by specialty. The Ambiguity Index is the mean of four binary proxy flags: multi-constraint structure, missing context, multi-step reasoning, and interpretation requirement.
Figure 6 : Preference vs. harmlessness. Each point denotes a model; Bradley–Terry preference strength is plotted against mean harmlessness.
Figure 7 : Preference vs. accuracy. Each point denotes a model; Bradley–Terry preference strength is plotted against mean accuracy.
Figure 8 : Model × criteria dispersion (std. dev.). Prompt-level dispersion varies sharply by criterion and model, with consistently higher variability on Accuracy than on Harmlessness .
Figure 9 : Inter-rater reliability across evaluation criteria, measured by Fleiss’ κ . Agreement is strongest for Fairness , Confidence , and Contextual Awareness , with Safety & Harmlessness showing moderate reliability. In contrast, Accuracy , Instruction Following , and Uncertainty Management exhibit lower agreement, suggesting that correctness-oriented judgments are less stable across raters than broader behavioral or safety-related dimensions.
Figure 10 : Bradley–Terry preference ranking with 95% bootstrap CIs. Horizontal bars indicate uncertainty in the estimated preference strength from resampling pairwise votes.
Specialty
n
Safety Share (%)
Surface Share (%)
Urology
556
33.6
66.4
All Practices
127
38.0
62.0
Surgery
380
42.2
57.8
Clinical Pharmacy
145
42.8
57.2
Pediatrics
311
45.2
54.8
Internal Medicine
519
45.9
54.1
Appendix
Table 11 : Preference decomposition within clinical specialties. Reported shares are normalized R2 values for the Safety Signal (Accuracy/Harmlessness) and Surface Signal (Length/Clarity/Completeness). In nearly all specialties, the surface signal explains more of the observed preference than objective safety criteria.
Large language models (LLMs) are increasingly used to detect unsafe content. A common approach is to combine judgments from a panel of models to correct individual mistakes, but this benefit may disappear when every model sees the same misleading context before voting. We study this risk in a controlled two-round experiment. Each model first judges an item alone, then judges it again after six simulated peers either assert the wrong label or abstain. We combine the final judgments by majority vote. Across six open-weight LLMs and six datasets, we find that the wrong-label peer message raises the average reviewer false-alarm rate from 56.5% under silent peers to 87.5%, and majority voting raises the panel false-alarm rate to 100%. Without an asserted label, the same panel outperforms its average member. The effect is strongly asymmetric: reviewers follow pushes toward "unsafe" far more than pushes toward "safe" (about 75% versus 17%), so the panel's false-alarm rate rises sharply while its harmful-miss rate changes little. The proprietary-model probe shows substantial variation across models. These results identify susceptibility to shared social cues as a failure mode of safety panels and provide a simple pre-deployment diagnostic.
Background: LLM judges increasingly score whether clinical language models give overconfident answers under incomplete evidence, yet whether a measured "safety gain" reflects real behavior change or the judge's calibration is unresolved. Using a structured evidence-sufficiency prompt as a test case, we asked whether it reduces unsafe overconfident answers, how far that effect depends on the scoring judge, and what it costs in helpfulness. Methods: In a retrospective public-data benchmark (Real-POCQi, HealthBench, MedRBench), four models (GPT-5.5, Claude Opus 4.8, Gemini 3.5 Flash, Grok 4.3) answered a fully paired common panel (1,200 cells) with a standard prompt and the wrapper. The pre-specified endpoint was the paired reduction in unsafe overconfidence scored by the primary judge (GPT-5.4-nano); secondary analyses added a different-family judge (Claude Sonnet 5), a correctness judge, matched scaffold controls, and a blinded three-clinician review. Results: Unsafe overconfidence fell from 49.3% to 24.7%, a paired reduction of 24.7 points (95% CI 21.8-27.7; p<0.001), robust in direction across models and paraphrases. Magnitude was judge-dependent: Sonnet agreed on direction but nearly halved the effect (+13.1 points), with one-directional disagreement. Blinded clinicians characterized the primary judge as a high-sensitivity (1.00), low-specificity (0.55) screen, not a calibrated rate. The gain carried a model-specific helpfulness cost (correct diagnosis 80.3% to 50.3%): near-free for GPT-5.5, near-total for Gemini (-58 points). Matched scaffold controls showed genuine behavior change, not judge circularity. Conclusions: LLM-judged clinical safety effects should be reported as directional and relative, anchored to human review and evaluated jointly with helpfulness, not as calibrated absolute rates. This does not establish clinical deployment readiness.
LLMs-as-judges are the primary way to evaluate safety at scale. Despite their importance, LLM-judges themselves are rarely evaluated beyond human agreement in simple, static benchmarks. We therefore investigate two under-explored but crucial properties of LLMs-as-judges: their sensitivity to in-context information, and their steerability to differing safety definitions, which may not align with their internal safety priors. We evaluate the safety judging abilities of 13 generalist LLMs and safety-specific judges, and investigate the impact of novel in-context information and changing safety definitions. We find that while LLM-judges can learn from new information, they are broadly unlikely to update their evaluations if the context or safety definition departs from their prior.