Language-model ratings of depression reflect the rater more than the patient
Authors: Baihan Lin
Organizations: Department of Artificial Intelligence and Human Health, Icahn School of Medicine at Mount Sinai, New York, NY, USA · Department of Psychiatry, Icahn School of Medicine at Mount Sinai, New York, NY, USA · Department of Neuroscience, Icahn School of Medicine at Mount Sinai, New York, NY, USA · Mental Illness Research, Education and Clinical Center, James J. Peters VA Medical Center, Bronx, NY, USA · Berkman Klein Center for Internet & Society, Harvard University, Cambridge, MA, USA
Depression has no diagnostic blood test. Language models promise tireless, consistent assessment, but can accurate raters disagree about individuals? We pre-registered 880 language-model raters, crossing 11 open models with prompting and scoring choices, and applied them to 189 interviews against the eight-item Patient Health Questionnaire. Model choice explained 30.0% of summed-symptom score variance, stable participant differences 10.5%. Two randomly drawn raters with area under the receiver operating characteristic curve (AUC) >= 0.70 disagreed on screening decisions for 40% of participants, on average. Average over-rating governed how many were flagged, yet equal-capacity raters chose differently for about one participant in five. A locked analysis of 86 new interviews reproduced the main pre-registered findings. Exploratory recalibration with 40 labelled participants raised accuracy from about 60% to 75% and halved disagreement, leaving one participant in five decided differently. Calibration repaired much of the rater dependence without securing agreement about individuals.
Figures & tables
Figure 1: The same interview, opposite verdicts. a , The choices that make a rater; each of the 880 combinations was applied to every interview. b , Scores given to one participant by all 880 raters, on the PHQ-8 scale; black line, the participant’s self-reported PHQ-8; dashed line, the screening threshold of 10; orange, scores of 10 or more. c , For every participant, the share of raters screening them positive, against their self-reported PHQ-8 (jittered horizontally); grey band, 20–80% of raters; dashed line, between self-reported totals of 9 and 10. d , The chance that two different raters with AUC ≥ 0.70, drawn at random, give a participant opposite decisions, by self-reported severity; dark grey bars, discovery; white bars, confirmation.
Figure 2: Accuracy does not decide who is flagged. a , AUC of each model’s 80 raters: dot, median; thick line, interquartile range; thin line, range. Dotted line, chance; dashed line, 0.70. b , Share of participants that each model’s raters screen positive; dashed line, self-reported prevalence. c , Share screened positive against the rater’s mean over-rating of the PHQ-8, for every rater in discovery (filled) and confirmation (open); line, normal-ogive curve fitted to the discovery raters, with its R2 in discovery and, applied unchanged, in confirmation. d , Decisions for 100 people with the sample’s prevalence by the typical configuration (median AUC) of two models with nearly the same AUC; shown is the pair of models whose typical configurations differ in AUC by at most 0.02 and differ most in the share screened positive.
Figure 3: The rater outweighs the patient. a , Shares of the variance of the expected item sum in the generalizability analysis of the 330 expected-value item-sum raters (discovery); black, participant main effect; dark grey, interactions involving the participant; light grey, components of the rater alone. b , Dependability of a score averaged over models, for one or all wordings and formats, with both views of the conversation averaged; dashed line, 0.80. c , Participants’ mean score over the same 330 raters against their self-reported PHQ-8; orange, PHQ-8 of 10 or more.
Figure 4: Small choices move the threshold, and size buys ranking but not calibration or specificity. a – c , For each model, the median share screened positive with the other wordings (filled) and with the instruction to rate unmentioned symptoms as absent (open) ( a ), and with answer letters A–D meaning ‘not at all’ to ‘nearly every day’ (filled) and the reverse (open) ( b ); and the median over-rating with digit answers when symptoms are summed (filled) or the total is rated directly (open) ( c ). d – f , Against model size, one point per model: median AUC ( d ), interquartile range across a model’s raters of the share screened positive ( e ), and the median difference between a rater’s correlation with the PHQ-8 and with PTSD severity ( f ); grey band, differences within ± 0.10. Lines, least-squares fits against log size; filled markers and solid lines, discovery; open markers and dashed lines, confirmation; ρ , Spearman correlations across models.
Figure 5: What the raters read. a , Among participants below the screening threshold, the share of raters with AUC ≥ 0.70 screening them positive beyond what their PHQ-8 predicts (residual of a linear fit, percentage points), against PTSD symptoms (PCL-C). b , The same for all participants, against the amount of speech (logarithmic scale). Points and lines show residuals and linear fits on the original scales; the statistics in the panel titles are partial Spearman correlations on ranks (discovery; confirmation). Filled markers and solid lines, discovery; open markers and dashed lines, confirmation. c , AUC of each model from the participant’s speech (filled) and from the interviewer’s speech alone (open), median over wordings and formats of the expected-value item sum (discovery). d , Female–male standardised difference estimated by every rater (discovery); black line, the difference in self-reported PHQ-8.
Figure 6: Checks that need no labels. a , b , AUC of each item-sum rater against its internal consistency ( a ) and against its agreement with the same model and format under the other wordings ( b ); filled, discovery; open, confirmation; Spearman correlations across all raters and medians of the correlations within models, for discovery and then confirmation. c , In the confirmation set, the AUC of each model’s median item-sum configuration, of the configuration chosen without labels by agreement across rewordings, of a fixed neutral configuration (the questionnaire’s wording, digit answers, the participant’s speech, expected value) and of its best configuration; the rule was specified before the confirmation analysis.
Figure 7: How many versus who. a , The chance that two different raters disagree about a participant at the threshold of 10, split by the identity P(A=B)=∣pA−pB∣+2min{P(A=1,B=0),P(A=0,B=1)} into the part forced by their different positive rates (solid) and the rest, from flagging different people (hatched), for all raters, raters with AUC ≥ 0.70 and one fixed configuration per model; dark grey, discovery; white, confirmation. b , Participants decided differently when every rater selects the same share of participants, its highest-scoring ones, by two expected-value raters with AUC ≥ 0.70 or two fixed configurations (one per model); diamonds, participants decided differently by a rater and the self-reported criterion when the rater selects as many as screen positive. Filled markers and solid lines, discovery; open markers and dashed lines, confirmation. c , Participants decided differently by configurations that differ in one choice only (discovery; changes of wording are relative to the questionnaire’s own wording): dot, median over pairs of configurations; bar, interquartile range; filled, at the threshold of 10; open, selecting as many as screen positive (expected-value scores only); deployment-stack and transcript comparisons are for decisions only. d , For raters admitted by their AUC on 40 labelled participants and recalibrated with the same 40 labels, the chance that two raters disagree against the share of correct decisions, both on the held-out participants (means over 200 splits); filled, discovery; open, confirmation. Exploratory analyses.
Regularity
Criterion
Discovery
Confirmation
Held
The share screened positive follows a configuration’s mean over-rating (probit R2 ; referrals per 100 per point)
R2≥ 0.75
0.91; 13
0.91; 14
yes
Configurations with the same AUC ( ± 0.02) differ in share screened positive (median, points)
≥ 15
28
28
yes
Participants reporting few symptoms (PHQ-8 ≤ 4) are screened positive (mean share of configurations with AUC ≥ 0.70)
≥ 20%
35%
44%
yes
Larger models discriminate better (Spearman, size vs median AUC), yet no model’s ratings correlate with the PHQ-8 more than 0.10 above their correlation with PTSD severity (median difference; not a test of diagnostic separation)
> 0.5; none > 0.10
0.82; max 0.04
0.77; max 0.01
yes
Below the threshold, more PTSD symptoms go with a larger share screened positive at equal PHQ-8 (partial Spearman)
> 0, P< 0.05
0.36
0.23
yes
Rating only explicitly reported symptoms lowers referrals; reversing the answer letters moves them (median points)
8 and 6 of 11 models
27; 21
31; 24
yes
Extended Data Table 1: Regularities found in the discovery set and checked in new interviews. Discovery values are exploratory; the criteria were specified before the confirmation analysis. π , a participant’s share of raters with AUC ≥ 0.70 screening them positive. For the first regularity, confirmation values refit the curve; applied unchanged, the discovery curve gives R2 = 0.90. P values for partial correlations use n−3 degrees of freedom ( P = 0.049 for the PTSD regularity in confirmation). With digit answers only, item sums over-rate in 8 of 11 models in both sets. The regularities share participants and raters, so they are not independent tests.
Extended Data Fig. 1: Format compliance and rated severity. a , Median probability that each model placed on the answer options, by wording (participant view). b , Mean expected item sum by model and wording; the self-reported mean PHQ-8 is given in the panel title. c , Mean expected item sum by model and answer format.
Extended Data Fig. 2: Agreement with individual symptoms. Spearman correlation between each model’s expected item rating and the participant’s self-reported score on the same item, for the 141 participants with complete item scores (participant view), per symptom: dot, median model; thick line, interquartile range over models; thin line, range over models; each model’s value is its median over wordings and formats. Exploratory analysis.
Extended Data Fig. 3: Decision studies. Generalizability ( Eρ2 ; a ) and dependability ( Φ ; b ) of the expected item sum (discovery), when scores are averaged over the given number of models and over one or all wordings and formats, with both views of the conversation averaged. Dashed line, 0.80.
Extended Data Fig. 4: What averaging and labels do (discovery). a , AUC of single raters, of scores averaged over models for each configuration of the other choices, and of the average over all raters; lines, medians; dashed line, 0.70. b , Share screened positive by each rater before and after isotonic calibration of its total score against 20 or 40 labelled participants (mean over 50 draws); dashed line, self-reported prevalence. c , Mean width of 95% intervals for the prevalence from 40 labelled participants (2,000 draws): prediction-powered intervals with the registered formula and for the finite cohort, and intervals from the labels alone; observed coverage above the bars.
Extended Data Fig. 5: Discovery and confirmation. Discovery (DAIC-WOZ, dark grey) and confirmation (E-DAIC sessions 600–718, automatic transcripts, white). a , AUC of every rater; lines, medians. b , Share screened positive; dashed lines, self-reported prevalence. c , Participants receiving both decisions, contested decisions (20–80% of raters), and both decisions among raters with AUC ≥ 0.70. d , Shares of the variance of the item-sum score: participant, model, wording, participant × model (P × M) and participant × wording (P × W).
Extended Data Fig. 6: All raters, ordered by accuracy (discovery). a , AUC of every rater, ordered, with 95% bootstrap intervals; dotted line, chance; dashed line, 0.70. b , Share screened positive by the same raters; dashed line, self-reported prevalence. c , The choices that make up each rater.
Extended Data Fig. 7: How disagreement depends on which raters are counted. a , Expected share of participants who receive both decisions from k raters with AUC ≥ 0.70 drawn at random without replacement (exact), against k ; filled markers and solid line, discovery; open markers and dashed line, confirmation. b , Chance that two different raters disagree about a participant, among all raters, raters with AUC ≥ 0.70, those that also screen 5–95% of participants positive, and those with AUC, sensitivity and specificity all ≥ 0.70; dark grey, discovery; white, confirmation. Exploratory analyses.
Extended Data Fig. 8: Recalibration in detail. a , The chance that two recalibrated raters disagree about a held-out participant against the number of labelled participants, for raters admitted by their AUC in the labelled sample and recalibrated to match the labelled prevalence, to keep at least 80% of labelled positives, or by logistic regression; solid and filled, discovery; dashed and open, confirmation. b , Median absolute error of the share flagged with the prevalence-matched threshold; bars, 2.5th to 97.5th percentiles across splits (not confidence intervals). c , Disagreement after matching the prevalence with 40 labels for three rater sets: one fixed configuration per model, raters admitted by their AUC in the labelled sample, and raters admitted by their AUC in the whole cohort, a selection that uses the held-out labels. Means over 200 splits; exploratory analyses.
Extended Data Fig. 9: Deferral at matched coverage. For raters admitted by their AUC on 40 labelled participants and recalibrated to the labelled prevalence, accuracy ( a , discovery; b , confirmation) and sensitivity ( c , discovery; d , confirmation) on the held-out participants who are decided rather than deferred, against the share decided: the majority of raters for all participants (50% consensus) and after setting aside participants on whom fewer than 60, 70, 80 or 90% of raters agree (black; unanimous consensus is not shown because it left fewer than 3% of participants); single raters on the same retained participants (grey, dashed); and single raters setting aside their own closest calls to reach the same coverage (grey, solid). Large circle, the 80% consensus rule. Means over 200 splits; exploratory analyses.
Transcript view (first line of every prompt)
participant
Below is a transcript of a clinical research interview. Only the participant’s speech is shown.
full
Below is a transcript of a clinical research interview between an interviewer and a participant.
interviewer
Below is a transcript of a clinical research interview. Only the interviewer’s speech is shown.
Template
instrument
Task: complete the PHQ-8 depression questionnaire on behalf of the participant, using the interview as your only source of information.
clinician
You are an experienced psychiatrist. Using your clinical judgement about this interview, rate the participant on the PHQ-8 depression scale.
Supplementary Table 1: Prompt texts. The transcript follows the first line; the template, the format and the item follow the transcript.
Supplementary Table 2: Models and weights. † Numerical-precision check only; not part of the multiverse.
Termination rule
Re-run against saved
Model
Totals
max ∣Δ∣
Flips
median ∣Δ∣
max ∣Δ∣
Flips
gpt-oss 20B
125
0.0000
0
0.006
0.68
0
Llama 3.1 8B
125
0.0000
0
0.001
0.01
0
Mistral 7B
125
0.0080
0
0.001
0.01
0
Qwen2.5 7B
125
0.0028
0
0.090
0.45
1
Supplementary Table 3: Direct totals re-scored for 25 participants (five wordings each). Termination rule: the factor applied to every candidate against the frozen rule, from the same forward pass; flips, screening decisions that change under either scoring rule. Re-run against saved: the frozen rule in the new pass against the saved outputs (different batch composition).
Set
Rule
Disagree (%)
Acc. (%)
Sens. (%)
Spec. (%)
Error (points)
Target (%)
Discovery
Threshold of 10
40
61
83
67
28 (21–32)
Mean over-rating
18
75
39
92
14 (7–25)
Linear
16
75
42
91
12 (5–27)
Logistic
17
75
45
89
11 (4–28)
Labelled prevalence
20
75
56
84
7 (3–19)
88
≥ 80% of positives
27
68
78
68
17 (7–33)
100
Supplementary Table 4: Recalibration with 40 labelled participants, evaluated on the held-out participants (raters admitted by their AUC in the labelled sample; means over 200 splits). Disagreement, the chance that two recalibrated raters decide a held-out participant differently; error, median absolute error of the share flagged (2.5th–97.5th percentile across splits); target, the share of raters and splits whose labelled positive rate came within 5 points of the labelled prevalence (prevalence rule) or that kept at least 80% of labelled positives (sensitivity rule).
Set
Raters
K
At 10 (%)
Forced (%)
Equal no. (%)
Recal. (%)
Acc. (%)
Discovery
Full grid
880
44
92
27
–
–
AUC ≥ 0.70
540
40
94
18
19
75
Defensible core
44
31
95
21
21
74
One per model
11
32
98
25
24
72
Confirmation
Full grid
440
43
94
26
–
–
AUC ≥ 0.70
319
40
95
20
21
76
Supplementary Table 5: How the decision results change under restrictions of the grid. K , raters; at 10, the chance that two raters decide a participant differently at the threshold; forced, the share of that disagreement forced by different positive rates; equal no., the chance at equal capacity among expected-value raters of the set (all expected-value raters for the full grid); recal. and acc., held-out disagreement and accuracy after matching the labelled prevalence with 40 labels (no selection for the defensible core and the fixed configurations; whole-cohort AUC for the AUC ≥ 0.70 row).
At 10
Equal numbers
Set
PHQ-8
n
Mean (%)
Share (%)
Mean (%)
Share (%)
Discovery
0–4
86
44
48
9
24
5–9
46
43
25
25
33
10–14
30
40
15
27
24
15–24
27
33
11
25
20
Confirmation
0–4
36
47
48
13
28
Supplementary Table 6: Where the disagreement falls, by self-reported PHQ-8, for expected-value raters with AUC ≥ 0.70: each participant’s chance of being decided differently by two raters, averaged within the band (mean) and as the band’s share of all such disagreement (share), at the threshold of 10 and at equal capacity.
Set
Rule
Coverage (%)
Accuracy (%)
Sens. (%)
Spec. (%)
Discovery
Single rater, all
100
75
55
83
Majority, all
100
78
58
88
Majority, consensus ≥ 0.8 (24+, 82 − )
71
86
56
95
single rater, same participants
71
83
54
91
single rater, own closest calls deferred
71
82
48
92
Majority, consensus ≥ 0.9 (13+, 70 − )
56
90
55
97
Supplementary Table 7: Deferral at matched coverage on held-out participants (raters admitted on 40 labels and recalibrated to the labelled prevalence; means over 200 splits). Coverage, participants decided rather than deferred; counts in parentheses, mean retained participants with and without a positive screen.
Discovery
Confirmation
Choice changed
At 10
Equal numbers
At 10
Equal numbers
Letter–answer mapping reversed
20 (7–55)
21 (12–43)
26 (6–49)
21 (12–36)
Digits vs letters
13 (5–25)
15 (7–24)
12 (5–22)
14 (9–22)
Psychiatrist persona
7 (2–13)
8 (6–15)
6 (2–12)
10 (7–16)
Unmentioned symptoms rated absent
21 (8–35)
14 (8–23)
23 (9–36)
14 (9–21)
Role-play as the participant
8 (2–14)
11 (7–17)
8 (3–14)
10 (7–19)
Supplementary Table 8: Participants decided differently by configurations that differ in one choice only: median (interquartile range) over pairs, at the threshold of 10 and, for expected-value pairs, selecting as many participants as screen positive by self-report.
Model
AUC, median (range)
AUC ≥ 0.70 (%)
Positive, % (range)
Sens. (%)
Spec. (%)
Bias
Option mass
Qwen2.5 1.5B
0.59 (0.41–0.77)
8
50 (0–100)
55
52
+3.2
0.99
Qwen2.5 3B
0.70 (0.49–0.82)
49
84 (0–100)
96
21
+5.0
1.00
Qwen2.5 7B
0.81 (0.64–0.87)
96
67 (17–98)
93
45
+5.0
1.00
Qwen2.5 14B
0.85 (0.78–0.88)
100
33 (1–82)
71
83
+0.7
1.00
R1-Distill 7B
0.67 (0.35–0.81)
32
80 (1–100)
89
23
+5.3
0.99
Llama 3.2 1B
0.61 (0.34–0.78)
24
99 (0–100)
100
2
+7.3
0.89
Supplementary Table 9: Results by model (discovery set; all 80 configurations of each model). Sens. and Spec., median sensitivity and specificity for PHQ-8 ≥ 10 over the model’s configurations (descriptive). Bias, median over configurations of the mean difference between the configuration’s score and the PHQ-8 total. Option mass, median over participants, templates and formats.
Source
σ^2 (raw)
σ^2
Share (%)
model
6.943
6.943
30.0
model × format
3.645
3.645
15.8
person
2.437
2.437
10.5
person × model
1.872
1.872
8.1
template
1.779
1.779
7.7
format
1.363
1.363
5.9
Supplementary Table 10: Variance components of the expected item sum (person × model × template × format × view).
Models
Templates
Formats
Eρ2
Φ
1
1
1
0.37
0.11
3
1
1
0.58
0.21
11
1
1
0.73
0.32
1
1
3
0.45
0.14
3
1
3
0.67
0.27
11
1
3
0.82
0.40
Supplementary Table 11: Decision studies (discovery; both transcript views averaged). For one view, averaging over all models, templates and formats gives Eρ2 = 0.82 and Φ = 0.53; a single configuration with one view gives Φ = 0.11.
Prediction
Criterion
Discovery
Confirmation
H1
Configuration contributes more variance to the AUC than participant sampling
ratio > 1
10.6
supported
5.6
supported
H2
Participants with discordant decisions among acceptable configurations
≥ 25%
100%
supported
100%
supported
H3
Range of implied prevalence across configurations
≥ 30 pp
100 pp
supported
100 pp
supported
H4
Person share of item-sum variance
< 50%
10.5%
supported
13.0%
supported
H5
Rank correlation of Cronbach’s α with the AUC
< 0.30
0.02
supported
-0.06
supported
H6
Median-model AUC from interviewer speech only
> 0.60
0.59
not supported
–
–
Supplementary Table 12: Directional predictions specified before the analysis, and their outcomes.
Division of Computational & Data Sciences (DCDS), Washington University in St. Louis, St. Louis, MO · Department of Computer Science, SMU · Department of Psychological & Brain Sciences, Washington University in St. Louis