Language-model ratings of depression reflect the rater more than the patient
Authors: Baihan Lin
Organizations: Department of Artificial Intelligence and Human Health, Icahn School of Medicine at Mount Sinai, New York, NY, USA · Department of Psychiatry, Icahn School of Medicine at Mount Sinai, New York, NY, USA · Department of Neuroscience, Icahn School of Medicine at Mount Sinai, New York, NY, USA · Mental Illness Research, Education and Clinical Center, James J. Peters VA Medical Center, Bronx, NY, USA · Berkman Klein Center for Internet & Society, Harvard University, Cambridge, MA, USA
Depression has no diagnostic blood test. Language models promise tireless, consistent assessment, but can accurate raters disagree about individuals? We pre-registered 880 language-model raters, crossing 11 open models with prompting and scoring choices, and applied them to 189 interviews against the eight-item Patient Health Questionnaire. Model choice explained 30.0% of summed-symptom score variance, stable participant differences 10.5%. Two randomly drawn raters with area under the receiver operating characteristic curve (AUC) >= 0.70 disagreed on screening decisions for 40% of participants, on average. Average over-rating governed how many were flagged, yet equal-capacity raters chose differently for about one participant in five. A locked analysis of 86 new interviews reproduced the main pre-registered findings. Exploratory recalibration with 40 labelled participants raised accuracy from about 60% to 75% and halved disagreement, leaving one participant in five decided differently. Calibration repaired much of the rater dependence without securing agreement about individuals.
Figures & tables
Figure 1: The same interview, opposite verdicts. a , The choices that make a rater; each of the 880 combinations was applied to every interview. b , Scores given to one participant by all 880 raters, on the PHQ-8 scale; black line, the participant’s self-reported PHQ-8; dashed line, the screening threshold of 10; orange, scores of 10 or more. c , For every participant, the share of raters screening them positive, against their self-reported PHQ-8 (jittered horizontally); grey band, 20–80% of raters; dashed line, between self-reported totals of 9 and 10. d , The chance that two different raters with AUC ≥ 0.70, drawn at random, give a participant opposite decisions, by self-reported severity; dark grey bars, discovery; white bars, confirmation.
Figure 2: Accuracy does not decide who is flagged. a , AUC of each model’s 80 raters: dot, median; thick line, interquartile range; thin line, range. Dotted line, chance; dashed line, 0.70. b , Share of participants that each model’s raters screen positive; dashed line, self-reported prevalence. c , Share screened positive against the rater’s mean over-rating of the PHQ-8, for every rater in discovery (filled) and confirmation (open); line, normal-ogive curve fitted to the discovery raters, with its R2 in discovery and, applied unchanged, in confirmation. d , Decisions for 100 people with the sample’s prevalence by the typical configuration (median AUC) of two models with nearly the same AUC; shown is the pair of models whose typical configurations differ in AUC by at most 0.02 and differ most in the share screened positive.
Figure 3: The rater outweighs the patient. a , Shares of the variance of the expected item sum in the generalizability analysis of the 330 expected-value item-sum raters (discovery); black, participant main effect; dark grey, interactions involving the participant; light grey, components of the rater alone. b , Dependability of a score averaged over models, for one or all wordings and formats, with both views of the conversation averaged; dashed line, 0.80. c , Participants’ mean score over the same 330 raters against their self-reported PHQ-8; orange, PHQ-8 of 10 or more.
Figure 4: Small choices move the threshold, and size buys ranking but not calibration or specificity. a – c , For each model, the median share screened positive with the other wordings (filled) and with the instruction to rate unmentioned symptoms as absent (open) ( a ), and with answer letters A–D meaning ‘not at all’ to ‘nearly every day’ (filled) and the reverse (open) ( b ); and the median over-rating with digit answers when symptoms are summed (filled) or the total is rated directly (open) ( c ). d – f , Against model size, one point per model: median AUC ( d ), interquartile range across a model’s raters of the share screened positive ( e ), and the median difference between a rater’s correlation with the PHQ-8 and with PTSD severity ( f ); grey band, differences within ± 0.10. Lines, least-squares fits against log size; filled markers and solid lines, discovery; open markers and dashed lines, confirmation; ρ , Spearman correlations across models.
Figure 5: What the raters read. a , Among participants below the screening threshold, the share of raters with AUC ≥ 0.70 screening them positive beyond what their PHQ-8 predicts (residual of a linear fit, percentage points), against PTSD symptoms (PCL-C). b , The same for all participants, against the amount of speech (logarithmic scale). Points and lines show residuals and linear fits on the original scales; the statistics in the panel titles are partial Spearman correlations on ranks (discovery; confirmation). Filled markers and solid lines, discovery; open markers and dashed lines, confirmation. c , AUC of each model from the participant’s speech (filled) and from the interviewer’s speech alone (open), median over wordings and formats of the expected-value item sum (discovery). d , Female–male standardised difference estimated by every rater (discovery); black line, the difference in self-reported PHQ-8.
Figure 6: Checks that need no labels. a , b , AUC of each item-sum rater against its internal consistency ( a ) and against its agreement with the same model and format under the other wordings ( b ); filled, discovery; open, confirmation; Spearman correlations across all raters and medians of the correlations within models, for discovery and then confirmation. c , In the confirmation set, the AUC of each model’s median item-sum configuration, of the configuration chosen without labels by agreement across rewordings, of a fixed neutral configuration (the questionnaire’s wording, digit answers, the participant’s speech, expected value) and of its best configuration; the rule was specified before the confirmation analysis.
Figure 7: How many versus who. a , The chance that two different raters disagree about a participant at the threshold of 10, split by the identity P(A=B)=∣pA−pB∣+2min{P(A=1,B=0),P(A=0,B=1)} into the part forced by their different positive rates (solid) and the rest, from flagging different people (hatched), for all raters, raters with AUC ≥ 0.70 and one fixed configuration per model; dark grey, discovery; white, confirmation. b , Participants decided differently when every rater selects the same share of participants, its highest-scoring ones, by two expected-value raters with AUC ≥ 0.70 or two fixed configurations (one per model); diamonds, participants decided differently by a rater and the self-reported criterion when the rater selects as many as screen positive. Filled markers and solid lines, discovery; open markers and dashed lines, confirmation. c , Participants decided differently by configurations that differ in one choice only (discovery; changes of wording are relative to the questionnaire’s own wording): dot, median over pairs of configurations; bar, interquartile range; filled, at the threshold of 10; open, selecting as many as screen positive (expected-value scores only); deployment-stack and transcript comparisons are for decisions only. d , For raters admitted by their AUC on 40 labelled participants and recalibrated with the same 40 labels, the chance that two raters disagree against the share of correct decisions, both on the held-out participants (means over 200 splits); filled, discovery; open, confirmation. Exploratory analyses.
Regularity
Criterion
Discovery
Confirmation
Held
The share screened positive follows a configuration’s mean over-rating (probit R2 ; referrals per 100 per point)
R2≥ 0.75
0.91; 13
0.91; 14
yes
Configurations with the same AUC ( ± 0.02) differ in share screened positive (median, points)
≥ 15
28
28
yes
Participants reporting few symptoms (PHQ-8 ≤ 4) are screened positive (mean share of configurations with AUC ≥ 0.70)
≥ 20%
35%
44%
yes
Larger models discriminate better (Spearman, size vs median AUC), yet no model’s ratings correlate with the PHQ-8 more than 0.10 above their correlation with PTSD severity (median difference; not a test of diagnostic separation)
> 0.5; none > 0.10
0.82; max 0.04
0.77; max 0.01
yes
Below the threshold, more PTSD symptoms go with a larger share screened positive at equal PHQ-8 (partial Spearman)
> 0, P< 0.05
0.36
0.23
yes
Rating only explicitly reported symptoms lowers referrals; reversing the answer letters moves them (median points)
8 and 6 of 11 models
27; 21
31; 24
yes
Extended Data Table 1: Regularities found in the discovery set and checked in new interviews. Discovery values are exploratory; the criteria were specified before the confirmation analysis. π , a participant’s share of raters with AUC ≥ 0.70 screening them positive. For the first regularity, confirmation values refit the curve; applied unchanged, the discovery curve gives R2 = 0.90. P values for partial correlations use n−3 degrees of freedom ( P = 0.049 for the PTSD regularity in confirmation). With digit answers only, item sums over-rate in 8 of 11 models in both sets. The regularities share participants and raters, so they are not independent tests.
Extended Data Fig. 1: Format compliance and rated severity. a , Median probability that each model placed on the answer options, by wording (participant view). b , Mean expected item sum by model and wording; the self-reported mean PHQ-8 is given in the panel title. c , Mean expected item sum by model and answer format.
Extended Data Fig. 2: Agreement with individual symptoms. Spearman correlation between each model’s expected item rating and the participant’s self-reported score on the same item, for the 141 participants with complete item scores (participant view), per symptom: dot, median model; thick line, interquartile range over models; thin line, range over models; each model’s value is its median over wordings and formats. Exploratory analysis.
Extended Data Fig. 3: Decision studies. Generalizability ( Eρ2 ; a ) and dependability ( Φ ; b ) of the expected item sum (discovery), when scores are averaged over the given number of models and over one or all wordings and formats, with both views of the conversation averaged. Dashed line, 0.80.
Extended Data Fig. 4: What averaging and labels do (discovery). a , AUC of single raters, of scores averaged over models for each configuration of the other choices, and of the average over all raters; lines, medians; dashed line, 0.70. b , Share screened positive by each rater before and after isotonic calibration of its total score against 20 or 40 labelled participants (mean over 50 draws); dashed line, self-reported prevalence. c , Mean width of 95% intervals for the prevalence from 40 labelled participants (2,000 draws): prediction-powered intervals with the registered formula and for the finite cohort, and intervals from the labels alone; observed coverage above the bars.
Extended Data Fig. 5: Discovery and confirmation. Discovery (DAIC-WOZ, dark grey) and confirmation (E-DAIC sessions 600–718, automatic transcripts, white). a , AUC of every rater; lines, medians. b , Share screened positive; dashed lines, self-reported prevalence. c , Participants receiving both decisions, contested decisions (20–80% of raters), and both decisions among raters with AUC ≥ 0.70. d , Shares of the variance of the item-sum score: participant, model, wording, participant × model (P × M) and participant × wording (P × W).
Extended Data Fig. 6: All raters, ordered by accuracy (discovery). a , AUC of every rater, ordered, with 95% bootstrap intervals; dotted line, chance; dashed line, 0.70. b , Share screened positive by the same raters; dashed line, self-reported prevalence. c , The choices that make up each rater.
Extended Data Fig. 7: How disagreement depends on which raters are counted. a , Expected share of participants who receive both decisions from k raters with AUC ≥ 0.70 drawn at random without replacement (exact), against k ; filled markers and solid line, discovery; open markers and dashed line, confirmation. b , Chance that two different raters disagree about a participant, among all raters, raters with AUC ≥ 0.70, those that also screen 5–95% of participants positive, and those with AUC, sensitivity and specificity all ≥ 0.70; dark grey, discovery; white, confirmation. Exploratory analyses.
Extended Data Fig. 8: Recalibration in detail. a , The chance that two recalibrated raters disagree about a held-out participant against the number of labelled participants, for raters admitted by their AUC in the labelled sample and recalibrated to match the labelled prevalence, to keep at least 80% of labelled positives, or by logistic regression; solid and filled, discovery; dashed and open, confirmation. b , Median absolute error of the share flagged with the prevalence-matched threshold; bars, 2.5th to 97.5th percentiles across splits (not confidence intervals). c , Disagreement after matching the prevalence with 40 labels for three rater sets: one fixed configuration per model, raters admitted by their AUC in the labelled sample, and raters admitted by their AUC in the whole cohort, a selection that uses the held-out labels. Means over 200 splits; exploratory analyses.
Extended Data Fig. 9: Deferral at matched coverage. For raters admitted by their AUC on 40 labelled participants and recalibrated to the labelled prevalence, accuracy ( a , discovery; b , confirmation) and sensitivity ( c , discovery; d , confirmation) on the held-out participants who are decided rather than deferred, against the share decided: the majority of raters for all participants (50% consensus) and after setting aside participants on whom fewer than 60, 70, 80 or 90% of raters agree (black; unanimous consensus is not shown because it left fewer than 3% of participants); single raters on the same retained participants (grey, dashed); and single raters setting aside their own closest calls to reach the same coverage (grey, solid). Large circle, the 80% consensus rule. Means over 200 splits; exploratory analyses.
Transcript view (first line of every prompt)
participant
Below is a transcript of a clinical research interview. Only the participant’s speech is shown.
full
Below is a transcript of a clinical research interview between an interviewer and a participant.
interviewer
Below is a transcript of a clinical research interview. Only the interviewer’s speech is shown.
Template
instrument
Task: complete the PHQ-8 depression questionnaire on behalf of the participant, using the interview as your only source of information.
clinician
You are an experienced psychiatrist. Using your clinical judgement about this interview, rate the participant on the PHQ-8 depression scale.
Supplementary Table 1: Prompt texts. The transcript follows the first line; the template, the format and the item follow the transcript.
Supplementary Table 2: Models and weights. † Numerical-precision check only; not part of the multiverse.
Termination rule
Re-run against saved
Model
Totals
max ∣Δ∣
Flips
median ∣Δ∣
max ∣Δ∣
Flips
gpt-oss 20B
125
0.0000
0
0.006
0.68
0
Llama 3.1 8B
125
0.0000
0
0.001
0.01
0
Mistral 7B
125
0.0080
0
0.001
0.01
0
Qwen2.5 7B
125
0.0028
0
0.090
0.45
1
Supplementary Table 3: Direct totals re-scored for 25 participants (five wordings each). Termination rule: the factor applied to every candidate against the frozen rule, from the same forward pass; flips, screening decisions that change under either scoring rule. Re-run against saved: the frozen rule in the new pass against the saved outputs (different batch composition).
Set
Rule
Disagree (%)
Acc. (%)
Sens. (%)
Spec. (%)
Error (points)
Target (%)
Discovery
Threshold of 10
40
61
83
67
28 (21–32)
Mean over-rating
18
75
39
92
14 (7–25)
Linear
16
75
42
91
12 (5–27)
Logistic
17
75
45
89
11 (4–28)
Labelled prevalence
20
75
56
84
7 (3–19)
88
≥ 80% of positives
27
68
78
68
17 (7–33)
100
Supplementary Table 4: Recalibration with 40 labelled participants, evaluated on the held-out participants (raters admitted by their AUC in the labelled sample; means over 200 splits). Disagreement, the chance that two recalibrated raters decide a held-out participant differently; error, median absolute error of the share flagged (2.5th–97.5th percentile across splits); target, the share of raters and splits whose labelled positive rate came within 5 points of the labelled prevalence (prevalence rule) or that kept at least 80% of labelled positives (sensitivity rule).
Set
Raters
K
At 10 (%)
Forced (%)
Equal no. (%)
Recal. (%)
Acc. (%)
Discovery
Full grid
880
44
92
27
–
–
AUC ≥ 0.70
540
40
94
18
19
75
Defensible core
44
31
95
21
21
74
One per model
11
32
98
25
24
72
Confirmation
Full grid
440
43
94
26
–
–
AUC ≥ 0.70
319
40
95
20
21
76
Supplementary Table 5: How the decision results change under restrictions of the grid. K , raters; at 10, the chance that two raters decide a participant differently at the threshold; forced, the share of that disagreement forced by different positive rates; equal no., the chance at equal capacity among expected-value raters of the set (all expected-value raters for the full grid); recal. and acc., held-out disagreement and accuracy after matching the labelled prevalence with 40 labels (no selection for the defensible core and the fixed configurations; whole-cohort AUC for the AUC ≥ 0.70 row).
At 10
Equal numbers
Set
PHQ-8
n
Mean (%)
Share (%)
Mean (%)
Share (%)
Discovery
0–4
86
44
48
9
24
5–9
46
43
25
25
33
10–14
30
40
15
27
24
15–24
27
33
11
25
20
Confirmation
0–4
36
47
48
13
28
Supplementary Table 6: Where the disagreement falls, by self-reported PHQ-8, for expected-value raters with AUC ≥ 0.70: each participant’s chance of being decided differently by two raters, averaged within the band (mean) and as the band’s share of all such disagreement (share), at the threshold of 10 and at equal capacity.
Set
Rule
Coverage (%)
Accuracy (%)
Sens. (%)
Spec. (%)
Discovery
Single rater, all
100
75
55
83
Majority, all
100
78
58
88
Majority, consensus ≥ 0.8 (24+, 82 − )
71
86
56
95
single rater, same participants
71
83
54
91
single rater, own closest calls deferred
71
82
48
92
Majority, consensus ≥ 0.9 (13+, 70 − )
56
90
55
97
Supplementary Table 7: Deferral at matched coverage on held-out participants (raters admitted on 40 labels and recalibrated to the labelled prevalence; means over 200 splits). Coverage, participants decided rather than deferred; counts in parentheses, mean retained participants with and without a positive screen.
Discovery
Confirmation
Choice changed
At 10
Equal numbers
At 10
Equal numbers
Letter–answer mapping reversed
20 (7–55)
21 (12–43)
26 (6–49)
21 (12–36)
Digits vs letters
13 (5–25)
15 (7–24)
12 (5–22)
14 (9–22)
Psychiatrist persona
7 (2–13)
8 (6–15)
6 (2–12)
10 (7–16)
Unmentioned symptoms rated absent
21 (8–35)
14 (8–23)
23 (9–36)
14 (9–21)
Role-play as the participant
8 (2–14)
11 (7–17)
8 (3–14)
10 (7–19)
Supplementary Table 8: Participants decided differently by configurations that differ in one choice only: median (interquartile range) over pairs, at the threshold of 10 and, for expected-value pairs, selecting as many participants as screen positive by self-report.
Model
AUC, median (range)
AUC ≥ 0.70 (%)
Positive, % (range)
Sens. (%)
Spec. (%)
Bias
Option mass
Qwen2.5 1.5B
0.59 (0.41–0.77)
8
50 (0–100)
55
52
+3.2
0.99
Qwen2.5 3B
0.70 (0.49–0.82)
49
84 (0–100)
96
21
+5.0
1.00
Qwen2.5 7B
0.81 (0.64–0.87)
96
67 (17–98)
93
45
+5.0
1.00
Qwen2.5 14B
0.85 (0.78–0.88)
100
33 (1–82)
71
83
+0.7
1.00
R1-Distill 7B
0.67 (0.35–0.81)
32
80 (1–100)
89
23
+5.3
0.99
Llama 3.2 1B
0.61 (0.34–0.78)
24
99 (0–100)
100
2
+7.3
0.89
Supplementary Table 9: Results by model (discovery set; all 80 configurations of each model). Sens. and Spec., median sensitivity and specificity for PHQ-8 ≥ 10 over the model’s configurations (descriptive). Bias, median over configurations of the mean difference between the configuration’s score and the PHQ-8 total. Option mass, median over participants, templates and formats.
Source
σ^2 (raw)
σ^2
Share (%)
model
6.943
6.943
30.0
model × format
3.645
3.645
15.8
person
2.437
2.437
10.5
person × model
1.872
1.872
8.1
template
1.779
1.779
7.7
format
1.363
1.363
5.9
Supplementary Table 10: Variance components of the expected item sum (person × model × template × format × view).
Models
Templates
Formats
Eρ2
Φ
1
1
1
0.37
0.11
3
1
1
0.58
0.21
11
1
1
0.73
0.32
1
1
3
0.45
0.14
3
1
3
0.67
0.27
11
1
3
0.82
0.40
Supplementary Table 11: Decision studies (discovery; both transcript views averaged). For one view, averaging over all models, templates and formats gives Eρ2 = 0.82 and Φ = 0.53; a single configuration with one view gives Φ = 0.11.
Prediction
Criterion
Discovery
Confirmation
H1
Configuration contributes more variance to the AUC than participant sampling
ratio > 1
10.6
supported
5.6
supported
H2
Participants with discordant decisions among acceptable configurations
≥ 25%
100%
supported
100%
supported
H3
Range of implied prevalence across configurations
≥ 30 pp
100 pp
supported
100 pp
supported
H4
Person share of item-sum variance
< 50%
10.5%
supported
13.0%
supported
H5
Rank correlation of Cronbach’s α with the AUC
< 0.30
0.02
supported
-0.06
supported
H6
Median-model AUC from interviewer speech only
> 0.60
0.59
not supported
–
–
Supplementary Table 12: Directional predictions specified before the analysis, and their outcomes.
Large Language Model (LLM) studies that use language responses elicited from depression assessments to predict scores on those same assessments often report near-perfect prediction of depression. We refer to these as "Mirror" evaluations and demonstrate an applied case of criterion contamination. N = 110 participants completed both structured diagnostic depression interviews (Mirror condition) and life history interviews ("Non-Mirror" condition). LLMs were prompted to predict depression scores in each condition. As expected, Mirror evaluations were near-perfect. However, Non-Mirror evaluations also displayed prediction sizes considered outstanding in psychology. Further, both Mirror and Non-Mirror predictions correlated with Patient Health Questionnaire-9 scores at similar sizes, suggesting the Mirror condition's advantage collapses when predicting an independent depression measurement. Topic modeling revealed differing depression-related themes across interview types. Mirror evaluations are better considered as reliability evaluations than as validity evaluations. Incorporating Non-Mirror approaches in LLM depression assessment may support more valid and clinically-relevant applications. Keywords: large language models, psychological assessment, psychopathology, depression, reliability, validity, criterion contamination
Tong Li, Rasiq Hussain, Mehak Gupta +1
Division of Computational & Data Sciences (DCDS), Washington University in St. Louis, St. Louis, MO · Department of Computer Science, SMU · Department of Psychological & Brain Sciences, Washington University in St. Louis
A large language model (LLM) can rate depression severity directly from a social media post or mark which clinical criteria the post shows and let code turn the count into a label. The latter is easier to audit because a clinician can check each marked criterion. We compare these approaches on two Reddit corpora using three LLMs (from 9B to frontier scale) and two questionnaires (PHQ-9, BDI-II), and measure agreement with quadratic weighted kappa. For the two frontier models, criteria extraction scores above chain-of-thought on one corpus only when its decision thresholds are fitted on labeled data. Neither model's gain is significant, with or without recalibrating chain-of-thought on the same labels. With thresholds fixed a priori from PHQ-9's criteria, extraction shows no gain on either corpus, even where models mark over two criteria per post. The 9B model behaves differently on a corpus from depression communities. It labels most posts severe, whether prompted directly or with chain-of-thought, while the a priori rule beats both without labels. After chain-of-thought is recalibrated on the same labels, no significant gap remains, consistent with a calibration effect. Yet higher ordinal agreement does not ensure better detection of severe cases. PHQ-9 criteria extraction misses most severe posts, and moving from direct prompting to chain-of-thought and then to extraction increases misses in nearly all comparisons. On the primary corpus, a relabeled stress dataset, a model using that dataset's own features, including word counts from the text, is not significantly different from frontier criteria extraction under the a priori rule.
Automatic depression detection from doctor-patient conversations has gained momentum thanks to the availability of public corpora and advances in language modeling. However, interpretability remains limited: strong performance is often reported without revealing what drives predictions. We analyze three datasets: ANDROIDS, DAIC-WOZ, E-DAIC and identify a systematic bias from interviewer prompts in semi-structured interviews. Models trained on interviewer turns exploit fixed prompts and positions to distinguish depressed from control subjects, often achieving high classification scores without using participant language. Restricting models to participant utterances distributes decision evidence more broadly and reflects genuine linguistic cues. While semi-structured protocols ensure consistency, including interviewer prompts inflates performance by leveraging script artifacts. Our results highlight a cross-dataset, architecture-agnostic bias and emphasize the need for analyses that localize decision evidence by time and speaker to ensure models learn from participants' language.
Hasindri Watawana, Sergio Burdisso, Diego A. Moreno-Galván +4
Idiap Research Institute, Switzerland · EPFL, Switzerland · Centro de Investigación en Matemáticas (CIMAT), Mexico +1