cs.CLOct 6, 2026

Language-model ratings of depression reflect the rater more than the patient

Authors: Baihan Lin

Organizations: Department of Artificial Intelligence and Human Health, Icahn School of Medicine at Mount Sinai, New York, NY, USA · Department of Psychiatry, Icahn School of Medicine at Mount Sinai, New York, NY, USA · Department of Neuroscience, Icahn School of Medicine at Mount Sinai, New York, NY, USA · Mental Illness Research, Education and Clinical Center, James J. Peters VA Medical Center, Bronx, NY, USA · Berkman Klein Center for Internet & Society, Harvard University, Cambridge, MA, USA

Abstract

Depression has no diagnostic blood test. Language models promise tireless, consistent assessment, but can accurate raters disagree about individuals? We pre-registered 880 language-model raters, crossing 11 open models with prompting and scoring choices, and applied them to 189 interviews against the eight-item Patient Health Questionnaire. Model choice explained 30.0% of summed-symptom score variance, stable participant differences 10.5%. Two randomly drawn raters with area under the receiver operating characteristic curve (AUC) >= 0.70 disagreed on screening decisions for 40% of participants, on average. Average over-rating governed how many were flagged, yet equal-capacity raters chose differently for about one participant in five. A locked analysis of 86 new interviews reproduced the main pre-registered findings. Exploratory recalibration with 40 labelled participants raised accuracy from about 60% to 75% and halved disagreement, leaving one participant in five decided differently. Calibration repaired much of the rater dependence without securing agreement about individuals.

Figures & tables

Explore similar work

Date pendingcs.CL

"Mirror" Large Language Model Evaluations of Depression are Criterion Contaminated

Large Language Model (LLM) studies that use language responses elicited from depression assessments to predict scores on those same assessments often report near-perfect prediction of depression. We refer to these as "Mirror" evaluations and demonstrate an applied case of criterion contamination. N = 110 participants completed both structured diagnostic depression interviews (Mirror condition) and life history interviews ("Non-Mirror" condition). LLMs were prompted to predict depression scores in each condition. As expected, Mirror evaluations were near-perfect. However, Non-Mirror evaluations also displayed prediction sizes considered outstanding in psychology. Further, both Mirror and Non-Mirror predictions correlated with Patient Health Questionnaire-9 scores at similar sizes, suggesting the Mirror condition's advantage collapses when predicting an independent depression measurement. Topic modeling revealed differing depression-related themes across interview types. Mirror evaluations are better considered as reliability evaluations than as validity evaluations. Incorporating Non-Mirror approaches in LLM depression assessment may support more valid and clinically-relevant applications. Keywords: large language models, psychological assessment, psychopathology, depression, reliability, validity, criterion contamination
Sep 30, 2026cs.CL

Structure vs. Chain-of-Thought: Evaluating LLM Criteria Extraction for Depression Severity

A large language model (LLM) can rate depression severity directly from a social media post or mark which clinical criteria the post shows and let code turn the count into a label. The latter is easier to audit because a clinician can check each marked criterion. We compare these approaches on two Reddit corpora using three LLMs (from 9B to frontier scale) and two questionnaires (PHQ-9, BDI-II), and measure agreement with quadratic weighted kappa. For the two frontier models, criteria extraction scores above chain-of-thought on one corpus only when its decision thresholds are fitted on labeled data. Neither model's gain is significant, with or without recalibrating chain-of-thought on the same labels. With thresholds fixed a priori from PHQ-9's criteria, extraction shows no gain on either corpus, even where models mark over two criteria per post. The 9B model behaves differently on a corpus from depression communities. It labels most posts severe, whether prompted directly or with chain-of-thought, while the a priori rule beats both without labels. After chain-of-thought is recalibrated on the same labels, no significant gap remains, consistent with a calibration effect. Yet higher ordinal agreement does not ensure better detection of severe cases. PHQ-9 criteria extraction misses most severe posts, and moving from direct prompting to chain-of-thought and then to extraction increases misses in nearly all comparisons. On the primary corpus, a relabeled stress dataset, a model using that dataset's own features, including word counts from the text, is not significantly different from frontier criteria extraction under the a priori rule.
Mar 25, 2026cs.CL

When Consistency Becomes Bias: Interviewer Effects in Semi-Structured Clinical Interviews

Automatic depression detection from doctor-patient conversations has gained momentum thanks to the availability of public corpora and advances in language modeling. However, interpretability remains limited: strong performance is often reported without revealing what drives predictions. We analyze three datasets: ANDROIDS, DAIC-WOZ, E-DAIC and identify a systematic bias from interviewer prompts in semi-structured interviews. Models trained on interviewer turns exploit fixed prompts and positions to distinguish depressed from control subjects, often achieving high classification scores without using participant language. Restricting models to participant utterances distributes decision evidence more broadly and reflects genuine linguistic cues. While semi-structured protocols ensure consistency, including interviewer prompts inflates performance by leveraging script artifacts. Our results highlight a cross-dataset, architecture-agnostic bias and emphasize the need for analyses that localize decision evidence by time and speaker to ensure models learn from participants' language.