COMPASS 2.0: psychometric representational similarity analysis distinguishes symptom structure from personal signal
Organizations: Department of Artificial Intelligence and Human Health, Icahn School of Medicine at Mount Sinai, New York, NY, USA · Department of Psychiatry, Icahn School of Medicine at Mount Sinai, New York, NY, USA · Department of Neuroscience, Icahn School of Medicine at Mount Sinai, New York, NY, USA · Mental Illness Research, Education and Clinical Center, James J. Peters VA Medical Center, Bronx, NY, USA · Berkman Klein Center for Internet & Society, Harvard University, Cambridge, MA, USA
Abstract
Language models can score psychiatric questionnaires from speech, but agreement with self-report may reflect the questionnaire rather than the person. We introduce psychometric representational similarity analysis, a framework for comparing the structure of speech-derived scores, self-report, item wording and theory, and implement it alongside person-level construct scoring in COMPASS 2.0. We show how similarly worded items induce covariance without psychological signal. In pre-registered discovery and confirmation analyses of clinical interviews from 275 participants, language-derived symptom geometry resembled wording more than self-report, with no structure beyond wording detected by the registered tests. Geometric agreement with self-report survived assigning participants someone else's answers, whereas person-paired scores captured distress more than specific symptoms. Complementary analyses examined counselling quality and wording structure across 34 instruments and the Research Domain Criteria (RDoC) framework. These findings distinguish agreement about psychological structure from evidence that language-derived assessments track individual people.
Figures & tables
| Corpus | Content | People or sessions | Units | Use |
|---|---|---|---|---|
| DAIC-WOZ | Clinical interviews by a human-operated virtual interviewer; human transcripts | 189 | 10,533 | Discovery (symptoms) |
| E-DAIC | Interviews by an autonomous agent; automatic transcripts | 86 | 5,214 | Confirmation (symptoms) |
| Alexander Street | Transcribed psychotherapy sessions, four clinical collections | 786 | 180,547 | Alliance re-examination |
| AnnoMI | Motivational-interviewing demonstrations, high/low quality (110/23) | 133 | 9,687 | Discovery (alliance) |
| High/low-quality counselling | Counselling sessions not in AnnoMI, high/low quality (119/65) | 184 | 7,877 | Confirmation (alliance) |
| Sample | Scorer | H1 | H2 | H3 |
|---|---|---|---|---|
| Discovery | MiniLM | 0.68 / 0.42 (0.15 to 0.47) | 0.009 ( = 0.93) | 0.004 ( 0.016 to 0.023) |
| bge | 0.69 / 0.47 (0.15 to 0.40) | 0.028 ( = 0.52) | 0.014 ( 0.002 to 0.031) | |
| Entailment | – | 0.051 ( = 0.27) | 0.038 (0.023 to 0.056) | |
| Confirmation | MiniLM | 0.63 / 0.36 (0.11 to 0.55) | 0.008 ( = 0.93) | 0.038 (0.007 to 0.072) |
| bge | 0.65 / 0.21 (0.33 to 0.58) | 0.096 ( = 0.15) | 0.042 (0.013 to 0.073) | |
| Entailment | – | 0.118 ( = 0.15) | 0.005 ( 0.022 to 0.035) |
| Partial after whitening by ( ), controlling for | |||||||
|---|---|---|---|---|---|---|---|
| Sample | Scorer | H2 interval | Block | Corr | registered | both | |
| D | MiniLM | [ 0.137, 0.129] | 0.925 | 0.904 | 0.027 (0.722) | 0.003 (0.972) | 0.017 (0.823) |
| D | bge | [ 0.067, 0.109] | 0.520 | 0.877 | 0.031 (0.522) | 0.060 (0.190) | 0.059 (0.282) |
| D | entailment | [ 0.012, 0.107] | 0.260 | 0.216 | 0.030 (0.567) | 0.005 (0.923) | 0.030 (0.551) |
| C | MiniLM | [ 0.098, 0.123] | 0.962 | 0.914 | 0.029 (0.697) | 0.046 (0.522) | 0.076 (0.291) |
| C | bge | [ 0.134, 0.001] | 0.057 | 0.889 | 0.079 (0.135) | 0.045 (0.367) | 0.004 (0.950) |
| Analysis | Sessions | Videos | Stance AUC [95% CI] | Keyed cosine AUC | Difference [95% CI] |
|---|---|---|---|---|---|
| Discovery (AnnoMI) | 133 | 119 | 0.547 [0.420, 0.671] | 0.662 | [ 0.275, 0.040] |
| Registered confirmation sample | 184 | 156 | 0.682 [0.583, 0.773] | 0.564 | [ 0.009, 0.243] |
| Registered sample, screen-negative only | 176 | 150 | 0.695 [0.594, 0.785] | 0.550 | [0.013, 0.275] |
| All screen-negative sessions without ID match | 192 | 166 | 0.678 [0.585, 0.765] | 0.554 | [ 0.0001, 0.246] |
| Registered sample without over-limit segments | 184 | 156 | 0.676 [0.581, 0.767] | 0.564 | [ 0.019, 0.239] |
| Registered | Identifier | Transcript | |
|---|---|---|---|
| analysed | ID distinct | screen-negative | 176 |
| analysed | ID distinct | text match | 8 |
| excluded | ID distinct | screen-negative | 1 |
| excluded | ID match | screen-negative | 4 |
| excluded | ID match | text match | 49 |
| excluded | ID unresolved | screen-negative | 15 |
| COMPASS PHQ-8 score with PHQ-8 total | General factors | ||||
| Sample | Scorer | partial | partial | ||
| D | MiniLM | 0.15 [0.00, 0.30] | 0.07 [ 0.08, 0.21] | 0.07 [ 0.06, 0.22] | 0.04 [ 0.16, 0.11] |
| D | bge | 0.26 [0.13, 0.37] | 0.16 [0.03, 0.29] | 0.26 [0.13, 0.38] | 0.14 [ 0.01, 0.27] |
| D | entailment | 0.54 [0.44, 0.62] | 0.52 [0.43, 0.61] | 0.49 [0.38, 0.59] | 0.46 [0.35, 0.56] |
| C | MiniLM | 0.39 [0.24, 0.53] | 0.12 [ 0.06, 0.28] | 0.33 [0.14, 0.50] | 0.02 [ 0.26, 0.24] |
| C | bge | 0.53 [0.41, 0.65] | 0.40 [0.23, 0.56] | 0.52 [0.38, 0.65] | 0.36 [0.16, 0.52] |
| Analysis | Estimand | Unit of resampling | Status |
|---|---|---|---|
| H1 | Agreement of the naive population geometry with item wording minus its agreement with self-reported geometry | Participants | Registered (OSF) |
| H2 | Partial correlation of the whitened population geometry with self-reported geometry, beyond wording | Items (permutation) | Registered (OSF), primary |
| H3 | Person-paired item correspondence: mean same-item minus different-item correlation | Participants | Registered (OSF) |
| A1, A2 | Separation of high- from low-quality sessions by stance scores, and its advantage over keyed cosine scores | Sessions | Pre-specified (fingerprinted plan) |
| Psychometric RDM comparisons (Fig. 1) | Second-order agreement of population geometries across sources; split-half reliabilities | Participants (reliability only) | Descriptive |
| General factors, COMPASS PHQ-8 score | Person-paired agreement of language and self-report, alone and beyond length and topic baselines | Participants | Exploratory |
| Bank | Domain | Items | Text used | Source |
|---|---|---|---|---|
| CAPE-42 | symptoms | 42 | items, verbatim | 78 |
| SAPS | symptoms | 34 | item names | 3 |
| C-SSRS | symptoms | 31 | items, verbatim | 68 |
| PANSS | symptoms | 30 | item names | 42 |
| SANS | symptoms | 25 | item names | 2 , 3 |
| BDI-II | symptoms | 21 | item titles | 5 |
| Item | Hypotheses |
|---|---|
| PHQ-8 1 NoInterest | I have little interest in doing things. I have little pleasure in doing things. |
| PHQ-8 2 Depressed | I feel down. I feel depressed. I feel hopeless. |
| PHQ-8 3 Sleep | I have trouble falling asleep. I have trouble staying asleep. I sleep too much. |
| PHQ-8 4 Tired | I feel tired. I have little energy. |
| PHQ-8 5 Appetite | I have a poor appetite. I overeat. |
| PHQ-8 6 Failure | I feel bad about myself. I feel like a failure. I have let myself or my family down. |