Structure vs. Chain-of-Thought: Evaluating LLM Criteria Extraction for Depression Severity
Organizations: Independent Researcher
Abstract
A large language model (LLM) can rate depression severity directly from a social media post or mark which clinical criteria the post shows and let code turn the count into a label. The latter is easier to audit because a clinician can check each marked criterion. We compare these approaches on two Reddit corpora using three LLMs (from 9B to frontier scale) and two questionnaires (PHQ-9, BDI-II), and measure agreement with quadratic weighted kappa. For the two frontier models, criteria extraction scores above chain-of-thought on one corpus only when its decision thresholds are fitted on labeled data. Neither model's gain is significant, with or without recalibrating chain-of-thought on the same labels. With thresholds fixed a priori from PHQ-9's criteria, extraction shows no gain on either corpus, even where models mark over two criteria per post. The 9B model behaves differently on a corpus from depression communities. It labels most posts severe, whether prompted directly or with chain-of-thought, while the a priori rule beats both without labels. After chain-of-thought is recalibrated on the same labels, no significant gap remains, consistent with a calibration effect. Yet higher ordinal agreement does not ensure better detection of severe cases. PHQ-9 criteria extraction misses most severe posts, and moving from direct prompting to chain-of-thought and then to extraction increases misses in nearly all comparisons. On the primary corpus, a relabeled stress dataset, a model using that dataset's own features, including word counts from the text, is not significantly different from frontier criteria extraction under the a priori rule.
Figures & tables
| DepSeverity | DepSign | |
| lowest class | 513 minimum | 184 not dep. |
| middle class(es) | 58 mild / 79 mod. | 472 mod. |
| highest class | 56 severe | 50 severe |
| majority acc. | 0.727 | 0.669 |
| majority | 0.000 | 0.000 |
| median words | 80 | 104 |
| Model | Cond. | MAE | Acc | F1 M | sev | |
|---|---|---|---|---|---|---|
| DepSeverity (4 classes, 56 severe ) | ||||||
| Qwen3.5-9B | C1 | 0.288 | 0.948 | 0.445 | 0.335 | 29/56 |
| C2 | 0.340 | 0.877 | 0.456 | 0.327 | 34/56 | |
| C3 | 0.489 / 0.322 | 0.465 | 0.728 | 0.367 | 43/56 | |
| C4 | 0.495 /0.470/0.057 | 0.517 | 0.643 | 0.420 | 40/56 | |
| DeepSeek | C1 | 0.219 | 1.280 | 0.329 | 0.251 | 16/56 |
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
| Corpus | Model | Comparison | [95% CI] | |
|---|---|---|---|---|
| DepSeverity | Qwen3.5-9B | C2 C1 | +0.052 [+0.009, +0.098] | C |
| DepSeverity | Qwen3.5-9B | C3 fitted C2 | +0.149 [+0.072, +0.223] | C,H |
| DepSeverity | Qwen3.5-9B | C3 a priori C2 | 0.018 [ 0.089, +0.054] | C |
| DepSeverity | Qwen3.5-9B | C4 fitted C2 | +0.155 [+0.082, +0.227] | |
| DepSeverity | Qwen3.5-9B | C4 a priori C2 | +0.130 [+0.050, +0.208] | |
| DepSeverity | Qwen3.5-9B | C3 fitted C1 | +0.201 [+0.127, +0.276] |
| Corpus | Model | Cond. | fitted | a priori |
|---|---|---|---|---|
| DepSeverity | Qwen3.5-9B | C3 | [0.5, 1, 1.5] | [0.5, 2.5, 4.5] |
| DepSeverity | Qwen3.5-9B | C4 | [0.5, 1.5, 3.5] | [0.5, 2.5, 4.5] |
| DepSeverity | DeepSeek | C3 | [0.5, 1.5, 2.5] | [0.5, 2.5, 4.5] |
| DepSeverity | DeepSeek | C4 | [0.5, 1.5, 3.5] | [0.5, 2.5, 4.5] |
| DepSeverity | Claude | C3 | [0.5, 1.5, 2.5] | [0.5, 2.5, 4.5] |
| DepSeverity | Claude | C4 | [0.5, 2.5, 5.5] | [0.5, 2.5, 4.5] |
| Corpus | Model | raw | +cal | map | C3 fitted +cal | |
|---|---|---|---|---|---|---|
| DepSeverity | Qwen3.5-9B | C1 | 0.288 | 0.320 | 0012 | +0.169 [+0.090, +0.246] |
| DepSeverity | Qwen3.5-9B | C2 | 0.340 | 0.379 | 0013 | +0.109 [+0.027, +0.189] |
| DepSeverity | DeepSeek | C1 | 0.219 | 0.330 | 0002 | +0.196 [+0.113, +0.277] |
| DepSeverity | DeepSeek | C2 | 0.504 | 0.504 | 0123 | +0.022 [ 0.039, +0.086] |
| DepSeverity | Claude | C1 | 0.299 | 0.352 | 0013 | +0.147 [+0.067, +0.226] |
| DepSeverity | Claude | C2 | 0.462 | 0.462 | 0123 | +0.037 [ 0.029, +0.103] |
| Corpus | Model | Comparison | [95% CI] |
|---|---|---|---|
| DepSeverity | Qwen3.5-9B | C3 a priori C2 | 0.058 [ 0.136, +0.023] |
| DepSeverity | Qwen3.5-9B | C4 fitted C2 | +0.116 [+0.024, +0.202] |
| DepSeverity | Qwen3.5-9B | C4 a priori C2 | +0.091 [ 0.001, +0.181] |
| DepSeverity | Qwen3.5-9B | C3 a priori C1 | +0.002 [ 0.072, +0.079] |
| DepSeverity | Qwen3.5-9B | C4 fitted C1 | +0.176 [+0.085, +0.259] |
| DepSeverity | Qwen3.5-9B | C4 a priori C1 | +0.150 [+0.060, +0.235] |
| Model | Cond. | minimum | mild | moderate | severe |
|---|---|---|---|---|---|
| Qwen3.5-9B | C1 | 0.46 | 0.29 | 0.43 | 0.48 |
| C2 | 0.50 | 0.26 | 0.38 | 0.39 | |
| C3 fitted | 0.92 | 0.00 | 0.35 | 0.23 | |
| C3 a priori | 0.92 | 0.36 | 0.03 | 0.00 | |
| C4 fitted | 0.77 | 0.29 | 0.35 | 0.29 | |
| C4 a priori | 0.77 | 0.53 | 0.14 | 0.12 |
| Model | Cond. | not dep. | moderate | severe |
|---|---|---|---|---|
| Qwen3.5-9B | C1 | 0.22 | 0.24 | 0.74 |
| C2 | 0.13 | 0.16 | 0.80 | |
| C3 fitted | 0.40 | 0.85 | 0.04 | |
| C3 a priori | 0.40 | 0.78 | 0.06 | |
| C4 fitted | 0.38 | 0.70 | 0.24 | |
| C4 a priori | 0.38 | 0.52 | 0.42 |
| Corpus | Model | C1 | C2 | C3 | C4 |
|---|---|---|---|---|---|
| DepSeverity | Qwen3.5-9B | 179 / 3 | 212 / 519 | 535 / 186 | 767 / 459 |
| DepSeverity | DeepSeek | 162 / 2 | 194 / 363 | 499 / 152 | 712 / 371 |
| DepSeverity | Claude | 238 / 5 | 282 / 545 | 737 / 227 | 1021 / 541 |
| DepSign | Qwen3.5-9B | 276 / 3 | 309 / 492 | 633 / 200 | 865 / 476 |
| DepSign | DeepSeek | 262 / 2 | 294 / 429 | 600 / 171 | 813 / 404 |
| DepSign | Claude | 369 / 4 | 413 / 614 | 871 / 251 | 1155 / 579 |
| Variant | pres./post | absent | fitted | a pr. | vs C2, fitted | vs C2, a pr. |
|---|---|---|---|---|---|---|
| original | 0.45 | 0.30% | 0.526 | 0.423 | +0.022 [ 0.039, +0.086] | 0.081 [ 0.147, 0.010] |
| permissive | 0.72 | 0.25% | 0.507 | 0.425 | +0.002 [ 0.059, +0.067] | 0.079 [ 0.145, 0.012] |
| symmetric evidence | 0.47 | 0.35% | 0.555 | 0.433 | +0.051 [ 0.011, +0.116] | 0.071 [ 0.138, +0.000] |