No Transformer Beats Six Covariates: Long-Horizon Prediction of Depressive Symptoms from Childhood Essays
Organizations: Department of Statistics, University of British Columbia, Vancouver, Canada · Vector Institute, Toronto, Canada · Department of Psychiatry, University of British Columbia, Vancouver, Canada · Department of Psychology, University of British Columbia, Vancouver, Canada
Abstract
Natural language processing (NLP) models can detect depression-related language in text written near the time symptoms are measured, but whether pretrained transformers can predict depressive symptoms from text written twelve years earlier is largely untested. In the National Child Development Study, a British birth cohort, we predict probable depressive symptoms at age 23 from essays the same people wrote at age 11. Our baseline, a logistic regression on six childhood covariates, outperforms every text model that sees only the essay: seven fine-tuned transformers, a bag-of-words model, frozen embeddings and four zero-shot large language models. Its area under the receiver operating characteristic curve (AUC-ROC) is 0.737 against 0.670 for the best transformer on the primary seed, and no added text score detectably raises the baseline's AUC-ROC. None of the five domain-pretrained transformers detectably beats its general-domain control after Bonferroni correction. For long-horizon prediction, the baseline remains the model to beat.
Figures & tables
| Step | N |
|---|---|
| Essays parsed, non-empty (UKDA-8313) | 10,509 |
| Plus sex available (UKDA-5565) | 10,507 |
| Plus all six childhood covariates | 9,333 |
| Plus Malaise at 23 (UKDA-5566) | 7,226 |
| Below the cut (Malaise ) | 6,729 (93.1%) |
| At or above the cut (Malaise ) | 497 (6.9%) |
| Model | AUC-ROC [95% CI] | AUC-ROC | AUC-PR | Brier |
|---|---|---|---|---|
| Covariate baseline (logistic regression) | 0.737 [0.688, 0.783] | – | 0.188 | 0.060 |
| TF-IDF + logistic regression (non-neural) | 0.644 [0.591, 0.697] | – | 0.115 | 0.063 |
| General-domain language models | ||||
| BERT-base-uncased | 0.669 [0.617, 0.718] | 0.663 0.026 | 0.128 | 0.064 |
| RoBERTa-base | 0.620 [0.567, 0.670] | 0.661 0.005 | 0.097 | 0.063 |
| Domain-specific language models | ||||
| Score added | AUC [95% CI] | |
| BERT-base | 0.74 / 0.32 | |
| RoBERTa-base | 0.60 / 0.68 | |
| ClinicalBERT | 0.29 / 0.61 | |
| BioBERT | 0.20 / 0.24 | |
| MentalBERT | 0.73 / 0.39 | |
| MentalRoBERTa | 0.92 / 0.66 |
| AUC-ROC of the score | ||||
| Text model | cov. | whole | cov.-fitted | residual |
| MentalBERT | 0.401 | 0.671 | 0.720 | 0.549 |
| BERT-base | 0.343 | 0.663 | 0.721 | 0.542 |
| RoBERTa-base | 0.415 | 0.661 | 0.725 | 0.522 |
| MentalRoBERTa | 0.427 | 0.654 | 0.717 | 0.525 |
| BioBERT | 0.289 | 0.651 | 0.733 | 0.506 |
| Age-11 target | BERT-base | MentalBERT | TF-IDF | Cov. ref. |
|---|---|---|---|---|
| Sex | 0.991 | 0.991 | 0.985 | 0.583 |
| Ability | 0.864 | 0.864 | 0.827 | 0.723 |
| Internalising | 0.682 | 0.693 | 0.658 | 0.742 |
| Externalising | 0.699 | 0.693 | 0.623 | 0.769 |
| Encoder | Stock | Adapted |
|---|---|---|
| BERT-base | 0.663 0.026 | 0.673 0.012 |
| RoBERTa-base | 0.661 0.005 | 0.687 0.017 |
| MentalBERT | 0.671 0.007 | 0.678 0.012 |
| bert-base-cased | 0.665 0.012 | 0.682 0.015 |
| DeBERTa-v3-base † | 0.657 0.009 | 0.678 0.022 |
| DeBERTa-v3-large † | 0.665 0.036 | 0.673 0.026 |
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
| Model | Parse | Odds ratio per SD [95% CI] |
|---|---|---|
| Qwen2.5-7B | 0.906 [0.803, 1.022] | |
| Qwen2.5-32B | 0.881 [0.782, 0.994] | |
| Gemma-2-9B | 0.930 [0.825, 1.048] | |
| Llama-3.1-8B | 0.948 [0.839, 1.070] |
| Variable | What it measures (NCDS code) | Reported by | Range | Age-11 target |
|---|---|---|---|---|
| Malaise total (outcome) | derived total mal of 24 yes/no questions on emotional and bodily symptoms ( n6016 to n6039 ) | participant, age 23 | 0–20 | outcome: positive if |
| Sex | female or male ( n622 ) | cohort records | 0 or 1 | female |
| Father’s social class | father’s or male head’s occupation: non-manual, manual or no male head ( n1685 , 1966 General Register Office scheme) | parent, age 11 | 1–3 | not used |
| Parental psychiatric history | psychiatric condition coded for the mother ( n1406 , n1407 ) or father ( n1415 , n1416 ) among chronic or serious illnesses since age 7, other codes 0 | parent, age 11 | 0 or 1 | not used |
| Cognitive ability | items correct of 80 (40 verbal, 40 non-verbal) on a 30-minute general ability test ( n920 ) | test, age 11 | 0–79 | low: (bottom third) |
| BSAG internalising | sum of depression ( n980 ), withdrawal ( n977 ), unforthcomingness ( n974 ) and writing off adults ( n989 ) | teacher, age 11 | 0–31 | high: (top tenth) |
| Learning rate | Weight decay | 0.01 | |
|---|---|---|---|
| Batch size | 16 | Warmup | 10% |
| Train epochs | 6 | Rate decay | linear |
| Early-stop patience | 3 epochs | Max gradient norm | 1.0 |
| Hugging Face model |
| Fine-tuned encoders (Table 2 ) |
| bert-base-uncased |
| roberta-base |
| emilyalsentzer/Bio_ClinicalBERT |
| dmis-lab/biobert-base-cased-v1.1 |
| mental/mental-bert-base-uncased |
| Outcome cut | 6 | 7 | 8 | 9 | 10 |
|---|---|---|---|---|---|
| Test positives | 218 | 154 | 99 | 59 | 37 |
| Baseline, relabelled | 0.690 | 0.712 | 0.737 | 0.733 | 0.739 |
| Baseline, refitted | 0.693 | 0.714 | 0.737 | 0.734 | 0.746 |
| Baseline minus model AUC-ROC, both relabelled | |||||
| BERT-base | 0.043 | 0.051 | 0.068 | 0.069 | 0.048 † |
| RoBERTa-base | 0.068 | 0.086 | 0.117 | 0.104 | 0.096 |