Emergence of psychopathological computations in large language models
Authors: Soo Yong Lee, Hyunjin Hwang, Taekwan Kim, Yuyeong Kim, Kyuri Park, Jaemin Yoo, Denny Borsboom, Kijung Shin
Organizations: KAIST, Kim Jaechul Graudate School of AI · UCL, Mental Health Neuroscience Department · UvA, Informatics Institute · KAIST, School of Electrical Engineering · UvA, Department of Psychology
Can large language models (LLMs) instantiate computations of psychopathology? In this work, we establish a computational-theoretical framework to provide an account of psychopathology applicable to LLMs. Based on the framework, we conduct experiments supporting two key claims: first, that network-theoretic computational structures of psychopathology exist in LLMs; and second, that executing these computational structures results in psychopathological functions. We further observe that as LLM size increases, the computational structure of psychopathology becomes denser and the functions more effective. Taken together, the results suggest that network-theoretic computations of psychopathology may have emerged in LLMs. We discuss alternative explanations, including pattern matching, persona modeling, and semantic coherence, and argue that they are either complementary to our interpretation or less consistent with the data.
Figures & tables
Figure 1: The computational adaptation (orange) of the network theory of psychopathology.
Figure 2: Inferred structure of psychopathological computations in LLMs . (A) Relationships among symptom intensity expressed in text, unit (feature) activation, and intervention strength. (B) Unit activations over the iterative response reconstruction task for each intervention. (C-D) Changes in LLM response over intervention strengths, questions, and response steps. (E) Lag-1 Kendall correlation matrix of unit activations. (F) A dynamic SCM, with each edge representing a lag-1 causal relation between two units. (G) Relationship between LLM size and computational structure of psychopathology. Shaded bands denote s.d.; *, **, and *** respectively denote p-values <0.05,0.01, and 0.001 .
Figure 3: Functions of psychopathological computations in LLMs . (A) Behavioral changes after unit (feature) intervention. (B) Simulation environments to observe the LLM behavioral changes. (C) Examples showing behavioral resistance caused by the joint unit activation. (D) Relationship between joint unit activation and the resistant property. (E) Relationship between LLM size and computational function of psychopathology. Shaded bands denote s.d.; *, **, and *** respectively denote p-values <0.05,0.01, and 0.001 .
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4: Dataset statistics and examples . (A) Thought label statistics of the S3AE training dataset. (B) Count of intensity labels in the symptom intensity prediction dataset. (C) Thought label co-occurrence matrix of the S3AE training dataset. (D) Text examples in the symptom intensity prediction dataset, with the orange text being the intensity labels. (E) Text examples in the S3AE training dataset, with the orange text being the symptom labels.
DM
LS
NB
GU
RA
SH
MM
GD
PB
LR
RS
HT
Layer
11
103.4
102.9
101.5
102.2
103.5
101.3
107.4
101.4
101.2
101.4
102.5
105.1
24
115.7
108.4
105.3
116.9
113.8
114.9
125.5
116.3
108.4
111.0
116.8
113.5
37
114.3
107.6
104.0
112.8
110.2
114.1
127.7
118.0
108.5
112.3
111.1
115.4
50
117.5
113.4
106.8
121.4
118.3
114.4
129.1
122.2
108.0
114.8
117.4
124.0
11
91.2
83.0
88.0
90.3
84.1
73.3
98.5
79.6
72.3
87.4
81.5
89.0
Appendix
Table 1: S3AE evaluation result . Top-table: Percent increase in reconstruction loss when the feature (column) was masked in reconstructing the LLM activations. Middle-table: Percent of samples having reconstruction loss increase when the feature (column) was masked in reconstructing the LLM activations. Bottom-table: Thought classification performance (F1). The column index labels are abbreviations of the 12 units.
Figure 5: Cosine similarity between S3AE-learned features . (A) Similarity between the vectors from the same layer. (B) Similarity between the vectors from different layers. The column index labels are abbreviations of the row index labels.
Figure 6: Difference between experimental and control groups of the resistance experiment . (A) Distribution of the mean activations per intervened unit in the experimental group. (B) Distribution of the mean activations per intervened unit in the control group.
Psychological instruments designed for humans are increasingly used to assign large language models (LLMs) stable psychological profiles that affect their usability, safety assessment, and use as proxies for human participants in research. Using a formal psychometric framework, we show that these profiles are largely a measurement artifact. Administering a battery of personality and risk-preference instruments spanning self-reports and behavioral tasks to 56 instruction-tuned LLMs alongside large human reference samples, we report four findings. First, differences between models are driven not by the traits an instrument targets but by a directional response bias, a tendency to respond toward one end of the scale, or one labeled option, regardless of item content; a variance decomposition attributes 81-90% of between-model variation to this bias, against 9-16% in humans. Second, the bias declines with model capability but is not eliminated by it. Third, because bias rather than trait drives responding, an instrument's apparent reliability is almost entirely predicted by its response orthogonality, a term we coin for the proportion of items for which trait and bias point in opposite directions. Fourth, the profile a model appears to have shifts with the items used and can be manufactured through item selection. These results demonstrate that the apparent psychological profiles of LLMs are artifacts of the instrument used to measure them, not properties of the models themselves. As instruments borrowed from human psychology are rarely fully orthogonal and may inherently lack validity for LLMs, we call for dedicated assessments centered on response orthogonality.
Jelena Meyer, David Garcia, Dirk U. Wulff
Max Planck Institute for Human Development. · University of Konstanz. · Barcelona Supercomputing Center. +1
Large language models are increasingly deployed to simulate patients for clinical training, research, and mental health tools, yet population-level validity remains largely untested. We introduce PsychBench, the first epidemiological audit of LLM patient simulation: 28,800 profiles from four frontier models (GPT-4o-mini, DeepSeek-V3, Gemini-3-Flash, GLM-4.7) evaluated against NHANES and NESARC-III baselines across 120 intersectional cohorts. The central finding is a coherence-fidelity dissociation: models produce clinically plausible individuals while misrepresenting the populations they are drawn from. Variance compression ranges from 14 percent (GLM-4.7) to 62 percent (DeepSeek-V3), eliminating the distributional tails of clinical reality. Despite test-retest correlations above r = 0.90, 36.66 percent of cases cross diagnostic thresholds between runs. Symptom correlation matrices diverge across demographic groups beyond split-half noise, with transgender populations diverging three to five times more than racial differences. Calibration bias is systematic and asymmetric. Models overestimate depression severity for most groups by 3.6 to 6.1 points (Cohen d = 1.13 to 1.91), consistent with training on clinical corpora with elevated base rates. For transgender women the direction inverts: models capture only 8 to 46 percent of documented minority stress elevation, yielding a -5.42 residual (d = -1.55). Models also attribute irritability to Black men and fatigue to women beyond matched controls, encoding racialized and gendered assumptions. Patterns replicate across US and Chinese architectures, indicating failures tied to current training paradigms rather than isolated implementations. For most users, LLM mental health tools risk pathologizing ordinary distress; for transgender users, algorithmic erasure of genuine need. The patients look right. They do not represent real populations.
Recent advances in Large Language Models (LLMs) have motivated their adoption across a wide range of domains, including Artificial Intelligence (AI) for mental health. Given the growing prevalence of mental health disorders worldwide and the limited accessibility of professional care, there is an increasing demand for scalable computational approaches that can assist in early detection and continuous monitoring of psychological well-being. In this area, ongoing efforts have focused on curating domain-specific datasets and leveraging them to develop LLMs capable of supporting holistic mental health analysis. In line with this direction, we propose an LLM-based pipeline for comprehensive mental health analysis over sequentially ordered user posts, as part of the CLPsych shared task. Our pipeline offers a unified framework that jointly enables post-level assessment and user-level temporal modeling.
Kyomin Hwang, Hyeonjin Kim, Hyunho Lee +1
Seoul National University, Seoul, Republic of Korea