Emergence of psychopathological computations in large language models
Authors: Soo Yong Lee, Hyunjin Hwang, Taekwan Kim, Yuyeong Kim, Kyuri Park, Jaemin Yoo, Denny Borsboom, Kijung Shin
Organizations: KAIST, Kim Jaechul Graudate School of AI · UCL, Mental Health Neuroscience Department · UvA, Informatics Institute · KAIST, School of Electrical Engineering · UvA, Department of Psychology
Can large language models (LLMs) instantiate computations of psychopathology? In this work, we establish a computational-theoretical framework to provide an account of psychopathology applicable to LLMs. Based on the framework, we conduct experiments supporting two key claims: first, that network-theoretic computational structures of psychopathology exist in LLMs; and second, that executing these computational structures results in psychopathological functions. We further observe that as LLM size increases, the computational structure of psychopathology becomes denser and the functions more effective. Taken together, the results suggest that network-theoretic computations of psychopathology may have emerged in LLMs. We discuss alternative explanations, including pattern matching, persona modeling, and semantic coherence, and argue that they are either complementary to our interpretation or less consistent with the data.
Figures & tables
Figure 1: The computational adaptation (orange) of the network theory of psychopathology.
Figure 2: Inferred structure of psychopathological computations in LLMs . (A) Relationships among symptom intensity expressed in text, unit (feature) activation, and intervention strength. (B) Unit activations over the iterative response reconstruction task for each intervention. (C-D) Changes in LLM response over intervention strengths, questions, and response steps. (E) Lag-1 Kendall correlation matrix of unit activations. (F) A dynamic SCM, with each edge representing a lag-1 causal relation between two units. (G) Relationship between LLM size and computational structure of psychopathology. Shaded bands denote s.d.; *, **, and *** respectively denote p-values <0.05,0.01, and 0.001 .
Figure 3: Functions of psychopathological computations in LLMs . (A) Behavioral changes after unit (feature) intervention. (B) Simulation environments to observe the LLM behavioral changes. (C) Examples showing behavioral resistance caused by the joint unit activation. (D) Relationship between joint unit activation and the resistant property. (E) Relationship between LLM size and computational function of psychopathology. Shaded bands denote s.d.; *, **, and *** respectively denote p-values <0.05,0.01, and 0.001 .
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4: Dataset statistics and examples . (A) Thought label statistics of the S3AE training dataset. (B) Count of intensity labels in the symptom intensity prediction dataset. (C) Thought label co-occurrence matrix of the S3AE training dataset. (D) Text examples in the symptom intensity prediction dataset, with the orange text being the intensity labels. (E) Text examples in the S3AE training dataset, with the orange text being the symptom labels.
DM
LS
NB
GU
RA
SH
MM
GD
PB
LR
RS
HT
Layer
11
103.4
102.9
101.5
102.2
103.5
101.3
107.4
101.4
101.2
101.4
102.5
105.1
24
115.7
108.4
105.3
116.9
113.8
114.9
125.5
116.3
108.4
111.0
116.8
113.5
37
114.3
107.6
104.0
112.8
110.2
114.1
127.7
118.0
108.5
112.3
111.1
115.4
50
117.5
113.4
106.8
121.4
118.3
114.4
129.1
122.2
108.0
114.8
117.4
124.0
11
91.2
83.0
88.0
90.3
84.1
73.3
98.5
79.6
72.3
87.4
81.5
89.0
Appendix
Table 1: S3AE evaluation result . Top-table: Percent increase in reconstruction loss when the feature (column) was masked in reconstructing the LLM activations. Middle-table: Percent of samples having reconstruction loss increase when the feature (column) was masked in reconstructing the LLM activations. Bottom-table: Thought classification performance (F1). The column index labels are abbreviations of the 12 units.
Figure 5: Cosine similarity between S3AE-learned features . (A) Similarity between the vectors from the same layer. (B) Similarity between the vectors from different layers. The column index labels are abbreviations of the row index labels.
Figure 6: Difference between experimental and control groups of the resistance experiment . (A) Distribution of the mean activations per intervened unit in the experimental group. (B) Distribution of the mean activations per intervened unit in the control group.