Too Categorical to be Human: Emotion Concepts in LLMs and Humans
Authors: Sree Bhattacharyya, Evgenii Kuriabov, Lucas Craig, Tharun Dilliraj, Reginald B. Adams,, Jia Li, James Z. Wang
Organizations: Department of Informatics and Intelligent Systems, College of Information Sciences and Technology, The Pennsylvania State University · Department of Statistics, Eberly College of Science, The Pennsylvania State University · Department of Computer Science and Engineering, College of Engineering, The Pennsylvania State University · Department of Psychology, College of the Liberal Arts, The Pennsylvania State University
Understanding human emotions is central to user-facing AI applications, safety alignment, and the simulation of human behavior. As emotional stimuli shape high-stakes behavior in Large Language Models (LLMs), there is increasing interest in how models represent emotion concepts internally. Mechanistic accounts of these representations, however, cannot be compared directly against humans: emotion processing in humans is highly distributed and yields no equivalent neural representation. To understand whether LLMs internalize emotion concepts in a way similar to humans, we propose characterizing the abstract concept of an emotion using external behavioral signatures, which we term behavioral representations. Using the theory of cognitive appraisals, which enables representing emotional situations along interpretable evaluative dimensions, we create a benchmark dataset of emotional scenarios spanning 15 emotion categories. We elicit behavioral representations of emotion concepts from LLMs and humans using our benchmark, and study their structural similarity. We find that LLMs represent emotion concepts more categorically, homogeneously, and determinately than humans, representing a single emotion concept with less internal diversity, and place different emotions further apart. The categorical structure of representations in LLMs is further robust to contextual variation, including with different task framing and demographic personas. Analyzing model checkpoints across different training stages, we also find that the discretized nature of representations appears after the mid-training stage itself and is unaffected by different post-training strategies. Through our results, we highlight a key difference in how LLMs behaviorally represent emotion concepts, curbing the subjectivity inherent to the human experience of emotions.
Figures & tables
Figure 1 : Overview. (1) CoRE contains 270 scenarios spanning 15 emotion categories, each rated on 17 appraisal dimensions. (2) Human participants and LLMs answer identical first-person appraisal prompts. For each agent, the scenarios of one emotion form a matrix Ee∈Rne×17 , that is, a cloud of ne points in appraisal space. (3) Comparing these clouds across emotions shows that LLMs arrange emotions similarly to humans but with sharper boundaries: emotion categories are further apart (categorical), vary less within a category (determinate), and are more similar across models (convergent). Clouds are schematic and not computed with the actual data.
Dataset
Perspective
Ratings
x-enVent ( Troiano et al., 2022 )
3rd
–
crowd-enVent ( Troiano et al., 2023 )
dual
1
Covid-ET ( Zhan et al., 2023 )
3rd
–
PEACE-Reviews ( Yeo and Jaidka, 2023 )
3rd
–
Yeo and Jaidka (2025)
3rd
–
FGE ( Skerry and Saxe, 2015 )
3rd
–
Table 1 : Existing appraisal datasets compared with CoRE. Perspective is the view the appraisal is made from (appraising self, or others). Ratings is the number of independent self-appraisals collected per scenario.
Figure 2 : Between-emotion and within-emotion distance measures, computed for each agent individually, using the W2 Wasserstein distance metric. Statistical significance is tested across each pair of emotions ( DBetween ) and for each emotion category ( DWithin ), comparing each model with humans, using a two-sided Wilcoxon signed-rank test, and rank-biserial correlation is also computed as effect sizes. Holm-Bonferroni method ( Holm, 1979 ) is used for correction of multiple comparisons, and all comparisons are significant ( p<0.001 ).
Figure 3 : 2-Wasserstein distance of emotion representations, comparing two different sources. For humans, to match the number of models, eight random groups are created (sampled 300 times) with representations averaged within each group to represent a single source. The 95% range of human distances is presented.
Figure 4 : Effect of: (blue) removing the emotion label from the prompt, (red) eliciting responses with a third-person framing (purple) inducing nationality-based cultural personas, and (green) conditioning on personality traits on the separation ratio and its distance components.
Figure 5 : (a) Separation ratio across the Olmo model checkpoints, across different training stages, (b) Normalised entropy of the probability distribution over the rating scale at the answer position.
Appendix figures & tables26 assets
Supplementary material from the paper’s appendix.
Appendix
anger
frustration
contempt
fear
disgust
boredom
sadness
guilt
Scenarios
22
16
17
20
17
17
16
18
Valence
neg.
neg.
neg.
neg.
neg.
neg.
neg.
neg.
shame
challenge
surprise
hope
interest
pride
happiness
Scenarios
18
25
17
18
20
15
14
Valence
neg.
amb.
amb.
pos.
pos.
pos.
pos.
Appendix
Table 2: The fifteen emotion categories and the number of scenarios rated for each, over the 270 scenarios in the benchmark.
Family
Dimensions
Scale as presented
Valence
pleasantness, enjoyment
−5 to 5
Attention
attention, consideration
−5 to 5
Comprehension
understand, certainty
1 – 11
Control
self, other, situational
1 – 11
Goal obstruction
problem, obstacles
1 – 11
Legitimacy
cheated, fairness
1 – 11
Appendix
Table 3: The seventeen appraisal dimensions, grouped by what each question asks about. Every dimension is a single rated question. Pleasantness and attention were presented on a −5 to 5 scale, following the original study ( Smith and Ellsworth, 1985 ) and shifted to 1 – 11 for analysis, so that all seventeen dimensions share one scale.
Figure 6 : Overview of the dataset curation process, including the four independent stages.
Emotion
hope
happiness
disgust
anger
interest
fear
sadness
challenge
guilt †
surprise †
pride †
boredom †
Profile r
0.997
0.989
0.953
0.944
0.920
0.888
0.769
0.758
0.534
0.476
0.465
0.304
Appendix
Table 4 : Directional alignment of the human mean appraisal profile with meta-analytic theory: Pearson correlation between each emotion’s mean deviation profile and the theory’s signed- r vector. Higher is better; † marks p>0.05 (positive but non-significant).
Emotion
shame
happiness
contempt
hope
interest
guilt
pride
fear
boredom
disgust
anger
sadness
surprise
challenge
Agreement (%)
100.0
91.7
88.2
81.9
78.9
73.0
71.7
71.1
69.1
69.1
63.6
63.2
60.8
47.6
Appendix
Table 5 : Per-scenario directional sign agreement with theory: percentage of mapped appraisal dimensions on which individual human scenario ratings fall on the theory-predicted side of the grand mean. Chance is 50%.
Figure 7 : Coarse correspondence in appraisal-based representation space for LLMs and humans. (a) shows the cumulative variance explained as a function of the number of principal components, (b) shows the effective number of dimensions used. For Fig. (a), beyond showing the variance explained from averaged human ratings, we also show an approximation of the variance if each scenario was rated by a single rater (gray band). For this, for each scenario, we sample a single rater at random to compute PCA and repeat the process 100 times to show the 95% interval.
Figure 8 : The complete loadings of each original appraisal dimension ( n=17 ) onto the top two principal components recovered (left = PC1, right = PC2).
Separation Ratio
s=30
s=50
s=70
s=100
Humans
1.293
1.298
1.296
1.298
GPT o4-mini
2.084
2.108
2.083
2.092
Gemini 2.5 Flash
2.196
2.198
2.190
2.193
DeepSeek R1
2.275
2.282
2.269
2.266
Phi-4
2.130
2.160
2.133
2.151
Appendix
Table 6: Robustness of the separation ratio with different random seeds, and different numbers of random halves used to estimate the within-emotion spread. Separation ratio is recomputed at s=30,50,70,100 .
raters drawn
mean depth
between-emotion
within-emotion
separation
per cell, k
achieved
W2
W2
ratio
1
1.00
4.745
4.356
1.089 ± 0.007
2
1.97
4.736
4.014
1.180 ± 0.009
3
2.83
4.674
3.762
1.243 ± 0.004
4
3.36
4.661
3.646
1.279 ± 0.005
5
3.66
4.654
3.600
1.293 ± 0.001
Appendix
Table 7: The human–model gap in the separation ratio is not an artefact of rater noise. Human appraisals were recomputed from k raters drawn at random per scenario and dimension, five draws at each depth; where a cell held fewer than k raters all of them were used, so the depth actually achieved is reported beside the requested one. Averaging more raters removes measurement error and compresses the within-emotion spread, raising the ratio from 1.089 at a single rater to 1.293 at the full panel. Extrapolating to a perfectly reliable panel by fitting W22(k)=A+B/k (within-emotion: A=11.17 , B=8.09 , R2=0.95 ; between-emotion: A=21.47 , B=1.15 , R2=0.71 ) yields 1.387 , still well below the least discrete model. Values after ± are standard deviations across the five draws.
Between-emotion
Within-emotion
Ratio
Significant
Default
Noise-matched
Default
Noise-matched
Default
Noise-matched
draws (B / W)
Humans
4.65
–
3.60
–
1.29
–
–
Gemini 3.7 Flash
5.44
5.31
2.15
2.96
2.53
1.80
5 / 5
GPT-5.6 Luna
5.38
5.23
2.22
3.16
2.43
1.66
5 / 5
Muse Glimmer 30B
5.43
5.28
2.26
3.07
2.40
1.72
5 / 5
DeepSeek R1
5.42
5.30
2.38
3.06
2.28
1.73
5 / 5
Appendix
Table 8: Separation of emotion concepts with the models measured as noisily as the human ratings, averaged across five noise draws. Between-emotion: mean W2 distance between two emotion categories. Within-emotion: mean W2 spread within one category. Ratio: between divided by within. Significant draws: the number of the five draws in which the model differs from humans at Holm-corrected p<0.05 (Wilcoxon signed-rank test) on the between-emotion distances (B) and on the within-emotion spreads (W).
Figure 9 : Visualization of multivariate W2 distances between representations of different emotion concepts. In humans, lesser clear categories are observable, with lower distances between different emotions. On the other hand, for all LLMs, clearer, discrete clusters of emotions are observed.
Figure 10 : Complete F1 results. (a) Per-emotion F1 scores across different entities, where for each of the 15 emotion categories, F1 scores are higher for each model compared to humans. (b) Macro-F1 scores shown for each entity.
Macro F1
Difference from humans
Significant
Default
Noise-matched
Default
Noise-matched
draws
Humans
0.26
–
–
–
–
Gemini 3.7 Flash
0.80
0.69
0.55
0.43
5
Muse Glimmer 30B
0.80
0.66
0.55
0.40
5
GPT-5.6 Luna
0.74
0.57
0.48
0.32
5
DeepSeek R1
0.73
0.62
0.47
0.36
5
Appendix
Table 9: Recovery of the emotion label from the appraisal ratings when the same measurement error from humans is added to the model ratings. Macro F1 : the average of the fifteen per-emotion F1 scores of the cross-validated classifier, chance is 0.07 . Significant draws: the number of the five draws in which the model’s per-emotion scores are higher than the human ones at Holm-corrected p<0.05 (Wilcoxon signed-rank test over the fifteen emotion categories).
Humans
GPT o4-mini
Gemini 2.5 Flash
DeepSeek R1
Phi-4
QwQ 32B
GPT-5.6 Luna
Gemini 3.7 Flash
Muse Glimmer 30B
mean ρ
0.058
0.401
0.508
0.496
0.289
0.240
0.356
0.554
0.423
median ρ
0.077
0.425
0.508
0.490
0.330
0.293
0.396
0.548
0.445
Appendix
Table 10: Consistency of emotion representations across two different elicitation formats. The table reports the Spearman correlation between the two rankings, averaged over the 90 pairs. Every model is above zero (Holm-corrected p<<0.001 ) and exceeds humans when paired on the same emotion pairs (all Holm-corrected p<<0.001 ; rank-biserial 0.55 to 0.96 ). The human value is not distinguishable from zero ( p=0.069 ). The human correspondence computed separately for each individual participant, is 0.017 on average ( p=0.61 ).
Noise-matched
Default
ρ
Gap kept
Humans
0.02
–
–
Gemini 3.7 Flash
0.55
0.55
98%
Gemini 2.5 Flash
0.51
0.44
86%
DeepSeek R1
0.50
0.48
96%
Muse Glimmer 30B
0.42
0.42
101%
Appendix
Table 11: Agreement between the explicit and the implicit ranking of the appraisal dimensions, with the model appraisal ratings given the same measurement error as the human ratings. ρ : the mean Spearman correlation over the emotion pairs between the two rankings. Gap kept: the share of the model’s clean gap to the human value that remains under matched error. The human value differs here from Table 10 as this is computed by pooling all ratings, instead of averaging.
Figure 11 : Mean balanced accuracy across all 105 emotion pairs, providing a measure of how distinguishable the emotion concepts are in an agent’s appraisal space.
Figure 12 : Appraisal-dimension importance across all emotion pairs. For each sub-plot, representing a single agent, each column denotes the 105 possible emotion pairs, and is left unlabelled for brevity. Discriminators fitted on LLMs’ behavioral representations places most of its importance on fewer dimensions for any given pair, compared to humans. Table 12 reports the same contrast numerically, as the concentration of the importance profile.
Figure 13 : Fitted discriminators for guilt against shame , for: humans, GPT-5.6 Luna and Muse Glimmer 30B, each labelled inside the respective sub-plot. A split node shows the appraisal dimension it tests and the threshold it tests at. A leaf reports the majority true class of the training rows routed to it, that class’s share of the leaf, and the corresponding number of rows. A leaf that no training row reaches is shown in grey. The trees for visualization are fit on all rows across the entire data comparing these two emotions, and not cross-validated. The percentages here thus describe the data fit rather than accuracy on held-out data.
Agent
SynthTree acc.
STFI concentration
ρ (STFI, explicit)
ρ (STFI, implicit)
Humans
0.76
0.55
0.04
0.58
GPT o4-mini
0.96
0.84
0.40
0.63
Gemini 2.5 Flash
0.96
0.83
0.43
0.67
DeepSeek R1
0.97
0.86
0.42
0.58
Phi-4
0.96
0.84
0.28
0.59
QwQ 32B
0.94
0.81
0.29
0.68
Appendix
Table 12: Interpretable discriminators of emotion pairs, shown per agent. SynthTree acc.: mean balanced accuracy on held-out folds. STFI concentration: the Gini coefficient of the importance the fitted discriminator assigns to the 17 appraisal dimensions, averaged over pairs, where a higher value means the discriminator rests on fewer dimensions. ρ (STFI, explicit): the rank correlation between the importance profile and the agent’s own statement of which dimensions separate the pair. This is computed across the 90 psychologically-informed, subsampled emotion pairs. ρ (STFI, implicit): the rank correlation between that importance profile and the per-dimension discriminability of the ratings themselves.
Pair
Agent
Acc.
Bal. acc.
Macro F1
Top 1
Top 2
Top 3
anger vs boredom
Humans
0.69
0.69
0.69
legit-cheat
attention
legit-fair
anger vs boredom
GPT o4-mini
0.97
0.97
0.97
attention
legit-cheat
pleasantness
anger vs boredom
Gemini 2.5 Flash
1.00
1.00
1.00
attention
consider
legit-cheat
anger vs boredom
DeepSeek R1
0.98
0.97
0.98
attention
pleasantness
ctrl-situ
anger vs boredom
Phi-4
1.00
1.00
1.00
attention
legit-fair
exert
anger vs boredom
QwQ 32B
0.94
0.94
0.94
legit-cheat
pleasantness
legit-fair
Appendix
Table 26
Between two models
Between two
Emotion
Default
Noise-matched
human groups
anger
1.95
3.62
3.97 [3.76, 4.19]
boredom
2.68
3.86
4.08 [3.82, 4.31]
challenge
1.94
3.40
3.95 [3.73, 4.15]
contempt
2.80
3.93
4.30 [4.06, 4.50]
disgust
3.19
4.39
4.59 [4.31, 4.85]
Appendix
Table 15: Distance between the emotion representations of two agents, with the model ratings given the same measurement error as the human ratings. Between two models: the mean W2 distance over the 28 model pairs. Between two human groups: the median over 300 random divisions of the participants into eight groups, with the 2.5 th and 97.5 th percentiles in brackets.
Between two models
Between two human groups
Emotion
Full clouds
Size-matched
All scenarios
Shared scenarios
anger
1.96
2.17
3.95
4.27
boredom
2.70
3.04
4.11
4.50
challenge
1.94
2.27
4.04
4.26
contempt
2.82
3.03
4.33
4.68
disgust
3.17
3.71
4.52
4.79
Appendix
Table 16: Distance between the emotion representations of two agents, under two controls for the number of scenarios in a cloud. Full clouds: the mean W2 distance over the 28 model pairs, as in Table 15 . Size-matched: each model’s cloud reduced at random to the size of the matched human group’s cloud. All scenarios: the median over 300 random divisions of the participants into eight groups, as in Table 15 . Shared scenarios: the same divisions, with each pair of groups restricted to the scenarios both groups cover. Model clouds hold 17.4 scenarios on average, human group clouds 6.6 , and the shared subset of two human groups 4.0 .
Correctly interpreted scenarios
Same size
All scenarios
Model
%
Between
Within
Ratio
Humans
Ratio
Ratio
Humans
Muse Glimmer 30B
57%
5.24
2.84
1.85
1.33
1.70
1.84
1.31
Gemini 3.7 Flash
57%
5.29
2.89
1.83
1.32
1.72
1.87
1.31
GPT-5.6 Luna
55%
5.29
2.93
1.81
1.27
1.71
1.87
1.31
DeepSeek R1
53%
5.13
3.22
1.60
1.25
1.48
1.57
1.31
Phi-4
49%
4.93
3.13
1.57
1.24
1.52
1.66
1.31
Appendix
Table 17: Separation of emotion concepts in the without-label condition, on the scenarios each model interpreted as the intended emotion. Between: the mean W2 distance between two emotion categories; Within: the mean spread inside one category; Ratio: their quotient. Humans: the human ratio computed on that model’s own subset. Same size: the mean ratio over twenty random subsets with the same number of scenarios per emotion, which isolates the effect of the smaller sample. All scenarios: the same quantities on every scenario the model rated without the label.
Dimensions
Emotions
Emotion × dimension
Trait
High
Low
High
Low
High
Low
Neuroticism
0.77
0.88
0.94
0.93
0.47
0.49
Conscientiousness
0.57
0.84
0.63
0.98
0.17
0.52
Agreeableness
0.43
0.85
0.55
0.97
0.15
0.53
Openness
0.11
0.83
0.16
0.96
0.04
0.47
Extraversion
0.49
0.59
0.56
0.78
0.20
0.26
Appendix
Table 18: Share of tests that show a significant change from the neutral condition, averaged over the eight models, for each trait at its high and low level. Dimensions: the 17 per-dimension tests, with emotions pooled. Emotions: the 15 whole-profile tests. Emotion × dimension: the 255 directional tests. All shares are after Benjamini-Hochberg correction within the level of each comparison.
Dimension
Share
δ
Dimension
Share
δ
Consideration
0.58
+0.19
Obstacles
0.32
+0.02
Understanding
0.58
+0.01
Enjoyment
0.32
+0.05
Attention
0.48
+0.22
Cheated
0.24
+0.03
Exertion
0.45
+0.12
Certainty
0.22
+0.02
Self-control
0.39
+0.01
Situational control
0.21
−0.03
Problem
0.38
+0.05
Other-responsibility
0.19
+0.01
Appendix
Table 19: Change by appraisal dimension, averaged over the eight models, the ten personality settings and the fifteen emotions. Share: the fraction of emotion-and-dimension tests that are significant. δ : the mean Cliff’s delta, written so that a positive value means the neutral rating is higher than the rating under a persona.
Distance to the human ratings
Agreement with theory
Trait
High
Low
High
Low
Agreeableness
+0.05
+0.81
−0.00
−0.18
Conscientiousness
+0.07
+0.47
−0.01
−0.08
Neuroticism
+0.44
+0.11
+0.01
−0.04
Openness
+0.09
+0.41
+0.01
−0.06
Extraversion
+0.15
−0.07
−0.01
−0.01
Appendix
Table 20: Change from the neutral condition under each personality setting, averaged over the eight models. Distance to the human ratings: the mean W2 distance between the model’s and the human profiles of an emotion, averaged over the fifteen emotions; the neutral value is 3.45 and a positive change means the model moved further from the human ratings. Agreement with theory: the correlation between the model’s appraisal-emotion associations and the meta-analytic reference values; the neutral value is 0.69 and a negative change means weaker agreement.
Faculty of Business and Commerce, Kansai University, Osaka, Japan · Faculty of Business Data Science, Kansai University, Osaka, Japan · RIKEN Center for Advanced Intelligence Project, Tokyo, Japan +2