cs.CLAug 7, 2025

Too Categorical to be Human: Emotion Concepts in LLMs and Humans

Authors: Sree Bhattacharyya, Evgenii Kuriabov, Lucas Craig, Tharun Dilliraj, Reginald B. Adams,, Jia Li, James Z. Wang

Organizations: Department of Informatics and Intelligent Systems, College of Information Sciences and Technology, The Pennsylvania State University · Department of Statistics, Eberly College of Science, The Pennsylvania State University · Department of Computer Science and Engineering, College of Engineering, The Pennsylvania State University · Department of Psychology, College of the Liberal Arts, The Pennsylvania State University

Abstract

Understanding human emotions is central to user-facing AI applications, safety alignment, and the simulation of human behavior. As emotional stimuli shape high-stakes behavior in Large Language Models (LLMs), there is increasing interest in how models represent emotion concepts internally. Mechanistic accounts of these representations, however, cannot be compared directly against humans: emotion processing in humans is highly distributed and yields no equivalent neural representation. To understand whether LLMs internalize emotion concepts in a way similar to humans, we propose characterizing the abstract concept of an emotion using external behavioral signatures, which we term behavioral representations. Using the theory of cognitive appraisals, which enables representing emotional situations along interpretable evaluative dimensions, we create a benchmark dataset of emotional scenarios spanning 15 emotion categories. We elicit behavioral representations of emotion concepts from LLMs and humans using our benchmark, and study their structural similarity. We find that LLMs represent emotion concepts more categorically, homogeneously, and determinately than humans, representing a single emotion concept with less internal diversity, and place different emotions further apart. The categorical structure of representations in LLMs is further robust to contextual variation, including with different task framing and demographic personas. Analyzing model checkpoints across different training stages, we also find that the discretized nature of representations appears after the mid-training stage itself and is unaffected by different post-training strategies. Through our results, we highlight a key difference in how LLMs behaviorally represent emotion concepts, curbing the subjectivity inherent to the human experience of emotions.

Figures & tables

Appendix figures & tables26 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. CAREBench: Evaluating LLMs' Emotion Understanding by Assessing Cognitive Appraisal Reasoning

    May 16, 2026Zhaoyue Sun, Hainiu Xu, Andero Uusberg +3Large Language Model EvaluationCognitive Science

  2. Do Large Language Models Have Emotions?

    Jun 3, 2026Amit Goldenberg, James J. GrossEmotionNeuroscience

  3. LLMs Capture Emotion Labels, Not Emotion Uncertainty: Distributional Analysis and Calibration of Human-LLM Judgment Gaps

    Apr 30, 2026Keito Inoshita, Xiaokang Zhou, Akira Kawai +1Large Language Model AnnotationsEmotion Recognition