cs.CVSep 28, 2026

Do Emotion Concepts Generalize Across Sources, Modalities, and Architectures in Vision-Language Models?

Authors: Bohao Xing, Xin Liu, Kaishen Yuan, Deng Li, Rong Gao, Guoying Zhao, Xiaolan Fu, Heikki Kälviäinen

Organizations: Lappeenranta-Lahti University of Technology LUT · Shanghai Jiao Tong University · The Hong Kong University of Science and Technology (Guangzhou) · University of Oulu · ELLIS Institute Finland · Brno University of Technology

Abstract

Recent studies suggest that large language models encode emotion concepts as structured internal representations, but most existing work focuses on text and a single architecture. Therefore, we ask, do emotion concepts generalize across sources, modalities, and architectures in vision--language models (VLMs)? To address this, we construct CMES (Cross-Modal Emotion Stimuli), a multi-source collection of emotion-conditioned stories, real facial expressions, synthetic portraits, and synthetic emotion-evoking scenes. For each stimulus source, we extract a separate set of six Ekman emotion vectors from each of three VLMs. We report four main findings as follows: 1) Image-derived emotion vectors form a low-dimensional geometry similar to that of text-derived vectors. Valence is relatively stable across sources, while arousal varies more. 2) Text- and image-derived emotion vectors have modest cosine similarity but still show held-out cross-modal correspondence. Text-derived vectors can also steer image interpretation. 3) Cross-architecture correspondence remains even when native cosine is near zero. Transformations estimated from generic ImageNet activations recover both correspondence and causal transfer without using the six emotion vectors or their labels. 4) After aligning representations across architectures, we construct a shared emotion subspace that preserves affective geometry and selective steering effects. The corresponding consensus emotion vectors also generalize to a held-out fourth architecture at two model sizes. These results suggest that emotion representations can share relational structure and causal effects across sources, modalities, and architectures, even when individual vector directions differ.

Figures & tables

Appendix figures & tables48 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Where Do Models Find Happiness? Emotion Vectors in Open-Source LLMs

    Jun 25, 2026Sinie van der Ben, Raphaël Baur, Yannick Metz +1Valence-Arousal EstimationEmotion

  2. Why Do Vision Language Models Struggle To Recognize Human Emotions?

    Apr 16, 2026Madhav Agarwal, Sotirios A. Tsaftaris, Laura Sevilla-Lara +1Vision-Language Foundation ModelsEmotion

  3. Interpreting and Enhancing Emotional Circuits in Large Vision-Language Models via Cross-Modal Information Flow

    May 21, 2026Chengsheng Zhang, Chenghao Sun, Zhining Xie +1Emotional RegulationCross-Modal