Do Emotion Concepts Generalize Across Sources, Modalities, and Architectures in Vision-Language Models?
Organizations: Lappeenranta-Lahti University of Technology LUT · Shanghai Jiao Tong University · The Hong Kong University of Science and Technology (Guangzhou) · University of Oulu · ELLIS Institute Finland · Brno University of Technology
Abstract
Recent studies suggest that large language models encode emotion concepts as structured internal representations, but most existing work focuses on text and a single architecture. Therefore, we ask, do emotion concepts generalize across sources, modalities, and architectures in vision--language models (VLMs)? To address this, we construct CMES (Cross-Modal Emotion Stimuli), a multi-source collection of emotion-conditioned stories, real facial expressions, synthetic portraits, and synthetic emotion-evoking scenes. For each stimulus source, we extract a separate set of six Ekman emotion vectors from each of three VLMs. We report four main findings as follows: 1) Image-derived emotion vectors form a low-dimensional geometry similar to that of text-derived vectors. Valence is relatively stable across sources, while arousal varies more. 2) Text- and image-derived emotion vectors have modest cosine similarity but still show held-out cross-modal correspondence. Text-derived vectors can also steer image interpretation. 3) Cross-architecture correspondence remains even when native cosine is near zero. Transformations estimated from generic ImageNet activations recover both correspondence and causal transfer without using the six emotion vectors or their labels. 4) After aligning representations across architectures, we construct a shared emotion subspace that preserves affective geometry and selective steering effects. The corresponding consensus emotion vectors also generalize to a held-out fourth architecture at two model sizes. These results suggest that emotion representations can share relational structure and causal effects across sources, modalities, and architectures, even when individual vector directions differ.
Figures & tables
Appendix figures & tables48 assets
Supplementary material from the paper’s appendix.
Appendix
| Label | Source | Stimulus type | Per class | Total |
| Text sources | ||||
| T1 | Gemini 3.1 Pro Preview | Generated text | ||
| T2 | Qwen3-Max stories | Generated text | ||
| T3 | Qwen3-VL stories | Generated text | ||
| Image sources | ||||
| AffectNet | Real faces | |||
| Emotion | Retrieval P@50 (%) | Classification Recall (%) | ||||
|---|---|---|---|---|---|---|
| Qwen3-VL | InternVL3.5 | Ministral-3 | Qwen3-VL | InternVL3.5 | Ministral-3 | |
| anger | 58 | 60 | 38 | 41 | 52 | 29 |
| disgust | 80 | 66 | 68 | 34 | 37 | 30 |
| fear | 72 | 70 | 72 | 52 | 49 | 45 |
| happiness | 92 | 84 | 88 | 93 | 89 | 93 |
| sadness | 82 | 88 | 90 | 54 | 47 | 50 |
| Emotion | Pool share (%) | Retrieval P@50 (%) | Classification Recall (%) | ||||
|---|---|---|---|---|---|---|---|
| Qwen3-VL | InternVL3.5 | Ministral-3 | Qwen3-VL | InternVL3.5 | Ministral-3 | ||
| anger | 5.3 | 72 | 66 | 64 | 69 | 51 | 54 |
| disgust | 5.2 | 8 | 16 | 6 | 60 | 48 | 53 |
| fear | 2.4 | 28 | 64 | 48 | 62 | 62 | 64 |
| happiness | 38.6 | 100 | 94 | 98 | 60 | 63 | 65 |
| sadness | 15.6 | 86 | 80 | 96 | 66 | 61 | 61 |
| Emotion vector | Model | Disgust | Fear | Anger | Happiness | Neutral | Surprise | Sadness |
| AffectNet validation pool | ||||||||
| disgust | Qwen3-VL | 80 | 2 | 18 | – | – | – | – |
| InternVL3.5 | 66 | 4 | 26 | – | – | – | 4 | |
| Ministral-3 | 68 | 6 | 18 | – | – | – | 8 | |
| fear | Qwen3-VL | 2 | 72 | 8 | – | – | 18 | – |
| InternVL3.5 | 2 | 70 | 6 | – | – | 22 | – | |
| Stimulus source | PC1 | PC2 | PC3 | PC4 | PC5 | PC1 PC2 |
|---|---|---|---|---|---|---|
| Qwen3-VL | ||||||
| T1 | 30.7 | 25.2 | 19.2 | 13.9 | 11.1 | 55.9 |
| T2 | 38.2 | 27.6 | 20.3 | 8.4 | 5.5 | 65.8 |
| T3 | 40.2 | 25.9 | 19.9 | 8.0 | 6.1 | 66.0 |
| AffectNet | 34.0 | 29.8 | 18.8 | 11.6 | 5.8 | 63.8 |
| AffectNet (filtered) | 32.8 | 23.0 | 21.4 | 15.2 | 7.5 | 55.9 |
| Stimulus source | Qwen3-VL | InternVL3.5 | Ministral-3 | Mean | Spread |
|---|---|---|---|---|---|
| Valence agreement | |||||
| RAF-DB | |||||
| T3 | |||||
| Synth-S | |||||
| T1 | |||||
| AffectNet | |||||
| Arousal (PC2) | Valence (PC1) | ||||
|---|---|---|---|---|---|
| Source | Model | Unfiltered | Filtered | Unfiltered | Filtered |
| AffectNet | Qwen3-VL | ||||
| InternVL3.5 | |||||
| Ministral-3 | |||||
| RAF-DB | Qwen3-VL | ||||
| InternVL3.5 | |||||
| R&M | NRC-VAD | vs. T2 text | ||||
| AffectNet subset | Valence | Arousal | Valence | Arousal | CKA | RSA |
| Qwen3-VL | ||||||
| Full set | 0.85 | 0.23 | 0.83 | 0.20 | 0.93 | 0.61 |
| Random subset | ||||||
| Success-filtered | 0.97 | 0.41 | 0.99 | 0.28 | 0.95 | 0.86 |
| InternVL3.5 | ||||||
| Mean cosine | Permutation | ||||
| Model | Source pair | Before alignment | Full-fit | LOEO | |
| Text sources | |||||
| Qwen3-VL | T1–T2 | ||||
| T1–T3 | |||||
| T2–T3 | |||||
| InternVL3.5 | T1–T2 | ||||
| Mean cosine | Permutation | |||
| Text image source | Before alignment | Full-fit | LOEO | |
| Qwen3-VL | ||||
| T1 AffectNet | ||||
| T1 RAF-DB | ||||
| T1 Synth-P | ||||
| T1 Synth-S | ||||
| Mean cosine | Permutation | |||
| Source / control | Before alignment | Full-fit | LOEO | |
| T1 | ||||
| T2 | ||||
| T3 | ||||
| AffectNet | ||||
| RAF-DB | ||||
| Stimulus source | Anger | Disgust | Fear | Happiness | Sadness | Surprise |
|---|---|---|---|---|---|---|
| T1 | ||||||
| T2 | ||||||
| T3 | ||||||
| AffectNet | ||||||
| RAF-DB | ||||||
| Synth-P |
| Mean cosine | ||||
| Initial condition | Direct | After LOEO | Held-out range | Permutation |
| Image vectors (AffectNet) | ||||
| No mapping | – | |||
| Neutral-face map | – | |||
| ImageNet map | – | |||
| Random orthogonal map | – | |||
| Injected vector | diag | off | diag | off | diag | off | diag | off |
|---|---|---|---|---|---|---|---|---|
| RAF-DB | ||||||||
| Image vector | ||||||||
| Random | ||||||||
| Orthogonal | ||||||||
| AffectNet | ||||||||
| RAF-DB | AffectNet | |||||||
| Injected vector | diag | off | diag | off | diag | off | diag | off |
| Qwen3-VL | ||||||||
| Image vector | ||||||||
| Random | ||||||||
| Orthogonal | ||||||||
| Evaluation set | diag | off | diag | off | diag | off | diag | off |
|---|---|---|---|---|---|---|---|---|
| RAF-DB | ||||||||
| AffectNet | ||||||||
| RAF-DB | AffectNet | |||||||
| Injected vector | diag | off | diag | off | diag | off | diag | off |
| Qwen3-VL | ||||||||
| Text vector (T3) | ||||||||
| Image vector | ||||||||
| InternVL3.5 | ||||||||
| Vector source | Mapping / control | Cosine | diag | off | diag | off |
|---|---|---|---|---|---|---|
| Ministral-3 Qwen3-VL | ||||||
| Image | No mapping | |||||
| Random orthogonal map | ||||||
| Target-orthogonal directions | ||||||
| Neutral-face map | ||||||
| Emotion readout | Gender readout | |||
|---|---|---|---|---|
| Injected emotion | Flip rate | Accuracy | Flip rate | Accuracy |
| No intervention | — | — | ||
| Anger | ||||
| Disgust | ||||
| Fear | ||||
| Happiness | ||||
| RAF-DB | AffectNet | |||
|---|---|---|---|---|
| Injection positions | diag | off | diag | off |
| Image tokens (default) | ||||
| Question tokens | ||||
| Answer prefix | ||||
| All positions | ||||
| RAF-DB | AffectNet | |||
|---|---|---|---|---|
| Removed direction or span | ||||
| No removal ( ) | ||||
| Single emotion vector (rank 1) | ||||
| Random direction | ||||
| Centered emotion-vector span (rank 5) | ||||
| Random subspace | ||||
| RAF-DB | AffectNet | |||
|---|---|---|---|---|
| Injected emotion | ||||
| Anger | ||||
| Disgust | ||||
| Fear | ||||
| Happiness | ||||
| Sadness | ||||
| RAF-DB | AffectNet | |||||||
|---|---|---|---|---|---|---|---|---|
| Emotion attributed to | diag | off | diag | off | diag | off | diag | off |
| the depicted person | ||||||||
| the assistant | ||||||||
| a third party | ||||||||
| Layer | Relative depth | diag | off | Selectivity (diag off) |
| Qwen3-VL | ||||
| 9 | ||||
| 14 | ||||
| 18 | ||||
| 24 | ||||
| 31 | ||||
| Model / alignment | Emotion | Modality | Model | Residual / interactions |
| Within each model | ||||
| Qwen3-VL | — | |||
| InternVL3.5 | — | |||
| Ministral-3 | — | |||
| Across all three aligned models | ||||
| Neutral-face alignment | ||||
| Mean | ||
| Injected vector | diag | off |
| Equal-norm directions | ||
| Original vector | ||
| Shared projection | ||
| Remainder | ||
| Random outside span | ||
| Source model Gemma-3 | Gemma-3 target model | |||||
| Source / control | Qwen3-VL | Ministral-3 | InternVL3.5 | Qwen3-VL | Ministral-3 | InternVL3.5 |
| Gemma-3-12B | ||||||
| T1 | ||||||
| T2 | ||||||
| T3 | ||||||
| AffectNet | ||||||
| Injected vector | Cosine | diag | off | diag | off |
|---|---|---|---|---|---|
| Gemma-3-12B | |||||
| Original Gemma-3 vector | |||||
| Neutral-face alignment | |||||
| Consensus vector | |||||
| Qwen3-VL vector | |||||