Affective computing has progressed from categorical emotion recognition to open-ended affective analysis with large multimodal models. Yet affective science describes emotion as an unfolding process shaped by appraisal, regulation, and social interpretation, which remains underexplored computationally. We propose TRACE, a cognition-oriented framework that formalizes an affective episode through three interrelated stages: Condition, Affect, and Effect, integrating observable cues with cognitive factors such as internal stance and regulation of emotional display. Based on this formulation, TRACE-Bench evaluates multimodal models in real-world social scenes through five tasks spanning grounded affect recognition, regulation decoding, cause reasoning, effect reasoning, and full-chain reconstruction, with 3,746 structured question-answer pairs over 646 videos. A matched human-model comparison reveals a substantial performance gap, while affect-specialized models also generally lag behind general-purpose MLLMs. Model outputs show recurring failures, including treating displayed behavior as genuine feeling and fabricating unsupported events during long-chain generation. We further propose TRACER, a cognition-grounded structured reasoning method that couples each inference with explicit premises from factual observations, cognitive appraisals, and established upstream conclusions, forming a traceable graph of intermediate and target conclusions. TRACER outperforms all evaluated model baselines on each of the five tasks. Project page: https://cogaffc.github.io/TRACE
Figures & tables
Fig. 1: Existing affective benchmarks often evaluate local targets such as affective states or causes. TRACE-Bench evaluates the full cognitive-affective chain: from the conditions that give rise to an emotion, through the regulation that reshapes its expression, to the consequences it produces for the subject or other participants.
Fig. 2: The Affective Blueprint . A subject-centered affective episode is represented through three causally connected stages: Condition , Affect , and Effect . Each stage is decomposed into several components, each represented by a set of fields.
Fig. 3: TRACE-Bench construction pipeline and dataset statistics, with the three-stage construction workflow on the left.
Fig. 4: Overview of TRACER . Factual observations are interpreted through character-centered appraisal and used in Blueprint-guided derivation. (a) Gray arrows denote intrinsic relations, while black arrows specify inference dependencies and premise requirements. Blue, green, and purple construct nodes denote known, intermediate, and queried constructs, respectively. (b) An abridged shoe-store example illustrates how observations, appraisals, and intermediate conclusions jointly support subsequent inferences.
Fig. 5: Model profiles by (a) task component and (b) family. Human scores average two participants.
Fig. 6: Evidence ablation. Q includes the question and supplied fields. Frames denotes 8 frames for GPT-5 and 60 for Qwen. Both (16f): GPT-5 with subtitles and 16 frames.
Fig. 7: T1 confusions from true (left) to predicted (right) labels. Band widths show proportions within each true class.
EmoQ*
OV-70B
Q3-VL
GPT5-N
Gem-3
GPT5-T
TRACER
Affective state
36.8
57.3
60.1
59.7
63.3
59.2
64.9
Category
27.2
40.1
38.8
36.1
48.1
41.5
40.5
Description
11.7
36.8
48.1
55.6
53.7
57.3
60.0
Polarity
57.5
80.1
81.4
69.1
81.4
71.3
77.5
Intensity
35.4
61.1
64.2
64.5
58.2
63.3
66.5
Object
51.8
67.7
67.8
73.0
74.7
62.5
79.9
TABLE II: T1 field-level performance. Shaded rows report construct scores aggregated within each instance.
Table 9
Fig. 8: T2 tactic confusion for GPT-5 (thinking) and TRACER . Rows: reference; columns: prediction; values: row percentages. Fabrication is omitted from both axes.
Figure 11
Cond.
State
Manif.
Reg.
Effect
Emotion-Qwen *
14.09
45.69
23.20
12.09
11.18
LLaVA-OV-70B
33.49
64.10
39.28
32.08
20.39
Qwen3-VL-32B
51.59
69.61
55.83
40.64
38.21
GPT-5 (non-thinking)
66.91
70.34
62.01
42.55
43.80
Gemini-3-Pro
63.67
73.22
58.72
21.46
43.25
GPT-5 (thinking)
70.29
70.54
63.14
41.09
47.74
TABLE V: T5 full-chain reconstruction scores by construct.
Fig. 11: Case studies for T1–T4, with T1 and T2 in the top row and T3 and T4 in the bottom row. Panels show task inputs, selected observations and appraisal readings, construct outputs, and reference and baseline comparisons. Text is condensed from the source records. The T1 example comes from a separate illustrative run. The layouts follow the corresponding task-specific reasoning topologies.
Figure 14
Think
Frames
Subs
Obs
App
CoT
RT
PDD
T1
T2
T3
T4
T5
Avg.
GPT-5 (non-thinking)
∘
⚫
⚫
∘
∘
∘
∘
∘
56.92
39.53
70.84
55.00
57.62
55.98
GPT-5 (thinking)
⚫
⚫
⚫
∘
∘
∘
∘
∘
56.84
43.49
75.39
57.73
60.89
58.87
w/ Appraisal-based CoT
⚫
⚫
⚫
∘
∘
⚫
∘
∘
55.62
47.39
80.01
58.86
61.15
60.61
w/ Observations
⚫
∘
∘
⚫
∘
∘
∘
∘
50.63
46.84
77.52
47.07
59.02
56.22
w/ Observations & Appraisals
⚫
∘
∘
⚫
⚫
∘
∘
∘
52.56
48.37
79.73
54.20
62.83
59.54
w/ RT
⚫
∘
∘
⚫
⚫
∘
⚫
∘
53.77
51.12
80.38
55.45
61.42
60.43
TABLE VI: Component ablation. Filled circles indicate enabled components (open Think: minimal reasoning). RT and PDD denote task-conditioned reasoning topology and premise-declared derivation, respectively. T2 uses the revised regulation task, as in Table 5 .
Appendix figures & tables19 assets
Supplementary material from the paper’s appendix.
Appendix
Affective State
Manifestation
Model
Cat.
Desc.
Pol.
Int.
Obj.
All
Facial
Body
Verbal
Vocal
All
Affect-specialized
Emotion-LLaMA
1.4
0.1
1.1
1.4
0.0
0.8
0.1
0.1
0.1
0.1
0.1
AffectGPT
21.8
23.5
71.5
50.2
56.4
44.8
27.7
22.8
15.6
31.8
24.4
Emotion-Qwen
27.2 *
11.7
57.5
35.4
51.8
36.8
30.0
20.4
16.0
23.3
22.1
Open-source general-purpose
Appendix
TABLE S1: T1 affect recognition. State and Manifestation totals are computed within each item.
T2 Regulation
T3 Cause
T4 Effect
Model
Tactic
Target
Goal
Evidence
External
Internal
Mental
Affective
Physical
Affect-specialized
Emotion-LLaMA
10.2 *
5.6
4.3
1.5
2.5 †
0.2
0.0 †
0.6
0.2
AffectGPT
7.0 †
1.9
6.2
1.6
19.6 *
10.8
12.4 *
22.1
12.6
Emotion-Qwen
36.7 *
28.0
13.5
10.0
24.4 *
9.1
2.1
19.6
15.5
Open-source general-purpose
Appendix
TABLE S2: Field-level results for T2 regulation decoding, T3 cause reasoning, and T4 effect reasoning.
Configuration
Calls
Input
Output
Hidden reasoning
GPT-5 (thinking)
1
1.7k
1.25k
1.06k
GPT-5 (thinking) + Obs
1
2.2k
1.10k
0.90k
GPT-5 (thinking) + Obs + App
1
2.9k
1.13k
0.93k
TRACER
1.4
17.3k
0.71k
0
Appendix
TABLE S3: Answer-generation calls and token usage, averaged over five tasks.
Fig. S1: Observation extraction. (a) Abridged observer prompt with simplified input labels; […] marks omissions and braces denote inserted inputs. (b) Output format. (c) One observation of a record.
Fig. S2: Core appraisal instructions. Wording is retained from the recorded prompt, with formatting adjusted.
Model
Cond.
State
Manif.
Reg.
Effect
Full
Affect-specialized
Emotion-LLaMA
0.00 †
1.38
0.00
0.00
0.00
0.20
AffectGPT
31.36 *
47.28
0.34
1.32
9.89
20.87
Emotion-Qwen
14.09 *
45.69
23.20
12.09
11.18
19.13
Open-source general-purpose
LLaVA-OneVision-7B
1.95 †
15.86
6.65
3.17
3.99
5.47
Appendix
TABLE S4: T5 full-chain reconstruction. Full is the overall T5 score.
Fig. S3: Rubric of the unsupported-content audit. Slots s1–s6 are the audited fields. Labels: S supported, L misplaced, F fabricated, U unverifiable. The off-topic denominator refers to the answer-level off-topic rate, not the statement-level fabrication rate.
Fig. S4: Prompts of the surface-reading diagnostic. (a) Description prompt. (b) Block appended in the appraisal-records setting. (c) Instruction prepended in the appraisal-CoT setting.
Configuration
Cond.
Affect
Effect
All
Unver.
Mispl.
GPT-5 (thinking)
3.1
12.2
17.3
11.0
12.6
9.5
+ Observations
1.8
7.7
7.1
6.3
12.6
6.5
+ RT
1.6
7.4
3.7
5.5
13.0
5.0
+ RT + PDD ‡
2.4
7.0
4.0
5.0
13.5
3.6
TRACER
3.0
6.2
3.3
4.7
12.9
3.9
Appendix
TABLE S5: T5 fabrication rates by stage (%, lower is better).
Fig. S5: Appraisal intervention on a regulated moment. Top: frames and dialogue of the annotated episode, with the focal utterance in bold. Bottom: the reference annotation and excerpts of the model descriptions without and with appraisal records. Without the records, both models take the laughter at face value. With them, both identify the annotated State.
Table 26
Fig. S6: Human–judge calibration on 100 free-text fields. Points show mean human ratings, with group sizes below the axis. The dashed line indicates equal scores.
Model / Human
T1
T2
T3
T4
T5
Avg.
Human
83.60
72.32
85.55
78.90
75.40
79.15
GPT-5 (thinking)
57.97
41.87
75.80
57.61
60.48
58.75
GPT-5 (non-thinking)
55.21
36.00
68.12
61.06
59.09
55.89
Gemini-3-Pro
55.59
41.53
68.00
60.19
54.18
55.90
Qwen3-VL-32B
50.56
34.26
63.66
49.00
50.68
49.63
Qwen3-Omni-30B
49.32
29.37
54.32
42.68
41.55
43.45
Appendix
TABLE S8: Human and model performance on the same 50 questions per task. Human scores average two participants; Avg. averages the five tasks.
Score
Criterion
0.00
No meaningful match, an opposite claim, or the wrong character or episode.
0.25
Only a broad or generic relation to the reference.
0.50
Partial match, with an important detail or role missing or incorrect.
0.75
Correct central meaning, with a secondary omission or harmless addition.
1.00
Correct content, characters, and episode, regardless of wording.
Appendix
TABLE S9: Semantic adequacy d1 : higher is better.
Score
Criterion
0.00
Focused answer with almost no unnecessary content.
0.25
Minor repetition, filler, or unnecessary hedging.
0.50
About half the answer is uninformative or lists competing guesses.
0.75
Most of the answer is repetition, generic content, or competing guesses.
1.00
Almost entirely redundant or noncommittal.
Appendix
TABLE S10: Redundancy d2 : higher is worse.
Model
Task
Pars. (%)
All
Pars. only
Emotion-LLaMA *
T2
61
5.22
8.61
Emotion-LLaMA †
T3
37
2.26
6.11
Emotion-LLaMA †
T4
49.5
0.31
0.63
Emotion-LLaMA †
T5
7
0.20
2.81
AffectGPT †
T2
17
4.64
27.38
AffectGPT *
T3
79
17.99
22.84
Appendix
TABLE S11: Output parsing rates below 90%. Bold values are reported in Table 2 of the main paper.
Fig. S7: Field-level rewrite rates after human review of retained MLLM-proposed chains.
Fig. S8: Effect boundaries. (a) Consecutive actions belong to one Physical Effect. (b) Barrett’s later surprise follows the dismissal, not Mel’s preceding Affect. Checks mark supported links, and the cross marks the incorrect attribution.
Fig. S9: Agreement between two independent re-annotations. Closed fields: Cohen’s κ (linear weights for Intensity). Open fields: proportion of pairs judged to state the same content. Only κ corrects for chance agreement; the two measures are not directly comparable.
Recent multimodal large language models (MLLMs) increasingly incorporate explainable reasoning for emotion understanding. However, reasoning based mainly on observable affective cues can reduce emotion understanding to superficial cue-label associations, giving rise to the Clever Hans effect. Such shortcuts become unreliable when affective cues are implicit, conflicting across modalities, linguistically misleading, or obscured by redundant details. In contrast, human emotions are shaped by how individuals interpret and evaluate surrounding events beyond observable cues. Inspired by appraisal theories of emotion, we formulate multimodal emotion understanding as a progression from perception to cognitive appraisal, and introduce a dataset, a model, and a benchmark to support this novel paradigm. CogEmo-40K is a large-scale instruction-tuning dataset constructed through a perception-to-appraisal pipeline to elicit evidence-grounded reasoning across six cognitive appraisal dimensions underlying emotion. CogEmo-MoE is a compact sparse MLLM that introduces interleaved MoE blocks for appraisal-specific adaptation, enabling effective appraisal reasoning at a substantially smaller scale than typical emotion MLLMs. CogEmo-Bench introduces an Appraisal Evidence Quality Score (AEQS) to assess cognitive-affective understanding across six complementary appraisal dimensions, addressing the limitation of conventional emotion metrics that evaluate what emotion is predicted but not why it arises. Extensive experiments show that our paradigm not only leads CogEmo-Bench, but also exhibits strong cross-domain generalization. Our findings suggest that perception-to-appraisal reasoning can move beyond surface-level cue-label associations toward more reliable multimodal emotion understanding and closer cognitive alignment between MLLMs and humans.
Jia Li, Yichao He, Yangchen Yu +6
Hefei University of Technology · Nanyang Technological University
Humans infer emotions by integrating observed multimodal cues with expectations about how affective states may unfold. Existing multimodal large language models (MLLMs), however, often treat emotion recognition as static fusion over complete audiovisual-text inputs, leaving affective dynamics implicit. We propose AffectVerse, a Qwen2.5-Omni-based model equipped with an Emotion World Module (EWM), an action-free representation-level module for short-horizon latent affective prediction. \rev{EWM contains three modules: 1) Cross-Modal Temporal Imagination predicts future video/audio representations from past tokens with multi-step rollout. 2) MAMA(Modality-Aware Multi-step Attention) Belief Aggregation compresses imagined tokens into modality-aware belief tokens. 3) Belief Injection inserts these belief tokens into the LLM for affective reasoning.} AffectVerse uses future prediction as a past-conditioned self-supervised signal: it does not replace modeling observed history or require unseen signals at inference, but forces the current belief state to encode transition cues that are predictive of subsequent affective change. Across nine benchmarks, AffectVerse improves at least 2.57% over other models, while controlled ablations show additive gains from temporal imagination, cross-modal rollout, and belief aggregation. These results suggest predictive belief-state modeling is a practical alternative for affective computing.
Bo Zhao, Fanghua Ye, Yixin Ji +3
1Great Bay University · 2Tencent · 3Tsinghua University +1
Emotion understanding is a core capability for LLMs to interact effectively with humans, yet existing evaluation paradigms rely on discrete emotion label prediction and fail to capture the cognitive processes underlying emotion generation. Grounded in appraisal theory, we introduce CAREBench, the first benchmark with complete inferential chain annotations from both first- and third-person perspectives on real-world narratives, spanning appraisal reasoning, appraisal ratings, and multi-label emotion annotation. We propose a process-level evaluation framework and conduct systematic experiments across six LLMs organized around four research questions. We find that stronger models match or surpass human observers on certain tasks, yet fall short on appraisal reasoning and positive emotion recognition; performance across chain steps and sensitivity to appraisal interventions exhibit dissociations across models; and current models have not internalized the mechanisms needed to capture human subjective heterogeneity. These findings suggest that downstream emotion prediction metrics may overestimate LLMs' true emotion understanding, and CAREBench provides a foundation for more diagnostically informative evaluation of LLMs' affective cognitive capabilities.
Zhaoyue Sun, Hainiu Xu, Andero Uusberg +3
Department of Informatics King’s College London · Institute of Psychology University of Tartu · Department of Psychology Stanford University +1