Affective computing has progressed from categorical emotion recognition to open-ended affective analysis with large multimodal models. Yet affective science describes emotion as an unfolding process shaped by appraisal, regulation, and social interpretation, which remains underexplored computationally. We propose TRACE, a cognition-oriented framework that formalizes an affective episode through three interrelated stages: Condition, Affect, and Effect, integrating observable cues with cognitive factors such as internal stance and regulation of emotional display. Based on this formulation, TRACE-Bench evaluates multimodal models in real-world social scenes through five tasks spanning grounded affect recognition, regulation decoding, cause reasoning, effect reasoning, and full-chain reconstruction, with 3,746 structured question-answer pairs over 646 videos. A matched human-model comparison reveals a substantial performance gap, while affect-specialized models also generally lag behind general-purpose MLLMs. Model outputs show recurring failures, including treating displayed behavior as genuine feeling and fabricating unsupported events during long-chain generation. We further propose TRACER, a cognition-grounded structured reasoning method that couples each inference with explicit premises from factual observations, cognitive appraisals, and established upstream conclusions, forming a traceable graph of intermediate and target conclusions. TRACER outperforms all evaluated model baselines on each of the five tasks. Project page: https://cogaffc.github.io/TRACE
Figures & tables
Fig. 1: Existing affective benchmarks often evaluate local targets such as affective states or causes. TRACE-Bench evaluates the full cognitive-affective chain: from the conditions that give rise to an emotion, through the regulation that reshapes its expression, to the consequences it produces for the subject or other participants.
Fig. 2: The Affective Blueprint . A subject-centered affective episode is represented through three causally connected stages: Condition , Affect , and Effect . Each stage is decomposed into several components, each represented by a set of fields.
Fig. 3: TRACE-Bench construction pipeline and dataset statistics, with the three-stage construction workflow on the left.
Fig. 4: Overview of TRACER . Factual observations are interpreted through character-centered appraisal and used in Blueprint-guided derivation. (a) Gray arrows denote intrinsic relations, while black arrows specify inference dependencies and premise requirements. Blue, green, and purple construct nodes denote known, intermediate, and queried constructs, respectively. (b) An abridged shoe-store example illustrates how observations, appraisals, and intermediate conclusions jointly support subsequent inferences.
Fig. 5: Model profiles by (a) task component and (b) family. Human scores average two participants.
Fig. 6: Evidence ablation. Q includes the question and supplied fields. Frames denotes 8 frames for GPT-5 and 60 for Qwen. Both (16f): GPT-5 with subtitles and 16 frames.
Fig. 7: T1 confusions from true (left) to predicted (right) labels. Band widths show proportions within each true class.
EmoQ*
OV-70B
Q3-VL
GPT5-N
Gem-3
GPT5-T
TRACER
Affective state
36.8
57.3
60.1
59.7
63.3
59.2
64.9
Category
27.2
40.1
38.8
36.1
48.1
41.5
40.5
Description
11.7
36.8
48.1
55.6
53.7
57.3
60.0
Polarity
57.5
80.1
81.4
69.1
81.4
71.3
77.5
Intensity
35.4
61.1
64.2
64.5
58.2
63.3
66.5
Object
51.8
67.7
67.8
73.0
74.7
62.5
79.9
TABLE II: T1 field-level performance. Shaded rows report construct scores aggregated within each instance.
Table 9
Fig. 8: T2 tactic confusion for GPT-5 (thinking) and TRACER . Rows: reference; columns: prediction; values: row percentages. Fabrication is omitted from both axes.
Figure 11
Cond.
State
Manif.
Reg.
Effect
Emotion-Qwen *
14.09
45.69
23.20
12.09
11.18
LLaVA-OV-70B
33.49
64.10
39.28
32.08
20.39
Qwen3-VL-32B
51.59
69.61
55.83
40.64
38.21
GPT-5 (non-thinking)
66.91
70.34
62.01
42.55
43.80
Gemini-3-Pro
63.67
73.22
58.72
21.46
43.25
GPT-5 (thinking)
70.29
70.54
63.14
41.09
47.74
TABLE V: T5 full-chain reconstruction scores by construct.
Fig. 11: Case studies for T1–T4, with T1 and T2 in the top row and T3 and T4 in the bottom row. Panels show task inputs, selected observations and appraisal readings, construct outputs, and reference and baseline comparisons. Text is condensed from the source records. The T1 example comes from a separate illustrative run. The layouts follow the corresponding task-specific reasoning topologies.
Figure 14
Think
Frames
Subs
Obs
App
CoT
RT
PDD
T1
T2
T3
T4
T5
Avg.
GPT-5 (non-thinking)
∘
⚫
⚫
∘
∘
∘
∘
∘
56.92
39.53
70.84
55.00
57.62
55.98
GPT-5 (thinking)
⚫
⚫
⚫
∘
∘
∘
∘
∘
56.84
43.49
75.39
57.73
60.89
58.87
w/ Appraisal-based CoT
⚫
⚫
⚫
∘
∘
⚫
∘
∘
55.62
47.39
80.01
58.86
61.15
60.61
w/ Observations
⚫
∘
∘
⚫
∘
∘
∘
∘
50.63
46.84
77.52
47.07
59.02
56.22
w/ Observations & Appraisals
⚫
∘
∘
⚫
⚫
∘
∘
∘
52.56
48.37
79.73
54.20
62.83
59.54
w/ RT
⚫
∘
∘
⚫
⚫
∘
⚫
∘
53.77
51.12
80.38
55.45
61.42
60.43
TABLE VI: Component ablation. Filled circles indicate enabled components (open Think: minimal reasoning). RT and PDD denote task-conditioned reasoning topology and premise-declared derivation, respectively. T2 uses the revised regulation task, as in Table 5 .
Appendix figures & tables19 assets
Supplementary material from the paper’s appendix.
Appendix
Affective State
Manifestation
Model
Cat.
Desc.
Pol.
Int.
Obj.
All
Facial
Body
Verbal
Vocal
All
Affect-specialized
Emotion-LLaMA
1.4
0.1
1.1
1.4
0.0
0.8
0.1
0.1
0.1
0.1
0.1
AffectGPT
21.8
23.5
71.5
50.2
56.4
44.8
27.7
22.8
15.6
31.8
24.4
Emotion-Qwen
27.2 *
11.7
57.5
35.4
51.8
36.8
30.0
20.4
16.0
23.3
22.1
Open-source general-purpose
Appendix
TABLE S1: T1 affect recognition. State and Manifestation totals are computed within each item.
T2 Regulation
T3 Cause
T4 Effect
Model
Tactic
Target
Goal
Evidence
External
Internal
Mental
Affective
Physical
Affect-specialized
Emotion-LLaMA
10.2 *
5.6
4.3
1.5
2.5 †
0.2
0.0 †
0.6
0.2
AffectGPT
7.0 †
1.9
6.2
1.6
19.6 *
10.8
12.4 *
22.1
12.6
Emotion-Qwen
36.7 *
28.0
13.5
10.0
24.4 *
9.1
2.1
19.6
15.5
Open-source general-purpose
Appendix
TABLE S2: Field-level results for T2 regulation decoding, T3 cause reasoning, and T4 effect reasoning.
Configuration
Calls
Input
Output
Hidden reasoning
GPT-5 (thinking)
1
1.7k
1.25k
1.06k
GPT-5 (thinking) + Obs
1
2.2k
1.10k
0.90k
GPT-5 (thinking) + Obs + App
1
2.9k
1.13k
0.93k
TRACER
1.4
17.3k
0.71k
0
Appendix
TABLE S3: Answer-generation calls and token usage, averaged over five tasks.
Fig. S1: Observation extraction. (a) Abridged observer prompt with simplified input labels; […] marks omissions and braces denote inserted inputs. (b) Output format. (c) One observation of a record.
Fig. S2: Core appraisal instructions. Wording is retained from the recorded prompt, with formatting adjusted.
Model
Cond.
State
Manif.
Reg.
Effect
Full
Affect-specialized
Emotion-LLaMA
0.00 †
1.38
0.00
0.00
0.00
0.20
AffectGPT
31.36 *
47.28
0.34
1.32
9.89
20.87
Emotion-Qwen
14.09 *
45.69
23.20
12.09
11.18
19.13
Open-source general-purpose
LLaVA-OneVision-7B
1.95 †
15.86
6.65
3.17
3.99
5.47
Appendix
TABLE S4: T5 full-chain reconstruction. Full is the overall T5 score.
Fig. S3: Rubric of the unsupported-content audit. Slots s1–s6 are the audited fields. Labels: S supported, L misplaced, F fabricated, U unverifiable. The off-topic denominator refers to the answer-level off-topic rate, not the statement-level fabrication rate.
Fig. S4: Prompts of the surface-reading diagnostic. (a) Description prompt. (b) Block appended in the appraisal-records setting. (c) Instruction prepended in the appraisal-CoT setting.
Configuration
Cond.
Affect
Effect
All
Unver.
Mispl.
GPT-5 (thinking)
3.1
12.2
17.3
11.0
12.6
9.5
+ Observations
1.8
7.7
7.1
6.3
12.6
6.5
+ RT
1.6
7.4
3.7
5.5
13.0
5.0
+ RT + PDD ‡
2.4
7.0
4.0
5.0
13.5
3.6
TRACER
3.0
6.2
3.3
4.7
12.9
3.9
Appendix
TABLE S5: T5 fabrication rates by stage (%, lower is better).
Fig. S5: Appraisal intervention on a regulated moment. Top: frames and dialogue of the annotated episode, with the focal utterance in bold. Bottom: the reference annotation and excerpts of the model descriptions without and with appraisal records. Without the records, both models take the laughter at face value. With them, both identify the annotated State.
Table 26
Fig. S6: Human–judge calibration on 100 free-text fields. Points show mean human ratings, with group sizes below the axis. The dashed line indicates equal scores.
Model / Human
T1
T2
T3
T4
T5
Avg.
Human
83.60
72.32
85.55
78.90
75.40
79.15
GPT-5 (thinking)
57.97
41.87
75.80
57.61
60.48
58.75
GPT-5 (non-thinking)
55.21
36.00
68.12
61.06
59.09
55.89
Gemini-3-Pro
55.59
41.53
68.00
60.19
54.18
55.90
Qwen3-VL-32B
50.56
34.26
63.66
49.00
50.68
49.63
Qwen3-Omni-30B
49.32
29.37
54.32
42.68
41.55
43.45
Appendix
TABLE S8: Human and model performance on the same 50 questions per task. Human scores average two participants; Avg. averages the five tasks.
Score
Criterion
0.00
No meaningful match, an opposite claim, or the wrong character or episode.
0.25
Only a broad or generic relation to the reference.
0.50
Partial match, with an important detail or role missing or incorrect.
0.75
Correct central meaning, with a secondary omission or harmless addition.
1.00
Correct content, characters, and episode, regardless of wording.
Appendix
TABLE S9: Semantic adequacy d1 : higher is better.
Score
Criterion
0.00
Focused answer with almost no unnecessary content.
0.25
Minor repetition, filler, or unnecessary hedging.
0.50
About half the answer is uninformative or lists competing guesses.
0.75
Most of the answer is repetition, generic content, or competing guesses.
1.00
Almost entirely redundant or noncommittal.
Appendix
TABLE S10: Redundancy d2 : higher is worse.
Model
Task
Pars. (%)
All
Pars. only
Emotion-LLaMA *
T2
61
5.22
8.61
Emotion-LLaMA †
T3
37
2.26
6.11
Emotion-LLaMA †
T4
49.5
0.31
0.63
Emotion-LLaMA †
T5
7
0.20
2.81
AffectGPT †
T2
17
4.64
27.38
AffectGPT *
T3
79
17.99
22.84
Appendix
TABLE S11: Output parsing rates below 90%. Bold values are reported in Table 2 of the main paper.
Fig. S7: Field-level rewrite rates after human review of retained MLLM-proposed chains.
Fig. S8: Effect boundaries. (a) Consecutive actions belong to one Physical Effect. (b) Barrett’s later surprise follows the dismissal, not Mel’s preceding Affect. Checks mark supported links, and the cross marks the incorrect attribution.
Fig. S9: Agreement between two independent re-annotations. Closed fields: Cohen’s κ (linear weights for Intensity). Open fields: proportion of pairs judged to state the same content. Only κ corrects for chance agreement; the two measures are not directly comparable.