VISTA: Value-Informed Event Appraisal for Multimodal Emotion Conflict
Organizations: State Key Laboratory of General Artificial Intelligence, School of Intelligence Science and Technology, Peking University · College of Artificial Intelligence, Tsinghua University · China United Network Communications Group Co., Ltd.
Abstract
Conflicting emotional cues can be individually valid: a subdued voice may reflect a blocked goal while a smile satisfies a social obligation. Their interpretation depends on what the event means to the person. We introduce VISTA (Value-Informed Semantic Trust Arbitration), a learned seven-field appraisal interface that conditions modality arbitration on concerns, event relations, and expression conditions while retaining a joint-evidence residual. A log-odds decomposition separates emotion expectation from cue diagnosticity, motivating an interface that lets appraisal change how evidence is interpreted. With a shared Qwen2.5-Omni-7B backbone and matched training examples and steps, VISTA reaches 64.5% conflict accuracy on CA-MER, improving on modality gating by 2.5 percentage points on conflict and 0.2 on consistency. Shuffling appraisal across scenes or removing its decision connection reduces this benefit. A common frozen-backbone probe reaches 0.600 macro CCC for appraisal readout, compared with 0.505 for emotion-only fine-tuning. Evaluations across five benchmarks connect recognition under increasing conflict with appraisal readout and downstream decision use. Together, the analyses and experiments support scene-specific appraisal as an intermediate representation that helps interpret conflicting emotional evidence.
Figures & tables
| Method | Video-aligned | Audio-aligned | Conflict | Consistent | Overall |
|---|---|---|---|---|---|
| Qwen2.5-Omni Base | 49.0 | 60.0 | 54.5 | 68.0 | 59.0 |
| Emotion-SFT | 55.0 | 65.0 | 60.0 | 73.0 | 64.3 |
| Generic-CoT-SFT | 57.0 | 66.0 | 61.5 | 74.0 | 65.7 |
| Modality-Gate-SFT | 58.0 | 66.0 | 62.0 | 74.0 | 66.0 |
| VISTA | 61.0 | 68.0 | 64.5 | 74.2 | 67.7 |
| Method | Direction | Class change | log-odds | Stability | |
|---|---|---|---|---|---|
| Emotion-SFT | 62.0 | 24.0 | 0.050 | 0.201 | 91.0 |
| Generic-CoT-SFT | 66.0 | 28.0 | 0.070 | 0.281 | 90.0 |
| Modality-Gate-SFT | 63.0 | 25.0 | 0.055 | 0.221 | 91.0 |
| VISTA | 75.0 | 36.0 | 0.120 | 0.484 | 92.0 |
| EmoMM accuracy | CH-SIMS v2 Acc2 | MELD | |||
|---|---|---|---|---|---|
| Method | Conflict | Conflict + missing | Overall | Q4 | wF1 |
| Qwen2.5-Omni Base | 46.5 | 38.0 | 79.4 | 70.0 | 59.55 |
| Emotion-SFT | 47.5 | 38.0 | 84.0 | 77.0 | 65.49 |
| Generic-CoT-SFT | 48.0 | 38.6 | 84.3 | 77.8 | 65.69 |
| Modality-Gate-SFT | 49.0 | 40.0 | 84.8 | 79.2 | 65.74 |
| VISTA | 52.0 | 43.5 | 85.9 | 81.5 | 66.94 |
Appendix figures & tables42 assets
Supplementary material from the paper’s appendix.
Appendix
| Field | Type | Values |
|---|---|---|
| Short text | Goal or concern supported by the scene | |
| Categorical | promotes , obstructs , irrelevant | |
| Categorical | expected , violated , uncertain | |
| Categorical | self , other , environment , shared , unknown | |
| Continuous | coping/control value | |
| Four binary entries | politeness , identity , status , relationship |
| Field | Evidence anchor | Distinction to preserve |
|---|---|---|
| Goal / concern | Requests, commitments, prior choices, or stated priorities in the available context | The current object of concern, such as obtaining this role, is more specific than a general value such as achievement. |
| Goal congruence | The observed outcome together with evidence about the relevant goal | Whether the event supports or obstructs the goal; the same external outcome can have different signs for different concerns. |
| Expectation | Prior plans, predictions, promises, or established patterns | Expectedness concerns anticipation. An undesired event can be expected, and a desirable event can be surprising. |
| Agency | Actions, decisions, attributions, and identified participants | Who or what produced the outcome; identifying a responsible agent does not by itself establish intent or blame. |
| Coping / control | Available remedies, remaining choices, resources, or finality of the decision | The person’s capacity to alter or manage the consequences, distinct from responsibility for causing them. |
| Norm / social relevance | Roles, relationships, agreed procedures, audience, and obligations | Which interpersonal standards bear on this situation, including whether an outcome or an expression is socially expected. |
| Approach | Organizing signal | Decision role | Evaluation focus |
|---|---|---|---|
| Wang et al. (2026) | Expectations and their violations | Predict emotion labels and shifts | Emotion labels and transition dynamics |
| AG-CTR 2 ( Chu et al., 2026 ) | Event understanding, appraisal, and coping | Retrieve support using appraisal chains | Support generation and retrieval-query quality |
| CHASE ( Sun et al., 2026a ) | Modality hidden states and attention preferences | Detect conflict and steer head-level attention | Recognition under conflict and missingness |
| VISTA | Concerns, event relations, and expression conditions | Condition modality arbitration and affect prediction | Conflict gains, correspondence, and decision access |
| Contrast | Reference retained | Dependency examined |
|---|---|---|
| Cross-sample shuffle | The current input and recognition task | Does another sample’s appraisal supply the same useful interpretation as the scene’s own appraisal? |
| Within-emotion shuffle | The current input and the donor’s emotion identity | Does event-specific correspondence matter beyond the donor appraisal’s association with the same class? |
| Appraisal with no decision connection | Auxiliary appraisal supervision | Does appraisal improve prediction when its forward connection to the decision is removed? |
| Generic semantic bottleneck | An intermediate semantic route to the same recognition task | Does organizing content as event appraisal contribute beyond the supplied semantic alternative? |
| Rationale / internal-state corruption | One of verbal presentation and internal appraisal is retained while the other is altered | Is the observed dependence stronger on displayed explanation text or on the internal appraisal information? |
| Relevant change / irrelevant paraphrase | The pairing identifies what event meaning should change or remain invariant | Does the response follow a relevant change while remaining stable to a reformulation of the same meaning? |
| Dataset | Public context | Role in this paper |
|---|---|---|
| CA-MER | 1,500 examples: 500 video-aligned, 500 audio-aligned, 500 consistent | Primary comparison of conflict and consistent performance |
| EmoMM | 4,000 base examples from Chinese CH-SIMS v2.0 and English CMU-MOSI; multimodal and unimodal sentiment annotations | Alignment, conflict, missingness, and their combination |
| CH-SIMS v2.0 | 4,402 labeled and 10,161 unlabeled Chinese clips, with multimodal and unimodal sentiment information | Binary accuracy (Acc2) across four conflict groups |
| THERADIA | 2,735 affect-annotated clips from interactions during cognitive exercises, with appraisal annotations | Frozen-representation appraisal probes and emotion-intensity regression |
| MELD | 13,708 utterances across 1,433 dialogues; official train/development/test sizes of 9,989/1,109/2,610 | Ordinary seven-class conversational emotion recognition |
| Input | Canonical | Input | Canonical |
|---|---|---|---|
| anger | angry | sadness | sad |
| happiness | happy | worried | worry |
| surprised | surprise | doubtful | doubt |
| fearful; afraid | fear | contemptuous | contempt |
| Comparison / population | Disagreement | Eligible clips | Rate (%) |
|---|---|---|---|
| Any text–audio–video pair | 1,117 | 2,281 | 48.97 |
| Text–audio | 800 | 2,281 | 35.07 |
| Text–video | 945 | 2,281 | 41.43 |
| Audio–video | 602 | 2,281 | 26.39 |
| Any pair: train | 676 | 1,368 | 49.42 |
| Any pair: validation | 231 | 456 | 50.66 |
| Quantity | Reported evidence | Population and interpretation |
|---|---|---|
| Three-modal disagreement | (1,117/2,281) | Our public-label audit of CH-SIMS: at least two unimodal polarity labels differ; Table 10 . |
| Image–text inconsistency | / | Filtered MVSA-Single / Multiple; neutral versus positive or negative modality labels ( Pan & Meng, 2024 ) . |
| Text–joint disagreement | UniC video reviews: complement of 63.69% text-only versus multimodal label agreement ( Du et al., 2025 ) . | |
| Strong-conflict prevalence | (661/4,402) | Labeled CH-SIMS v2.0 clips screened at ; corpus frequency ( Wang & Wu, 2025 ) . |
| Conflict difficulty | pp | MulT binary accuracy: 89.13% aligned, 65.60% conflict; selected DiffEmo tests, 173 examples per group ( Wang & Wu, 2025 ) . |
| Primary benchmark mix | 1,000/1,500 conflict | CA-MER’s equal allocation across two conflict directions and consistency; evaluation design ( Han et al., 2025 ) . |
| Comparator | |||
|---|---|---|---|
| Emotion-SFT | 4.5 | 1.2 | 3.3 |
| Generic-CoT-SFT | 3.0 | 0.2 | 2.8 |
| Modality-Gate-SFT | 2.5 | 0.2 | 2.3 |
| Comparator | ||||||
|---|---|---|---|---|---|---|
| Emotion-SFT | 6.0 | 3.0 | 1.2 | 4.8 | 1.8 | 3.0 |
| Generic-CoT-SFT | 4.0 | 2.0 | 0.2 | 3.8 | 1.8 | 2.0 |
| Modality-Gate-SFT | 3.0 | 2.0 | 0.2 | 2.8 | 1.8 | 1.0 |
| Method | Visual | Audio | Conflict Avg | Consistent | Overall | Invalid |
|---|---|---|---|---|---|---|
| Base | 49.0 | 60.0 | 54.5 | 68.0 | 59.0 | 0.5 |
| Emotion-SFT | 55.0 | 65.0 | 60.0 | 73.0 | 64.3 | 0.2 |
| Generic-CoT-SFT | 57.0 | 66.0 | 61.5 | 74.0 | 65.7 | 0.4 |
| Modality-Gate-SFT | 58.0 | 66.0 | 62.0 | 74.0 | 66.0 | 0.4 |
| VISTA | 61.0 | 68.0 | 64.5 | 74.2 | 67.7 | 0.7 |
| Method | Conflict Avg (%) | VISTA gain |
|---|---|---|
| MoSEAR | 61.5 | 3.0 |
| CHASE | 61.0 | 3.5 |
| VISTA | 64.5 | — |
| Method | Consistent | Conflict | Missing | Conflict + missing | Macro Avg |
|---|---|---|---|---|---|
| Base | 52.1 | 46.5 | 43.3 | 38.0 | 45.0 |
| Emotion-SFT | 55.0 | 47.5 | 44.0 | 38.0 | 46.1 |
| Generic-CoT-SFT | 55.5 | 48.0 | 44.2 | 38.6 | 46.6 |
| Modality-Gate-SFT | 55.6 | 49.0 | 45.1 | 40.0 | 47.4 |
| VISTA | 56.3 | 52.0 | 46.8 | 43.5 | 49.7 |
| CHASE (reevaluated) | 53.2 | 51.3 | 49.3 | 44.5 | 49.6 |
| Method | Q1 | Q2 | Q3 | Q4 | Overall Acc2 | Q4 MAE |
|---|---|---|---|---|---|---|
| Base | 87.5 | 82.5 | 77.5 | 70.0 | 79.4 | 0.460 |
| Emotion-SFT | 90.0 | 86.5 | 82.5 | 77.0 | 84.0 | 0.365 |
| Generic-CoT-SFT | 89.8 | 86.7 | 82.8 | 77.8 | 84.3 | 0.355 |
| Modality-Gate-SFT | 89.7 | 86.8 | 83.5 | 79.2 | 84.8 | 0.337 |
| VISTA | 90.2 | 87.2 | 84.5 | 81.5 | 85.9 | 0.315 |
| Comparator | ||||||
|---|---|---|---|---|---|---|
| Emotion-SFT | 0.2 | 0.7 | 2.0 | 4.5 | 4.3 | 1.42 |
| Generic-CoT-SFT | 0.4 | 0.5 | 1.7 | 3.7 | 3.3 | 1.11 |
| Modality-Gate-SFT | 0.5 | 0.4 | 1.0 | 2.3 | 1.8 | 0.60 |
| Condition | Conflict | Consistent | ||
|---|---|---|---|---|
| Full VISTA | 64.5 | 74.2 | 0.0 | 3.3 |
| Within-emotion appraisal shuffle | 63.7 | 74.0 | 2.7 | |
| Cross-sample appraisal shuffle | 62.0 | 73.5 | 1.5 | |
| Field-name randomization | 64.2 | 74.2 | 3.0 | |
| Auxiliary only; decision connection cut | 62.5 | 74.0 | 1.5 | |
| No expression regulation | 63.6 | 74.0 | 2.6 |
| Method | Conflict subset | Visual weight | Audio weight |
|---|---|---|---|
| Modality-Gate-SFT | Visual-aligned | 0.37 | — |
| Modality-Gate-SFT | Audio-aligned | — | 0.43 |
| VISTA | Visual-aligned | 0.42 | 0.30 |
| VISTA | Audio-aligned | 0.25 | 0.47 |
| Prompt | Conflict Avg | Consistent | Invalid | Output tokens |
|---|---|---|---|---|
| Direct | 54.5 | 68.0 | 0.5 | 8 |
| Modality decomposition | 55.8 | 69.0 | 0.8 | 210 |
| Generic conflict CoT | 56.4 | 69.0 | 1.0 | 290 |
| Appraisal prompt | 56.0 | 68.5 | 1.8 | 440 |
| Method | Direction | Label change | log-odds | Rewrite stability | |
|---|---|---|---|---|---|
| Emotion-SFT | 62.0 | 24.0 | 0.050 | 0.201 | 91.0 |
| Generic-CoT-SFT | 66.0 | 28.0 | 0.070 | 0.281 | 90.0 |
| Modality-Gate-SFT | 63.0 | 25.0 | 0.055 | 0.221 | 91.0 |
| VISTA | 75.0 | 36.0 | 0.120 | 0.484 | 92.0 |
| Intervention type | Pairs | Direction compliance (%) |
|---|---|---|
| Goal congruence | 25 | 84.0 |
| Expression regulation | 25 | 81.0 |
| Coping / control | 25 | 68.0 |
| Social norm | 25 | 69.0 |
| Method | Novelty | Pleasantness | Goal | Coping | Macro CCC |
|---|---|---|---|---|---|
| Base | 0.400 | 0.470 | 0.490 | 0.520 | 0.470 |
| Emotion-SFT | 0.440 | 0.500 | 0.530 | 0.550 | 0.505 |
| Generic-CoT-SFT | 0.470 | 0.530 | 0.560 | 0.570 | 0.533 |
| Modality-Gate-SFT | 0.470 | 0.530 | 0.560 | 0.580 | 0.535 |
| VISTA | 0.540 | 0.590 | 0.620 | 0.650 | 0.600 |
| Macro CCC | MAE | RMSE | Pearson | Spearman |
| 0.600 | 0.086 | 0.111 | 0.658 | 0.625 |
| Appraisal condition | Macro CCC | Gain over none |
|---|---|---|
| No appraisal | 0.390 | 0.000 |
| Auxiliary only; decision connection cut | 0.415 | 0.025 |
| Model-generated appraisal | 0.450 | 0.060 |
| Human appraisal (oracle) | 0.480 | 0.090 |
| Candidate disposition | Count | Retained source | Count |
|---|---|---|---|
| Automatically accepted | 2,872 | CH-SIMS v2.0 | 2,123 |
| Accepted after human adjudication | 1,231 | THERADIA | 866 |
| Retained with partial-field masking | 3,897 | MELD | 5,011 |
| Rejected | 2,256 | ||
| All candidates | 10,256 | All retained | 8,000 |
| Field | Valid | Coverage | Sampling consistency | Grounded |
|---|---|---|---|---|
| Goal / concern | 240 | 60.0 | 80.0 | 76.0 |
| Goal congruence | 296 | 74.0 | 85.0 | 81.0 |
| Expectation | 172 | 43.0 | 76.0 | 67.0 |
| Agency | 260 | 65.0 | 86.0 | 84.0 |
| Coping / control | 192 | 48.0 | 81.0 | 73.0 |
| Social norm | 136 | 34.0 | 78.0 | 68.0 |
| Teacher condition | Consistency | Grounded | Copying | MI (bits) | Dev Acc2 | Conflict Acc2 |
|---|---|---|---|---|---|---|
| Label-blind | 0.53 | 73.0 | 3.0 | 0.120 | 84.5 | 80.5 |
| Label-visible, isolated | 0.65 | 67.0 | 21.0 | 0.240 | 85.0 | 81.2 |
| Method | Params (M) | GPU-h | Peak GiB | P50 (s) | P95 (s) | Tokens |
|---|---|---|---|---|---|---|
| Emotion-SFT | 10.12 | 32.0 | 42.0 | 1.2 | 4.0 | 8 |
| Generic-CoT-SFT | 10.12 | 40.0 | 45.0 | 13.5 | 26.0 | 360 |
| Modality-Gate-SFT | 12.19 | 64.0 | 52.0 | 6.0 | 13.0 | 96 |
| VISTA | 12.39 | 72.0 | 56.0 | 15.5 | 31.0 | 360 |
| Error group | Reported error category | Share (%) |
|---|---|---|
| Only Emotion-SFT wrong | Conflict-source identification | 24.0 |
| Only Emotion-SFT wrong | Expression regulation / sarcasm | 22.0 |
| Only VISTA wrong | Goal / concern | 11.0 |
| Only VISTA wrong | Goal congruence / Expectation | 12.0 |
| Both models wrong | Insufficient context | 20.0 |
| Method | Weighted-F1 | Macro-F1 | Accuracy | wF1 |
|---|---|---|---|---|
| Base | 59.55 | 43.73 | 60.59 | |
| Emotion-SFT | 65.49 | 50.65 | 66.48 | 0.00 |
| Generic-CoT-SFT | 65.69 | 51.12 | 66.64 | |
| Modality-Gate-SFT | 65.74 | 50.84 | 66.73 | |
| VISTA | 66.94 | 52.89 | 67.85 |
| Emotion | Emotion-SFT | VISTA |
|---|---|---|
| Disgust | 18.0 | 21.0 |
| Fear | 20.0 | 23.0 |
| Sadness | 39.0 | 41.5 |
| Neutral | — | 81.7 |
| Setting | Final configuration |
|---|---|
| Backbone | Qwen2.5-Omni-7B Thinker |
| Frozen components | Audio encoder, visual encoder, and Talker |
| Retained training examples | 8,000 |
| Epochs / optimization steps | 3 / 375 |
| Effective batch size | 64 |
| LoRA targets | q_proj , k_proj , v_proj , o_proj in 28 layers |
| Evaluation | Base | Four trained methods |
|---|---|---|
| CA-MER | Frozen Base | Common-stage trained models; normalized output interface |
| EmoMM | Frozen Base | Common-stage checkpoint; no EmoMM adaptation |
| CH-SIMS v2.0 | Zero-shot | Separate task adaptation; train-defined conflict groups |
| MELD | Zero-shot | Separate task adaptation for seven emotion classes |
| THERADIA probe | Frozen backbone + common probe | Frozen backbone + common probe |
| Method | CA conflict | CA consistent | EmoMM conflict | SIMS Q4 | CCC | MELD wF1 |
|---|---|---|---|---|---|---|
| Base | 54.5 | 68.0 | 46.5 | 70.0 | 0.470 | 59.55 |
| Emotion-SFT | 60.0 | 73.0 | 47.5 | 77.0 | 0.505 | 65.49 |
| Generic-CoT-SFT | 61.5 | 74.0 | 48.0 | 77.8 | 0.533 | 65.69 |
| Modality-Gate-SFT | 62.0 | 74.0 | 49.0 | 79.2 | 0.535 | 65.74 |
| VISTA | 64.5 | 74.2 | 52.0 | 81.5 | 0.600 | 66.94 |
| Evidence family | Question and complete contents | Location |
|---|---|---|
| Core settings and definitions | Backbone, shared training, task adaptation, units, and six-metric overview | Tables 34 – 36 ; this section |
| CA-MER recognition | All subsets, overall accuracy, invalid output, external reevaluations | Tables 14 , 15 |
| Conflict specificity | All three control contrasts and directional decomposition | Tables 12 , 13 |
| EmoMM | All four conditions and CHASE reevaluation | Table 16 |
| CH-SIMS v2.0 | All conflict groups, overall Acc2, Q4 MAE | Table 17 |
| MELD | Weighted-F1, macro-F1, accuracy, unrounded differences, class recall | Tables 32 , 33 |