Current approaches for Multimodal Sentiment Analysis (MSA) primarily leverage the knowledge and reasoning capabilities of parameter-heavy (Multimodal) LLMs for classification, overlooking autonomous multimodal sentiment reasoning generation in resource-constrained environments. In this paper, we focus on the Resource-Limited Joint Multimodal Sentiment Reasoning and Classification task, JMSRC, which simultaneously performs multimodal sentiment reasoning chain generation and sentiment classification only with a lightweight model. We propose a Multimodal Chain-of-Thought Reasoning Distillation model, MulCoT-RD, designed for JMSRC that employs a "Teacher-Assistant-Student" distillation paradigm to address deployment constraints in resource-limited environments. We first leverage a high-performance Multimodal Large Language Model (MLLM) to generate the initial reasoning dataset and train a medium-sized assistant model with a multi-task learning mechanism. A lightweight student model is jointly trained to perform efficient multimodal sentiment reasoning generation and classification. Extensive experiments on four datasets demonstrate that MulCoT-RD, with only 3B parameters, achieves strong performance on JMSRC while exhibiting robust generalization and enhanced interpretability.
Figure 2: MulCoT-RD comprises two core modules, i.e., (1) Multimodal CoT Enhancement Module, (2) Reasoning Distillation Module (Assistant Model with Multi-Task Learning, Student Model with Joint Learning).
Dataset
Train
Dev
Test
Train g+
Train q+
MVSA-Single
3608
451
452
6483
6350
MVSA-Multiple
13619
1702
1702
23424
23697
Twitter-2015
3179
1122
1037
6166
6218
Twitter-2017
3562
1176
1234
6652
6871
Table 1: Statistics of datasets. g+ and q+ represent the teacher models GPT-4o-mini Hurst et al. (2024) and Qwen2.5-VL-72B Bai et al. (2025) , respectively.
ID
Teacher Model
Assistant Model
Student Model
1
GPT-4o-mini
Qwen3-VL-8B
Qwen3-VL-2B
2
Qwen2.5-VL-7B
Qwen2.5-VL-3B
3
Qwen2.5-VL-72B
Qwen3-VL-8B
Qwen3-VL-2B
4
Qwen2.5-VL-7B
Qwen2.5-VL-3B
Table 2: Four reasoning distillation architectures.
Model
Venue
MVSA-S
MVSA-M
Acc
w-F1
Acc
w-F1
MultiSentiNet
CIKM’17
69.8
69.8
68.9
68.1
HSAN
ISI’17
69.9
66.9
68.0
67.8
CoMN-Hop6
SIGIR’18
70.5
70.0
68.9
68.8
MGNNS
ACL’21
73.8
72.7
72.5
69.3
CLMLF
NAACL’21
75.3
73.5
72.0
69.8
Table 3: Results for coarse-grained MSA. Models above the middle line are small models fully fine-tuned, while those below are (M)LLMs fine-tuned with LoRA. † denotes the results reproduced by us using models retrained on our datasets. ‡ indicates the 16-shot performance under In-Context Learning (ICL). The best results are bold-typed and the second best ones are underlined. ∗ means the zero-shot performance.
Model
Venue
Twitter-15
Twitter-17
Acc
m-F1
Acc
m-F1
ESAFN
TASLP’20
73.4
67.4
67.8
64.2
TomBERT
IJCAI’19
77.2
71.8
70.5
68.0
CapTrBERT
ACM MM’21
78.0
73.2
72.3
70.2
JML
EMNLP’21
78.7
-
72.7
-
VLP-MABSA
ACL’22
78.6
73.8
73.8
71.8
Table 4: Results of different methods for MASC. “-” means it does not exist in the original paper.
Model
Dataset
Sim
Meteor
Bleu
Rouge-L
Dist-1
Dist-2
ELLA
MVSA-S
87.6
35.9
14.6
35.1
49.8
80.2
MVSA-M
84.7
36.0
15.9
35.9
52.5
83.7
Twitter-15
86.3
38.6
18.3
39.3
42.7
72.9
Twitter-17
86.6
38.1
17.6
38.2
43.0
73.1
Asst
MVSA-S
92.6
59.8
47.8
55.0
49.8
80.2
MVSA-M
93.0
57.4
48.1
57.2
48.6
79.4
Table 5: Evaluation results of generated reasoning from Emotion-LLaMA, assistant and student models.
Method
MVSA-S
MVSA-M
Twitter-15
Twitter-17
Acc
w-F1
Acc
w-F1
Acc
m-F1
Acc
w-F1
MulCoT-RD
83.2
82.1
76.9
73.8
80.8
77.2
75.0
75.1
w/o Img
79.4
77.7
73.7
73.0
78.4
72.5
73.5
73.5
w/o Txt
77.9
77.1
66.2
67.7
65.6
56.6
64.6
59.4
w/o CoT
79.9
79.7
74.2
73.1
79.9
75.5
74.2
73.4
w/o Asst
81.9
81.3
75.2
74.1
79.3
72.3
73.7
73.3
Table 6: The performance comparison of our full model and its ablated methods under the setting where Qwen2.5-VL-72B serves as the teacher model.
Figure 3: Efficiency comparison on MVSA-Single and Twitter-2015. Metrics are measured on the batch size as 1 and all samples are from the test set. Note that the models are trained under the paradigm where Qwen2.5-VL-72B serves as the teacher model. Detailed efficiency comparison results are provided in Appendix H .
Qwen3-VL
MVSA-S
MVSA-M
Twitter-15
Twitter-17
Acc
w-F1
Acc
w-F1
Acc
m-F1
Acc
m-F1
8B (asst) 1
83.6
82.9
76.0
73.5
80.5
76.3
76.1
75.1
2B (stu) 1
84.1
83.7
76.7
74.1
80.0
75.2
74.8
74.7
8B (asst) 2
83.2
82.6
76.7
73.2
80.9
77.4
75.0
73.4
2B (stu) 2
83.4
82.4
76.8
74.2
80.4
76.7
73.8
73.0
Table 7: Performance of Qwen3-VL-based models on coarse-grained MSA and MASC. The best results are bold-typed, and the second best ones are underlined. 1 and 2 indicate the Teacher models, with 1 being GPT-4o-mini and 2 being Qwen2.5-VL-72B.
Figure 4: Visualization of two samples, using the MulCoT-RD architecture with ID 4 from Table 2 .
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
Samples
GPT-4o-mini
Qwen2.5-VL-72B
Acc
w-F1
m-F1
Acc
w-F1
m-F1
MVSA-S
3608
79.7
79.7
69.9
76.0
77.1
66.9
MVSA-M
13619
72.0
68.0
55.3
74.0
70.5
60.6
Twitter-15
3179
94.0
94.1
92.6
95.6
95.6
94.6
Twitter-17
3562
86.8
86.7
86.4
92.9
92.9
93.4
Appendix
Table 8: Performance of the Assistant Model (Qwen-2.5-VL-7B) on Training Sets During Data Expansion, Guided by Different Teacher Models.
Role
Model
Access
Release Date
Teacher
GPT-4o-mini
Closed
2024.07
Qwen2.5-VL-72B
Open
2025.02
Assistant
Qwen3-VL-8B
Open
2025.10
Qwen2.5-VL-7B
Open
2025.02
Student
Qwen3-VL-2B
Open
2025.10
Qwen2.5-VL-3B
Open
2025.02
Appendix
Table 9: Model Selection and Characteristics.
Figure 5: Effect of Loss Weight on Convergence for CoT Generation and Sentiment Classification Tasks.
Model
Total Parameters
Trainable
Ratio (%)
Qwen3-VL-8B
8,810,770,672
43,646,976
0.4954
Qwen3-VL-2B
2,144,964,608
17,432,576
0.8127
Qwen2.5-VL-7B
8,339,756,032
47,589,376
0.5706
Qwen2.5-VL-3B
3,791,775,744
37,152,768
0.9798
Appendix
Table 10: Total and trainable parameter counts under LoRA fine-tuning.
Figure 6: Two-stage reasoning prompt template.
Figure 7: Accuracy comparison of teacher (GPT-3.5-Turbo), assistant (Flan-T5-Large with 783M parameters) and student (Flan-T5-Base) models.
Figure 8: Weighted-F1 comparison of teacher(GPT-3.5-Turbo), assistant(Flan-T5-Large with 783M parameters) and student(Flan-T5-Base with 248M parameters) models.
Figure 9: Macro-F1 comparison of teacher(GPT-3.5-Turbo), assistant(Flan-T5-Large with 783M parameters) and student(Flan-T5-Base with 248M parameters) models.
Model
Params
TFLOPs
Latency
Memory
(B)
-
(ms)
(GB)
Qwen2.5-VL-72B
73.41
38.24
17849.37
149.69
Emotion-LLaMA
14.58
4.11
11342.73
31.61
Qwen3-VL-8B
8.81
3.13
10708.24
16.61
Qwen3-VL-2B
2.14
3.15
8302.95
4.11
Qwen2.5-VL-7B
8.34
3.02
7002.99
15.81
Appendix
Table 11: Efficiency comparison on the MVSA-Single dataset.
Model
Params
TFLOPs
Latency
Memory
(B)
-
(ms)
(GB)
Qwen2.5-VL-72B
73.41
38.41
17524.94
149.76
Emotion-LLaMA
14.58
4.31
11582.16
31.66
Qwen3-VL-8B
8.81
3.39
9930.38
16.62
Qwen3-VL-2B
2.14
3.60
8710.76
4.12
Qwen2.5-VL-7B
8.34
3.21
6209.65
15.83
Appendix
Table 12: Efficiency comparison on the Twitter-2015 dataset.
Recent multimodal large language models (MLLMs) have shown strong chain-of-thought (CoT) reasoning ability on vision-language tasks, but their direct deployment in real-world systems is often limited by latency and resource constraints. In practice, smaller MLLMs are preferred for online serving, yet their reasoning performance is bottlenecked by the lack of large-scale, high-quality multimodal CoT supervision. In this paper, we present OmniThoughtVis, a scalable data curation and distillation pipeline for transferring multimodal reasoning capabilities from high-capacity teacher models to smaller, deployment-oriented MLLMs. Starting from a diverse open-source seed pool, our pipeline generates structured CoT traces and performs joint annotation of reasoning difficulty, answer quality, and semantic task tags. To maintain data quality at scale, we combine rule-based filtering, difficulty-aware selection, and tag-based diversity sampling, resulting in a curated corpus of 1.8M samples that supports controllable subset construction for downstream training. We use OmniThoughtVis to distill Qwen3-VL models from 2B to 8B parameters and evaluate them on nine multimodal reasoning benchmarks. The resulting distilled models show consistent gains across model scales, including improvements of up to +16.8 points on MathVerse and +5.6 points on MMMU-Pro for the 4B model. Notably, the distilled 4B model matches or surpasses the undistilled 8B baseline on several tasks, highlighting the practical value of scalable reasoning distillation for deployment-oriented MLLMs.
Recent advances in multimodal large language models (MLLMs) have significantly improved the performance of multimodal emotion recognition (MER) and enabled interpretable description generation by jointly modeling video, audio, and language, etc. However, these performance improvements are often accompanied by an increase in model parameter size (e.g, at least 7B), which simultaneously incurs high computational costs and reduces inference efficiency, thereby hindering real-time deployment on resource-constrained platforms such as robots and mobile devices. This raises a fundamental question: do we really need the multimodal MER model larger than 1B parameters for high-quality MER? In this paper, we challenge the assumption that larger models are inherently necessary and proposes a lightweight MER framework (called Light-MER), which achieves better and faster multimodal sentiment understanding and recognition through knowledge distillation. It can transfer knowledge from a strong, large-scale teacher model to a lightweight sub-billion-parameter student model, aiming to preserve rich multimodal emotion reasoning and recognition while substantially improving deployment efficiency. Specifically, we introduce two new optimization strategies to enhance knowledge transfer: (1) a new optimal transport loss that combines Sliced Wasserstein Distance with hidden-state alignment, and (2) a new multi-reward optimization strategy based on GRPO that balances MER performance and efficiency, aimed at further enhancing the learning capabilities of student models. Extensive experiments on nine benchmark datasets demonstrate that Light-MER achieves state-of-the-art performance while significantly improving inference efficiency. This highlights the strong potential of small multimodal emotion language models for future research. Code is available at https://github.com/GAIR-Lab/Light-MER.
Kaiwen Zheng, Junchen Fu, Wenhao Deng +3
University of Glasgow United Kingdom · Institute of Computing Technology, Chinese Academy of Sciences China · School of Artificial Intelligence, Shandong University China +1
Open multimodal reasoning models have benefited from large-scale reasoning supervision, yet reliable post-training remains challenging due to uneven data quality, inefficient supervision construction, imbalanced difficulty, and cross-domain interference. We introduce MMVistaReason (MVR), an open-data post-training recipe with three components: (1) broader capability coverage across complementary Analytical and Real-World reasoning groups, emphasizing structured reasoning versus visual perception and spatial grounding; (2) efficient SFT and RL data construction, standardizing heterogeneous open data through staged cleaning and annotation, combining difficulty-aware cascaded teacher distillation with answer-likelihood-based trajectory selection to construct MVR-SFT-528K, and applying scale-specific frontier filtering for MVR-RL-63K; and (3) specialize-then-integrate training, which trains complementary RL experts and consolidates their capabilities through multi-teacher on-policy distillation (MOPD). Our analyses reveal a capacity-dependent interaction between supervision difficulty, trajectory quality, and model capacity: smaller students benefit more from selected supervision, while larger students are robust to trajectory variation and mixed-domain interference. Mixed-domain RL introduces benchmark-level negative transfer, whereas MOPD provides consistent capability integration, with the preferred KL direction varying across model scales. Across 15 multimodal benchmarks, MVR-4B achieves an average score of 72.8, outperforming Qwen3.5-9B (Instruct) and MMFineReason-8B while using about 70% fewer samples than MMFineReason. Scaling to 9B improves the average to 74.4, surpassing Qwen3.5-35B-A3B (Instruct). Overall, MMVistaReason demonstrates that systematic open-data construction and capacity-aware post-training provide a practical and scalable path toward reliable multimodal reasoning.
Juekai Lin, Honglin Lin, Yuqian Yuan +7
Zhejiang University · Shanghai Artificial Intelligence Laboratory, OpenDataLab · Shanghai Jiao Tong University +1