Resource-Limited Joint Multimodal Sentiment Reasoning and Classification via Chain-of-Thought Enhancement and Distillation
Organizations: School of Computer Science and Engineering, Northeastern University, Shenyang, China
Abstract
Current approaches for Multimodal Sentiment Analysis (MSA) primarily leverage the knowledge and reasoning capabilities of parameter-heavy (Multimodal) LLMs for classification, overlooking autonomous multimodal sentiment reasoning generation in resource-constrained environments. In this paper, we focus on the Resource-Limited Joint Multimodal Sentiment Reasoning and Classification task, JMSRC, which simultaneously performs multimodal sentiment reasoning chain generation and sentiment classification only with a lightweight model. We propose a Multimodal Chain-of-Thought Reasoning Distillation model, MulCoT-RD, designed for JMSRC that employs a "Teacher-Assistant-Student" distillation paradigm to address deployment constraints in resource-limited environments. We first leverage a high-performance Multimodal Large Language Model (MLLM) to generate the initial reasoning dataset and train a medium-sized assistant model with a multi-task learning mechanism. A lightweight student model is jointly trained to perform efficient multimodal sentiment reasoning generation and classification. Extensive experiments on four datasets demonstrate that MulCoT-RD, with only 3B parameters, achieves strong performance on JMSRC while exhibiting robust generalization and enhanced interpretability.
Figures & tables
| Dataset | Train | Dev | Test | Train g+ | Train q+ |
|---|---|---|---|---|---|
| MVSA-Single | 3608 | 451 | 452 | 6483 | 6350 |
| MVSA-Multiple | 13619 | 1702 | 1702 | 23424 | 23697 |
| Twitter-2015 | 3179 | 1122 | 1037 | 6166 | 6218 |
| Twitter-2017 | 3562 | 1176 | 1234 | 6652 | 6871 |
| ID | Teacher Model | Assistant Model | Student Model |
|---|---|---|---|
| 1 | GPT-4o-mini | Qwen3-VL-8B | Qwen3-VL-2B |
| 2 | Qwen2.5-VL-7B | Qwen2.5-VL-3B | |
| 3 | Qwen2.5-VL-72B | Qwen3-VL-8B | Qwen3-VL-2B |
| 4 | Qwen2.5-VL-7B | Qwen2.5-VL-3B |
| Model | Venue | MVSA-S | MVSA-M | ||
|---|---|---|---|---|---|
| Acc | w-F1 | Acc | w-F1 | ||
| MultiSentiNet | CIKM’17 | 69.8 | 69.8 | 68.9 | 68.1 |
| HSAN | ISI’17 | 69.9 | 66.9 | 68.0 | 67.8 |
| CoMN-Hop6 | SIGIR’18 | 70.5 | 70.0 | 68.9 | 68.8 |
| MGNNS | ACL’21 | 73.8 | 72.7 | 72.5 | 69.3 |
| CLMLF | NAACL’21 | 75.3 | 73.5 | 72.0 | 69.8 |
| Model | Venue | Twitter-15 | Twitter-17 | ||
|---|---|---|---|---|---|
| Acc | m-F1 | Acc | m-F1 | ||
| ESAFN | TASLP’20 | 73.4 | 67.4 | 67.8 | 64.2 |
| TomBERT | IJCAI’19 | 77.2 | 71.8 | 70.5 | 68.0 |
| CapTrBERT | ACM MM’21 | 78.0 | 73.2 | 72.3 | 70.2 |
| JML | EMNLP’21 | 78.7 | - | 72.7 | - |
| VLP-MABSA | ACL’22 | 78.6 | 73.8 | 73.8 | 71.8 |
| Model | Dataset | Sim | Meteor | Bleu | Rouge-L | Dist-1 | Dist-2 |
|---|---|---|---|---|---|---|---|
| ELLA | MVSA-S | 87.6 | 35.9 | 14.6 | 35.1 | 49.8 | 80.2 |
| MVSA-M | 84.7 | 36.0 | 15.9 | 35.9 | 52.5 | 83.7 | |
| Twitter-15 | 86.3 | 38.6 | 18.3 | 39.3 | 42.7 | 72.9 | |
| Twitter-17 | 86.6 | 38.1 | 17.6 | 38.2 | 43.0 | 73.1 | |
| Asst | MVSA-S | 92.6 | 59.8 | 47.8 | 55.0 | 49.8 | 80.2 |
| MVSA-M | 93.0 | 57.4 | 48.1 | 57.2 | 48.6 | 79.4 |
| Method | MVSA-S | MVSA-M | Twitter-15 | Twitter-17 | ||||
|---|---|---|---|---|---|---|---|---|
| Acc | w-F1 | Acc | w-F1 | Acc | m-F1 | Acc | w-F1 | |
| MulCoT-RD | 83.2 | 82.1 | 76.9 | 73.8 | 80.8 | 77.2 | 75.0 | 75.1 |
| w/o Img | 79.4 | 77.7 | 73.7 | 73.0 | 78.4 | 72.5 | 73.5 | 73.5 |
| w/o Txt | 77.9 | 77.1 | 66.2 | 67.7 | 65.6 | 56.6 | 64.6 | 59.4 |
| w/o CoT | 79.9 | 79.7 | 74.2 | 73.1 | 79.9 | 75.5 | 74.2 | 73.4 |
| w/o Asst | 81.9 | 81.3 | 75.2 | 74.1 | 79.3 | 72.3 | 73.7 | 73.3 |
| Qwen3-VL | MVSA-S | MVSA-M | Twitter-15 | Twitter-17 | ||||
|---|---|---|---|---|---|---|---|---|
| Acc | w-F1 | Acc | w-F1 | Acc | m-F1 | Acc | m-F1 | |
| 8B (asst) 1 | 83.6 | 82.9 | 76.0 | 73.5 | 80.5 | 76.3 | 76.1 | 75.1 |
| 2B (stu) 1 | 84.1 | 83.7 | 76.7 | 74.1 | 80.0 | 75.2 | 74.8 | 74.7 |
| 8B (asst) 2 | 83.2 | 82.6 | 76.7 | 73.2 | 80.9 | 77.4 | 75.0 | 73.4 |
| 2B (stu) 2 | 83.4 | 82.4 | 76.8 | 74.2 | 80.4 | 76.7 | 73.8 | 73.0 |
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
| Dataset | Samples | GPT-4o-mini | Qwen2.5-VL-72B | ||||
|---|---|---|---|---|---|---|---|
| Acc | w-F1 | m-F1 | Acc | w-F1 | m-F1 | ||
| MVSA-S | 3608 | 79.7 | 79.7 | 69.9 | 76.0 | 77.1 | 66.9 |
| MVSA-M | 13619 | 72.0 | 68.0 | 55.3 | 74.0 | 70.5 | 60.6 |
| Twitter-15 | 3179 | 94.0 | 94.1 | 92.6 | 95.6 | 95.6 | 94.6 |
| Twitter-17 | 3562 | 86.8 | 86.7 | 86.4 | 92.9 | 92.9 | 93.4 |
| Role | Model | Access | Release Date |
|---|---|---|---|
| Teacher | GPT-4o-mini | Closed | 2024.07 |
| Qwen2.5-VL-72B | Open | 2025.02 | |
| Assistant | Qwen3-VL-8B | Open | 2025.10 |
| Qwen2.5-VL-7B | Open | 2025.02 | |
| Student | Qwen3-VL-2B | Open | 2025.10 |
| Qwen2.5-VL-3B | Open | 2025.02 |
| Model | Total Parameters | Trainable | Ratio (%) |
|---|---|---|---|
| Qwen3-VL-8B | 8,810,770,672 | 43,646,976 | 0.4954 |
| Qwen3-VL-2B | 2,144,964,608 | 17,432,576 | 0.8127 |
| Qwen2.5-VL-7B | 8,339,756,032 | 47,589,376 | 0.5706 |
| Qwen2.5-VL-3B | 3,791,775,744 | 37,152,768 | 0.9798 |
| Model | Params | TFLOPs | Latency | Memory |
|---|---|---|---|---|
| (B) | - | (ms) | (GB) | |
| Qwen2.5-VL-72B | 73.41 | 38.24 | 17849.37 | 149.69 |
| Emotion-LLaMA | 14.58 | 4.11 | 11342.73 | 31.61 |
| Qwen3-VL-8B | 8.81 | 3.13 | 10708.24 | 16.61 |
| Qwen3-VL-2B | 2.14 | 3.15 | 8302.95 | 4.11 |
| Qwen2.5-VL-7B | 8.34 | 3.02 | 7002.99 | 15.81 |
| Model | Params | TFLOPs | Latency | Memory |
|---|---|---|---|---|
| (B) | - | (ms) | (GB) | |
| Qwen2.5-VL-72B | 73.41 | 38.41 | 17524.94 | 149.76 |
| Emotion-LLaMA | 14.58 | 4.31 | 11582.16 | 31.66 |
| Qwen3-VL-8B | 8.81 | 3.39 | 9930.38 | 16.62 |
| Qwen3-VL-2B | 2.14 | 3.60 | 8710.76 | 4.12 |
| Qwen2.5-VL-7B | 8.34 | 3.21 | 6209.65 | 15.83 |