Learning What to Trust in Multimodal Learning under Noisy Supervision
Organizations: University of Science and Technology of China
Abstract
Multimodal classification processes and relates information from multiple modalities to achieve more accurate predictions. However, existing methods typically rely on high-quality ground-truth labels, which are difficult to obtain in real-world scenarios. While sample-selection methods for learning with noisy labels aim to identify correctly labeled examples from noisy data, traditional methods primarily focus on unimodal settings and fail to exploit multimodal information fully. This motivates us to build a more reliable noise detector in multimodal learning. To this end, we theoretically analyze the relationship between representation structure and noise detection capability. Based on this analysis, we propose REFINE, which is a multimodal label-noise detection framework that jointly uses fused and unimodal representations for label-noise detection. Specifically, REFINE constructs discriminative eigenvectors through discriminative analysis of the target and background classes and selects trusted representation spaces with better noise detection capability for each class. Within each trusted space, REFINE measures the alignment between each instance representation and the discriminative eigenvectors. It then combines the subsets selected from these spaces. The combined set provides cleaner supervision for updating the multimodal classifier, thereby reducing the influence of mislabeled examples during training and improving model generalization. Extensive experiments across diverse tasks demonstrate REFINE's superiority compared to baseline methods. The source code will be publicly available.
Figures & tables
| Dataset | Method | Noise | Average | |||||
| Sym 50% | Asym 40% | F-IDN 40% | F-IDN 50% | P-IDN 40% | P-IDN 50% | |||
| UPMC-Food101 | Standard | 82.75 ± 0.33 | 67.86 ± 0.22 | 84.25 ± 0.10 | 82.50 ± 0.11 | 84.09 ± 0.08 | 82.11 ± 0.07 | 80.59 ± 0.05 |
| CRUST | 84.81 ± 0.16 | 81.93 ± 0.23 | 83.11 ± 0.25 | 81.75 ± 0.04 | 82.19 ± 0.33 | 80.79 ± 0.20 | 82.43 ± 0.08 | |
| AUM | 83.18 ± 0.12 | 66.08 ± 0.56 | 83.80 ± 0.13 | 82.09 ± 0.30 | 82.01 ± 0.23 | 79.80 ± 0.22 | 79.49 ± 0.15 | |
| L2D | 84.49 ± 0.22 | 72.89 ± 0.37 | 85.75 ± 0.21 | 84.06 ± 0.11 | 85.46 ± 0.22 | 83.83 ± 0.35 | 82.75 ± 0.07 | |
| DIST | 85.56 ± 0.20 | 80.97 ± 0.38 | 85.61 ± 0.23 | 84.29 ± 0.15 | 84.37 ± 0.17 | 82.29 ± 0.15 | 83.85 ± 0.10 | |
| Method | Standard | AUM | CRUST | L2D | DIST | DIST+CT | FINE | REFINE |
| Accuracy | 60.37 ± 0.65 | 58.90 ± 0.73 | 58.70 ± 0.71 | – | 60.27 ± 0.74 | 60.28 ± 0.72 | 60.71 ± 0.98 | 61.03 ± 0.67 |
| Method | Noise | |||
| Sym 50% | Asym 40% | F-IDN 50% | P-IDN 50% | |
| Standard | 87.54 ± 0.20 | 73.77 ± 1.06 | 86.72 ± 0.15 | 86.05 ± 0.25 |
| CRUST | 88.41 ± 0.18 | 81.92 ± 0.79 | 87.10 ± 0.21 | 86.36 ± 0.14 |
| AUM | 88.52 ± 0.12 | 74.13 ± 1.31 | 87.02 ± 0.21 | 86.09 ± 0.28 |
| L2D | 87.91 ± 0.17 | 80.39 ± 0.66 | 86.90 ± 0.21 | 86.15 ± 0.30 |
| DIST | 87.53 ± 0.10 | 74.15 ± 1.29 | 86.73 ± 0.16 | 86.10 ± 0.27 |
| Method | Sym 50% | Asym 40% | F-IDN 50% | P-IDN 50% | ||||||||
| P | R | F1 | P | R | F1 | P | R | F1 | P | R | F1 | |
| CRUST | 93.76 ± 0.08 | 93.67 ± 0.09 | 93.72 ± 0.08 | 94.24 ± 0.24 | 78.46 ± 0.20 | 85.63 ± 0.21 | 85.68 ± 0.24 | 85.61 ± 0.23 | 85.65 ± 0.23 | 84.85 ± 0.18 | 84.78 ± 0.17 | 84.81 ± 0.17 |
| AUM | 96.40 ± 0.74 | 90.50 ± 0.82 | 93.35 ± 0.10 | 60.62 ± 0.09 | 91.79 ± 0.29 | 73.01 ± 0.06 | 82.36 ± 0.70 | 93.67 ± 0.41 | 87.65 ± 0.25 | 78.39 ± 0.56 | 93.50 ± 0.32 | 85.28 ± 0.20 |
| L2D | 87.64 ± 0.25 | 87.63 ± 0.25 | 87.63 ± 0.25 | 66.40 ± 0.09 | 66.39 ± 0.09 | 66.40 ± 0.09 | 81.42 ± 0.09 | 81.42 ± 0.09 | 81.42 ± 0.09 | 82.40 ± 0.13 | 82.40 ± 0.13 | 82.40 ± 0.13 |
| DIST | 94.80 ± 0.11 | 95.98 ± 0.13 | 95.39 ± 0.07 | 93.22 ± 0.37 | 85.92 ± 0.30 | 89.42 ± 0.31 | 81.45 ± 0.17 | 97.04 ± 0.15 | 88.56 ± 0.08 | 80.66 ± 0.35 | 95.89 ± 0.12 | 87.61 ± 0.18 |
| DIST+CT | 94.80 ± 0.11 | 95.98 ± 0.15 | 95.39 ± 0.07 | 93.16 ± 0.39 | 85.97 ± 0.34 | 89.42 ± 0.33 | 81.49 ± 0.15 | 97.02 ± 0.16 | 88.58 ± 0.06 | 80.74 ± 0.38 | 95.88 ± 0.11 | 87.66 ± 0.21 |
| Method | Noise | ||||
| Sym 50% | Asym 40% | F-IDN 50% | P-IDN 50% | ||
| DivideMix | Best | 86.98 ± 0.16 | 75.36 ± 0.46 | 86.48 ± 0.07 | 86.41 ± 0.12 |
| Last | 86.97 ± 0.17 | 72.30 ± 0.18 | 86.44 ± 0.06 | 86.39 ± 0.10 | |
| Mean Teacher | Best | 87.23 ± 0.23 | 75.26 ± 0.31 | 86.49 ± 0.11 | 86.36 ± 0.19 |
| Last | 87.23 ± 0.23 | 71.32 ± 0.41 | 86.49 ± 0.11 | 86.36 ± 0.19 | |
| ICT | Best | 87.09 ± 0.21 | 75.30 ± 0.37 | 86.41 ± 0.11 | 86.21 ± 0.13 |
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
| Dataset | Method | Noise | Average | |||||
| Sym 50% | Asym 40% | F-IDN 40% | F-IDN 50% | P-IDN 40% | P-IDN 50% | |||
| UPMC-Food101 | Standard | 82.63 ± 0.34 | 67.76 ± 0.24 | 84.30 ± 0.10 | 82.62 ± 0.10 | 84.39 ± 0.15 | 82.75 ± 0.06 | 80.74 ± 0.05 |
| CRUST | 84.71 ± 0.17 | 81.93 ± 0.23 | 83.11 ± 0.23 | 82.08 ± 0.05 | 82.60 ± 0.19 | 81.74 ± 0.18 | 82.69 ± 0.05 | |
| AUM | 83.04 ± 0.12 | 65.89 ± 0.59 | 83.73 ± 0.14 | 82.06 ± 0.29 | 82.08 ± 0.27 | 80.14 ± 0.16 | 79.49 ± 0.14 | |
| L2D | 84.39 ± 0.22 | 72.80 ± 0.37 | 85.70 ± 0.20 | 84.04 ± 0.12 | 85.47 ± 0.20 | 83.90 ± 0.32 | 82.72 ± 0.06 | |
| DIST | 85.45 ± 0.22 | 80.78 ± 0.38 | 85.54 ± 0.22 | 84.23 ± 0.16 | 84.37 ± 0.18 | 82.40 ± 0.13 | 83.80 ± 0.11 | |