Dynamic Alignment and Calibration for Multimodal Learning
Authors: Jinghao Xu, Zhenhua Guo, Xiaofeng Zhu, Xiaoshuang Shi
Organizations: School of Computer Science and Engineering, University of Electronic Science and Technology of China, Chengdu,China · Tianyijiaotong Technology Ltd.,China · School of Computer Science and Technology, Hainan University, Haikou, China
Dynamic multimodal learning aims to learn robust representations by adaptively modeling information discrepancies across modalities. However, existing methods still suffer from two limitations: (i) static cross-modal alignment strategies usually impose uniform constraints on all samples while overlooking sample-wise variations, potentially leading to unreasonable over-alignment; and (ii) confidence- or uncertainty-aware fusion methods often fail to adequately account for feature magnitude and confidence differences across modalities. For modality pairs with significant feature magnitude differences or small confidence gaps, it might be unreliable to strictly align fusion weights according to confidence. To address these issues, we propose an Alignment- and Calibration-driven Multimodal Learning framework (ACML). Specifically, ACML incorporates a dynamic cross-modal triplet alignment module, which enforces strong semantic consistency for high-confidence positive pairs while encouraging diverse representation learning between high- and low-confidence positive pairs according to their confidence gaps. Additionally, ACML introduces a difference-aware attention calibration strategy that adaptively adjusts attention regularization based on feature magnitude and confidence differences across modalities, thereby mitigating biases caused by unreasonable fusion constraints. Extensive experiments on multiple multimodal benchmark datasets demonstrate that ACML consistently achieves superior performance and robustness over recent state-of-the-art methods.
Figures & tables
Figure 1: Overview of the proposed framework ACML. In LDCTA , △ and □ denote different categories, colors represent modalities, and color intensity reflects confidence.
\toprule \multirow 2*Dataset
\multirow 2*Method
ϵ=0.0
ϵ=5.0
ϵ=10.0
Avg±Std
Worst
Avg±Std
Worst
Avg±Std
Worst
\midrule \multirow 10* \makecell NYU Depth V2
Depth
61.97 ±0.99
60.40
49.65 ±3.22
43.27
42.42 ±2.42
36.85
RGB
62.57 ±0.82
61.32
52.56 ±2.28
47.55
46.16 ±3.10
38.84
LateFusion
69.22 ±0.98
67.43
56.35 ±2.84
51.22
48.65 ±3.40
42.36
Concat
68.70 ±0.67
67.89
55.76 ±2.23
49.54
48.17 ±3.36
40.21
TMC
71.42 ±0.84
70.03
61.44 ±2.19
56.58
53.36 ±3.36
45.87
Table 1: Classification comparison when 50% of the modalities are corrupted with Gaussian noise, i.e., zero mean with variance of ϵ . We highlight the optimal results using bold text and indicate the suboptimal results with underlining for clarity.
Table 2: Ablation studies on NYU Depth V2. We report both the average performance and the worst-case performance under different noise levels.
Figure 2: 2D t -SNE visualization of multimodal features. △ indicates the image modality, and □ denotes the text modality. Each color corresponds to a distinct class, with darker shades indicating higher confidence.
Figure 3: Performance sensitivity to hyperparameters λ1 and λ2 across two datasets.