Multimodal joint training often suffers from modality imbalance, where a dominant modality suppresses the optimization of others. Existing methods mainly balance modality learning by modulating gradient magnitudes or directions, modifying optimization objectives, or adjusting training strategies, with most interventions focusing on the current update. However, when combined with widely used momentum-based optimizers, the update also incorporates accumulated information from previous gradients, which is not explicitly addressed by current-step modulation alone. To address this issue, we propose Variance-Calibrated MomentuM (VCMM), which adapts gradient memory to modality-specific gradient dynamics. Specifically, VCMM estimates minibatch noise and temporal drift online and uses their relative strength to determine modality-specific momentum through a Kalman-inspired controller. We further center the control signal across modalities and apply exact bias correction for the time-varying first moment, enabling adaptive gradient memory without extra network passes or explicit learning-rate scaling. Experiments on four multimodal benchmarks demonstrate consistent improvements with modest training overhead.
Figures & tables
Method
CREMA-D
KSounds
Twitter
NVGesture
Acc.
mAP
Acc.
mAP
Acc.
F1
Acc.
F1
Unimodal-1
.6317
.5879
.5412
.5669
.5863
.4333
.7822
.7833
Unimodal-2
.4583
.6861
.5562
.5837
.7367
.6849
.7863
.7865
Unimodal-3
–
–
–
–
–
–
.8154
.8183
Concat
.6361
.6841
.6455
.7130
.7011
.6386
.8237
.8270
G-Blend [ 4 ]
.6465
.7392
.6722
.7274
.7309
.6799
.8299
.8305
Table 1: Comparison with representative multimodal learning methods. The best results are highlighted in bold . The underline denotes the second-best performance. Gray-shaded results indicate that multimodal performance is lower than that of the best unimodal approach.