Multimodal learning beyond two modalities commonly leverages a specific modality (e.g., text) to bind other modalities. However, how to establish a more balanced representation space that approximates shared semantics while respecting the holistic geometry of n-modal data remains challenging. In this work, we present BaryBind, which aims to transport the specific modality towards the Wasserstein barycenter (WB) optimized across all modalities and introduces a volumetric alignment objective to establish a unified semantic space around the WB embedding. Specifically, we project specific modalities to the WB, which minimizes the average Wasserstein distances to multimodal distributions and serves as the anchor for subsequent alignment. We then construct a barycenter simplex, whose volume is taken as a similarity metric for global alignment centered at the WB. Experiments show that BaryBind achieves competitive performance in text-video-audio retrieval, classification, videoQA, and cross-modal generation tasks, along with robustness under modality absence and scalability to more than three modalities. Code is released at https://github.com/xl-tang3/BaryBind.
Figures & tables
Figure 1 : BaryBind learns the multimodal Wasserstein Barycenter (WB) of all modalities via WB optimization, and the WB serves as the anchor to bind modalities via volumetric alignment based on the barycenter simplex volume. Notably, BaryBind yields more balanced multimodal retrieval results and enables TA2I generation that well preserves both text and audio semantics.
Figure 2 : The t-SNE comparison of multimodal embeddings on the zero-shot VGGSound dataset [ 8 ] . Compared to VAST [ 9 ] , BaryBind induces a latent space where class clusters are clearly separated, and multimodal embeddings are grouped around the WB embeddings.
Figure 3 : BaryBind establishes unified representation space through two processes: 1) The WB optimization process transforms modality-specific initializer m0 into the WB anchor b , in which the WB map Tθ is optimized by minimizing the average of Wasserstein distances to all multimodal distributions Pk . 2) The volumetric alignment process utilizes the WB anchor to induce a barycenter simplex, whose volume is contrasted to achieve global alignment to the WB embedding.
Figure 4 : Visualization of multimodal embeddings before alignment (left) and after applying TV + TA cosine-based alignment (middle) and the proposed BVC alignment (right). The BVC loss promotes convergence of embeddings towards the WB anchors, forming compact clusters around the WB embeddings, which highlights improved balanced multimodal alignment.
Method
Modality
Acc@1
Acc@5
ImageBind [ 16 ]
A
31.6
58.7
ImageBind [ 16 ]
V
37.9
65.9
LanguageBind [ 67 ]
A
34.1
62.8
LanguageBind [ 67 ]
V
39.6
64.5
GRAM [ 12 ]
V
41.6
74.3
GRAM [ 12 ]
A+V
42.3
76.4
Table 1: Zero-shot video/audio classification results on VGGSound5K. Results from our baseline VAST and the proposed BaryBind , and the performance gains are highlighted accordingly.
Methods
Modality
MSR-VTT
DiDeMo
ActivityNet
VATEX
T2V
V2T
T2V
V2T
T2V
V2T
T2V
V2T
X-CLIP [ 35 ]
T-V
46.1
46.8
45.2
42.3
44.3
42.6
-
-
ImageBind [ 16 ]
T-V
36.8
-
-
-
-
-
-
-
ViCLIP [ 57 ]
T-V
42.4
41.3
18.4
27.9
15.1
24.0
-
-
VideoPrism-b [ 66 ]
T-V
51.4
50.2
-
-
49.6
47.9
62.5
77.1
LanguageBind [ 67 ]
T-V
44.8
40.9
39.9
39.8
41.0
39.1
-
-
Table 2 : Zero-shot text-to-video (T2V) and video-to-text (V2T) retrieval results in terms of Recall at 1 score (R@1). Results from our baseline VAST and BaryBind are highlighted accordingly.
Method
Training modality
T2A
A2T
VAST
32.1
26.1
GRAM
33.2
27.4
Triangle
T-VA
32.2
28.1
BaryBind
35.7
32.5
vs. baseline (VAST)
+3.6
+6.4
VAST
10.4
6.7
Table 3 : Text-to-audio (T2A) and audio-to-text (A2T) R@1 retrieval results on AudioCaps w/ and w/o audio during training .
Method
Training modality
T2A
A2T
VAST
32.1
26.1
GRAM
33.2
27.4
Triangle
T-VA
32.2
28.1
BaryBind
35.7
32.5
vs. baseline (VAST)
+3.6
+6.4
VAST
10.4
6.7
Table 3 : Text-to-audio (T2A) and audio-to-text (A2T) R@1 retrieval results on AudioCaps w/ and w/o audio during training .
Method
Inference modality
Acc@1
Acc@5
VAST
48.1
79.6
GRAM
42.3
76.4
Triangle
A+V
44.8
80.0
BaryBind
55.6
83.4
vs. baseline (VAST)
+7.5
+3.8
VAST
40.8
71.6
Table 4 : Multimodal event classification on VGGSound5K w/ and w/o video during inference when training with audio and video.
Figure 5 : Training dynamics comparison with cosine-based VAST: (a) Training loss comparison between BaryBind and VAST. (b) Recall@1 (R@1) evolution on MSR-VTT under T2V and V2T retrieval tasks. BaryBind achieves higher retrieval accuracy and better modal balance (smaller T2V/V2T R@1 gap). (c) V2T/T2V retrieval gaps throughout training.
Loss function components
Classification
Retrieval
VGGSound
MSR-VTT
TV+TA CL
MWB
BVC
DAM
A
V
V+A
T2V
V2T
✓
✗
✗
✗
38.1
42.8
44.5
46.8
40.1
✓
✗
✗
✓
40.3
43.6
46.3
49.3
43.7
✓
✓
✗
✗
43.6 (+5.5)
45.6
48.6
49.1
47.3 (+7.2)
✓
✓
✗
✓
44.3 (+4.0)
46.8
49.8
50.6
48.8 (+5.1)
Table 5 : Ablation study on loss functions. We report the zero-shot top-1 classification accuracy (Acc@1) on VGGSound5K and top-1 retrieval recall at 1 score (R@1) on MSR-VTT. The performance gains over the baseline are highlighted accordingly.
Batch Size
T2V/V2T R@1
A+C Acc@1
16
55.2/52.6
54.5
32
55.8/53.4
55.1
64
55.9/53.2
55.4
128
56.1/53.5
55.8
256
56.1/53.6
55.6
Table 6 : Ablation on batch size.
Appendix figures & tables20 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6 : The simplex’s volume is correlated ( ρ=−0.9844 ) with the downstream retrieval performance.
Methods
Modality
MSR-VTT
DiDeMo
ActivityNet
VATEX
T2V
V2T
T2V
V2T
T2V
V2T
T2V
V2T
CLIP4Clip [ 32 ]
T-V
45.6
45.9
43.0
43.6
40.3
41.6
63.0
78.3
InternVideo-L [ 59 ]
T-V
53.1
54.4
57.9
59.1
62.2
62.8
69.8
80.6
mPLUG-2 [ 64 ]
T-V
53.1
-
56.4
-
-
-
-
-
ViCLIP [ 57 ]
T-V
52.5
51.8
49.4
50.2
49.8
48.1
-
-
T-MASS [ 54 ]
T-VA
52.7
-
53.3
-
-
-
65.6
-
Appendix
Table 8 : Finetuning text-to-video (T2V) and video-to-text (V2T) retrieval results in terms of Recall at 1 score (R@1). Results from our baseline VAST and BaryBind are highlighted accordingly.
Number of modality n
2
3
4
5
10
20
Pairwise cosine similarity
3.0×10−7
7.0×10−7
1.0×10−6
1.3×10−6
2.9×10−6
4.9×10−6
Barycenter simplex volume
4.9×10−6
6.8×10−6
5.9×10−6
9.8×10−6
3.6×10−5
8.8×10−5
Appendix
Table 9 : Computation time (in seconds) of similarity metrics vs. number of modalities, measured on an NVIDIA A100 GPU with batch size B=64 and embedding dimension D=512 .
Figure 7 : Zero-shot text-to-video retrieval R@1 results as scaling from 2 (T-V) modalities to 5 (T-VASD) modalities.
Anchor Modality
VGGSound
MSR-VTT
V
A
V + A
T2V
V2T
Acc@1
Acc@1
Acc@1
R@1
R@5
R@1
R@5
Video (w/o WB opt.)
45.8
39.6
46.5
50.9
72.6
47.4
74.8
Text (w/o WB opt.)
46.3
40.3
48.1
51.5
72.2
46.8
73.1
T+V+A (w/o WB opt.)
46.1
41.8
50.5
51.2
73.5
46.9
74.5
Video-initialized WB
47.9 (+2.8)
44.8 (+5.2)
55.2 (+8.7)
54.9 (+4.0)
74.6 (+2.0)
52.9 (+5.5)
78.5 (+3.7)
Appendix
Table 10 : Ablation on the anchor selection under the T-VA setting.
Anchor
Modality
Retrieval (R@1)
VideoQA (Acc)
T2V
V2T
MSRVTT-QA
MSVD-QA
T+V+A (w/o WB opt.)
T-VA
51.2
46.9
47.6
55.1
T+V+A-initialized WB
T-VA
56.1
53.6
52.0
60.3
Text-initialized WB
T-VA
56.0
53.2
52.2
60.3
T+V+A+S (w/o WB opt.)
T-VAS
51.7
47.3
50.3
57.9
T+V+A+S-initialized WB
T-VAS
57.1
53.8
55.7
63.8
Appendix
Table 11 : Comparison of arithmetic fusion and WB-based models on MSR-VTT and VideoQA under different modalities
Normalization
Setting
T2V
V2T
(V/2!)1/2
T-V
53.9
52.2
V
T-V
54.1
52.0
(V/3!)1/3
T-VA
54.8
52.9
V
T-VA
55.1
52.7
(V/4!)1/4
T-VAS
57.3
54.2
V
T-VAS
57.2
54.5
Appendix
Table 12 : Zero-shot generalization on MSR-VTT under different normalization strategies for the simplex volume.
α1
0.3
0.6
1
1.5
1.8
T2V (R@1)
56.2
55.7
56.5
56.0
55.3
V2T (R@1)
58.0
57.5
58.3
57.9
56.5
Appendix
Table 13 : Sensitivity analysis of α1 on MSR-VTT validation set.
α1
0.3
0.6
1
1.5
1.8
T2V (R@1)
56.2
55.7
56.5
56.0
55.3
V2T (R@1)
58.0
57.5
58.3
57.9
56.5
Appendix
Table 13 : Sensitivity analysis of α1 on MSR-VTT validation set.
α2
0.02
0.06
0.1
0.14
0.18
T2V (R@1)
55.6
55.8
56.5
56.2
56.0
V2T (R@1)
57.6
57.2
58.3
57.9
57.7
Appendix
Table 14 : Sensitivity analysis of α2 on MSR-VTT validation set.
Methods
MSR-VTT
ActivityNet
T2V
V2T
T2V
V2T
VAST [ 9 ]
35.0
39.1
26.7
25.2
LanguageBind [ 67 ]
31.8
36.4
29.3
26.1
GRAM [ 12 ]
35.7
39.4
27.2
29.1
Triangle [ 11 ]
39.4
40.1
27.8
29.3
BaryBind
44.8
43.2
33.6
34.2
Appendix
Table 15 : Training-from-scratch T2V and V2T results on MSR-VTT and ActivityNet. The training follows the T-VA setting.
Iterations
Mean Lf
Range
Δ
0–500
-0.376
[-1.000,-0.154]
0.846
500–1000
-0.237
[-0.309,-0.171]
0.138
1000–2000
-0.181
[-0.242,-0.134]
0.108
2000–3000
-0.098
[-0.138,-0.045]
0.093
3000–10000
-0.030
[-0.049,-0.011]
0.038
Appendix
Table 16 : Stability analysis of potential optimization during alternating updates.
Constraint
T2V/V2T R@1
A+C Acc@1
Hard constraint
56.1/53.6
55.6
Soft constraint
55.9/53.2
55.1
Appendix
Table 17 : Comparison between hard and soft congruence constraints.
Iterations
Mean LT
Range
Δ
0–500
0.971
[0.913,1.000]
0.087
500–1000
0.602
[0.317,0.913]
0.597
1000–2000
0.136
[0.073,0.330]
0.257
2000–3000
0.052
[0.020,0.095]
0.075
3000–10000
0.029
[0.000,0.022]
0.022
Appendix
Table 18 : Loss analysis of WB map Tθ adaptation during encoder fine-tuning.
Model
Params
Batch size
Forward+Backward
Steps/Epoch
Time/Epoch
VAST
1.28B
64
∼ 8.8s
2344
∼ 5.7h
GRAM
1.30B
64
∼ 9.4s
2344
∼ 6.1h
BaryBind
1.34B
64
∼ 9.7s
2344
∼ 6.3h
Appendix
Table 19 : Computational overhead comparison during training.
Modalities
MSR-VTT
VATEX
VAST
BaryBind
VAST
BaryBind
Text
Video
Audio
Sub.
Depth
T2V
V2T
T2V
V2T
T2V
V2T
T2V
V2T
✓
✓
✗
✗
✗
48.7
43.2
54.2
52.0
78.8
77.0
82.3
79.8
✓
✓
✓
✗
✗
49.3
43.7
56.0
53.2
80.0
77.3
84.2
81.3
✓
✓
✓
✓
✗
50.9
47.9
57.2
53.6
82.1
78.7
84.6
83.5
✓
✓
✓
✓
✓
51.2
49.3
57.9
54.4
82.4
79.2
84.9
83.8
Appendix
Table 22 : Downstream performance with increasing number of modalities.
Method
MSRVTT-QA
MSVD-QA
mPLUG-2
48.0
58.1
VALOR-L
49.2
60.0
VAST
50.1
60.2
GRAM
52.3
61.5
Triangle
52.0
61.9
BaryBind
55.4
63.6
Appendix
Table 23 : Evaluation on VideoQA benchmarks. Methods utilizing audio or subtitle modalities are bolded .
Method
MSRVTT-QA
MSVD-QA
mPLUG-2
48.0
58.1
VALOR-L
49.2
60.0
VAST
50.1
60.2
GRAM
52.3
61.5
Triangle
52.0
61.9
BaryBind
55.4
63.6
Appendix
Table 23 : Evaluation on VideoQA benchmarks. Methods utilizing audio or subtitle modalities are bolded .
Method
T2I
A2I
TA2I
ImageBind
50.1
53.6
46.8
VAST
48.1
57.2
50.6
GRAM
46.7
56.5
45.2
Triangle
46.4
54.5
44.3
BaryBind
43.6
50.3
38.6
Appendix
Table 24 : Cross-modal generation results on VGGSound (FID ↓ ). T2I, A2I, and TA2I denote text-to-image, audio-to-image, and joint text-audio-to-image generation, respectively.
Figure 8 : Visual results of text-to-video retrieval. We display 5 frames from the top-1 video.
Hong Kong University of Science and Technology, Hong Kong · Vrije Universiteit Amsterdam, Netherlands · UiT – The Arctic University of Norway, Norway +3