DiscoVL: Unveiling Disentangled C ross-Modal Representation Learning via Orthogonal Adversarial Regularization for V ision-Language Models
Authors: Mengping Dong, Jinbao Li, Fei Li
Organizations: Shandong Artificial Intelligence Institute, Qilu University of Technology (Shandong Academy of Sciences), Jinan, China · University of Florida, Florida, USA
Pre-trained vision-language models excel across varied perception tasks, but adapting them to novel downstream settings without sacrificing generalization remains non-trivial. Existing parameter-efficient prompt learning method often yields inconsistent representations and fails to account for semantic distribution shifts. In this work, we present DiscoVL, a disentangled cross-modal representation learning framework that couples orthogonal adversarial regularization with structured cross-modal alignment for vision-language models. To address the insufficient cross-modal interaction, our DiscoVL designs a multi-branch low-rank residual aligner that decomposes representations into subspaces and enables bidirectional cross-modal feedback between visual and textual streams at each layer. Furthermore, while conventional triplet constraints overfit features to class centroids, we design an orthogonal regularization for adversarial triplet loss, which prevents centroid collapse and substantially boosts generalization. Evaluations on 15 benchmarks demonstrate that DiscoVL delivers consistent improvements over state-of-the-art methods for base-to-novel generalization, cross-dataset evaluation, and few-shot learning
Figures & tables
Figure 1 : Comparison of the harmonic mean between previous state-of-the-art method and our DiscoVL across 11 diverse datasets for base-to-novel generalization.
Figure 2 : Overview of DiscoVL, which disentangles representations into shared and task-specific subspaces. DCRA separates and aligns cross-modal cues via multi-branch residual structure and bidirectional feedback, while OARL enforces orthogonal adversarial constraints to prevent representation tokens from collapsing onto class centroids.
Figure 3 : Main component. (a) Disentangled cross-modal representation aligner, which contains a multi-branch residual aligner and bi-directional feedback between different image-text modalities. (b) Adversarial triplet regularization. The red dashed box is vanilla adversarial triplet regularization.
Table 1 : Base-to-novel generalization results. HM indicates the harmonic mean. The best and second best results are marked in bold and underline. DiscoVL consistently improves base class performance while preserving generalization to novel classes.
Method
Source
Target dataset
ImageNet
Caltech101
Pets
Cars
Flowers102
Food101
Aircraft
SUN397
DTD
EuroSAT
UCF101
Average
CoOp [ 41 ]
71.51
93.70
89.14
64.51
68.71
85.30
18.47
64.15
41.90
46.39
66.55
63.88
CoCoOp [ 40 ]
71.02
94.43
90.14
65.32
71.88
86.06
22.94
67.36
45.73
45.37
68.21
65.74
MaPLe [ 15 ]
70.72
93.53
90.49
65.57
72.23
86.20
24.74
67.01
46.49
48.06
68.69
66.30
PromptSRC [ 16 ]
71.27
93.60
90.25
65.70
70.25
86.15
23.90
67.10
46.87
45.50
68.75
65.81
TCP [ 37 ]
71.40
93.97
91.25
64.69
71.21
86.69
23.45
67.15
44.35
51.45
68.73
66.29
Table 2 : Comparisons with state-of-the-art methods on cross-dataset evaluation. Bold values indicate the best results. Overall, DiscoVL provides the highest average accuracy, indicating better generalization.
Method
Source
Target dataset
ImageNet
ImageNetV2
ImageNet-Sk
ImageNet-A
ImageNet-R
Average
CLIP [ 26 ]
66.73
60.83
46.15
47.77
73.96
57.18
CoOp [ 41 ]
71.51
64.20
47.99
49.71
75.21
59.28
CoCoOp [ 40 ]
71.02
64.07
48.75
50.63
76.18
59.91
MaPLe [ 15 ]
70.72
64.07
49.15
50.90
76.98
60.27
PromptSRC [ 16 ]
71.27
64.35
49.55
50.90
77.80
60.65
Table 3 : Comparisons with state-of-the-art methods on domain generalization. On an average, DiscoVL achieves consistent improvement.
Figure 4 : Comparison with previous few-shot fine-tuning methods on 11 datasets under different shot settings. Our DiscoVL achieves new state-of-the-art performance.
Table 4 : Ablation study of DiscoVL components (average over 11 datasets) and loss function contributions (StanfordCars).
Figure 5 : t-SNE visualization. Different colors denote different classes.
Figure 6 : Grad-CAM visualization.
Method
ImageNet
EuroSAT
Base
Novel
HM
Base
Novel
HM
Coprompt ⋆
76.97±0.57
71.10±0.00
73.92
93.67±1.08
77.73±8.16
84.96
MMRL
77.90±0.08
71.30±0.28
74.45
95.60±0.33
80.17±5.05
87.21
SkipT. ⋆
77.77±0.09
70.40±0.22
73.87
92.60±1.28
82.10±2.91
87.03
DiscoVL (Ours)
78.07±0.13
71.63±0.05
74.71
96.80±0.21
81.20±2.99
88.32
Table 5 : Training stability comparison, measured by mean ± standard deviation (%). ⋆ means reproduced by ourselves.
Method
Param. (M)
Lat.(ms)
FPS
FLOPs (G)
Train time (min)
HM (%)
MMRL
4.99
6.31
158.44
21.58
156
74.45
DiscoVL (Ours)
2.21 (-2.78)
6.11 (-0.20)
163.68 (+5.24)
21.50 (-0.08)
147 (-9)
74.71 (+0.26)
Table 6 : Efficiency on a single NVIDIA RTX 3060 GPU.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Table S1 : Ablation studies on γ , α and β .
Table 14
Layer index
1
2
3
4
5
6
7
mean
Fusion weights
0.7276
0.7273
0.7228
0.7294
0.7276
0.7276
0.7306
0.7276
Appendix
Table S4 : Fusion weights analysis.
Layer
Cos. Sim. ( ↓ )
Shared Probe Acc. (%)
Task-Specific Probe Acc. (%)
6
0.08
67.42
81.35
7
0.07
68.11
82.04
8
0.09
67.89
82.47
9
0.06
68.53
82.71
10
0.08
68.20
82.13
11
0.07
68.74
81.96
Appendix
Table S5 : Separability across different layers (averaged over 11 datasets). ’Cos. Sim.’ is the mean absolute cosine similarity between shared and task-specific branch features. ’Probe Acc.’ is the base-class accuracy of a frozen linear probe trained on each branch independently.
ImageNet
EuroSAT
UCF101
Setting
Base
Novel
HM
Base
Novel
HM
Base
Novel
HM
Full
78.07
71.63
74.71
96.80
81.20
88.32
88.67
80.63
84.46
\ Task
71.90
70.47
71.18
89.53
79.80
84.39
81.20
79.43
80.31
\ Shared
77.13
67.20
71.82
96.07
76.53
85.19
87.93
76.10
81.59
Appendix
Table S6 : Branch intervention results. ’Full’ denotes the proposed DiscoVL. ’ \ Task’ and ’ \ Shared’ denote the task-specific and shared branches, respectively.