DiscoVL: Unveiling Disentangled C ross-Modal Representation Learning via Orthogonal Adversarial Regularization for V ision-Language Models
Authors: Mengping Dong, Jinbao Li, Fei Li
Organizations: Shandong Artificial Intelligence Institute, Qilu University of Technology (Shandong Academy of Sciences), Jinan, China · University of Florida, Florida, USA
Pre-trained vision-language models excel across varied perception tasks, but adapting them to novel downstream settings without sacrificing generalization remains non-trivial. Existing parameter-efficient prompt learning method often yields inconsistent representations and fails to account for semantic distribution shifts. In this work, we present DiscoVL, a disentangled cross-modal representation learning framework that couples orthogonal adversarial regularization with structured cross-modal alignment for vision-language models. To address the insufficient cross-modal interaction, our DiscoVL designs a multi-branch low-rank residual aligner that decomposes representations into subspaces and enables bidirectional cross-modal feedback between visual and textual streams at each layer. Furthermore, while conventional triplet constraints overfit features to class centroids, we design an orthogonal regularization for adversarial triplet loss, which prevents centroid collapse and substantially boosts generalization. Evaluations on 15 benchmarks demonstrate that DiscoVL delivers consistent improvements over state-of-the-art methods for base-to-novel generalization, cross-dataset evaluation, and few-shot learning
Figures & tables
Figure 1 : Comparison of the harmonic mean between previous state-of-the-art method and our DiscoVL across 11 diverse datasets for base-to-novel generalization.
Figure 2 : Overview of DiscoVL, which disentangles representations into shared and task-specific subspaces. DCRA separates and aligns cross-modal cues via multi-branch residual structure and bidirectional feedback, while OARL enforces orthogonal adversarial constraints to prevent representation tokens from collapsing onto class centroids.
Figure 3 : Main component. (a) Disentangled cross-modal representation aligner, which contains a multi-branch residual aligner and bi-directional feedback between different image-text modalities. (b) Adversarial triplet regularization. The red dashed box is vanilla adversarial triplet regularization.
Table 1 : Base-to-novel generalization results. HM indicates the harmonic mean. The best and second best results are marked in bold and underline. DiscoVL consistently improves base class performance while preserving generalization to novel classes.
Method
Source
Target dataset
ImageNet
Caltech101
Pets
Cars
Flowers102
Food101
Aircraft
SUN397
DTD
EuroSAT
UCF101
Average
CoOp [ 41 ]
71.51
93.70
89.14
64.51
68.71
85.30
18.47
64.15
41.90
46.39
66.55
63.88
CoCoOp [ 40 ]
71.02
94.43
90.14
65.32
71.88
86.06
22.94
67.36
45.73
45.37
68.21
65.74
MaPLe [ 15 ]
70.72
93.53
90.49
65.57
72.23
86.20
24.74
67.01
46.49
48.06
68.69
66.30
PromptSRC [ 16 ]
71.27
93.60
90.25
65.70
70.25
86.15
23.90
67.10
46.87
45.50
68.75
65.81
TCP [ 37 ]
71.40
93.97
91.25
64.69
71.21
86.69
23.45
67.15
44.35
51.45
68.73
66.29
Table 2 : Comparisons with state-of-the-art methods on cross-dataset evaluation. Bold values indicate the best results. Overall, DiscoVL provides the highest average accuracy, indicating better generalization.
Method
Source
Target dataset
ImageNet
ImageNetV2
ImageNet-Sk
ImageNet-A
ImageNet-R
Average
CLIP [ 26 ]
66.73
60.83
46.15
47.77
73.96
57.18
CoOp [ 41 ]
71.51
64.20
47.99
49.71
75.21
59.28
CoCoOp [ 40 ]
71.02
64.07
48.75
50.63
76.18
59.91
MaPLe [ 15 ]
70.72
64.07
49.15
50.90
76.98
60.27
PromptSRC [ 16 ]
71.27
64.35
49.55
50.90
77.80
60.65
Table 3 : Comparisons with state-of-the-art methods on domain generalization. On an average, DiscoVL achieves consistent improvement.
Figure 4 : Comparison with previous few-shot fine-tuning methods on 11 datasets under different shot settings. Our DiscoVL achieves new state-of-the-art performance.
Table 4 : Ablation study of DiscoVL components (average over 11 datasets) and loss function contributions (StanfordCars).
Figure 5 : t-SNE visualization. Different colors denote different classes.
Figure 6 : Grad-CAM visualization.
Method
ImageNet
EuroSAT
Base
Novel
HM
Base
Novel
HM
Coprompt ⋆
76.97±0.57
71.10±0.00
73.92
93.67±1.08
77.73±8.16
84.96
MMRL
77.90±0.08
71.30±0.28
74.45
95.60±0.33
80.17±5.05
87.21
SkipT. ⋆
77.77±0.09
70.40±0.22
73.87
92.60±1.28
82.10±2.91
87.03
DiscoVL (Ours)
78.07±0.13
71.63±0.05
74.71
96.80±0.21
81.20±2.99
88.32
Table 5 : Training stability comparison, measured by mean ± standard deviation (%). ⋆ means reproduced by ourselves.
Method
Param. (M)
Lat.(ms)
FPS
FLOPs (G)
Train time (min)
HM (%)
MMRL
4.99
6.31
158.44
21.58
156
74.45
DiscoVL (Ours)
2.21 (-2.78)
6.11 (-0.20)
163.68 (+5.24)
21.50 (-0.08)
147 (-9)
74.71 (+0.26)
Table 6 : Efficiency on a single NVIDIA RTX 3060 GPU.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Table S1 : Ablation studies on γ , α and β .
Table 14
Layer index
1
2
3
4
5
6
7
mean
Fusion weights
0.7276
0.7273
0.7228
0.7294
0.7276
0.7276
0.7306
0.7276
Appendix
Table S4 : Fusion weights analysis.
Layer
Cos. Sim. ( ↓ )
Shared Probe Acc. (%)
Task-Specific Probe Acc. (%)
6
0.08
67.42
81.35
7
0.07
68.11
82.04
8
0.09
67.89
82.47
9
0.06
68.53
82.71
10
0.08
68.20
82.13
11
0.07
68.74
81.96
Appendix
Table S5 : Separability across different layers (averaged over 11 datasets). ’Cos. Sim.’ is the mean absolute cosine similarity between shared and task-specific branch features. ’Probe Acc.’ is the base-class accuracy of a frozen linear probe trained on each branch independently.
ImageNet
EuroSAT
UCF101
Setting
Base
Novel
HM
Base
Novel
HM
Base
Novel
HM
Full
78.07
71.63
74.71
96.80
81.20
88.32
88.67
80.63
84.46
\ Task
71.90
70.47
71.18
89.53
79.80
84.39
81.20
79.43
80.31
\ Shared
77.13
67.20
71.82
96.07
76.53
85.19
87.93
76.10
81.59
Appendix
Table S6 : Branch intervention results. ’Full’ denotes the proposed DiscoVL. ’ \ Task’ and ’ \ Shared’ denote the task-specific and shared branches, respectively.
Vision-Language Models (VLMs) achieve strong cross-modal performance, yet recent evidence suggests they over-rely on textual descriptions while under-utilizing visual evidence -- a phenomenon termed ``text shortcut learning.'' We propose an adversarial evaluation framework that quantifies this cross-modal dependency by measuring accuracy degradation (Drop) when semantically conflicting text is paired with unchanged images. Four adversarial strategies -- shape_swap, color_swap, position_swap, and random_text -- are applied to a controlled geometric-shapes dataset (n=1,000). We compare three configurations: Baseline CLIP (ViT-B/32), LoRA fine-tuning, and LoRA Optimized (integrating Hard Negative Mining, Label Smoothing, layer-wise learning rates, Cosine Restarts, curriculum learning, and data augmentation). The optimized model reduces average Drop from 27.5% to 9.8% (64.4% relative improvement, p<0.001) while maintaining 97% normal accuracy. Attention visualization and embedding-space analysis confirm that the optimized model attends more to visual features and achieves tighter cross-modal alignment.
Lijie Zhou
School of Computer Science University of Nottingham Ningbo China Ningbo, China
Prompt learning has become a dominant paradigm for adapting vision-language models (VLMs) such as CLIP to downstream tasks without modifying pretrained weights. While extending prompts to both vision and text encoders across multiple transformer layers significantly boosts performance, it dramatically increases the number of trainable parameters, with state-of-the-art methods requiring millions of parameters and abandoning the parameter efficiency that makes prompt tuning attractive. In this work, we propose MMLoP (Multi-Modal Low-Rank Prompting), a framework that achieves deep multi-modal prompting with only 11.5K trainable parameters, comparable to early text-only methods like CoOp. MMLoP parameterizes vision and text prompts at each transformer layer through a low-rank factorization that constrains prompts to a compact subspace, providing parameter efficiency while motivating the need for our complementary regularization components. To further close the accuracy gap with state-of-the-art methods, we introduce three complementary components: a self-regulating consistency loss that anchors prompted representations to frozen zero-shot CLIP features at both the feature and logit levels, a uniform drift correction that removes the global embedding shift induced by prompt tuning to preserve class-discriminative structure, and a shared up-projection that couples vision and text prompts through a common low-rank factor to enforce cross-modal alignment. Extensive experiments across three benchmarks and 11 diverse datasets demonstrate that MMLoP achieves a highly favorable accuracy-efficiency tradeoff, outperforming the majority of existing methods including those with orders of magnitude more parameters, while achieving a harmonic mean of 79.70% on base-to-novel generalization. Code is available at https://github.com/sajjad-ucsb/MMLoP.
Recent advances in pre-training vision-language models (VLMs), e.g., contrastive language-image pre-training (CLIP) methods, have shown great potential in learning out-of-distribution (OOD) representations. Despite showing competitive performance, the prompt-based CLIP methods still suffer from: i) inaccurate text descriptions, which leads to degraded accuracy and robustness, and poses a challenge for zero-shot CLIP methods. ii) limited vision-language embedding alignment, which is one important factor affecting generalization performance. To tackle the above issues, this paper proposes a novel Conditional Domain prompt Learning (CoDoL) method, which utilizes readily-available domain information to form prompts and contributes to improved vision-language embedding alignment, which we identify as one factor underlying the observed OOD generalization gains. To capture both instance-specific and domain-specific information, we further propose a lightweight Domain Meta Network (DMN) to generate input-conditional tokens for images in each domain. Extensive experiments on four OOD benchmarks (PACS, VLCS, OfficeHome, and DigitDG) validate the effectiveness of our proposed CoDoL method in terms of empirically improves vision-language embedding alignment across four DG benchmarks, which we present as a contributing factor (rather than the sole cause) of the observed OOD gains.
Min Zhang, Yuyin Wang, Zhongxiang Dai +4
East China Normal University · Xidian University · Zhejiang Key Laboratory of Artificial Intelligence of Things (AIoT) Network and Data Security +3