Contrastive vision-language models learn shared embedding spaces by aligning matched image-text pairs, yet their representations remain separated by a modality gap. Prior work reports divergent effects of modifying this gap: reducing it can improve zero-shot classification and cross-modal alignment, whereas removing gap-related structure can degrade image-text retrieval. In this paper, we provide a unified geometric explanation for these task-dependent effects. Across CLIP and SigLIP encoders, we find that a single dominant direction captures 94.4-99.9% of the squared norm of the image-text mean separation, revealing that the mean-separation component is approximately rank-one. A decomposition of the similarity score then identifies three task-specific roles. In zero-shot classification, query-side fixed gap-offset subtraction is exactly equivalent to an additive class bias. In standard cross-modal retrieval, projecting out the gap direction and renormalising residuals discards candidate-specific norm information, inducing a multiplicative ranking distortion; a geometry-derived exponent tracks the grid-search optimum (Spearman rho = 0.93) and restores performance in some settings, although the gains transfer unevenly. In mixed-modal retrieval, the gap direction sorts candidates by modality; its removal can improve cross-modal ranking, unlike random or non-gap controls. Residual semantic structure after removal defines the limits of the rank-one account. Together, these results explain why gap modification can improve, degrade, or restore performance across downstream settings. By clarifying when and why gap modification changes model behavior, this account provides a principled basis for selecting gap interventions in similarity-based vision-language systems across evaluated downstream tasks.
Figures & tables
Figure 2: Score-term sufficiency is readout-dependent. Cells show five-seed mean prediction or ordering recovery, rather than numerical score error. The algebraic B/C labels are image/text-based: B is candidate-varying for image-to-text retrieval and C for text-to-image, so the retrieval row aggregates direction-specific functional roles. The full score gives exact numerical reconstruction within tolerance; near ties can still change strict decisions.
Figure 3: Geometry predicts configuration-specific scale calibration. A: predicted versus estimate-split fitted exponents for 16 configurations; Spearman ρ=0.930 with a 2,000-resample bootstrap interval. B: evaluation Mean Recall over five paired seeds; error bars show 95% Student- t intervals. The fitted exponent is diagnostic and the shared constant is configuration-invariant.
Method
Binary NDCG@10 ↑
Pool-adj. skew ↓
Within-mod. ρ↑
Standard ret. Δ MR
Unchanged
0.0002
0.9999
1.000
0.000
O2: remove v1
0.2160
0.7952
0.896
−0.071
Remove PC2
0.0013
0.9993
0.872
—
Remove seeded random
0.0002
0.9999
1.000
—
O4: full centering
0.2417
0.8175
0.796
−0.014
Table 2: Binary shared-index retrieval is direction-specific. Means over 48 cells across eight Flickr30k/COCO-5K encoder–dataset pairs, both query modalities, and image:text ratios 25:75, 50:50, and 75:25; five paired seeds per cell; MixBench is excluded. Near-zero unchanged NDCG accompanies severe modality-dominated ranking. Standard-retrieval Δ MR is from a separate matched 16-configuration evaluation.
Check
Evidence
A. Held-out checkpoint
Scalar effect directions
7/7
Operation-effect signs
10/12
Exact task ordering
1/3
Frozen interval coverage
1/7
Scalar / operation-effect MAE
0.257 / 0.092
Table 3: Transfer and boundary. Frozen held-out prediction (A) and residual structure after centering (B).
Vision-language models such as CLIP embed images and text in a shared space, where modality-specific distributions often remain separated. Existing accounts connect this modality gap to initialization, contrastive dynamics, and information imbalance, while its distributional and pairwise contributions to retrieval remain unresolved. We introduce UOT-Gap, a training-free variational diagnostic that models frozen image and text embeddings with unbalanced entropic optimal transport (UOT). The UOT optimum separates transport, coupling complexity, and marginal mass variation; a complementary pair-aware residual compares observed image-caption pairs with the UOT soft matching. On Flickr8K and COCO-1K with frozen CLIP, OpenCLIP, and SigLIP encoders, caption degradation reduces Flickr8K Recall@1 from 0.559 to 0.003. Across six dataset-model conditions, the pair-aware residual tracks retrieval degradation with mean absolute Spearman 0.973, compared with 0.392 for the mean gap. The association remains stable across five random COCO-1K subsets at 0.954±0.026, with a minimum of 0.943. UOT barycentric updates reduce the transport objective while degrading retrieval, distinguishing geometric objective descent from task improvement. These results establish UOT-Gap as a diagnostic for caption quality, modality alignment, and retrieval robustness.
Vision-Language Models (VLMs) achieve strong cross-modal performance, yet recent evidence suggests they over-rely on textual descriptions while under-utilizing visual evidence -- a phenomenon termed ``text shortcut learning.'' We propose an adversarial evaluation framework that quantifies this cross-modal dependency by measuring accuracy degradation (Drop) when semantically conflicting text is paired with unchanged images. Four adversarial strategies -- shape_swap, color_swap, position_swap, and random_text -- are applied to a controlled geometric-shapes dataset (n=1,000). We compare three configurations: Baseline CLIP (ViT-B/32), LoRA fine-tuning, and LoRA Optimized (integrating Hard Negative Mining, Label Smoothing, layer-wise learning rates, Cosine Restarts, curriculum learning, and data augmentation). The optimized model reduces average Drop from 27.5% to 9.8% (64.4% relative improvement, p<0.001) while maintaining 97% normal accuracy. Attention visualization and embedding-space analysis confirm that the optimized model attends more to visual features and achieves tighter cross-modal alignment.
Lijie Zhou
School of Computer Science University of Nottingham Ningbo China Ningbo, China
Vision-language models are evaluated by aggregate accuracy on multimodal benchmarks, a practice that implicitly assumes the model uses its visual input. We show this assumption fails on 40%--97% of samples across six VLMs and three perceptual benchmarks: blurring the question-relevant visual region leaves the next-token distribution nearly unchanged. We name this phenomenon the Visual Insensitivity Gap and quantify it with a per-sample Visual Sensitivity Index (VSI). The gap is a property of samples, not of models: VSI ranks correlate across models (grand-mean Spearman rho=+0.40, permutation p<10^-3), so the same samples are flagged insensitive by VLMs sharing no architectural detail beyond a contrastively pretrained vision tower. The mechanism is concrete: on the insensitive samples, a linear probe on each model's own vision tower distinguishes perturbed from clean images at 0.72--0.79 accuracy, yet the model's argmax token changes on only 2%--11% of the same samples, an encoder--LLM gap above 0.65 on every model. Mapping VSI's diagnostic utility cell by cell surfaces a strong regime (multi-choice reasoning on capable VLMs: AUROC=0.85--0.87) and a weak regime (well-calibrated factuality, where softmax confidence already leads). VSI is not a universal best abstention signal; it is a sample-intrinsic indicator of vision-ignoring failure, best used as a conditional ensemble component.