Multimodal retrieval, which seeks to retrieve relevant content across modalities such as text or image, supports applications from AI search to contents production. Despite the success of separate-encoder approaches like CLIP aligning modality-specific embeddings with contrastive learning, recent multimodal large language models (MLLMs) enable a unified encoder that directly processes composed inputs. While flexible and advanced, we identify that unified encoders trained with conventional contrastive learning are prone to learn modality shortcut, leading to poor robustness under distribution shifts. We propose a modality composition awareness framework to mitigate this issue. Concretely, it consists of a preference loss enforces multimodal embeddings to outperform their unimodal counterparts, and a composition regularization objective aligns multimodal embeddings with prototypes composed from its unimodal parts. These objectives explicitly model structural relationships between the composed representation and its unimodal counterparts. Experiments on various benchmarks show gains in out-of-distribution retrieval, highlighting modality composition awareness as a effective principle for robust composed multimodal retrieval when utilizing MLLMs as the unified encoder.
Figures & tables
Figure 1: Illustration of modality shortcut problem in multimodal retrieval. Given the inputs consists of an image and textual instruction, left : baseline model trained with conventional CL loss retrieve a visually similar image, focusing too much on vision modality; right : model trained with modality composition awareness retrieves the relevant image following the instruction.
Figure 2: Modality Composition Awareness (MCA, § 3.2 ). Unimodal parts are shown as dotted circles when not explicitly produced. (a) Training with vanilla CL: the composed embedding can be close to the target but still align disproportionately with one unimodal part, leading to modality shortcuts. (b) Training with CL and MCA: the composed embedding is explicitly constrained to be closer to the target than any of its unimodal parts by MCP (§ 3.2.1 ), and anchored to a compositional prototype mixed (Equation 5 ) from unimodal embeddings by MCR (§ 3.2.2 ) to mitigate modality shortcut.
In-domain Retrieval
OOD Retrieval
Zero-shot Grounding
Method
VisDial
CIRR
VisualNews_t2i
VisualNews_i2t
MSCOCO_t2i
MSCOCO_i2t
NIGHTS
WebQA
Avg.
OVEN
FashionIQ
EDIS
Avg.
MSCOCO
Visual7W-P
RefCOCO
RefCOCO-M
Avg.
Dual-Encoder Baselines
CLIP ( Radford et al., 2021 )
30.7
12.6
78.9
79.6
59.5
57.7
60.4
67.5
55.8
41.1
11.4
81.0
44.5
33.8
55.1
56.9
61.3
51.7
OpenCLIP ( Cherti et al., 2023 )
25.4
15.4
74.0
78.0
63.6
62.1
66.1
62.1
55.8
45.0
13.8
77.5
45.4
34.5
56.3
54.2
68.3
53.3
SigLIP ( Zhai et al., 2023 )
21.5
15.1
51.0
52.4
58.3
55.0
62.9
58.1
46.7
56.0
20.1
23.6
33.2
46.4
70.1
70.8
50.8
59.5
BLIP2 ( Li et al., 2023 )
18.0
9.8
48.1
13.5
53.7
20.3
56.5
55.4
39.5
39.3
9.3
54.4
34.4
28.9
52.0
47.4
59.5
46.9
Table 1: Overall results. † indicates the baselines reproduced with identical settings, backbone LLMs, and training data, and Δ indicates the difference compared to the completely comparable baselines.
Figure 3: Breakdown of MCA though loss ablation.
Figure 4: Convergence analysis of MCA. For a fair comparison, the y-axis range for the loss is set to 0–1, and the range for benchmark performance is consistently set to 20.
MCA Ratio
IND Avg.
OOD Avg.
OOD Retrieval Avg.
Zero-shot Grounding Avg.
Low resolution (128 × 128)
0
69.7
46.7
42.5
49.9
0.01
69.5 (-0.3%)
44.0 (-5.8%)
37.6 (-11.6%)
48.8 (-2.2%)
0.10
69.3 (-0.6%)
49.5 (+5.9%)
46.8 (+10.1%)
51.4 (+3.0%)
1.00
69.5 (-0.3%)
50.9 (+8.9%)
48.1 (+13.1%)
52.9 (+6.0%)
Mid resolution (672 × 672)
Table 2: Impact of weighting MCA on varying input image resolution. Darker colors denote larger deviations from the vanilla CL.
Figure 5: Qualitative examples. The central part of the retrieved image by MCA for #6 is zoomed in for a better visibility.
Mixer
In-domain Avg.
OOD Avg.
OOD Retrieval Avg.
Zero-shot Grounding Avg.
Vanilla CL
73.2
56.3
54.2
57.9
Mean Pooling
73.3 (+0.1%)
56.5 (+0.3%)
54.8 (+1.1%)
57.8 (-0.2%)
Gated Fusion
73.2 (+0.0%)
59.4 (+5.5%)
57.4 (+5.9%)
61.0 (+5.4%)
MFB
73.0 (-0.2%)
57.7 (+2.4%)
54.6 (+0.7%)
60.0 (+3.6%)
Table 3: Impact of mixer. Darker colors denote larger deviations from the vanilla CL.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6: A conceptual diagram showing how OOD benchmarks evaluate modality shortcut.
Dataset
MSCOCO
VisualNews
VisDial
CIRR
NIGHTS
WebQA
Input modalities
I → T
T → I
I → T
T → I
T → I
(I+T) → I
I → I
T → (I+T)
# of training pairs
113K
100K
100K
100K
123K
26K
16K
17K
Appendix
Table 4: Statistics of training datasets. I and T are the abbreviation of image and text, respectively.
Figure 7: Visualizations of queries and targets.
Method
In-domain Retrieval
OOD Retrieval
Zero-shot Grounding
VisDial
CIRR
VN-t2i
VN-i2t
COCO-t2i
COCO-i2t
NIGHTS
WebQA
Avg.
OVEN
FashionIQ
EDIS
Avg.
MSCOCO
Visual7W-P
RefCOCO
RefCOCO-M
Avg.
Baseline
81.3
51.6
75.7
78.6
73.9
71.9
65.6
86.8
73.2
63.0
15.6
84.0
54.2
41.0
60.7
60.3
69.6
57.9
CL+MCP
80.6
51.9
74.6
78.4
74.8
73.0
65.6
86.9
73.2
64.1
17.3
84.4
55.3
42.1
64.7
59.0
66.8
58.2
CL+MCR
80.6
49.2
76.5
78.7
75.1
71.6
66.2
88.0
73.2
63.5
18.0
84.6
55.4
41.2
61.1
62.9
69.9
58.8
MCA
80.3
50.5
76.3
78.7
74.6
72.1
66.1
86.8
73.2
66.7
19.3
86.1
57.4
42.7
65.4
65.0
70.9
61.0
Appendix
Table 5: Per-dataset comparison of different training objectives.
# Tasks
Vanilla CL
MCA
Δ
Overall
67
52.09
52.91
+0.82
Composed
28
55.78
56.96
+1.18
Non-composed
39
49.45
50.01
+0.56
Image composed
18
61.90
62.81
+0.91
Video composed
8
43.87
45.59
+1.72
Audiovisual composed
2
48.33
49.80
+1.47
Appendix
Table 6: Omni-modal generalization results on 67 tasks. We report the macro-average Accuracy@1 across tasks. Δ denotes the absolute improvement in percentage points.
Mixer
In-domain
OOD
OOD Retrieval
Zero-shot
Avg.
Avg.
Avg.
Grounding Avg.
Low resolution
Qwen2-VL-7B
62.8
43.6
38.2
47.7
MCA-7B
71.4 (+8.6)
55.2 (+11.6)
49.7 (+11.5)
59.3 (+11.6)
Mid resolution
Qwen2-VL-7B
76.1
62.6
60.1
64.5
Appendix
Table 7: Category-level average results on Qwen2-VL-7B. OOD Avg. is computed over all seven OOD retrieval and zero-shot grounding tasks. Values in parentheses denote absolute improvements over the corresponding Qwen2-VL-7B baseline.
Model / checkpoint
Composition margin ↑
Held-out MCP loss ↓
Value
Δ vs. base
Value
Δ vs. base
Base model, step 0
0.1744
–
0.0667
–
Vanilla CL, final
0.1248
−28.4%
0.0368
−44.8%
MCA, final
0.2125
+21.8%
0.0100
−85.0%
Appendix
Table 8: Held-out modality-shortcut diagnostics on 200 composed test examples from CIRR and NIGHTS. A larger composition margin indicates clearer separation from the strongest unimodal counterpart, while a lower held-out MCP loss indicates a stronger preference for the composed representation.
Composed Image Retrieval (CIR) is a multimodal retrieval task where a query consists of a reference image and a textual modification, and the goal is to retrieve a target image satisfying both. In principle, strong performance on CIR benchmarks is assumed to require multimodal composition, i.e., combining complementary information from reference image and textual modification. In this work, we show that this assumption does not always hold. Across four widely used CIR benchmarks and eleven Generalist Multimodal Embedding models, a large fraction of queries can be solved using a single modality (from 32.2% to 83.6%), revealing pervasive unimodal shortcuts. Thus, high CIR performance can arise from unimodal signals rather than true multimodal composition. To better understand this issue, we perform a two-stage audit. First, we identify shortcut-solvable queries through cross-model analysis. Second, we conduct human validation on 4,741 shortcut-free queries, of which only 1,689 are well-formed, with common issues including ambiguous edits and mismatched targets. Re-evaluating models on this validated subset reveals qualitatively different behaviour: queries can no longer be solved with a single modality, and successful retrieval requires combining both inputs. While accuracy decreases, reliance on multimodal information increases. Overall, current CIR benchmarks conflate shortcut-solvable, noisy, and genuinely compositional queries, leading to an overestimation of model capability in multimodal composition.
Matteo Attimonelli, Alessandro De Bellis, Aryo Pradipta Gema +8
Unified multimodal retrieval aims to identify candidates that satisfy complex user intent expressed through heterogeneous inputs. Although Large Vision-Language Model (LVLM)-based retrievers are efficient and scalable, directly encoding raw multimodal inputs often misses fine-grained discriminative cues, leading to confusion among semantically similar candidates. Recent methods mitigate this limitation by generating Chain-of-Thought (CoT) rationales to enrich the query representation. However, such reasoning is typically derived from the query alone: it explains what the query describes, but not what the retriever misunderstands. We argue that effective retrieval reasoning should instead be conditioned on retrieval feedback. Based on this insight, we introduce UniME-R1, an embedder-adviser framework that learns to reason over initially retrieved candidates and generate Retrieval-Centric Chain-of-Thought (RC-CoT). The adviser analyzes candidates individually to identify the discriminative cues confused by the embedder. If the target appears in the initial top-k set, UniME-R1 directly reranks the candidates; otherwise, it generates RC-CoT to refine the retrieval direction and performs full-corpus re-retrieval with a dual-mode embedder. To train the framework, we mine hard negatives to simulate realistic retrieval failures, jointly optimize direct retrieval and RC-CoT-augmented retrieval, and align the adviser with retrieval outcomes through supervised learning and retrieval-oriented reinforcement learning. Extensive experiments on MMEB-V2 and a diverse set of general multimodal retrieval benchmarks demonstrate that UniME-R1 consistently improves retrieval performance over strong baselines.
Multi-modal retrieval has become increasingly critical for handling the growing volume of integrated visual-textual data in real-world applications, but existing frameworks rely on implicit fusion via text encoder self-attention, limiting explicit cross-modal semantic alignment. To address this gap, this paper proposes UniCA (Unified Cross-Attention Encoder), a multi-modal retrieval model with four key innovations: 1) a bi-directional cross-attention (Bi-CA) block that enables active semantic exchange between visual and textual tokens prior to concatenation, capturing inter-modal correlations more efficiently. 2) a Positive Similarity Loss that optimizes absolute semantic proximity between query and positive candidate embeddings. 3) a streamlined dataset UMR-S10 (Universal Multimodal Retrieval Sample 10%) to reduce computational costs while retaining semantic diversity and task representativeness. 4) an experimental validation on the WebQA benchmark demonstrates that UniCA outperforms the baseline model across Hybrid and Image-Text tasks, achieving improvements of up to 4.09% in Recall@5, 3.28% in Recall@10, and 3.96% in MRR@1 for the hybrid task. UniCA provides an efficient and robust solution for multi-modal retrieval, lowering deployment barriers through its lightweight dataset and enhanced fusion mechanism.
Yini Huang, Wenlong Zhang
The Hong Kong University of Science and Technology (Guangzhou), Guangdong, China · Southern Medical University, Guangdong, China