MCA: Modality Composition Awareness for Robust Composed Multimodal Retrieval
Organizations: Sony Group Corporation · Sony AI
Abstract
Multimodal retrieval, which seeks to retrieve relevant content across modalities such as text or image, supports applications from AI search to contents production. Despite the success of separate-encoder approaches like CLIP aligning modality-specific embeddings with contrastive learning, recent multimodal large language models (MLLMs) enable a unified encoder that directly processes composed inputs. While flexible and advanced, we identify that unified encoders trained with conventional contrastive learning are prone to learn modality shortcut, leading to poor robustness under distribution shifts. We propose a modality composition awareness framework to mitigate this issue. Concretely, it consists of a preference loss enforces multimodal embeddings to outperform their unimodal counterparts, and a composition regularization objective aligns multimodal embeddings with prototypes composed from its unimodal parts. These objectives explicitly model structural relationships between the composed representation and its unimodal counterparts. Experiments on various benchmarks show gains in out-of-distribution retrieval, highlighting modality composition awareness as a effective principle for robust composed multimodal retrieval when utilizing MLLMs as the unified encoder.
Figures & tables
| In-domain Retrieval | OOD Retrieval | Zero-shot Grounding | ||||||||||||||||
| Method | VisDial | CIRR | VisualNews_t2i | VisualNews_i2t | MSCOCO_t2i | MSCOCO_i2t | NIGHTS | WebQA | Avg. | OVEN | FashionIQ | EDIS | Avg. | MSCOCO | Visual7W-P | RefCOCO | RefCOCO-M | Avg. |
| Dual-Encoder Baselines | ||||||||||||||||||
| CLIP ( Radford et al., 2021 ) | 30.7 | 12.6 | 78.9 | 79.6 | 59.5 | 57.7 | 60.4 | 67.5 | 55.8 | 41.1 | 11.4 | 81.0 | 44.5 | 33.8 | 55.1 | 56.9 | 61.3 | 51.7 |
| OpenCLIP ( Cherti et al., 2023 ) | 25.4 | 15.4 | 74.0 | 78.0 | 63.6 | 62.1 | 66.1 | 62.1 | 55.8 | 45.0 | 13.8 | 77.5 | 45.4 | 34.5 | 56.3 | 54.2 | 68.3 | 53.3 |
| SigLIP ( Zhai et al., 2023 ) | 21.5 | 15.1 | 51.0 | 52.4 | 58.3 | 55.0 | 62.9 | 58.1 | 46.7 | 56.0 | 20.1 | 23.6 | 33.2 | 46.4 | 70.1 | 70.8 | 50.8 | 59.5 |
| BLIP2 ( Li et al., 2023 ) | 18.0 | 9.8 | 48.1 | 13.5 | 53.7 | 20.3 | 56.5 | 55.4 | 39.5 | 39.3 | 9.3 | 54.4 | 34.4 | 28.9 | 52.0 | 47.4 | 59.5 | 46.9 |
| MCA Ratio | IND Avg. | OOD Avg. | OOD Retrieval Avg. | Zero-shot Grounding Avg. |
|---|---|---|---|---|
| Low resolution (128 128) | ||||
| 0 | 69.7 | 46.7 | 42.5 | 49.9 |
| 0.01 | 69.5 (-0.3%) | 44.0 (-5.8%) | 37.6 (-11.6%) | 48.8 (-2.2%) |
| 0.10 | 69.3 (-0.6%) | 49.5 (+5.9%) | 46.8 (+10.1%) | 51.4 (+3.0%) |
| 1.00 | 69.5 (-0.3%) | 50.9 (+8.9%) | 48.1 (+13.1%) | 52.9 (+6.0%) |
| Mid resolution (672 672) | ||||
| Mixer | In-domain Avg. | OOD Avg. | OOD Retrieval Avg. | Zero-shot Grounding Avg. |
|---|---|---|---|---|
| Vanilla CL | 73.2 | 56.3 | 54.2 | 57.9 |
| Mean Pooling | 73.3 (+0.1%) | 56.5 (+0.3%) | 54.8 (+1.1%) | 57.8 (-0.2%) |
| Gated Fusion | 73.2 (+0.0%) | 59.4 (+5.5%) | 57.4 (+5.9%) | 61.0 (+5.4%) |
| MFB | 73.0 (-0.2%) | 57.7 (+2.4%) | 54.6 (+0.7%) | 60.0 (+3.6%) |
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
| Dataset | MSCOCO | VisualNews | VisDial | CIRR | NIGHTS | WebQA | ||
|---|---|---|---|---|---|---|---|---|
| Input modalities | I T | T I | I T | T I | T I | (I+T) I | I I | T (I+T) |
| # of training pairs | 113K | 100K | 100K | 100K | 123K | 26K | 16K | 17K |
| Method | In-domain Retrieval | OOD Retrieval | Zero-shot Grounding | |||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| VisDial | CIRR | VN-t2i | VN-i2t | COCO-t2i | COCO-i2t | NIGHTS | WebQA | Avg. | OVEN | FashionIQ | EDIS | Avg. | MSCOCO | Visual7W-P | RefCOCO | RefCOCO-M | Avg. | |
| Baseline | 81.3 | 51.6 | 75.7 | 78.6 | 73.9 | 71.9 | 65.6 | 86.8 | 73.2 | 63.0 | 15.6 | 84.0 | 54.2 | 41.0 | 60.7 | 60.3 | 69.6 | 57.9 |
| CL+MCP | 80.6 | 51.9 | 74.6 | 78.4 | 74.8 | 73.0 | 65.6 | 86.9 | 73.2 | 64.1 | 17.3 | 84.4 | 55.3 | 42.1 | 64.7 | 59.0 | 66.8 | 58.2 |
| CL+MCR | 80.6 | 49.2 | 76.5 | 78.7 | 75.1 | 71.6 | 66.2 | 88.0 | 73.2 | 63.5 | 18.0 | 84.6 | 55.4 | 41.2 | 61.1 | 62.9 | 69.9 | 58.8 |
| MCA | 80.3 | 50.5 | 76.3 | 78.7 | 74.6 | 72.1 | 66.1 | 86.8 | 73.2 | 66.7 | 19.3 | 86.1 | 57.4 | 42.7 | 65.4 | 65.0 | 70.9 | 61.0 |
| # Tasks | Vanilla CL | MCA | ||
|---|---|---|---|---|
| Overall | 67 | 52.09 | 52.91 | +0.82 |
| Composed | 28 | 55.78 | 56.96 | +1.18 |
| Non-composed | 39 | 49.45 | 50.01 | +0.56 |
| Image composed | 18 | 61.90 | 62.81 | +0.91 |
| Video composed | 8 | 43.87 | 45.59 | +1.72 |
| Audiovisual composed | 2 | 48.33 | 49.80 | +1.47 |
| Mixer | In-domain | OOD | OOD Retrieval | Zero-shot |
|---|---|---|---|---|
| Avg. | Avg. | Avg. | Grounding Avg. | |
| Low resolution | ||||
| Qwen2-VL-7B | 62.8 | 43.6 | 38.2 | 47.7 |
| MCA-7B | 71.4 (+8.6) | 55.2 (+11.6) | 49.7 (+11.5) | 59.3 (+11.6) |
| Mid resolution | ||||
| Qwen2-VL-7B | 76.1 | 62.6 | 60.1 | 64.5 |
| Model / checkpoint | Composition margin | Held-out MCP loss | ||
|---|---|---|---|---|
| Value | vs. base | Value | vs. base | |
| Base model, step 0 | 0.1744 | – | 0.0667 | – |
| Vanilla CL, final | 0.1248 | 0.0368 | ||
| MCA, final | 0.2125 | 0.0100 | ||