Multimodal large language models (MLLMs) have shown strong potential for universal multimodal representation learning. However, existing methods either compress each input into a single vector, limiting fine-grained expressiveness, or retain long sequences of visual-token vectors, incurring substantial storage and interaction costs. To resolve this trade-off, we propose ResComEmb, a trainable framework for effective and efficient universal multi-vector multimodal embedding. ResComEmb first encodes each input at native dynamic resolution into ordered global, intermediate, and fine-grained views. After MLLM contextualization and embedding projection, a trainable Residual Homogeneity Compression (RHC) module reduces within-granularity redundancy and cross-granularity repetition under explicit visual token budgets. Then, ResComEmb introduces a length-adaptive Bidirectional Late-Interaction Matching mechanism for robust query-document scoring, which averages the strongest token-level matches in each direction and combines the two scores using a weight based on how many valid tokens each side has. Extensive experiments on MMEB, ViDoRe V1, and ViDoRe V2 show that ResComEmb produces higher-quality universal multimodal embeddings than VLM2Vec-V2, and outperforms ColQwen2.5 in visual document retrieval using only 37.5% of its full visual token budget, demonstrating a favorable effectiveness-efficiency trade-off.
Figures & tables
Figure 1: Overview of proposed ResComEmb framework. Ordered multi-granularity views are encoded by a shared MLLM, then compressed into nested representations. The Bidirectional Late-Interaction Matching is applied to compute relevance scores between multimodal queries and targets.
Model
Backbone
Size
Per Meta-Task Score
Average Score
Cls.
VQA
Ret.
Gnd.
IND
OOD
Avg.
Single-Vector Embedding
CLIP
ViT-L
428M
55.2
19.7
53.2
62.2
47.6
42.8
45.4
UniIR
ViT-L
428M
44.3
16.2
61.8
65.3
47.1
41.7
44.7
MagicLens
ViT-L
613M
38.8
8.3
35.4
26.0
–
–
27.8
VLM2Vec
Qwen2-VL
2B
58.7
49.3
65.0
72.9
64.9
53.3
59.7
Table 1: MMEB Precision@1 (%). Cls., Ret., and Gnd. abbreviate classification, retrieval, and grounding; IND/OOD denote in-domain/out-of-domain tasks. Bold and underline mark the best and second-best scores in each column.
Model
Backbone
Size
Arxiv
Doc
Info
TabF
TATQ
Shift
AI
Ener.
Gov.
Hlth.
Avg.
Single-Vector Embedding
ONE-PEACE
ONE-PEACE
4B
43.9
23.4
59.9
57.0
13.4
17.0
45.4
53.2
55.9
59.5
42.9
E5-V
LLaVA-NeXT
8B
41.1
24.3
49.5
58.2
9.0
13.2
46.1
57.7
53.0
59.6
41.2
DSE
Phi-3-Vision
4B
78.1
45.8
82.0
79.2
49.0
69.8
96.8
92.6
92.0
96.3
78.2
GME
Qwen2-VL
2B
82.8
53.1
90.2
93.3
69.9
89.5
97.5
91.9
94.6
98.7
86.2
GME
Qwen2-VL
7B
86.9
57.5
91.6
94.6
74.1
96.8
99.6
95.3
98.8
99.3
89.5
Table 2: ViDoRe V1 NDCG@5 (%) across 10 visual document retrieval tasks. Arxiv, Doc, and Info denote ArxivQ, DocQ, and InfoQ; Ener. and Hlth. abbreviate Energy and Health.
Model
Backbone
Size
ESG Human
Eco Mul
Bio Mul
ESG Syn-Mul
Bio
ESG Syn
Eco
Avg.
Single-Vector Embedding
SigLIP
SigLIP
652M
28.8
14.0
18.2
21.9
33.8
19.8
29.8
23.8
VisRAG-Ret
MiniCPM-V2.0
3B
53.7
48.7
47.7
46.4
54.8
45.9
59.6
51.0
VLM2Vec
Qwen2-VL
7B
33.9
42.0
29.7
38.4
38.8
36.7
51.4
38.7
GME
Qwen2-VL
7B
65.8
56.2
55.1
56.7
64.0
54.3
62.9
59.3
mmE5
Llama-3.2-Vision
11B
52.8
44.3
46.8
54.7
51.3
55.1
48.6
50.5
Table 3: ViDoRe V2 NDCG@5 (%) across 7 tasks. Syn, Mul, and Bio denote synthetic, multilingual, and biomedical data.
Prefix
Variant
ViDoRe
MMEB
V1
V2
Avg.
Cls.
VQA
Ret.
Gnd.
Avg.
E(1)
w/ MRL
89.2
60.8
77.5
62.8
61.2
67.7
86.8
66.7
w/o MRL
86.4
53.2
72.7
58.9
60.0
63.6
85.1
63.7
Δ
(2.8 ↓ )
(7.6 ↓ )
(4.8 ↓ )
(3.9 ↓ )
(1.2 ↓ )
(4.1 ↓ )
(1.7 ↓ )
(3.0 ↓ )
E(2)
w/ MRL
89.6
61.4
78.0
62.8
61.4
68.2
86.8
66.9
w/o MRL
88.6
55.2
74.8
59.2
60.4
64.1
85.3
64.1
Table 4: Effect of MRL across nested prefixes. Prefixes E(1) , E(2) , and E(3) retain 128, 256, and 384 visual tokens, respectively, with identical text tokens; Δ denotes the drop without MRL.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Variant
ArxivQ
DocQ
InfoQ
TabF
TATQ
Shift
AI
Energy
Gov.
Health
Avg.
ResComEmb
88.8
61.7
94.2
95.2
80.4
90.4
99.6
96.6
97.9
99.3
90.4
w/o Importance Score
87.3
63.0
93.6
94.6
80.8
88.9
99.5
96.5
97.5
99.1
90.1
w/o Novelty Score
86.6
62.5
93.8
94.2
80.7
87.4
99.3
96.5
96.5
99.2
89.7
Appendix
Table 5: Ablation of RHC scoring signals on all ViDoRe V1 subsets (NDCG@5 %).
Variant
ESG Human
Eco Mul
Bio Mul
ESG Syn-Mul
Bio
ESG Syn
Eco
Avg.
ResComEmb
69.8
56.0
60.0
57.1
64.8
60.9
62.7
61.6
w/o Importance Score
68.4
54.7
57.3
55.9
60.5
59.6
60.0
59.5
w/o Novelty Score
67.5
51.2
56.8
48.8
62.3
57.3
62.7
58.1
Appendix
Table 6: Ablation of RHC scoring signals on all ViDoRe V2 subsets (NDCG@5 %). Syn, Mul, and Bio denote synthetic, multilingual, and biomedical data.
Variant
ViDoRe
MMEB
V1
V2
Avg.
Cls.
VQA
Ret.
Gnd.
Avg.
Vanilla MaxSim
90.0
59.8
77.6
41.3
15.1
62.7
71.8
44.5
ResComEmb
90.4
61.6
78.5
63.1
61.1
69.5
87.9
67.4
w/o Mean
90.0
59.1
77.3
54.6
41.3
64.5
89.3
58.1
w/o TopK
89.1
54.5
74.9
61.2
55.7
66.7
81.4
63.8
w/o Adaptive Weighting
89.0
54.0
74.6
62.8
61.3
67.4
89.3
66.9
Appendix
Table 7: Late-interaction ablations by benchmark and MMEB category. w/o Adaptive Weighting keeps both directions with fixed 0.5/0.5 weights; w/o Bidirectionality retains only query-to-document TopK-mean MaxSim. ViDoRe reports NDCG@5 (%); MMEB reports Precision@1 (%). ViDoRe Avg. weights V1/V2 by their 10/7 subsets.
Variant
Tokens
ViDoRe
MMEB
V1
V2
Avg.
Cls.
VQA
Ret.
Gnd.
Avg.
ResComEmb
384
90.4
61.6
78.5
63.1
61.1
69.5
87.9
67.4
w/o g1
384
90.0
59.1
77.3
62.1
61.2
66.4
89.5
66.3
w/o g2
384
90.4
61.9
78.7
62.8
61.3
67.5
89.6
66.9
w/o g3
384
90.2
61.5
78.4
62.9
61.2
67.8
89.2
67.0
Appendix
Table 8: Visual-granularity ablations at a fixed total budget of 384 visual tokens. ViDoRe reports NDCG@5 (%); MMEB reports Precision@1 (%). ViDoRe Avg. weights V1/V2 by their 10/7 subsets.
Stage budget
Tokens
ViDoRe
MMEB
V1
V2
Avg.
Cls.
VQA
Ret.
Gnd.
Avg.
16/16/16
48
86.2
55.5
73.6
62.0
60.9
66.7
89.6
66.3
32/32/32
96
88.9
58.6
76.4
62.3
61.3
67.7
89.8
66.9
64/64/64
192
89.5
60.6
77.6
62.6
61.4
68.6
90.2
67.3
128/128/128
384
90.4
61.6
78.5
63.1
61.1
69.5
87.9
67.4
256/256/256
768
90.6
61.7
78.7
63.2
61.1
66.9
88.5
66.7
Appendix
Table 9: Per-stage token-budget ablations for RHC. Stage budgets are written as b1/b2/b3 ; Tokens counts the total visual budget. ViDoRe reports NDCG@5 (%); MMEB reports Precision@1 (%). ViDoRe Avg. weights V1/V2 by their 10/7 subsets.
Multimodal Large Language Models (MLLMs) have shown immense promise in universal multimodal retrieval, which aims to find relevant items of various modalities for a given query. However, their practical application is often hindered by the substantial computational cost incurred from processing a large number of tokens from visual inputs. In this paper, we propose Magic-MM-Embedding, a series of novel models that achieve both high efficiency and state-of-the-art performance in universal multimodal embedding. Our approach is built on two synergistic pillars: (1) a highly efficient MLLM architecture incorporating visual token compression to drastically reduce inference latency and training time, and (2) a multi-stage progressive training strategy designed to not only recover but significantly boost performance. This coarse-to-fine training paradigm begins with extensive continued training to restore multimodal understanding and generation capabilities, progresses to large-scale contrastive pretraining and hard negative mining to enhance discriminative power, and culminates in a task-aware fine-tuning stage guided by an MLLM-as-a-Judge for precise data curation. Comprehensive experiments show that our model outperforms existing methods by a large margin while being more inference-efficient.
Although Multimodal Large Language Models (MLLMs) have shown remarkable potential in Visual Document Retrieval (VDR) through generating high-quality multi-vector embeddings, the substantial storage overhead caused by representing a page with thousands of visual tokens limits their practicality in real-world applications. To address this challenge, we propose an auto-regressive generation approach, CausalEmbed, for constructing multi-vector embeddings. By incorporating iterative margin loss during contrastive training, CausalEmbed encourages the embedding models to learn compact and well-structured representations. Our method enables efficient VDR tasks using only dozens of visual tokens, achieving a 30-155x reduction in token count while maintaining highly competitive performance across various backbones and benchmarks. Theoretical analysis and empirical results demonstrate the unique advantages of auto-regressive embedding generation in terms of training efficiency and scalability at test time. As a result, CausalEmbed introduces a flexible test-time scaling strategy for multi-vector VDR representations and sheds light on the generative paradigm within multimodal document retrieval. Our code is available at https://github.com/Z1zs/Causal-Embed.
Jiahao Huo, Yu Huang, Yibo Yan +7
The Hong Kong University of Science and Technology (Guangzhou) · Alibaba Cloud Computing · The Hong Kong University of Science and Technology +1
Universal multimodal retrieval requires compact embeddings that preserve task-relevant semantic information across diverse modalities. Prior works have incorporated latent reasoning into multimodal embedding learning to refine this information before embedding extraction. However, most existing approaches remain confined to deterministic latent paths, without exploring alternative trajectories to discover better embeddings. Thus, we propose VaME (Variational Multimodal Embeddings), a framework that models latent reasoning as a learnable distribution over trajectories. Specifically, we first introduce Variational Latent Reasoning (VLR) to enable autoregressive exploration in latent space, guided by answer reconstruction through a lightweight decoder. Meanwhile, we augment the original embedding-token readout with a latent-fused embedding to facilitate exploration during subsequent reinforcement learning. Finally, we optimize latent reasoning over stochastic variational trajectories through reinforcement learning, using Semantic Decoding Reward (SDR) to favor semantically meaningful trajectories with interpretable decoded outcomes. On the 78-task MMEB-V2 benchmark, spanning image, video, and visual-document retrieval, VaME outperforms most explicit CoT-based models and all latent-reasoning baselines. VaME also demonstrates robust performance on reasoning-intensive benchmarks such as MRMR, with substantial gains after reinforcement learning. Importantly, VaME achieves these gains with at least a 4.25x inference speedup over the deterministic latent autoregressive baselines. The code will be made publicly available.
Peixi Wu, Mingzhou Jiang, Feipeng Ma +11
University of Science and Technology of China · Tsinghua University · Kuaishou +1