Multimodal large language models (MLLMs) have shown strong potential for universal multimodal representation learning. However, existing methods either compress each input into a single vector, limiting fine-grained expressiveness, or retain long sequences of visual-token vectors, incurring substantial storage and interaction costs. To resolve this trade-off, we propose ResComEmb, a trainable framework for effective and efficient universal multi-vector multimodal embedding. ResComEmb first encodes each input at native dynamic resolution into ordered global, intermediate, and fine-grained views. After MLLM contextualization and embedding projection, a trainable Residual Homogeneity Compression (RHC) module reduces within-granularity redundancy and cross-granularity repetition under explicit visual token budgets. Then, ResComEmb introduces a length-adaptive Bidirectional Late-Interaction Matching mechanism for robust query-document scoring, which averages the strongest token-level matches in each direction and combines the two scores using a weight based on how many valid tokens each side has. Extensive experiments on MMEB, ViDoRe V1, and ViDoRe V2 show that ResComEmb produces higher-quality universal multimodal embeddings than VLM2Vec-V2, and outperforms ColQwen2.5 in visual document retrieval using only 37.5% of its full visual token budget, demonstrating a favorable effectiveness-efficiency trade-off.
Figures & tables
Figure 1: Overview of proposed ResComEmb framework. Ordered multi-granularity views are encoded by a shared MLLM, then compressed into nested representations. The Bidirectional Late-Interaction Matching is applied to compute relevance scores between multimodal queries and targets.
Model
Backbone
Size
Per Meta-Task Score
Average Score
Cls.
VQA
Ret.
Gnd.
IND
OOD
Avg.
Single-Vector Embedding
CLIP
ViT-L
428M
55.2
19.7
53.2
62.2
47.6
42.8
45.4
UniIR
ViT-L
428M
44.3
16.2
61.8
65.3
47.1
41.7
44.7
MagicLens
ViT-L
613M
38.8
8.3
35.4
26.0
–
–
27.8
VLM2Vec
Qwen2-VL
2B
58.7
49.3
65.0
72.9
64.9
53.3
59.7
Table 1: MMEB Precision@1 (%). Cls., Ret., and Gnd. abbreviate classification, retrieval, and grounding; IND/OOD denote in-domain/out-of-domain tasks. Bold and underline mark the best and second-best scores in each column.
Model
Backbone
Size
Arxiv
Doc
Info
TabF
TATQ
Shift
AI
Ener.
Gov.
Hlth.
Avg.
Single-Vector Embedding
ONE-PEACE
ONE-PEACE
4B
43.9
23.4
59.9
57.0
13.4
17.0
45.4
53.2
55.9
59.5
42.9
E5-V
LLaVA-NeXT
8B
41.1
24.3
49.5
58.2
9.0
13.2
46.1
57.7
53.0
59.6
41.2
DSE
Phi-3-Vision
4B
78.1
45.8
82.0
79.2
49.0
69.8
96.8
92.6
92.0
96.3
78.2
GME
Qwen2-VL
2B
82.8
53.1
90.2
93.3
69.9
89.5
97.5
91.9
94.6
98.7
86.2
GME
Qwen2-VL
7B
86.9
57.5
91.6
94.6
74.1
96.8
99.6
95.3
98.8
99.3
89.5
Table 2: ViDoRe V1 NDCG@5 (%) across 10 visual document retrieval tasks. Arxiv, Doc, and Info denote ArxivQ, DocQ, and InfoQ; Ener. and Hlth. abbreviate Energy and Health.
Model
Backbone
Size
ESG Human
Eco Mul
Bio Mul
ESG Syn-Mul
Bio
ESG Syn
Eco
Avg.
Single-Vector Embedding
SigLIP
SigLIP
652M
28.8
14.0
18.2
21.9
33.8
19.8
29.8
23.8
VisRAG-Ret
MiniCPM-V2.0
3B
53.7
48.7
47.7
46.4
54.8
45.9
59.6
51.0
VLM2Vec
Qwen2-VL
7B
33.9
42.0
29.7
38.4
38.8
36.7
51.4
38.7
GME
Qwen2-VL
7B
65.8
56.2
55.1
56.7
64.0
54.3
62.9
59.3
mmE5
Llama-3.2-Vision
11B
52.8
44.3
46.8
54.7
51.3
55.1
48.6
50.5
Table 3: ViDoRe V2 NDCG@5 (%) across 7 tasks. Syn, Mul, and Bio denote synthetic, multilingual, and biomedical data.
Prefix
Variant
ViDoRe
MMEB
V1
V2
Avg.
Cls.
VQA
Ret.
Gnd.
Avg.
E(1)
w/ MRL
89.2
60.8
77.5
62.8
61.2
67.7
86.8
66.7
w/o MRL
86.4
53.2
72.7
58.9
60.0
63.6
85.1
63.7
Δ
(2.8 ↓ )
(7.6 ↓ )
(4.8 ↓ )
(3.9 ↓ )
(1.2 ↓ )
(4.1 ↓ )
(1.7 ↓ )
(3.0 ↓ )
E(2)
w/ MRL
89.6
61.4
78.0
62.8
61.4
68.2
86.8
66.9
w/o MRL
88.6
55.2
74.8
59.2
60.4
64.1
85.3
64.1
Table 4: Effect of MRL across nested prefixes. Prefixes E(1) , E(2) , and E(3) retain 128, 256, and 384 visual tokens, respectively, with identical text tokens; Δ denotes the drop without MRL.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Variant
ArxivQ
DocQ
InfoQ
TabF
TATQ
Shift
AI
Energy
Gov.
Health
Avg.
ResComEmb
88.8
61.7
94.2
95.2
80.4
90.4
99.6
96.6
97.9
99.3
90.4
w/o Importance Score
87.3
63.0
93.6
94.6
80.8
88.9
99.5
96.5
97.5
99.1
90.1
w/o Novelty Score
86.6
62.5
93.8
94.2
80.7
87.4
99.3
96.5
96.5
99.2
89.7
Appendix
Table 5: Ablation of RHC scoring signals on all ViDoRe V1 subsets (NDCG@5 %).
Variant
ESG Human
Eco Mul
Bio Mul
ESG Syn-Mul
Bio
ESG Syn
Eco
Avg.
ResComEmb
69.8
56.0
60.0
57.1
64.8
60.9
62.7
61.6
w/o Importance Score
68.4
54.7
57.3
55.9
60.5
59.6
60.0
59.5
w/o Novelty Score
67.5
51.2
56.8
48.8
62.3
57.3
62.7
58.1
Appendix
Table 6: Ablation of RHC scoring signals on all ViDoRe V2 subsets (NDCG@5 %). Syn, Mul, and Bio denote synthetic, multilingual, and biomedical data.
Variant
ViDoRe
MMEB
V1
V2
Avg.
Cls.
VQA
Ret.
Gnd.
Avg.
Vanilla MaxSim
90.0
59.8
77.6
41.3
15.1
62.7
71.8
44.5
ResComEmb
90.4
61.6
78.5
63.1
61.1
69.5
87.9
67.4
w/o Mean
90.0
59.1
77.3
54.6
41.3
64.5
89.3
58.1
w/o TopK
89.1
54.5
74.9
61.2
55.7
66.7
81.4
63.8
w/o Adaptive Weighting
89.0
54.0
74.6
62.8
61.3
67.4
89.3
66.9
Appendix
Table 7: Late-interaction ablations by benchmark and MMEB category. w/o Adaptive Weighting keeps both directions with fixed 0.5/0.5 weights; w/o Bidirectionality retains only query-to-document TopK-mean MaxSim. ViDoRe reports NDCG@5 (%); MMEB reports Precision@1 (%). ViDoRe Avg. weights V1/V2 by their 10/7 subsets.
Variant
Tokens
ViDoRe
MMEB
V1
V2
Avg.
Cls.
VQA
Ret.
Gnd.
Avg.
ResComEmb
384
90.4
61.6
78.5
63.1
61.1
69.5
87.9
67.4
w/o g1
384
90.0
59.1
77.3
62.1
61.2
66.4
89.5
66.3
w/o g2
384
90.4
61.9
78.7
62.8
61.3
67.5
89.6
66.9
w/o g3
384
90.2
61.5
78.4
62.9
61.2
67.8
89.2
67.0
Appendix
Table 8: Visual-granularity ablations at a fixed total budget of 384 visual tokens. ViDoRe reports NDCG@5 (%); MMEB reports Precision@1 (%). ViDoRe Avg. weights V1/V2 by their 10/7 subsets.
Stage budget
Tokens
ViDoRe
MMEB
V1
V2
Avg.
Cls.
VQA
Ret.
Gnd.
Avg.
16/16/16
48
86.2
55.5
73.6
62.0
60.9
66.7
89.6
66.3
32/32/32
96
88.9
58.6
76.4
62.3
61.3
67.7
89.8
66.9
64/64/64
192
89.5
60.6
77.6
62.6
61.4
68.6
90.2
67.3
128/128/128
384
90.4
61.6
78.5
63.1
61.1
69.5
87.9
67.4
256/256/256
768
90.6
61.7
78.7
63.2
61.1
66.9
88.5
66.7
Appendix
Table 9: Per-stage token-budget ablations for RHC. Stage budgets are written as b1/b2/b3 ; Tokens counts the total visual budget. ViDoRe reports NDCG@5 (%); MMEB reports Precision@1 (%). ViDoRe Avg. weights V1/V2 by their 10/7 subsets.