Learning Multimodal Embeddings with Evidence-Aligned Readout
Authors: Zirong Chen, Fuda Ye, Enjun Du, Junfu Pu, Xinlei Wang, Xinyu Zuo, Lisheng Duan, Haijin Liang, +3 more
Organizations: The Hong Kong University of Science and Technology (Guangzhou) · Tencent Yuanbao · Tsinghua University · The University of Hong Kong · ARC Lab, Tencent · University of Tsukuba
Multimodal large language models can expose task-relevant evidence through generation, but producing useful evidence does not by itself determine how it enters a retrieval embedding. We study whether the semantic organization of that evidence can also specify where representations are read. To address this question, we introduce EviAlign, which couples Semantic Evidence Generation with Boundary Readout in a shared multimodal large language model. It organizes evidence into five semantic units, reads the contextualized state at each unit boundary, and aggregates these states into a single normalized embedding. Generation and contrastive retrieval objectives jointly train this shared structure. With the same trailing readout, semantic evidence and free-form CoT yield nearly identical retrieval performance, suggesting that evidence organization alone does not explain the full gain. A controlled 2×3 study compares consistent and permuted evidence organization across three readout strategies, using training targets with matched evidence spans. With five readout states and the same mean pooling, the advantage of consistent semantic organization grows from 0.65 points at length-based training positions to 2.39 at evidence boundaries, yielding a 1.74-point co-design interaction. Across 12 MMEB retrieval tasks, EviAlign achieves 76.9 average Recall@1 with 500K training pairs while retaining single-vector indexing and scoring.
Figures & tables
Figure 1 : Evidence–readout co-design in EviAlign. (a) Training targets share five evidence spans across readout conditions. Distributed uses length-based training positions; at inference, readouts follow the emitted tokens. (b) Mixed permutes labeled evidence spans while fixing boundary-token identities and order; Semantic preserves role-to-boundary correspondence. Boundary readout enlarges the semantic–mixed gap from 0.65 to 2.39 points, yielding a 1.74-point co-design interaction.
Figure 2 : Overview of EviAlign’s evidence–readout co-design. Semantic Evidence Generation defines evidence units and their boundaries; Boundary Readout extracts states at those boundaries and pools them into one normalized embedding.
Method
Backbone
Data
In-Domain
Out-of-Domain
Overall
Direct embedding
GME ( Zhang et al., 2025 )
Qwen2-VL-7B
∼ 8M
70.9
71.8
71.2
LamRA-Ret ( Liu et al., 2025b )
Qwen2-VL-7B
∼ 1.4M
70.0
69.9
70.0
VLM2Vec ( Jiang et al., 2025 )
Qwen2-VL-7B
∼ 662K
75.2
57.9
69.4
VLM2Vec-V2 ( Meng et al., 2026 )
Qwen2-VL-2B
∼ 1.7M
74.8
58.7
69.5
UniME-V2 ( Gu et al., 2026a )
Qwen2-VL-7B
∼ 662K
–
–
73.1 †
Table 1 : Comparison on the 12 MMEB retrieval tasks (Recall@1, %). In-Domain and Out-of-Domain average eight and four tasks, respectively. Split averages are computed from the corresponding per-task results. † UniME-V2 and LaME report the 12-task retrieval mean without the 8/4 split.
Figure 3 : Analysis of EviAlign’s representation construction under the 500K training budget. (a) Performance across Qwen-VL model generations and scales; darker regions show the absolute gains from the larger model within each generation. (b) Effect of the joint evidence-unit and boundary-readout configuration on MMEB Recall@1.
Figure 4 : A CIRR example. The query pairs a reference image with a text modification; the generated evidence describes the requested three-bottle target while retaining visual attributes of the reference.
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
Split
Dataset
Query
Target
Retrieval Task
In Domain
VisDial
T
I
Multi-turn dialogue resolution
CIRR
I+T
I
Composed image retrieval
VisualNews (t2i)
T
I
Entity-grounded news matching
VisualNews (i2t)
I
T
Visual-to-caption grounding
MSCOCO (t2i)
T
I
General visual-language alignment
MSCOCO (i2t)
I
T
General visual-language alignment
Appendix
Table 4 : Summary of the MMEB retrieval evaluation suite. “Query” and “Target” denote input modalities: T = text, I = image, I+T = image-text pair.
Hyperparameter
Value
Model Architecture
Base Backbone
Qwen3-VL-8B-Instruct
Vision Encoder
Frozen
LLM & Projector
Full Fine-Tuning
Boundary Token Initialization
[EOS] embedding
Min / Max Pixels
768 / 1,572,864
Appendix
Table 5 : Hyperparameter settings and hardware configurations for EviAlign.
Organization
Readout
Evidence order and readout
Semantic
Trailing
E <ENT> A <ATT> R <REL> D <DET> S <SUM> ; <SUM> state only
Semantic
Distributed
U(E∣A∣R∣D∣S) ; five inserted-token states
Semantic
Boundary
E <ENT> A <ATT> R <REL> D <DET> S <SUM>
Mixed
Trailing
D <ENT> S <ATT> E <REL> A <DET> R <SUM> ; <SUM> state only
Mixed
Distributed
U(D∣S∣E∣A∣R) ; five inserted-token states
Mixed
Boundary
D <ENT> S <ATT> E <REL> A <DET> R <SUM>
Appendix
Table 6 : Schematic of the 2×3 evidence order and readout for a CIRR query requesting cookies on a counter. The five evidence spans are abbreviated, and Mixed shows one example permutation.
Dataset
Queries
Valid outputs
Valid format (%)
In-domain
VisDial
1,000
998
99.8
CIRR
1,000
994
99.4
VisualNews-t2i
1,000
991
99.1
VisualNews-i2t
1,000
997
99.7
MSCOCO-t2i
1,000
993
99.3
Appendix
Table 7 : Query-side boundary-format validity of the final 500K EviAlign model. A valid output contains each of the five boundary tokens exactly once and in the prescribed order.
Setting
ENT
ATT
REL
DET
SUM
One-slot
76.12
75.60
74.94
73.52
74.37
Leave-one-out
76.58
76.63
76.71
76.66
76.68
Full (all five readouts): 76.94
Appendix
Table 8 : Readout contribution on MMEB (average Recall@1, %).
Task
ENT
ATT
REL
DET
SUM
Full
VisDial
85.5
85.7
84.3
84.3
81.4
85.5
CIRR
74.7
75.2
75.8
78.6
75.7
77.2
VisualNews-t2i
76.0
72.7
77.3
77.9
76.1
78.9
VisualNews-i2t
82.7
84.8
78.1
77.6
79.1
83.3
MSCOCO-t2i
81.0
79.9
80.1
79.3
79.2
82.1
MSCOCO-i2t
79.2
78.7
75.2
74.2
76.3
79.8
Appendix
Table 9 : Per-task single-readout results on MMEB (Recall@1, %; 500K model). Each readout is evaluated without retraining; Full averages all five.
Task
w/o ENT
w/o ATT
w/o REL
w/o DET
w/o SUM
Full
VisDial
85.1
85.1
85.4
85.4
85.4
85.5
CIRR
77.1
77.1
76.9
76.2
76.9
77.2
VisualNews-t2i
78.8
78.6
78.5
78.1
78.8
78.9
VisualNews-i2t
82.7
82.2
83.2
83.1
82.9
83.3
MSCOCO-t2i
81.9
82.0
82.0
82.0
81.9
82.1
MSCOCO-i2t
79.2
79.5
79.7
79.5
79.7
79.8
Appendix
Table 10 : Per-task leave-one-out results on MMEB (Recall@1, %; 500K model). Each column omits one readout without retraining; Full averages all five.
Setting
Positive-pair alignment
Within-input dispersion
Recall@1 (%)
Mixed + boundary
0.53
−3.18
74.55
Semantic + distributed
0.56
−3.34
74.85
Semantic + boundary (EviAlign)
0.63
−3.27
76.94
Appendix
Table 11 : Readout geometry under matched five-readout controls on MMEB. Dispersion is diagnostic (more negative means greater separation); Recall@1 is the 12-task average.
In-domain
Dataset
Qwen3-8B
Qwen2-7B
VisDial
85.5
84.7
CIRR
77.2
76.3
VisualNews-t2i
78.9
78.0
VisualNews-i2t
83.3
82.5
MSCOCO-t2i
82.1
81.2
Appendix
Table 12 : Per-task EviAlign results with Qwen2-VL-7B and Qwen3-VL-8B under the 500K training budget.
Figure 5 : LLM-as-a-Judge comparison of EviAlign structured training targets and free-form CoT training targets. Win rates compare independently assigned 1–5 scores and exclude ties; mean scores use all sampled inputs. Multipliers denote the ratio of EviAlign wins to CoT wins among non-tied comparisons.
Figure 6 : t-SNE visualization of query and target embeddings on four representative MMEB subsets.
Figure 7 : Attention patterns in the reasoning-prefix diagnostic. (a) Attention from the last evidence-boundary token ( <SUM> ) to preceding generated positions, compared with the trailing token of free-form CoT. (b) Layer-wise self-attention over sequence positions, with the five evidence-boundary positions marked. Panels (a) and (b) use separate color scales, shown above each panel.
Figure 8 : Data scaling curves of EviAlign on MMEB from 1K to 500K training pairs. We report Recall@1 across in-domain and out-of-domain retrieval benchmarks.
Figure 9 : VisualNews example: an image query, structured retrieval evidence, and the retrieved caption.
Figure 10 : OVEN example: the query combines an image and a place-identification question. The retrieved candidate is an image–text knowledge entry; its image is shown.