Learning Multimodal Embeddings with Evidence-Aligned Readout
Authors: Zirong Chen, Fuda Ye, Enjun Du, Junfu Pu, Xinlei Wang, Xinyu Zuo, Lisheng Duan, Haijin Liang, +3 more
Organizations: The Hong Kong University of Science and Technology (Guangzhou) · Tencent Yuanbao · Tsinghua University · The University of Hong Kong · ARC Lab, Tencent · University of Tsukuba
Multimodal large language models can expose task-relevant evidence through generation, but producing useful evidence does not by itself determine how it enters a retrieval embedding. We study whether the semantic organization of that evidence can also specify where representations are read. To address this question, we introduce EviAlign, which couples Semantic Evidence Generation with Boundary Readout in a shared multimodal large language model. It organizes evidence into five semantic units, reads the contextualized state at each unit boundary, and aggregates these states into a single normalized embedding. Generation and contrastive retrieval objectives jointly train this shared structure. With the same trailing readout, semantic evidence and free-form CoT yield nearly identical retrieval performance, suggesting that evidence organization alone does not explain the full gain. A controlled 2×3 study compares consistent and permuted evidence organization across three readout strategies, using training targets with matched evidence spans. With five readout states and the same mean pooling, the advantage of consistent semantic organization grows from 0.65 points at length-based training positions to 2.39 at evidence boundaries, yielding a 1.74-point co-design interaction. Across 12 MMEB retrieval tasks, EviAlign achieves 76.9 average Recall@1 with 500K training pairs while retaining single-vector indexing and scoring.
Figures & tables
Figure 1 : Evidence–readout co-design in EviAlign. (a) Training targets share five evidence spans across readout conditions. Distributed uses length-based training positions; at inference, readouts follow the emitted tokens. (b) Mixed permutes labeled evidence spans while fixing boundary-token identities and order; Semantic preserves role-to-boundary correspondence. Boundary readout enlarges the semantic–mixed gap from 0.65 to 2.39 points, yielding a 1.74-point co-design interaction.
Figure 2 : Overview of EviAlign’s evidence–readout co-design. Semantic Evidence Generation defines evidence units and their boundaries; Boundary Readout extracts states at those boundaries and pools them into one normalized embedding.
Method
Backbone
Data
In-Domain
Out-of-Domain
Overall
Direct embedding
GME ( Zhang et al., 2025 )
Qwen2-VL-7B
∼ 8M
70.9
71.8
71.2
LamRA-Ret ( Liu et al., 2025b )
Qwen2-VL-7B
∼ 1.4M
70.0
69.9
70.0
VLM2Vec ( Jiang et al., 2025 )
Qwen2-VL-7B
∼ 662K
75.2
57.9
69.4
VLM2Vec-V2 ( Meng et al., 2026 )
Qwen2-VL-2B
∼ 1.7M
74.8
58.7
69.5
UniME-V2 ( Gu et al., 2026a )
Qwen2-VL-7B
∼ 662K
–
–
73.1 †
Table 1 : Comparison on the 12 MMEB retrieval tasks (Recall@1, %). In-Domain and Out-of-Domain average eight and four tasks, respectively. Split averages are computed from the corresponding per-task results. † UniME-V2 and LaME report the 12-task retrieval mean without the 8/4 split.
Figure 3 : Analysis of EviAlign’s representation construction under the 500K training budget. (a) Performance across Qwen-VL model generations and scales; darker regions show the absolute gains from the larger model within each generation. (b) Effect of the joint evidence-unit and boundary-readout configuration on MMEB Recall@1.
Figure 4 : A CIRR example. The query pairs a reference image with a text modification; the generated evidence describes the requested three-bottle target while retaining visual attributes of the reference.
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
Split
Dataset
Query
Target
Retrieval Task
In Domain
VisDial
T
I
Multi-turn dialogue resolution
CIRR
I+T
I
Composed image retrieval
VisualNews (t2i)
T
I
Entity-grounded news matching
VisualNews (i2t)
I
T
Visual-to-caption grounding
MSCOCO (t2i)
T
I
General visual-language alignment
MSCOCO (i2t)
I
T
General visual-language alignment
Appendix
Table 4 : Summary of the MMEB retrieval evaluation suite. “Query” and “Target” denote input modalities: T = text, I = image, I+T = image-text pair.
Hyperparameter
Value
Model Architecture
Base Backbone
Qwen3-VL-8B-Instruct
Vision Encoder
Frozen
LLM & Projector
Full Fine-Tuning
Boundary Token Initialization
[EOS] embedding
Min / Max Pixels
768 / 1,572,864
Appendix
Table 5 : Hyperparameter settings and hardware configurations for EviAlign.
Organization
Readout
Evidence order and readout
Semantic
Trailing
E <ENT> A <ATT> R <REL> D <DET> S <SUM> ; <SUM> state only
Semantic
Distributed
U(E∣A∣R∣D∣S) ; five inserted-token states
Semantic
Boundary
E <ENT> A <ATT> R <REL> D <DET> S <SUM>
Mixed
Trailing
D <ENT> S <ATT> E <REL> A <DET> R <SUM> ; <SUM> state only
Mixed
Distributed
U(D∣S∣E∣A∣R) ; five inserted-token states
Mixed
Boundary
D <ENT> S <ATT> E <REL> A <DET> R <SUM>
Appendix
Table 6 : Schematic of the 2×3 evidence order and readout for a CIRR query requesting cookies on a counter. The five evidence spans are abbreviated, and Mixed shows one example permutation.
Dataset
Queries
Valid outputs
Valid format (%)
In-domain
VisDial
1,000
998
99.8
CIRR
1,000
994
99.4
VisualNews-t2i
1,000
991
99.1
VisualNews-i2t
1,000
997
99.7
MSCOCO-t2i
1,000
993
99.3
Appendix
Table 7 : Query-side boundary-format validity of the final 500K EviAlign model. A valid output contains each of the five boundary tokens exactly once and in the prescribed order.
Setting
ENT
ATT
REL
DET
SUM
One-slot
76.12
75.60
74.94
73.52
74.37
Leave-one-out
76.58
76.63
76.71
76.66
76.68
Full (all five readouts): 76.94
Appendix
Table 8 : Readout contribution on MMEB (average Recall@1, %).
Task
ENT
ATT
REL
DET
SUM
Full
VisDial
85.5
85.7
84.3
84.3
81.4
85.5
CIRR
74.7
75.2
75.8
78.6
75.7
77.2
VisualNews-t2i
76.0
72.7
77.3
77.9
76.1
78.9
VisualNews-i2t
82.7
84.8
78.1
77.6
79.1
83.3
MSCOCO-t2i
81.0
79.9
80.1
79.3
79.2
82.1
MSCOCO-i2t
79.2
78.7
75.2
74.2
76.3
79.8
Appendix
Table 9 : Per-task single-readout results on MMEB (Recall@1, %; 500K model). Each readout is evaluated without retraining; Full averages all five.
Task
w/o ENT
w/o ATT
w/o REL
w/o DET
w/o SUM
Full
VisDial
85.1
85.1
85.4
85.4
85.4
85.5
CIRR
77.1
77.1
76.9
76.2
76.9
77.2
VisualNews-t2i
78.8
78.6
78.5
78.1
78.8
78.9
VisualNews-i2t
82.7
82.2
83.2
83.1
82.9
83.3
MSCOCO-t2i
81.9
82.0
82.0
82.0
81.9
82.1
MSCOCO-i2t
79.2
79.5
79.7
79.5
79.7
79.8
Appendix
Table 10 : Per-task leave-one-out results on MMEB (Recall@1, %; 500K model). Each column omits one readout without retraining; Full averages all five.
Setting
Positive-pair alignment
Within-input dispersion
Recall@1 (%)
Mixed + boundary
0.53
−3.18
74.55
Semantic + distributed
0.56
−3.34
74.85
Semantic + boundary (EviAlign)
0.63
−3.27
76.94
Appendix
Table 11 : Readout geometry under matched five-readout controls on MMEB. Dispersion is diagnostic (more negative means greater separation); Recall@1 is the 12-task average.
In-domain
Dataset
Qwen3-8B
Qwen2-7B
VisDial
85.5
84.7
CIRR
77.2
76.3
VisualNews-t2i
78.9
78.0
VisualNews-i2t
83.3
82.5
MSCOCO-t2i
82.1
81.2
Appendix
Table 12 : Per-task EviAlign results with Qwen2-VL-7B and Qwen3-VL-8B under the 500K training budget.
Figure 5 : LLM-as-a-Judge comparison of EviAlign structured training targets and free-form CoT training targets. Win rates compare independently assigned 1–5 scores and exclude ties; mean scores use all sampled inputs. Multipliers denote the ratio of EviAlign wins to CoT wins among non-tied comparisons.
Figure 6 : t-SNE visualization of query and target embeddings on four representative MMEB subsets.
Figure 7 : Attention patterns in the reasoning-prefix diagnostic. (a) Attention from the last evidence-boundary token ( <SUM> ) to preceding generated positions, compared with the trailing token of free-form CoT. (b) Layer-wise self-attention over sequence positions, with the five evidence-boundary positions marked. Panels (a) and (b) use separate color scales, shown above each panel.
Figure 8 : Data scaling curves of EviAlign on MMEB from 1K to 500K training pairs. We report Recall@1 across in-domain and out-of-domain retrieval benchmarks.
Figure 9 : VisualNews example: an image query, structured retrieval evidence, and the retrieved caption.
Figure 10 : OVEN example: the query combines an image and a place-identification question. The retrieved candidate is an image–text knowledge entry; its image is shown.
Multimodal Large Language Models (MLLMs) have recently been applied to universal multimodal retrieval, where Chain-of-Thought (CoT) reasoning improves candidate reranking. However, existing approaches remain largely language-driven, relying on static visual encodings and lacking the ability to actively verify fine-grained visual evidence, which often leads to speculative reasoning in visually ambiguous cases. We propose V-Retrver, an evidence-driven retrieval framework that reformulates multimodal retrieval as an agentic reasoning process grounded in visual inspection. V-Retrver enables an MLLM to selectively acquire visual evidence during reasoning via external visual tools, performing a multimodal interleaved reasoning process that alternates between hypothesis generation and targeted visual verification.To train such an evidence-gathering retrieval agent, we adopt a curriculum-based learning strategy combining supervised reasoning activation, rejection-based refinement, and reinforcement learning with an evidence-aligned objective. Experiments across multiple multimodal retrieval benchmarks demonstrate consistent improvements in retrieval accuracy (with 23.0% improvements on average), perception-driven reasoning reliability, and generalization.
Dongyang Chen, Chaoyang Wang, Dezhao Su +6
Tsinghua University · University of Central Florida · Fudan University +4
Universal multimodal retrieval requires compact embeddings that preserve task-relevant semantic information across diverse modalities. Prior works have incorporated latent reasoning into multimodal embedding learning to refine this information before embedding extraction. However, most existing approaches remain confined to deterministic latent paths, without exploring alternative trajectories to discover better embeddings. Thus, we propose VaME (Variational Multimodal Embeddings), a framework that models latent reasoning as a learnable distribution over trajectories. Specifically, we first introduce Variational Latent Reasoning (VLR) to enable autoregressive exploration in latent space, guided by answer reconstruction through a lightweight decoder. Meanwhile, we augment the original embedding-token readout with a latent-fused embedding to facilitate exploration during subsequent reinforcement learning. Finally, we optimize latent reasoning over stochastic variational trajectories through reinforcement learning, using Semantic Decoding Reward (SDR) to favor semantically meaningful trajectories with interpretable decoded outcomes. On the 78-task MMEB-V2 benchmark, spanning image, video, and visual-document retrieval, VaME outperforms most explicit CoT-based models and all latent-reasoning baselines. VaME also demonstrates robust performance on reasoning-intensive benchmarks such as MRMR, with substantial gains after reinforcement learning. Importantly, VaME achieves these gains with at least a 4.25x inference speedup over the deterministic latent autoregressive baselines. The code will be made publicly available.
Peixi Wu, Mingzhou Jiang, Feipeng Ma +11
University of Science and Technology of China · Tsinghua University · Kuaishou +1
Unified multimodal retrieval aims to identify candidates that satisfy complex user intent expressed through heterogeneous inputs. Although Large Vision-Language Model (LVLM)-based retrievers are efficient and scalable, directly encoding raw multimodal inputs often misses fine-grained discriminative cues, leading to confusion among semantically similar candidates. Recent methods mitigate this limitation by generating Chain-of-Thought (CoT) rationales to enrich the query representation. However, such reasoning is typically derived from the query alone: it explains what the query describes, but not what the retriever misunderstands. We argue that effective retrieval reasoning should instead be conditioned on retrieval feedback. Based on this insight, we introduce UniME-R1, an embedder-adviser framework that learns to reason over initially retrieved candidates and generate Retrieval-Centric Chain-of-Thought (RC-CoT). The adviser analyzes candidates individually to identify the discriminative cues confused by the embedder. If the target appears in the initial top-k set, UniME-R1 directly reranks the candidates; otherwise, it generates RC-CoT to refine the retrieval direction and performs full-corpus re-retrieval with a dual-mode embedder. To train the framework, we mine hard negatives to simulate realistic retrieval failures, jointly optimize direct retrieval and RC-CoT-augmented retrieval, and align the adviser with retrieval outcomes through supervised learning and retrieval-oriented reinforcement learning. Extensive experiments on MMEB-V2 and a diverse set of general multimodal retrieval benchmarks demonstrate that UniME-R1 consistently improves retrieval performance over strong baselines.