Universal multimodal retrieval requires compact embeddings that preserve task-relevant semantic information across diverse modalities. Prior works have incorporated latent reasoning into multimodal embedding learning to refine this information before embedding extraction. However, most existing approaches remain confined to deterministic latent paths, without exploring alternative trajectories to discover better embeddings. Thus, we propose VaME (Variational Multimodal Embeddings), a framework that models latent reasoning as a learnable distribution over trajectories. Specifically, we first introduce Variational Latent Reasoning (VLR) to enable autoregressive exploration in latent space, guided by answer reconstruction through a lightweight decoder. Meanwhile, we augment the original embedding-token readout with a latent-fused embedding to facilitate exploration during subsequent reinforcement learning. Finally, we optimize latent reasoning over stochastic variational trajectories through reinforcement learning, using Semantic Decoding Reward (SDR) to favor semantically meaningful trajectories with interpretable decoded outcomes. On the 78-task MMEB-V2 benchmark, spanning image, video, and visual-document retrieval, VaME outperforms most explicit CoT-based models and all latent-reasoning baselines. VaME also demonstrates robust performance on reasoning-intensive benchmarks such as MRMR, with substantial gains after reinforcement learning. Importantly, VaME achieves these gains with at least a 4.25x inference speedup over the deterministic latent autoregressive baselines. The code will be made publicly available.
Figures & tables
Figure 1: From deterministic embedding to variational latent reasoning. (a) Discriminative embedders map each input directly to a single embedding. (b) Deterministic reasoning adds latent computation but follows one fixed path. (c) VaME samples and compares alternative latent trajectories during training and follows the conditional means for deterministic inference.
Figure 2: Overview of the VaME framework. Latent SFT aligns variational trajectories with retrieval and answer supervision. Latent RL compares sampled trajectories and refines the policy using retrieval feedback and the semantic decoding reward.
Model
Image
Video
VisDoc
All
CLS
QA
RET
GD
Avg.
CLS
QA
RET
MRET
Avg.
VDRv1
VDRv2
VR
OOD
Avg.
# of Datasets
10
10
12
4
36
5
5
5
3
18
10
4
6
4
24
78
∼ 2B Model Size
VLM2Vec
58.7
49.3
65.0
72.9
59.7
33.4
30.5
20.6
33.0
29.0
49.8
13.5
51.8
33.5
41.6
47.0
DUME
59.3
55.0
66.3
78.0
62.5
37.7
46.6
17.1
30.0
33.2
67.6
43.3
47.1
33.8
52.8
52.7
GME
54.4
29.9
66.9
55.5
51.9
34.9
42.0
25.6
32.4
33.9
86.1
54.0
82.5
43.1
72.7
54.1
Table 1: Results on the MMEB-V2 benchmark. Bold and underline indicate the best and second-best scores within each size group, respectively. CLS: classification, QA: question answer, RET: retrieval, GD: grounding, MRET: moment retrieval, VDR: ViDoRe, VR: VisRAG, OOD: out-of-distribution. Reported metrics adhere to the settings of VLM2Vec-V2.
Model
Backbone
Knowledge
Theorem
Contradiction
All
Art
Med.
Sci.
Hum.
Math
Phy.
Eng.
Bus.
Neg.
Des.
Tra.
E5-V
LLaVA-Next-8B
25.1
11.7
16.6
10.8
2.1
3.4
2.5
5.2
11.5
3.7
2.1
8.6
EVA-CLIP
EVA-ViT-0.4B
10.2
13.5
26.1
12.9
6.2
10.5
9.3
11.7
8.5
4.4
5.4
10.8
OpenCLIP
ViT-G/14-1B
56.0
17.9
33.2
22.0
5.7
5.0
7.0
9.7
13.0
8.1
12.4
17.3
VLM2Vec
Qwen2-VL-7B
53.5
22.4
36.7
24.0
2.1
2.8
2.8
2.9
11.5
5.6
18.3
18.1
VISTA
Qwen2-VL-2B
21.3
27.8
32.6
17.0
18.8
17.1
17.3
28.6
20.0
20.2
9.4
20.9
Table 2: Results on the reasoning-intensive MRMR benchmark including Art, Medicine (Med.), Science (Sci.), Humanities (Hum.), Math, Physics (Phy.), Engineering (Eng.), Business (Bus.), Negation (Neg.), Design (Des.), and Traffic (Tra.). Bold and underlined scores denote the best and second-best performance, respectively.
VLR
Latent SFT
Latent RL
Image
Video
VisDoc
All
×
×
×
68.5
43.9
71.3
63.8
×
✓
×
68.9
45.1
72.8
64.6
✓
×
×
67.8
43.3
70.9
63.1
✓
✓
×
69.0
45.5
73.3
64.9
✓
✓
✓
69.5
46.3
74.4
65.7
Table 3: Ablation on core components on MMEB-V2. All crosses ( ××× ) denote the deterministic AR baseline. All reports the overall MMEB-V2 score.
#
Model
Image
Video
VisDoc
All
MRMR
1
Latent SFT
69.0
45.5
73.3
64.9
38.3
2
w/ ret reward
69.0
45.8
73.8
65.2
38.5
3
w/ SDR
69.1
46.2
74.1
65.3
40.7
4
VaME-2B
69.5
46.3
74.4
65.7
40.5
Table 5: Ablation on latent RL rewards on MMEB-V2. Latent RL combines the retrieval reward and SDR. All reports the overall MMEB-V2 score.
Figure 3: Effect of latent tokens on gains (solid lines) and throughput (dashed lines).
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
Initial
Final SFT
Ratio
Retrieval direction
Image-based (MMEB-train)
A-OKVQA
37,929
22,685
59.81%
Image–text → text
CIRR
35,085
32,203
91.79%
Image–text → image–text
ChartQA
39,512
34,962
88.48%
Image–text → text
DocVQA
47,401
42,700
90.08%
Image–text → text
HatefulMemes
16,572
10,612
64.04%
Image–text → text
Appendix
Table 6: Statistics of the latent-SFT training data composition.
Grade
Reward
Criterion
a
0.00
No requested fact is correct, or the sole atomic answer is wrong, contradictory, irrelevant, or absent.
b
0.25
Only a peripheral fragment is correct; the requested answer is mostly wrong.
c
0.50
A meaningful subset of facts is correct, but another substantive fact is wrong or missing.
d
0.75
All core facts are correct, with only a minor non-core defect.
f
1.00
All answer facts are correct, with no contradiction.
Reasoning-driven universal multimodal embedding has advanced rapidly by introducing Chain-of-Thought (CoT) reasoning into the embedding pipeline. Despite the strong performance across both general and complex tasks, this paradigm suffers from two core limitations: (i) autoregressive CoT reasoning incurs high computational cost, making it impractical for low-latency retrieval; and (ii) embedding performance is heavily coupled with CoT annotation quality, making large-scale training unreliable. These raise fundamental questions: Is textual CoT the optimal form of reasoning for embedding, and can effective embedding reasoning be accomplished in latent space? To this end, we propose LaME (Latent Reasoning Multimodal Embedding), which formulates embedding-oriented latent reasoning as a weakly supervised information bottleneck. LaME employs K learnable reason tokens as a fixed-capacity bottleneck, completing all reasoning within a single forward pass. The two weak supervision signals structurally decouple contrastive from autoregressive objectives and eliminate dependence on CoT annotations, while a two-stage training pipeline ensures stable convergence. Experiments on MMEB-v2 and MRMR show that LaME achieves competitive performance, surpassing some explicit CoT-based models, while delivering 60x faster inference than explicit CoT methods and 2x faster than latent baselines with throughput comparable to discriminative embedding models. Code is available at https://github.com/PeppaWu/LaME.
Peixi Wu, Biao Yang, Feipeng Ma +7
University of Science and Technology of China · 2Kuaishou Technology · 3Zhejiang University +1
Universal multimodal embedding (UME) maps multimodal inputs into a shared embedding space for diverse retrieval tasks. Recent methods improve embeddings through Chain-of-Thought (CoT) reasoning optimized with GRPO using retrieval rewards. However, existing methods overlook the mismatch bettween candidate-aware retrieval supervision and input-only CoT generation: (1)trajectory-level rewards convey retrieval outcomes without explicitly identifying the input-supported evidence that distinguishes the positive from hard negatives; (2) input-only generation cannot directly assess whether further reasoning improves retrieval, potentially producing redundant CoTs with substantial latency. To bridge this gap, we propose Reason What Matters (ReWAM), a retrieval-grounded framework that aligns candidate-aware supervision with input-only generation. Specifically, we introduce Retrieval-Aware Self-Distillation (RASD), which extracts privileged guidance from input-supported facts and evidence distinguishing the positive from hard negatives. Conditioned on this guidance, an on-policy self-teacher provides token-level feedback to refine credit assignment, directing policy updates toward retrieval-relevant reasoning grounded in the input. We further propose Retrieval-Adaptive Inference (RAI), which learns a retrieval-aware stopping criterion from prefix-level retrieval feedback. It stops redundant reasoning without candidate access and uses speculative decoding to further reduce CoT latency. Extensive experiments on MMEB-V2 and MRMR demonstrate that ReWAM achieves state-of-the-art retrieval performance while delivering up to 5x the inference throughput of competitive explicit-CoT UME methods. ReWAM thus enables high-quality retrieval through efficient input-only reasoning, making explicit CoT practical for corpus-scale multimodal retrieval. The code will be publicly available.
Mingzhou Jiang, Peixi Wu, Hang Cheng +7
Tsinghua Shenzhen International Graduate School, Tsinghua University · School of Artificial Intelligence and Data Science, USTC · Kuaishou Technology +1
Multimodal Large Language Models (MLLMs) have emerged as a promising foundation for universal multimodal embeddings. Recent studies have shown that reasoning-driven generative multimodal embeddings can outperform discriminative embeddings on several embedding tasks. However, Chain-of-Thought (CoT) reasoning tends to generate redundant thinking steps and introduce semantic ambiguity in the summarized answers in broader retrieval scenarios. To address this limitation, we propose Rewrite-driven Multimodal Embedding (RIME), a unified framework that jointly optimizes generation and embedding through a retrieval-friendly rewrite. Meanwhile, we present the Cross-Mode Alignment (CMA) to bridge the generative and discriminative embedding spaces, enabling flexible mutual retrieval to trade off efficiency and accuracy. Based on this, we also introduce Refine Reinforcement Learning (Refine-RL) that treats discriminative embeddings as stable semantic anchors to guide the rewrite optimization. Extensive experiments on MMEB-V2, MRMR and UVRB demonstrate that RIME substantially outperforms prior generative embedding models while significantly reducing the length of thinking. Code is available at https://github.com/PeppaWu/RIME.
Peixi Wu, Ke Mei, Feipeng Ma +15
WeChat Vision, Tencent Inc. · Zhejiang University · Tsinghua University +1