Universal multimodal retrieval requires compact embeddings that preserve task-relevant semantic information across diverse modalities. Prior works have incorporated latent reasoning into multimodal embedding learning to refine this information before embedding extraction. However, most existing approaches remain confined to deterministic latent paths, without exploring alternative trajectories to discover better embeddings. Thus, we propose VaME (Variational Multimodal Embeddings), a framework that models latent reasoning as a learnable distribution over trajectories. Specifically, we first introduce Variational Latent Reasoning (VLR) to enable autoregressive exploration in latent space, guided by answer reconstruction through a lightweight decoder. Meanwhile, we augment the original embedding-token readout with a latent-fused embedding to facilitate exploration during subsequent reinforcement learning. Finally, we optimize latent reasoning over stochastic variational trajectories through reinforcement learning, using Semantic Decoding Reward (SDR) to favor semantically meaningful trajectories with interpretable decoded outcomes. On the 78-task MMEB-V2 benchmark, spanning image, video, and visual-document retrieval, VaME outperforms most explicit CoT-based models and all latent-reasoning baselines. VaME also demonstrates robust performance on reasoning-intensive benchmarks such as MRMR, with substantial gains after reinforcement learning. Importantly, VaME achieves these gains with at least a 4.25x inference speedup over the deterministic latent autoregressive baselines. The code will be made publicly available.
Figures & tables
Figure 1: From deterministic embedding to variational latent reasoning. (a) Discriminative embedders map each input directly to a single embedding. (b) Deterministic reasoning adds latent computation but follows one fixed path. (c) VaME samples and compares alternative latent trajectories during training and follows the conditional means for deterministic inference.
Figure 2: Overview of the VaME framework. Latent SFT aligns variational trajectories with retrieval and answer supervision. Latent RL compares sampled trajectories and refines the policy using retrieval feedback and the semantic decoding reward.
Model
Image
Video
VisDoc
All
CLS
QA
RET
GD
Avg.
CLS
QA
RET
MRET
Avg.
VDRv1
VDRv2
VR
OOD
Avg.
# of Datasets
10
10
12
4
36
5
5
5
3
18
10
4
6
4
24
78
∼ 2B Model Size
VLM2Vec
58.7
49.3
65.0
72.9
59.7
33.4
30.5
20.6
33.0
29.0
49.8
13.5
51.8
33.5
41.6
47.0
DUME
59.3
55.0
66.3
78.0
62.5
37.7
46.6
17.1
30.0
33.2
67.6
43.3
47.1
33.8
52.8
52.7
GME
54.4
29.9
66.9
55.5
51.9
34.9
42.0
25.6
32.4
33.9
86.1
54.0
82.5
43.1
72.7
54.1
Table 1: Results on the MMEB-V2 benchmark. Bold and underline indicate the best and second-best scores within each size group, respectively. CLS: classification, QA: question answer, RET: retrieval, GD: grounding, MRET: moment retrieval, VDR: ViDoRe, VR: VisRAG, OOD: out-of-distribution. Reported metrics adhere to the settings of VLM2Vec-V2.
Model
Backbone
Knowledge
Theorem
Contradiction
All
Art
Med.
Sci.
Hum.
Math
Phy.
Eng.
Bus.
Neg.
Des.
Tra.
E5-V
LLaVA-Next-8B
25.1
11.7
16.6
10.8
2.1
3.4
2.5
5.2
11.5
3.7
2.1
8.6
EVA-CLIP
EVA-ViT-0.4B
10.2
13.5
26.1
12.9
6.2
10.5
9.3
11.7
8.5
4.4
5.4
10.8
OpenCLIP
ViT-G/14-1B
56.0
17.9
33.2
22.0
5.7
5.0
7.0
9.7
13.0
8.1
12.4
17.3
VLM2Vec
Qwen2-VL-7B
53.5
22.4
36.7
24.0
2.1
2.8
2.8
2.9
11.5
5.6
18.3
18.1
VISTA
Qwen2-VL-2B
21.3
27.8
32.6
17.0
18.8
17.1
17.3
28.6
20.0
20.2
9.4
20.9
Table 2: Results on the reasoning-intensive MRMR benchmark including Art, Medicine (Med.), Science (Sci.), Humanities (Hum.), Math, Physics (Phy.), Engineering (Eng.), Business (Bus.), Negation (Neg.), Design (Des.), and Traffic (Tra.). Bold and underlined scores denote the best and second-best performance, respectively.
VLR
Latent SFT
Latent RL
Image
Video
VisDoc
All
×
×
×
68.5
43.9
71.3
63.8
×
✓
×
68.9
45.1
72.8
64.6
✓
×
×
67.8
43.3
70.9
63.1
✓
✓
×
69.0
45.5
73.3
64.9
✓
✓
✓
69.5
46.3
74.4
65.7
Table 3: Ablation on core components on MMEB-V2. All crosses ( ××× ) denote the deterministic AR baseline. All reports the overall MMEB-V2 score.
#
Model
Image
Video
VisDoc
All
MRMR
1
Latent SFT
69.0
45.5
73.3
64.9
38.3
2
w/ ret reward
69.0
45.8
73.8
65.2
38.5
3
w/ SDR
69.1
46.2
74.1
65.3
40.7
4
VaME-2B
69.5
46.3
74.4
65.7
40.5
Table 5: Ablation on latent RL rewards on MMEB-V2. Latent RL combines the retrieval reward and SDR. All reports the overall MMEB-V2 score.
Figure 3: Effect of latent tokens on gains (solid lines) and throughput (dashed lines).
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
Initial
Final SFT
Ratio
Retrieval direction
Image-based (MMEB-train)
A-OKVQA
37,929
22,685
59.81%
Image–text → text
CIRR
35,085
32,203
91.79%
Image–text → image–text
ChartQA
39,512
34,962
88.48%
Image–text → text
DocVQA
47,401
42,700
90.08%
Image–text → text
HatefulMemes
16,572
10,612
64.04%
Image–text → text
Appendix
Table 6: Statistics of the latent-SFT training data composition.
Grade
Reward
Criterion
a
0.00
No requested fact is correct, or the sole atomic answer is wrong, contradictory, irrelevant, or absent.
b
0.25
Only a peripheral fragment is correct; the requested answer is mostly wrong.
c
0.50
A meaningful subset of facts is correct, but another substantive fact is wrong or missing.
d
0.75
All core facts are correct, with only a minor non-core defect.
f
1.00
All answer facts are correct, with no contradiction.
Tsinghua Shenzhen International Graduate School, Tsinghua University · School of Artificial Intelligence and Data Science, USTC · Kuaishou Technology +1