Omnimodal embeddings naturally involve both shared representations and modality-specific features across heterogeneous inputs. However, existing omnimodal embedding methods often rely on a single shared parameter space over mixed-modality data, limiting structural separation between universal and modality-specific representations. To address this, we propose Syn-Omni, a unified framework for structured omnimodal adaptation with modality specialization and controlled cross-modal collaboration. Specifically, we introduce Orthogonal Modality-Expert LoRA (OME-LoRA), which decomposes adaptation into a shared LoRA path for universal semantics and modality-expert LoRA paths for modality-aware specialization. Furthermore, Progressive Synergy Routing (PSR) enables experts to first establish modality-specific priors, then gradually interact with other modality-experts for cross-modal synergy. Evaluated across 81 diverse tasks spanning image, video, audio, and audiovisual modalities, Syn-Omni consistently outperforms omnimodal baselines, demonstrating the effectiveness of structured specialization and cross-modal progressive collaboration.
Figures & tables
Figure 1: Omnimodal adaptation strategies. We compare three contrastive embedding setups: (1) modality-specific models, each trained independently on a single modality using a shared LoRA, (2) Uni-Omni (baseline), trained on mixed-modality data using a single shared LoRA, and (3) Syn-Omni (Ours), which combines a shared adaptation path with modality-specialized expert paths. Here, the overall score of the modality-specific setup averages each model across all four modalities, including those it was not trained on. While mixed-modality training enables strong unified representations, Syn-Omni further improves the Uni-Omni baseline across diverse modality scenarios.
Figure 2: Illustration of the Syn-Omni framework. (a) Overall Framework: Omnimodal inputs xm are encoded by modality-specific encoders and a pretrained LLM backbone with OME-LoRA to produce unified embeddings e(xm) . (b) Orthogonal Modality-Expert LoRA (OME-LoRA): Shared LoRA path captures universal semantics, while modality-specific expert LoRAs model modality-aware characteristics. The outputs of modality-experts are combined by the routing weight w . (c) Progressive Synergy Routing (PSR): PSR gradually shifts from initial modality-separate routing to later-stage collaborative routing, enabling progressive cross-modal synergy.
Figure 3: Conceptual comparison between BCE Loss and Synergy Margin Loss (SML). While BCE strictly suppresses absent modalities to 0, SML maintains a relative margin γ ( vpresentR−vabsentR≥γ ), which allows absent-modality experts to retain meaningful activations, preserving cross-modal synergy during training.
Image (36)
Video (23)
Audio (12)
Audiovisual (10)
All (81)
3B Models
Omni-Embed-Nemotron Xu et al. (2025b)
44.1
36.5
24.5
28.5
33.4
LCO-Emb Xiao et al. (2025)
58.1
43.7
42.1
35.5
44.8
e5-omni Chen et al. (2026)
64.8
40.6
37.4
31.6
43.6
Uni-Omni (Contrastive baseline)
66.8
40.4
43.1
40.4
47.7
Syn-Omni (Ours)
68.1
40.5
45.6
43.4
49.4
Table 1: Comparison of omnimodal embedding models using Hit@1 (%). The numbers in parentheses indicate the number of evaluation tasks for each modality and the total. Here, Uni-Omni corresponds to a strong baseline trained with a standard LoRA setting. Across both 3B- and 7B-scale models, Syn-Omni achieves higher overall scores than previous methods, demonstrating well-balanced performance across diverse modalities, while also improving over our baseline. For each model scale, the best results are bolded , and the second-best are underlined .
OME-LoRA
PSR
id
S-Exp.
Lortho
Soft
Prog.
I
V
A
AV
All
(a)
-
-
-
-
66.8
40.4
43.1
40.4
47.7
(b)
✓
-
-
-
67.6
39.6
44.9
42.6
48.7
(c)
✓
✓
-
-
67.4
39.9
45.2
41.7
48.6
(d)
✓
✓
✓
-
67.3
39.9
45.0
42.1
48.6
(e)
✓
-
✓
✓
68.2
40.1
45.2
42.6
49.0
Table 2: Component analysis of Syn-Omni framework. Starting from the Uni-Omni baseline in (a), each row incorporates individual components of our design. The results show that both OME-LoRA and PSR contribute to improved and balanced performance over the baseline, with the full model achieving the best overall results.
id
Variant
Rank
Params.
PSR
I
V
A
AV
All
(a)
Shr
16
29.9M
-
66.8
40.4
43.1
40.4
47.7
(b)
Shr
24
44.9M
-
67.3
38.7
43.0
41.6
47.6
(c)
Shr
36
59.9M
-
67.3
40.7
43.8
41.2
48.2
(d)
Exp
16
29.9M
-
66.1
39.4
42.5
40.1
47.0
(e)
Exp
16
33.3M
✓
65.3
39.5
42.1
40.5
46.9
(f)
Shr+Exp
24
44.9M
-
67.6
39.6
44.9
42.6
48.7
Table 3: Analysis of the adaptation structure. ‘Shr’ denotes a single shared LoRA path ( Uni-Omni ) with increasing rank. ‘Exp’ denotes modality-expert paths without a shared LoRA path, and ‘Shr+Exp’ denotes our OME-LoRA architecture. Neither path alone is sufficient: scaling the shared path yields limited gains, and removing it degrades performance further.
Routing Strategy (PSR)
id
Router
Progress
loss
I
V
A
AV
All
(a)
Hard
-
-
67.4
39.9
45.2
41.7
48.6
(b)
Soft
-
BCE
67.3
39.9
45.0
42.1
48.6
(c)
Soft
-
SML
67.6
40.4
45.1
41.9
48.8
(d)
Soft
✓
BCE
67.6
39.6
46.0
41.9
48.8
(e)
Soft
✓
SML
68.1
40.5
45.6
43.4
49.4
Table 4: Analysis on Progressive Synergy Routing (PSR) strategy. We evaluate the impact of routing types, progressive blending, and router loss functions in the OME-LoRA layers. We highlight that combining a progressive blending of routing weights with the Synergy Margin Loss (SML) is crucial for both expert specialization and cross-modal collaboration among experts.
Figure 4: Sensitivity analysis of hyperparameters on the orthogonal penalty weight λortho , the synergy margin loss weight λSML , and the margin parameter γ . Red stars indicate the optimal choice in our final model.
Figure 5: Pairwise cosine similarity between OME-LoRA pathways. Without progressive blending (left), high cross-path similarities reveal redundancy across different pathways. In contrast, Syn-Omni (right), which differs only in applying progressive blending, produces significantly lower similarities. This indicates effective decoupling and clear specialization across pathways.
Figure 6: Routing dynamics of PSR under video-audio inputs. Present-modality experts retain dominant routing weights, while absent-modality experts gradually acquire non-zero activations through progressive blending. This indicates the emergence of synergistic cross-modal interactions along with modality specialization.
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: Distribution of sample counts for each query-to-target modality composition in our training mixture. For each composition, ‘T’, ‘I’, ‘V’, and ‘A’ correspond to text, image, video, and audio modality, respectively. We also specify the total sample size per modality group.
Uni-Omni
Syn-Omni
3B
7B
3B
7B
Backbone
Qwen-2.5-Omni
LoRA rank (Shared)
16
8
LoRA rank (Expert)
0
4×4
LoRA rank (Total)
16
24
LoRA params.
29.9 M
40.4 M
44.9 M
60.6 M
Appendix
Table 5: Summary of LoRA configuration and trainable parameters for Uni-Omni and Syn-Omni model variants.
Figure 8: Visualization of layer-wise soft routing weights under different input modalities. The baseline (BCE loss) exhibits a collapsed hard routing behavior with nearly binary activations, whereas ours (synergy margin loss) dynamically allocates routing weights even to absent-modality experts to encourage cross-modal synergy.
I
V
A
AV
All
I-specific
68.0
43.0
35.4
33.3
44.9
V-specific
48.1
41.1
34.9
35.9
40.0
A-specific
40.3
36.5
43.5
30.9
37.8
AV-specific
38.0
32.4
34.6
37.1
35.5
Uni-Omni (3B)
66.8
40.4
43.1
40.4
47.7
Syn-Omni (3B)
68.1
40.5
45.6
43.4
49.4
Appendix
Table 6: Comprehensive evaluation of modality-specific models on all modalities. The results demonstrate intrinsic cross-modal transferability and highlight the necessity of unified multimodal training for cross-modal synergy. Here, I, V, A, and AV denote image, video, audio, and audiovisual modalities, respectively.
Model
Latency (ms)
Peak GPU Mem. (MB)
Uni-Omni (3B)
300.3 ± 25
9612
Syn-Omni (3B)
354.6 ± 20
9676
Uni-Omni (7B)
377.1 ± 26
17672
Syn-Omni (7B)
443.5 ± 30
17765
Appendix
Table 7: Runtime analysis on Uni-Omni and our Syn-Omni model variants. We report mean latency (ms) and peak GPU memory usage (MB), testing on VALOR-32k audiovisual input samples.
Dataset
Composition
Size
ImageNet 1K Deng et al. (2009)
I → T
15,000
N24News Wang et al. (2022)
TI → T
15,000
HatefulMemes Kiela et al. (2020)
I → T
8,500
VOC2007 Everingham et al. (2015)
I → T
7,844
SUN397 Xiao et al. (2010)
I → T
15,000
OK-VQA Marino et al. (2019)
TI → T
9,007
Appendix
Table 8: An overview of image-text datasets in the omnimodal training dataset. Each dataset is sourced from the MMEB-v1 Jiang et al. (2025) , where either the full original set or a randomly sampled subset was utilized.
Dataset
Composition
Size
LLaVAHound Retrieval Zhang et al. (2025)
T → V
30,000
V → T
30,000
LLaVAHound QA Zhang et al. (2025)
TV → T
30,000
PE-Video Bolya et al. (2025)
V → T
40,000
T → V
40,000
MSVD Chen and Dolan (2011)
V → T
1,200
Appendix
Table 9: An overview of video-centric datasets included in the omnimodal training dataset.
Dataset
Composition
Size
AudioCaps Kim et al. (2019)
A → T
45,000
T → A
45,000
WavCaps Mei et al. (2024)
A → T
45,000
T → A
45,000
AudioSet-SL Gemmeke et al. (2017)
A → T
20,000
MusicCaps Agostinelli et al. (2023)
A → T
2,580
Appendix
Table 10: An overview of audio-centric datasets included in the omnimodal training dataset.
Dataset
Composition
Size
AudioSet Gemmeke et al. (2017)
V → A
30,000
A → V
30,000
VAST Chen et al. (2023)
VA → T
30,000
T → VA
30,000
InternVideo2 Wang et al. (2024)
VA → T
30,000
T → VA
30,000
Appendix
Table 11: An overview of audiovisual-centric datasets included in the omnimodal training dataset.
Task
Dataset
Composition
Query Size
Corpus Size
I-CLS
ImageNet-1K Deng et al. (2009)
I2T
1000
1000
N24News Wang et al. (2022)
TI2T
1000
24
HatefulMemes Kiela et al. (2020)
I2T
1000
2
VOC2007 Everingham et al. (2015)
I2T
1000
20
SUN397 Xiao et al. (2010)
I2T
1000
397
Place365 Zhou et al. (2018a)
I2T
1000
365
Appendix
Table 12: Evaluation datasets for image modality, including task type and query-to-target composition.
Task
Dataset
Composition
Query Size
Corpus Size
V-CLS
SmthSmthV2 Goyal et al. (2017)
V2T
1000
174
HMDB51 Kuehne et al. (2011)
V2T
1000
51
UCF101 Soomro et al. (2012)
V2T
1000
101
K700 Carreira et al. (2019)
V2T
1000
700
Breakfast Kuehne et al. (2014)
V2T
433
10
T2V RET
MSR-VTT Xu et al. (2016)
T2V
1000
1000
Appendix
Table 13: Evaluation datasets for video modality, including task type and query-to-target composition.
Task
Dataset
Composition
Query Size
Corpus Size
A-CLS
ESC50 Piczak (2015)
A2T
1237
50
GTZAN Tzanetakis and Cook (2002)
A2T
290
10
NSynth Engel et al. (2017)
A2T
4096
10
T2A RET
AudioCaps Kim et al. (2019)
T2A
4411
883
Clotho Drossos et al. (2020)
T2A
5225
1045
MusicCaps Agostinelli et al. (2023)
T2A
2772
2772
Appendix
Table 14: Evaluation datasets for audio modality, including task type and query-to-target composition.
Task
Dataset
Composition
Query Size
Corpus Size
A2V RET
AVE Tian et al. (2018)
A2V
402
402
VALOR32k Liu et al. (2025)
A2V
3239
3239
V2A RET
AVE Tian et al. (2018)
V2A
402
402
VALOR32k Liu et al. (2025)
V2A
3239
3239
T2VA RET
AVHBench Kim et al. (2025)
T2VA
1105
1105
VALOR32k Liu et al. (2025)
T2VA
3239
3239
Appendix
Table 15: Evaluation datasets for audiovisual modality, including task type and query-to-target composition.
Nemotron (3B)
LCO-Emb (3B)
e5-omni (3B)
Uni-Omni (3B)
Syn-Omni (3B)
OmniEmbed (7B)
Multivent (7B)
LCO Emb (7B)
e5-omni (7B)
WAVE (7B)
Uni-Omni (7B)
Syn-Omni (7B)
Average
Overall (36)
44.1
58.1
64.8
66.8
68.1
44.2
51.8
61.6
72.5
42.7
71.3
71.7
I-CLS (10)
47.9
57.3
60.7
64.8
65.4
44.5
56.5
59.5
67.6
49.0
66.1
67.5
I-QA (10)
20.1
58.2
61.3
63.1
62.7
22.5
29.6
62.7
69.9
25.8
67.8
66.8
I-RET (12)
58.7
55.4
68.3
66.3
66.3
49.5
61.1
58.9
71.9
44.1
68.3
69.1
VG (4)
49.5
61.5
69.1
72.8
77.8
60.5
60.0
65.1
80.9
51.9
83.2
83.6
Appendix
Table 16: Individual performances on image modality tasks. The method order is identical to that in Tab. 1 .
Nemotron (3B)
LCO-Emb (3B)
e5-omni (3B)
Uni-Omni (3B)
Syn-Omni (3B)
OmniEmbed (7B)
Multivent (7B)
LCO Emb (7B)
e5-omni (7B)
WAVE (7B)
Uni-Omni (7B)
Syn-Omni (7B)
Average
Overall (23)
36.5
43.7
40.6
40.4
40.5
35.0
40.2
45.5
44.3
39.8
40.2
41.4
V-CLS (5)
43.1
44.6
37.6
47.6
44.7
36.4
50.7
47.6
49.4
48.9
49.5
47.0
T2V RET (5)
34.8
34.2
38.9
27.9
29.3
33.7
35.9
36.9
36.1
30.0
25.0
28.5
V2T RET (5)
32.0
34.3
33.2
38.3
39.2
27.9
34.7
35.9
40.8
35.1
40.5
41.6
M-RET (3)
25.6
47.2
41.0
36.9
37.5
28.7
28.1
47.1
34.1
40.1
32.4
36.1
Appendix
Table 17: Individual performances on video modality tasks. The method order is identical to that in Tab. 1 .
Nemotron (3B)
LCO-Emb (3B)
e5-omni (3B)
Uni-Omni (3B)
Syn-Omni (3B)
OmniEmbed (7B)
Multivent (7B)
LCO Emb (7B)
e5-omni (7B)
WAVE (7B)
Uni-Omni (7B)
Syn-Omni (7B)
Average
Overall (12)
24.5
42.1
37.4
43.1
45.6
31.6
39.5
45.2
44.3
33.4
47.5
48.7
A-CLS (3)
40.7
67.5
51.5
59.5
62.0
36.0
52.9
71.1
62.8
56.3
67.4
67.2
T2A RET (4)
6.4
19.7
26.0
28.7
30.8
25.5
28.6
24.5
30.2
22.5
32.5
34.9
A2T RET (3)
6.8
17.3
14.2
30.0
32.4
17.2
21.0
18.7
20.7
17.7
32.5
34.0
A-QA (2)
44.0
63.8
57.9
54.4
57.0
47.8
55.4
66.4
63.7
37.0
57.6
58.5
Appendix
Table 18: Individual performances on audio modality tasks. The method order is identical to that in Tab. 1 .
Nemotron (3B)
LCO-Emb (3B)
e5-omni (3B)
Uni-Omni (3B)
Syn-Omni (3B)
OmniEmbed (7B)
Multivent (7B)
LCO Emb (7B)
e5-omni (7B)
WAVE (7B)
Uni-Omni (7B)
Syn-Omni (7B)
Average
Overall (10)
28.5
35.5
31.6
40.4
43.4
29.0
34.4
35.5
42.3
38.1
43.9
44.6
A2V RET (2)
5.2
11.1
7.0
17.4
19.1
5.2
11.4
12.6
7.2
8.8
22.0
23.1
V2A RET (2)
7.3
10.5
9.2
14.3
17.9
13.5
16.1
13.8
14.8
19.9
20.1
18.9
T2VA RET (2)
46.9
51.1
64.4
63.3
64.2
57.2
48.9
51.7
65.7
52.6
64.7
64.9
VA2T RET (2)
41.4
46.6
30.5
62.5
63.0
27.3
49.1
46.3
62.4
56.9
64.3
63.9
Appendix
Table 19: Individual performances on audiovisual modality tasks. The method order is identical to that in Tab. 1 .
In this report, we introduce \textbf{Ovis-Embedding}, a state-of-the-art omni-modal embedding family built on native integration of text, image, video, and audio. Instead of assembling separate modality towers, Ovis-Embedding uses a shared multimodal backbone to encode different modalities in a common representation space. Specifically, we make \textbf{three key advances}: (1) \textbf{native omni-modal initialization}: we adopt a pretrained Qwen-omni model as the embedding backbone and adapt it through contrastive training with low-rank initialization; (2) \textbf{data-centric omni-modal training}: we construct a broad, high-quality corpus spanning text, images, video, audio, and interleaved multimodal data. To improve data efficiency, we introduce homogeneous-source sampling to form task-consistent batches with informative in-batch negatives; and (3) \textbf{embedding-specific training and inference optimization}: we use focal loss to emphasize hard examples and similarity-based Embedding Distillation to transfer fine-grained similarity structure from complementary experts. At inference time, low-rank feature decomposition enables compact embeddings with flexible dimensionality and minimal performance loss. Empirical evaluations show that the \textbf{Ovis-Embedding} family achieves state-of-the-art performance on \textbf{MMEB-v3}, \textbf{MMEB-v2}, \textbf{MVEB}, \textbf{MAEB}, and \textbf{RTEB}, demonstrating its effectiveness across text, image, video, and audio modalities. These results highlight the potential of unified omni-modal training to overcome modality fragmentation and advance universal embedding models for any-to-any retrieval.
Omni-modal Large Language Models (Omni-MLLMs) are designed to reason over diverse sensory streams within a unified model. However, we find that multimodal reasoning is not solely determined by sensory evidence, but also by how modality information is organized. Through systematic analysis across diverse reasoning scenarios, we show that different modality organizations exhibit distinct advantages and limitations, and no single topology is universally optimal. Rather than proposing a general-purpose improvement to Omni-MLLM accuracy, we ask a narrower question: can these topology-induced failures be systematically identified and corrected? Motivated by this, we propose Chain of Modality (CoM), a framework that dynamically reorganizes multimodal topology during inference and learns adaptive organization strategies. Experiments across five benchmarks, diverse architectures, and model scales show that CoM reliably recovers topology-sensitive failures, through both training-free planning and lightweight Planner-SFT, while preserving performance on the full benchmark.
Ziyang Luo, Nian Liu, Junwei Han
Northwestern Polytechnical University Xi’an, Shaanxi, China
Omnimodal language models (OLMs) enable unified audio-visual understanding, but processing long joint token sequences makes inference computationally prohibitive. While recent token compression methods attempt to alleviate this burden, compressing modalities in isolation often destroys the temporal cross-modal anchors necessary for coherent reasoning. We introduce Omni2LoRA, a two-stage framework for efficient parametric memory compression via coherence-preserving context distillation that bypasses the token bottleneck entirely. First, a Perceiver hypernetwork processes intermediate representations from a frozen OLM to encode the multimodal context into a full-rank Low-Rank Adaptation (LoRA) adapter in a single forward pass. To prevent the resulting parameter footprint from scaling linearly with recording length, we optimize a discrete rank allocation policy via Group Relative Policy Optimization (GRPO) that uses a modality-ablated counterfactual reward to explicitly penalize the loss of audio-visual coherence, forcing the model to allocate its fixed sub-linear rank budget to synergistic cross-modal anchors rather than isolated visual features. Across three omnimodal backbones, Omni2LoRA operating at a 30% rank budget outperforms direct full-context inference and strong token-compression baselines (OmniZip, OMAC, O-MARC) on four audio-visual question answering benchmarks, improving average accuracy by 8-12% over the strongest baseline and remaining stable under compression ratios as tight as 75%, where token-pruning methods degrade sharply. By converting multimodal memory into a fixed-budget, reusable parameter state, our method drives answer-time multimodal-token load to zero, cutting per-query Time to First Token (TTFT) by up to 12x relative to full-context inference and amortizing to under 0.5s after a handful of queries.