Omni-Embed-Mini: Binding Modalities Without Forgetting via Dense Distillation
Organizations: Mohamed Bin Zayed University of Artificial Intelligence (MBZUAI)
Abstract
Extending a text embedding model to new modalities typically degrades text retrieval quality, and existing omni-modal embedders compensate with multi-billion parameters. We present Omni-Embed-Mini, a 0.9B-parameter model that maps text, speech, audio, images, video, and visually-rich documents into a single shared cosine space without updating any text-side parameter. Our key insight is that the teacher signal requires no separate embedding model: each media sample is paired with a dense cascaded caption, and the teacher target is simply the frozen backbone's own embedding of that caption. Because teacher and student share the same backbone weights, they inhabit byte-identical geometry, and lightweight projectors plus phased LoRA adapters on the modality encoders suffice for alignment. Training combines a Matryoshka SigLIP contrastive loss with an online hybrid hard-negative miner whose negatives sharpen as the encoder improves. The recipe carries over to a 2.3B variant by swapping in a native vision-language backbone. Omni-Embed-Mini-0.9B keeps its text weights bit-identical to the backbone, so training cannot regress text retrieval (49.57 nDCG@10 on MTEB-v2 BEIR-8), while extending it to five additional modalities, and is ~2.7x to 9.5x smaller than every open omni embedder we compare against. The 2.3B variant is competitive with the closed gemini-embedding-2, edging ahead of it on the overall-modality average. Models, code, data and evaluation harness are on our project page: https://omniembed.cvmbzuai.com
Figures & tables
| Model | Backbone | Inference | Frozen | Trained | Trained % |
| Omni-Embed-Mini-0.9B | Qwen3-Embedding-0.6B | 935.35 M | 869.98 M | 68.32 M | 7.3% |
| Omni-Embed-Mini-2.3B | Qwen3-VL-Embedding-2B | 2.44 B | 2.30 B | 141.68 M | 5.8% |
| BidirLM-Omni-2.5B | Qwen3-1.7B merged | 2.45 B | 724 M | 1.72 B | 70.4% |
| omni-embed-nemotron-3B | Qwen2.5-Omni-3B Thinker | 4.70 B | 4.67 to 4.70 B | 7 to 30 M | 0.2 to 0.6% |
| e5-omni-3B | Qwen2.5-Omni-3B Thinker | 4.70 B | 4.58 to 4.70 B | 7 to 120 M | 0.2 to 2.5% |
| LCO-Embedding-Omni-3B | Qwen2.5-Omni-3B Thinker | 4.70 B | 4.58 B | 120 M | 2.5% |
| Style | Sp. | Aud. | Img. | Vid. | Overall |
| Original Captions | 41.75 | 31.02 | 17.63 | 4.74 | 23.78 |
| Dense Captions | 44.43 | 33.72 | 26.23 | 18.80 | 30.79 |
| 0.9B | 2.3B | |||||
| Modality | Frozen | +LoRA | Frozen | +LoRA | ||
| Text | 49.57 | 46.44 | 47.94 | 47.29 | ||
| Speech | 44.43 | 47.09 | 41.55 | 31.02 | ||
| Audio | 33.72 | 34.83 | 33.97 | 25.57 | ||
| Image | 26.23 | 30.04 | 64.80 | 64.34 | ||
| Video | 18.80 | 24.46 | 55.18 | 54.60 | ||
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
| Model | Params | ArgA | CQA-E | CQA-P | CQA-Pr | FiQA | NFC | SCID | SciF | Avg |
| Text embedders | ||||||||||
| voyage-3-m-exp | 6.9B | 89.79 | 69.98 | 73.71 | 65.87 | 78.01 | 46.99 | 34.53 | 85.09 | 68.00 |
| Conan-embedding-v2 | 1.5B | 77.79 | 57.50 | 60.58 | 60.66 | 71.86 | 55.74 | 31.63 | 84.27 | 62.50 |
| Qwen3-Embedding-8B | 7.6B | 76.85 | 57.44 | 58.31 | 53.49 | 64.57 | 41.45 | 32.74 | 78.46 | 57.91 |
| inf-retriever-v1 | 7.1B | 84.86 | 55.96 | 52.35 | 46.64 | 62.35 | 43.66 | 30.78 | 85.40 | 57.75 |
| Qwen3-Embedding-4B | 4.0B | 75.64 | 55.00 | 56.58 | 51.17 | 62.65 | 41.10 | 31.44 | 78.33 | 56.49 |
| Model | Params | Clotho HR@5 | MACS HR@5 | US8K HR@5 | BjOp. Acc | Bird Acc | GTZAN Acc | Mrd. Acc | Veh. V-meas. | FSD19 LRAP | GTZ-R MAP@1k | Avg |
| Audio–language models | ||||||||||||
| Qwen2-Audio-7B | 7.0B | 0.45 | 1.53 | 0.22 | 97.45 | 37.10 | 93.10 | 61.17 | 5.52 | 32.64 | 80.85 | 41.00 |
| CLAP-style dual encoders | ||||||||||||
| CLAP-LAION (general) | 194M | 33.72 | 33.08 | 0.96 | 93.62 | 17.00 | 84.50 | 49.45 | 3.37 | 80.63 | 66.78 | 46.31 |
| CLAP-LAION (music+speech) | 194M | 32.39 | 30.53 | 0.94 | 91.52 | 16.40 | 83.50 | 48.99 | 4.72 | 80.62 | 65.65 | 45.53 |
| MS-CLAP 2023 | 160M | 41.93 | 27.23 | 0.94 | 91.11 | 17.30 | 78.10 | 52.09 | 2.86 | 79.73 | 75.43 | 46.67 |
| Model | Params | GigaSp HR@5 | SpSQuAD HR@5 | CREMA Acc | AgeDet Acc | IEMOCAP Acc | VoxCel Acc | CREMA-C V-meas. | CREMA-P max-AP | NMSQA max-AP | VoxPop max-AP | Ravd ZS Acc | SpCmd ZS Acc | Avg |
| Audio–language models | ||||||||||||||
| Qwen2-Audio-7B | 7.0B | 0.09 | 1.00 | 73.99 | 17.59 | 92.96 | 29.54 | 32.37 | 68.87 | 48.85 | 52.93 | 14.37 | 10.38 | 36.91 |
| CLAP-style dual encoders | ||||||||||||||
| CLAP-LAION (general) | 194M | 0.16 | 2.00 | 39.83 | 20.51 | 89.28 | 30.47 | 13.17 | 54.71 | 47.42 | 53.55 | 17.29 | 12.44 | 31.74 |
| CLAP-LAION (music+speech) | 194M | 0.21 | 2.00 | 40.18 | 16.59 | 93.12 | 33.78 | 14.43 | 56.66 | 46.14 | 53.49 | 17.29 | 9.03 | 31.91 |
| MS-CLAP 2023 | 160M | 0.10 | 0.00 | 37.06 | 15.35 | 85.86 | 26.18 | 10.73 | 57.65 | 51.09 | 52.05 | 15.21 | 10.01 | 30.11 |
| Model | Params | VOC07 CLS | Ctry211 CLS | OKVQA VQA | EDIS T2I | COCO-T2I T2I | VN-T2I T2I | COCO-I2T I2T | VN-I2T I2T | NIGHTS I2I | FashIQ I2I | Avg |
| Vision–language embedders | ||||||||||||
| seed1.6-embedding-1215 | ? | 89.30 | 46.90 | 74.90 | 96.70 | 81.30 | 83.80 | 77.80 | 83.70 | 70.70 | 50.40 | 75.55 |
| seed-1.6-embedding | ? | 91.90 | 47.70 | 74.20 | 91.50 | 78.80 | 83.40 | 77.40 | 84.50 | 72.10 | 49.00 | 75.05 |
| Qwen3-VL-Embedding-8B | 8.1B | 93.40 | 27.30 | 77.80 | 96.30 | 81.10 | 81.10 | 79.10 | 85.70 | 72.70 | 44.50 | 73.90 |
| IFM-TTE-7B | 8.3B | 83.40 | 60.50 | 83.40 | 95.50 | 77.70 | 79.20 | 72.60 | 82.50 | 69.40 | 32.40 | 73.66 |
| WeMM-Embedding-8B | 8.8B | 94.50 | 32.20 | 73.40 | 94.90 | 81.20 | 81.70 | 78.20 | 84.90 | 69.00 | 43.80 | 73.38 |
| Model | Params | HMDB V-CLS | UCF V-CLS | MSRVTT V-T2V | MSVD V-T2V | DiDeMo V-T2V | VATEX V-T2V | Avg |
| Vision–language embedders | ||||||||
| seed1.6-embedding-1215 | ? | 93.50 | 98.70 | 60.20 | 74.48 | 66.43 | 54.87 | 74.70 |
| Qwen3-VL-Embedding-8B | 8.1B | 83.40 | 95.10 | 58.20 | 75.67 | 66.04 | 54.87 | 72.21 |
| WeMM-Embedding-8B | 8.8B | 61.50 | 84.20 | 55.30 | 73.43 | 69.22 | 55.05 | 66.45 |
| WeMM-Embedding-2B | 2.1B | 57.70 | 75.10 | 53.10 | 73.43 | 65.24 | 52.52 | 62.85 |
| IFM-TTE-7B | 8.3B | 65.40 | 79.60 | 52.70 | 73.13 | 49.70 | 51.45 | 62.00 |
| Model | Params | Fin | HR | Ind | Phr | CS | En | Phy | Avg |
| Visual-document retrievers | |||||||||
| colqwen3.5-4.5B-v3 | 4.6B | 68.69 | 66.09 | 59.35 | 67.31 | 79.87 | 66.16 | 50.08 | 65.36 |
| nomic-embed-multimodal-7b | 7.8B | 59.43 | 58.27 | 48.60 | 63.44 | 71.47 | 56.96 | 44.70 | 57.55 |
| nomic-embed-multimodal-3b | 3.8B | 57.84 | 54.88 | 46.63 | 60.35 | 70.29 | 54.96 | 45.17 | 55.73 |
| Vision–language embedders | |||||||||
| llama-nemotron-embed-vl-1b-v2 | 1.7B | 64.53 | 60.59 | 47.85 | 63.52 | 74.01 | 57.72 | 43.65 | 58.84 |
| Modality | Media (described) | Original caption | Dense caption |
| Audio | A woman speaks in a public space while a slow heartbeat-like thump and faint ticking play in the background. | A woman talks while a heartbeat and ticking play in the background. | A woman’s voice, speaking in a mid-to-low pitch, is present at the beginning of the clip, accompanied by faint, indistinct conversation from other people. Simultaneously, a deep, rhythmic thumping with a slow, steady tempo, resembling a heartbeat, begins and continues throughout. The background includes a persistent, low-level ambient hum and the general murmur of a public indoor space with some reverberation. |
| Speech | A young woman with a clear, neutral accent reads a short factual question in a calm, measured voice in a quiet indoor setting. | What are organic plants understood to be? | A young adult female with a high-pitched, clear, and articulate voice delivers the line in a standard, neutral accent. Her speaking style is measured and calm, with a moderate pace and consistent volume, conveying a neutral, informative affect. She says, “What are organic plants understood to be?” The audio is clean and crisp, likely recorded in a quiet indoor space with no discernible background noise or reverberation. |
| Image | Two women going down a conveyor belt with a cart full of groceries. | The image, taken from a high-angle perspective, shows two women on a descending escalator. The woman in the foreground, wearing a floral-patterned short-sleeved shirt, leans forward, gripping the handle of a metal shopping cart. Behind her, another woman with brown hair, wearing a white top over a dark shirt, also holds the cart’s handle. The cart is filled with various items, including a visible bag of potato chips and other grocery products. The setting appears to be indoors, with a neutral-toned, geometric-patterned wall in the background; lighting is even and diffuse, characteristic of an indoor public space. | |
| Video | There is a girl gymnast who is performing on the bars. | A gymnast, Cienna Samiley, performs on the uneven bars during the 2014 US Challenge warm-up. The scene is a large, indoor arena with sparse audience seating. The camera cuts to a podium where seven young female gymnasts stand, wearing leotards and medals, in front of a banner reading “2014 SECRET U.S. CLASSIC”. A scoreboard graphic is superimposed, displaying the top six scores. The gymnasts on the podium then raise their arms and cheer. | |
| Visual doc | (not available) | A detailed analysis of an infographic titled “Facts About Gender Inequality,” which uses a combination of text, icons, and data visualizations to present information on gender disparities. The infographic is set against a light teal background, structured into several sections, with a bold sans-serif title at the top … Icons use simple black silhouettes (gender symbols, a wheelbarrow for the workplace, a house for home, books for school); the colour scheme is teal/red/black; data is presented via a pie chart and a bar chart with a consistent red-for-women, blue-for-men encoding. |
| Source | Modality | Samples | License |
| TED-LIUM 3 Hernandez et al. (2018) | speech | 26,826 | CC BY-NC-ND 3.0 |
| HeySQuAD Wu et al. (2023a) | speech | 7,199 | CC BY 4.0 (HF declared); SQuAD text CC BY-SA 4.0 |
| Spoken Alpaca Shih et al. (2023) | speech | 5,135 | CC BY-NC 4.0 |
| Gemini Speech SB (2025) | speech | 4,726 | Apache 2.0 |
| LibriSpeech Panayotov et al. (2015) | speech | 2,854 | CC BY 4.0 |
| WavCaps (FreeSound) Mei et al. (2024) | audio | 26,217 | Academic / research only |
| Modality | s42 | s71 | s1234 | Mean | SD |
| Text | 49.57 | 49.57 | 49.57 | 49.57 | 0.00 |
| Speech | 43.93 | 41.95 | 43.41 | 43.10 | 1.03 |
| Audio | 33.28 | 34.22 | 34.37 | 33.96 | 0.59 |
| Image | 27.96 | 27.35 | 27.13 | 27.48 | 0.43 |
| Video | 20.86 | 19.61 | 18.99 | 19.82 | 0.95 |
| Visual Doc | 46.76 | 46.21 | 43.47 | 45.48 | 1.76 |
| Component | Params |
| Frozen text backbone (Qwen3-Embedding-0.6B) | 595.78 M |
| Vision encoder (ViT from Qwen3.5-0.8B) | 100.59 M |
| Whisper-small encoder | 88.15 M |
| Dasheng-base encoder | 85.46 M |
| Projectors and pooler | 65.37 M |
| Total at inference | 935.35 M |
| Component | Setting |
| Architecture | |
| Backbone, frozen | Qwen3-Embedding-0.6B |
| Hidden size | 1024 |
| MRL dimensions | |
| Vision encoder | ViT extracted from Qwen3.5-0.8B |
| Image resolution | |