cs.CVSep 29, 2026

ResComEmb: Effective and Efficient Multimodal Embedding via Residual Homogeneity Compression

Authors: Zijing Cai, Yuzhe Wang, Jingxian Zhu, Fengbin Zhu, Richang Hong

Organizations: University of Science and Technology of China · Hefei University of Technology · National University of Singapore

Abstract

Multimodal large language models (MLLMs) have shown strong potential for universal multimodal representation learning. However, existing methods either compress each input into a single vector, limiting fine-grained expressiveness, or retain long sequences of visual-token vectors, incurring substantial storage and interaction costs. To resolve this trade-off, we propose ResComEmb, a trainable framework for effective and efficient universal multi-vector multimodal embedding. ResComEmb first encodes each input at native dynamic resolution into ordered global, intermediate, and fine-grained views. After MLLM contextualization and embedding projection, a trainable Residual Homogeneity Compression (RHC) module reduces within-granularity redundancy and cross-granularity repetition under explicit visual token budgets. Then, ResComEmb introduces a length-adaptive Bidirectional Late-Interaction Matching mechanism for robust query-document scoring, which averages the strongest token-level matches in each direction and combines the two scores using a weight based on how many valid tokens each side has. Extensive experiments on MMEB, ViDoRe V1, and ViDoRe V2 show that ResComEmb produces higher-quality universal multimodal embeddings than VLM2Vec-V2, and outperforms ColQwen2.5 in visual document retrieval using only 37.5% of its full visual token budget, demonstrating a favorable effectiveness-efficiency trade-off.

Figures & tables

Appendix figures & tables5 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. CausalEmbed: Auto-Regressive Multi-Vector Generation in Latent Space for Visual Document Embedding

    Jan 29, 2026Jiahao Huo, Yu Huang, Yibo Yan +7Visual Document RetrievalMultimodal Embeddings

  2. VaME: Exploring Variational Latent Reasoning for Multimodal Embeddings

    Sep 27, 2026Peixi Wu, Mingzhou Jiang, Feipeng Ma +11Multimodal EmbeddingsMultimodal Retrieval