cs.CVOct 4, 2026

A Strong Baseline for Evaluating Vision Encoders in Multimodal Large Language Models

Authors: Yilin Yang, Jun-Tao Tang, Kengyi Wang, Siyuan Su, Gaoyong Luo, Mingda Chen

Organizations: School of Artificial Intelligence, Shanghai Jiao Tong University · Nanjing University · Fudan University · Independent Researcher

Abstract

Evaluating vision encoders requires metrics that reliably predict their downstream performance in multimodal large language models (MLLMs). Although recent studies have shown that cross-modal metrics can better capture such performance, unimodal metrics remain the dominant choice in practice. In this work, we revisit cross-modal evaluation of vision encoders through large-scale experiments. We identify important limitations in both the experimental design and methodological formulation of prior approaches. After addressing these limitations and introducing simple improvements, we propose RAVEL, a training-free method based on cross-modal nearest-neighbor retrieval. Despite its simplicity, RAVEL achieves state-of-the-art performance across our experiments, outperforming prior methods by a substantial margin. Our results demonstrate that simple cross-modal metrics, when evaluated under a careful and comprehensive setup, can provide a strong basis for evaluating vision encoders for MLLMs.

Figures & tables

Appendix figures & tables2 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. RAVE: Re-Allocating Visual Attention in Large Multimodal Models

    May 18, 2026Xi Leng, Xinhong Ma, Ziqiang Dong +4Cross-Modal AttentionLarge Multimodal Models

  2. Rethinking Model Selection in VLM Through the Lens of Gromov-Wasserstein Distance

    May 2, 2026Muyang Li, Yucheng Liu, Jianbo Ma +3Vision EncodersVision-Language Alignment