Evaluating vision encoders requires metrics that reliably predict their downstream performance in multimodal large language models (MLLMs). Although recent studies have shown that cross-modal metrics can better capture such performance, unimodal metrics remain the dominant choice in practice. In this work, we revisit cross-modal evaluation of vision encoders through large-scale experiments. We identify important limitations in both the experimental design and methodological formulation of prior approaches. After addressing these limitations and introducing simple improvements, we propose RAVEL, a training-free method based on cross-modal nearest-neighbor retrieval. Despite its simplicity, RAVEL achieves state-of-the-art performance across our experiments, outperforming prior methods by a substantial margin. Our results demonstrate that simple cross-modal metrics, when evaluated under a careful and comprehensive setup, can provide a strong basis for evaluating vision encoders for MLLMs.
Figures & tables
Figure 1: Effect of vision encoder candidate-pool size and heterogeneity on evaluation metrics. Cross-modal metrics outperform unimodal alternatives on moderately sized candidate pools, but their predictive performance drops substantially as more diverse vision encoders are added. In contrast, RAVEL, with improved representation conditioning and finer-grained similarity scoring, consistently outperforms competing metrics across different vision encoder pools.
Qwen3-1.7B
Qwen2.5-1.5B
SmolLM2-1.7B
Method
ρ↑
r↑
ρ↑
r↑
ρ↑
r↑
Vision-only metrics
Linear probing ( Caron et al., 2021 )
0.484
0.490
0.568
0.510
0.593
0.597
KNN ( Wu et al., 2018 )
0.551
0.513
0.534
0.598
0.603
0.567
Cross-modal metrics requiring additional model training
Mid-Training Loss
0.342
0.349
0.434
0.403
0.272
0.300
Table 1: FLOPs-controlled comparison of different evaluation metrics. All methods use matched computational budgets except AC Policy † , whose computational cost cannot be reduced to the matched FLOP budget due to its method design. Best and second-best results are shown in bold and underlined, respectively.
Table 3
Figure 2: Feature representation and patch-set matching in RAVEL. (a) Effect of the text feature extraction layer on ranking correlation. (b) Patch-set matching does not benefit long captions. (c) Patch differentiation under controlled feature intervention.
Figure 3: Representation post-processing choices. (a) Five post-processing methods and the unprocessed baseline. (b) Spearman change after removing high- or low-variance visual PCs. (c) Ablation of L2 normalization and spectral stabilization under conventional PCA whitening.
Figure 4: Patch-level matching preserves local structure. (a) Red lines link high-similarity patch pairs. (b) Color shows each patch’s best-match similarity (brighter = stronger).
Table 7
Figure 5: Robustness and scaling behavior.
Figure 6: Reconstruction quality does not determine MLLM utility. In the first example, the MLLM correctly identifies the author despite the unreadable reconstructed text. In the second, the reconstruction clearly preserves the book title, yet the MLLM produces an incorrect answer.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: Overview of RAVEL. Given paired image–text samples, RAVEL applies stabilized PCA whitening, constructs visual and textual neighborhoods using patch-level and global similarity, respectively, and measures their agreement through the known image–text correspondences.
Table 6: Complete vision encoder pool. All 70 evaluated architecture, checkpoint, and input-resolution configurations are listed and grouped by model family.
Large multimodal models (LMMs) inherit the self-attention mechanism of pretrained language backbones, yet standard attention can exhibit suboptimal allocation, including cross-modal misallocation between textual and visual evidence and intra-visual imbalance among visual tokens. We propose RAVE (Re-Allocating Visual Attention), a lightweight pair-gating mechanism that adds a learned query-key bias to pre-softmax attention scores over visual keys, derived from pre-RoPE query and key features. RAVE requires no architectural modification to the backbone and can be trained end-to-end with the rest of the model. Across a suite of multimodal benchmarks, RAVE improves over standard attention by an average of 3 points, with the largest gains on perception-intensive tasks -- including multilingual OCR, chart understanding, document VQA, and scene text VQA -- where accurate visual grounding is critical.
Xi Leng, Xinhong Ma, Ziqiang Dong +4
Qwen Business Unit of Alibaba · The Chinese University of Hong Kong, Shenzhen · Beijing Institute of Technology
Most modern multimodal large language models (MLLMs) build on a pretrained visual encoder that provides a strong visual prior. Encoder-free MLLMs instead learn visual representations directly from raw pixels, offering a simple and unified architecture, but their scaling behavior has not been systematically characterized. To fill this gap, we compare scaling laws for encoder-free and encoder-based MLLMs and report three main findings: (1) Removing the visual encoder shifts the compute-optimal allocation for the multimodal objective toward larger models, while leaving that for text nearly unchanged. (2) The two architectures exhibit nearly overlapping loss--compute frontiers on the text objective, but diverge on the multimodal objective: encoder-free models underperform at small scales yet are predicted to catch up at around 1022 FLOPs, well within practical pretraining budgets. (3) Without a visual encoder, the language model learns to take over its role via vision-specific adaptation: bidirectional interactions among visual tokens become increasingly beneficial as training compute grows, visual processing shifts toward earlier layers, and expert routing for visual tokens becomes more concentrated. Overall, our results indicate that the advantage of the visual prior provided by a pretrained encoder diminishes with scale, positioning encoder-free architectures as a promising direction for multimodal pretraining.
Lin Chen, Bolin Ni, Qi Yang +6
CASIA · UCAS · Foundation Model Department, Tencent
Vision-Language Models (VLMs) have enhanced traditional LLMs with visual capabilities through the integration of vision encoders. While recent works have explored various combinations of vision encoders and LLMs, there still lacks a principled understanding of what makes a vision encoder suitable for VLM alignment. In this paper, we systematically investigate this question via comprehensive experiments on a curated collection of 19 pre-trained vision encoders from diverse sources. We first demonstrate that common practices, such as choosing encoders with the largest size or highest zero-shot accuracy, consistently fail to identify optimal models. In fact, these metrics show only weak to moderate correlation with VLM performance. This intriguing finding begs a fundamental question: What factors of vision-encoders matter in VLM? Through comprehensive analysis, we identify that the structural similarity across modalities plays a crucial but previously overlooked role in vision-encoder selection, which we measure using the Gromov-Wasserstein distance as a proxy. From a theoretical perspective, we show that the learnability of cross-modality mapping can be provably associated with the Gromov-Wasserstein distance. Empirical verification on 60+ full VLM training runs shows that our proposed inference-only metric performs significantly better than alternative model selection strategies and exhibits a much stronger correlation with final VLM performance, thereby enabling efficient and effective prediction of VLM performance before full training.
Muyang Li, Yucheng Liu, Jianbo Ma +3
1Sydney AI Centre, The University of Sydney · 2Dolby Laboratories · 3TMLR Group, Hong Kong Baptist University