A Strong Baseline for Evaluating Vision Encoders in Multimodal Large Language Models
Organizations: School of Artificial Intelligence, Shanghai Jiao Tong University · Nanjing University · Fudan University · Independent Researcher
Abstract
Evaluating vision encoders requires metrics that reliably predict their downstream performance in multimodal large language models (MLLMs). Although recent studies have shown that cross-modal metrics can better capture such performance, unimodal metrics remain the dominant choice in practice. In this work, we revisit cross-modal evaluation of vision encoders through large-scale experiments. We identify important limitations in both the experimental design and methodological formulation of prior approaches. After addressing these limitations and introducing simple improvements, we propose RAVEL, a training-free method based on cross-modal nearest-neighbor retrieval. Despite its simplicity, RAVEL achieves state-of-the-art performance across our experiments, outperforming prior methods by a substantial margin. Our results demonstrate that simple cross-modal metrics, when evaluated under a careful and comprehensive setup, can provide a strong basis for evaluating vision encoders for MLLMs.
Figures & tables
| Qwen3-1.7B | Qwen2.5-1.5B | SmolLM2-1.7B | ||||
| Method | ||||||
| Vision-only metrics | ||||||
| Linear probing ( Caron et al., 2021 ) | 0.484 | 0.490 | 0.568 | 0.510 | 0.593 | 0.597 |
| KNN ( Wu et al., 2018 ) | 0.551 | 0.513 | 0.534 | 0.598 | 0.603 | 0.567 |
| Cross-modal metrics requiring additional model training | ||||||
| Mid-Training Loss | 0.342 | 0.349 | 0.434 | 0.403 | 0.272 | 0.300 |
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
| Family (#) | Architectures and input resolutions | Reference |
| Language-supervised continuous encoders | ||
| OpenAI CLIP (1) | ViT-L/14 (224) | Radford et al.,2021 |
| MetaCLIP (9) | ViT-B/16 (400M, 224; 2.5B, 224); ViT-B/32 (400M, 224; 2.5B, 224); ViT-L/14 (400M, 224; 2.5B, 224); ViT-H/14 (2.5B, 224; v1.2, 224); ViT-G/14 (2.5B, 224) | Xu et al.,2024 |
| MetaCLIP 2 (15) | ViT-B/16 (224, 384); ViT-B/32 (224, 384; mT5, 224); ViT-G/14 (224, 378); ViT-H/14 (378); ViT-L/14 (224); ViT-M/16 (224, 384; mT5, 224); ViT-S/16 (224, 384; mT5, 224) | Chuang et al.,2025 |
| SigLIP2 (15) | So400m/14 (224, 384); So400m/16 (256, 384, 512); ViT-B/16 (224, 256, 384, 512); ViT-B/32 (256); ViT-G/16 (256, 384); ViT-L/16 (256, 384, 512) | Tschannen et al.,2025 |
| Perception Encoder (3) | PE-Core-B/16 (224); PE-Core-G/14 (448); PE-Lang-L/14 (448) | Bolya et al.,2025 |