cs.CVDate pending

Solve the Missing First Step: Can VLMs Standardize Raw Heterogeneous Medical Data?

Authors: Xin ChenDongliang XuCunhao ZhuXudong LuoHaoyang LyuXiaoxiao SunSerena Yeung-LevyYue Yao

Organizations: Shandong University · Stanford University

Abstract

As vision-language models (VLMs) are increasingly applied to medical AI, existing benchmarks mainly focus on evaluating their diagnostic ability over given medical images and texts, implicitly assuming that standardized medical images, texts, or question-answer pairs are already prepared. However, this assumption does not hold when we apply VLMs in real clinical practice, where medical data is often raw, heterogeneous, and fragmented across different sources. In this paper, we study this missing step, i.e., raw medical data standardization. Specifically, models are given raw dataset folders and evaluated on their ability to identify source formats, convert raw medical images into VLM-compatible visual inputs, extract relevant textual information, and organize the results into structured image-text pairs. To construct this Medical Data Standardization Benchmark (MDS-Bench), we manually annotate 1,939 raw medical data standardization tasks covering diverse clinical practice, radiology modalities, annotation formats, and directory layouts. Extensive experiments show that even the best performing VLM, i.e., Gemini 3 Flash, achieves only a 48.6% end-to-end success rate. Our research highlights raw medical data standardization as a critical bottleneck for medical AI diagnosis in real practice.

Explore similar work

CardsList
  1. Aloe-Vision: Robust Vision-Language Models for Healthcare

    Jun 25, 2026Jaume Guasch-Martí, Enrique Lopez-Cuena, Martín Suárez-Fernández +3Medical Vision-Language ModelsLarge Vision Language Models