Medical Image Benchmarks

Latest papers 237

All topics
CardsList
  1. A Neuroimaging Simulation Framework for Developing and Evaluating Causal AI

    Jun 27, 2026Eryn Libert-Scott, Emma A. M. Stanley, Vibujithan Vigneshwaran +3Image GenerationCausal Discovery

  2. Aloe-Vision: Robust Vision-Language Models for Healthcare

    Jun 25, 2026Jaume Guasch-Martí, Enrique Lopez-Cuena, Martín Suárez-Fernández +3VLM EvaluationVLM Robustness

  3. CORTEX: A Structured Reasoning Benchmark for Trustworthy 3D Chest CT MLLMs

    Jun 25, 2026Hashmat Shadab Malik, Anees Ur Rehman Hashmi, Numan Saeed +3Radiology Report Generation3D Medical Imaging

  4. FunPiQ: A New Benchmark for Pixel-Level Quality Assessment in Fundus Images

    Jun 24, 2026Pengwei Wang, José Morano, Virginia Mares +1Fundus ImagingExplainable Medical Image Analysis

  5. Text Over Image: Auditing Multimodal Robustness in Synthetic Medical Image Detection

    Jun 24, 2026Ching-Hao Chiu, Hao-Wei Chung, Gelei Xu +7Multimodal RobustnessMedical VLMs

  6. Multilingual Hematology Visual Question Answering Dataset

    Jun 24, 2026Hajra Malik, Hafiza Tooba Aftab, Abdul Rehman +2Medical VQAMedical VLMs

  7. BenchX: Benchmarking AI Models for Cancer Detection and Localization with Demographic and Protocol Biases

    Jun 23, 2026Qi Chen, Wenxuan Li, Pedro R. A. S. Bassi +14Distribution ShiftMedical Image Analysis

  8. A Leakage-Aware Comparative Benchmark of Machine Learning, Deep Learning, and Transformer Models for Reliable Leukemia Detection

    Jun 23, 2026Nisreen AlbzourData LeakageMedical Image Classification

  9. Performance and Interpretability of Convolutional, Transformer, and Hybrid Deep Learning Models in Colorectal Histology Classification

    Jun 21, 2026Reza BozorgpourHybrid CNN-Transformer ArchitecturesVision Transformer

  10. Large Language Model-Assisted Cleaning of Report-Derived Labels in a Large-Scale Chest CT Dataset

    Jun 21, 2026Yosuke Yamagishi, Atsushi Takamatsu, Mototsugu Sato +4Training Data CurationLLM-Assisted Annotation

  11. SAGE: An Expert-Annotated South Asian GI Endoscopy Dataset for Multimodal Learning and Hallucination Analysis

    Jun 20, 2026Niyoj Oli, Sachin Acharya, Sandesh Pokhrel +7EndoscopyMedical Imaging

  12. Scaling up fine-grained intracranial vessel annotations in computed tomography angiography

    Jun 19, 2026Chu-Hsuan Lin, Alberto Mario Ceballos-Arroyo, Jisoo Kim +4Medical Image SegmentationComputed Tomography

  13. NoduLoCC2026: Lung Nodule Localization and Classification Contest from Chest X-Ray Images

    Jun 19, 2026Adnan Mustafic, Halim Benhabiles, Adnane Cabani +19Chest X-Ray ClassificationMedical Image Anomaly Detection

  14. MEDLAYXPLAIN: Benchmarking the Expert-Lay Gap in Medical Vision-Language Models

    Jun 19, 2026Han Jang, Junhyeok Lee, Songsoo Kim +4VLM EvaluationMedical VLMs

  15. MammoExpert: Benchmarking Chain-of-Thought Reasoning in Mammography Diagnosis

    Jun 19, 2026Di Dai, Bo Liu, Youcheng Li +9Medical ImagingCoT Reasoning

  16. CheXpercept: A Benchmark for Evaluating Expert-Level Lesion Perception in Chest X-rays

    Jun 19, 2026Geon Choi, Hangyul Yoon, Nalee Kim +5VLM EvaluationMedical Imaging

  17. GIM-ENDO: A Multimodal Endoscopic Image and Video Dataset for Gastric Intestinal Metaplasia Morphology and Pathology

    Jun 18, 2026Mojgan Forootan, Mahziar Setayeshfar, Ali Darvishi +2EndoscopyMedical Imaging

  18. HEad and neCK TumOR (HECKTOR) 2025: Benchmark of Segmentation, Diagnosis, and Prognosis in Multimodal PET/CT

    Jun 18, 2026Numan Saeed, Salma Hassan, Shahad Hardan +27Cancer Survival PredictionPET/CT Imaging

  19. A Controlled Benchmark of Quantum-Latent GAN Augmentation for Brain MRI

    Jun 17, 2026Syed Mujtaba Haider, Silvia FiginiSynthetic Data AugmentationQuantum Generative Modeling

  20. A Multi-Center Benchmark for Abdominal Disease Diagnosis and Report Generation from Non-Contrast CT

    Jun 15, 2026Mariam Elbakry, Aliaa Sayed Sheha, Salma Hassan Tantawy +5Radiology Report GenerationMedical Image Classification

  21. Federated Medical Image Segmentation under Real-World Label Noise: A Benchmark Suite for Noisy Label Learning Method Selection

    Jun 15, 2026Markus Bujotzek, Dimitrios Bounias, Stefan Denner +4Medical Image SegmentationMedical Image Benchmarks

  22. A Comprehensive Survey of Medical Image Segmentation: Challenges, Benchmarks, and Beyond

    Jun 15, 2026Pengyu Zhu, Xiaojing Zhang, Kunbo Zhang +2U-NetMedical Image Segmentation

  23. How Seemingly Inconsequential Design Choices Dictate Performance of LLMs in Pathology

    Jun 10, 2026Kian R. Weihrauch, Thomas A. Buckley, William Lotter +1Multimodal Large Language ModelsHistopathology Image Classification

  24. Atlas H&E-TME: Scalable AI-Based Tissue Profiling at Expert Pathologist-Level Accuracy

    Jun 10, 2026Kai Standvoss, Miriam Hägele, Rosemarie Krupar +25Computational PathologyPathology Foundation Models

  25. OpenMedReason: Scientific Reasoning Supervision for Medical Vision-Language Models

    Jun 10, 2026Negin Baghbanzadeh, Pritam Sarkar, Michael Colacci +6Medical VQAMedical VLMs

  26. From Patches to Patients: A study of the tile-to-slide performance transferability in Digital Pathology

    Jun 9, 2026Sofiène Boutaj, Leo Fillioux, Maria Vakalopoulou +2Pathology Foundation ModelsMedical Image Benchmarks

  27. A Controlled Audit of Pretraining Contamination in Public Medical Vision-Language Benchmarks

    Jun 8, 2026Bruce Changlong Xu, Lan Wu, Alexander RyuVLM EvaluationMedical VQA

  28. C3VD-DEFCOL: A Deformable Colonoscopy Dataset with Time-Resolved 3D Ground Truth and Realistic Appearance

    Jun 5, 2026Ethan Luk, Mayank V. Golhar, Anthony Song +5Endoscopy3D Reconstruction