cs.CLOct 19, 2024

MultiViewDx: Evidence-Linked Multi-View Clinical Diagnosis

Authors: Junda Wang, Zonghai Yao, Yujan Ting, Eric Z. Chen, Hieu Tran, Hong Yu, Weijing Huang, Terrence Chen

Organizations: United Imaging Intelligence, MA, USA · Manning College of Information and Computer Sciences, University of Massachusetts Amherst, MA, USA · Department of Medicine, University of Massachusetts Medical School, Worcester, MA, USA · Miner School of Computer and Information Sciences, University of Massachusetts Lowell, MA, USA

Abstract

Medical multimodal large language models (MLLMs) can perform well on existing medical visual question answering (MedVQA) benchmarks, but their training data often does not match clinical diagnosis. Most supervision is organized around isolated images or short QA pairs, leaving two structures weakly specified: how evidence leads to a decision, and how views, series, modalities, and patient context from the same case are linked. We introduce MultiViewDx, a partly physician-validated multimodal instruction dataset for evidence-linked multi-view medical imaging diagnosis. MultiViewDx uses the clinical case as the supervision unit. It links imaging studies with patient context, normalizes heterogeneous reports into an evidence-linked workflow (evidence -> findings -> differential discussion -> diagnosis), and uses a unified image-text retriever to constrain instruction synthesis to source-supported evidence. It covers X-ray, CT, MRI, ultrasound, histopathology, and other clinical visual sources. We fine-tune MultiViewDx-8B-AN and evaluate it on both existing MedVQA benchmarks and real-world case-based diagnostic reasoning. Across four MedVQA benchmarks, it achieves the best average accuracy among compared systems (79.0%), outperforming HuatuoGPT-Vision-34B (66.7%) and Claude3-Opus (55.7%). Beyond MedVQA, on JAMA Clinical Challenge cases, it receives the strongest overall rating under a physician-designed rubric for key clinical points, diagnostic inference, and evidence grounding. Controlled ablations and clinician evaluation show that both case-level multi-view organization and evidence-linked reasoning targets contribute to the gain.

Figures & tables

Appendix figures & tables11 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. OpenMedReason: Scientific Reasoning Supervision for Medical Vision-Language Models

    Jun 10, 2026Negin Baghbanzadeh, Pritam Sarkar, Michael Colacci +6Medical Vision-Language ModelsRecent Vision-Language Models

  2. DDX-TRACE: A Benchmark for Medical Diagnostic Trajectories in VLMs

    May 22, 2026Jiazhen Pan, Weixiang Shen, Jun Li +7Patient TrajectoriesTraces

  3. MedQA-MM: Shortcuts Behind Medical Visual Reasoning

    Sep 3, 2026Benlu Wang, Yifan Zhang, Jiaqing Yu +7Medical Visual Question AnsweringMultimodal Clinical Data