cs.LGSep 30, 2026

Ranking-Aware Prompt Optimization for Multimodal Clinical Diagnosis

Authors: Tian Xia, Minghao Liu, Yiqing Liang, Laixi Shi, Jiayun Wang

Organizations: Harvard University · University of California, Santa Cruz · Brown University · Johns Hopkins University · Georgia Institute of Technology

Abstract

Multimodal large language models (MLLMs) are rapidly advancing clinical diagnosis, yet their adaptation pipelines remain anchored to accuracy-based objectives. Clinical data are heavily class-imbalanced: a constant-majority predictor can score above 90% accuracy while being clinically useless. We therefore evaluate and optimize for AUROC, a threshold-free score that ranks positives above negatives and is invariant to class balance. We focus on prompt optimization in MLLMs. Reflective methods such as GEPA use a binary scores matrix with one row per evaluation instance and one column per candidate prompt; cells record per-instance correctness, so the column average is accuracy and drives candidate selection. We introduce pair-level Pareto prompt evolution (Ranking-PE), which replaces each correctness row with a pairwise-ordering row over (positive, negative) instance pairs: the cell is 1 if the candidate scores the positive higher than the paired negative. The column average then equals empirical AUROC (by the Wilcoxon-Mann-Whitney identity). We apply this swap at all three layers the prompt evolution search reads from - the scores matrix that decides Pareto dominance, the per-example feedback to the reflection LM, and final candidate selection - at no extra model calls and with no surrogate loss. Across three diseases on MIMIC, accuracy-based prompt evolution can degrade ranking; Ranking-PE reverses this, beating the accuracy-based recipe by +5.8 AUROC pp on fine-tuned Qwen3-VL-8B and +16.2 pp on MedGemma-4B. Ablations examine each design component and show that a medical-grade visual backbone - via vision-encoder-tuned SFT or medical pretraining - is a prerequisite that prompt search cannot replace - our recipe extends reflective prompt evolution from text-only data to multimodal clinical decision-making.

Figures & tables

Appendix figures & tables10 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. LLM-Guided Evolution for Medical Decision Pipelines

    Jun 5, 2026Ivan Sviridov, Artem Oskin, Ivan Panin +4TriageAlphaevolve

  2. Evaluating Multi-Turn Multimodal Diagnostic Reasoning on Challenging Real-World Clinical Cases

    Jul 28, 2026Rui Yang, Weihao Xuan, Yi Lin +23Multimodal Clinical DataClinical Reasoning

  3. Fathom-Vaidya: Advancing Medical Reasoning with Rubric-Based Rewards

    Sep 21, 2026Kalash Shah, Kunal Singh, Snehan J +1