Ranking-Aware Prompt Optimization for Multimodal Clinical Diagnosis
Organizations: Harvard University · University of California, Santa Cruz · Brown University · Johns Hopkins University · Georgia Institute of Technology
Abstract
Multimodal large language models (MLLMs) are rapidly advancing clinical diagnosis, yet their adaptation pipelines remain anchored to accuracy-based objectives. Clinical data are heavily class-imbalanced: a constant-majority predictor can score above 90% accuracy while being clinically useless. We therefore evaluate and optimize for AUROC, a threshold-free score that ranks positives above negatives and is invariant to class balance. We focus on prompt optimization in MLLMs. Reflective methods such as GEPA use a binary scores matrix with one row per evaluation instance and one column per candidate prompt; cells record per-instance correctness, so the column average is accuracy and drives candidate selection. We introduce pair-level Pareto prompt evolution (Ranking-PE), which replaces each correctness row with a pairwise-ordering row over (positive, negative) instance pairs: the cell is 1 if the candidate scores the positive higher than the paired negative. The column average then equals empirical AUROC (by the Wilcoxon-Mann-Whitney identity). We apply this swap at all three layers the prompt evolution search reads from - the scores matrix that decides Pareto dominance, the per-example feedback to the reflection LM, and final candidate selection - at no extra model calls and with no surrogate loss. Across three diseases on MIMIC, accuracy-based prompt evolution can degrade ranking; Ranking-PE reverses this, beating the accuracy-based recipe by +5.8 AUROC pp on fine-tuned Qwen3-VL-8B and +16.2 pp on MedGemma-4B. Ablations examine each design component and show that a medical-grade visual backbone - via vision-encoder-tuned SFT or medical pretraining - is a prerequisite that prompt search cannot replace - our recipe extends reflective prompt evolution from text-only data to multimodal clinical decision-making.
Figures & tables
| Atelectasis | Cardiomegaly | Consolidation | Avg | |||||
|---|---|---|---|---|---|---|---|---|
| Method | AUROC | Bal. Acc. | AUROC | Bal. Acc. | AUROC | Bal. Acc. | AUROC | Bal. Acc. |
| Qwen3-VL-8B | 65.4 | 64.0 | 61.3 | 51.2 | 74.0 | 64.4 | 66.9 | 59.9 |
| Gemma3-4B | 38.2 | 48.9 | 55.1 | 49.7 | 48.6 | 54.3 | 47.3 | 51.0 |
| GPT-4o-mini | 60.1 | 60.2 | 54.7 | 50.5 | 70.5 | 53.0 | 61.8 | 54.6 |
| GPT-5-mini | 50.6 | 58.2 | 71.5 | 66.8 | 75.9 | 64.6 | 66.0 | 63.2 |
| Gemini-2.5-flash | 46.5 | 54.3 | 69.0 | 58.3 | 75.1 | 66.4 | 63.5 | 59.7 |
| Model | Configuration | Atelectasis | Cardiomegaly | Consolidation | Avg |
|---|---|---|---|---|---|
| Qwen3-VL-8B | Base | 65.4 | 61.3 | 74.0 | 66.9 |
| Ranking-PE | 63.5 | 65.8 | 76.0 | 68.4 | |
| SFT | 63.3 | 68.0 | 72.6 | 68.0 | |
| SFT Ranking-PE | 70.7 | 71.3 | 79.3 | 73.8 | |
| Gemma3-4B | Base | 38.2 | 55.1 | 48.6 | 47.3 |
| Ranking-PE | 50.5 | 55.1 | 49.3 | 51.6 |
| Atelectasis | Cardiomegaly | Consolidation | Avg | |||||
|---|---|---|---|---|---|---|---|---|
| Method | AUROC | Bal. Acc. | AUROC | Bal. Acc. | AUROC | Bal. Acc. | AUROC | Bal. Acc. |
| Accuracy-PE | 63.6 | 45.4 | 68.1 | 58.0 | 72.4 | 61.8 | 68.0 | 55.1 |
| + AUROC final selection | 43.7 | 51.5 | 68.3 | 56.5 | 72.0 | 63.9 | 61.3 | 57.3 |
| + pair-level Pareto | 63.8 | 58.6 | 70.6 | 64.9 | 77.6 | 64.6 | 70.7 | 62.7 |
| + ranking feedback (full) | 70.7 | 70.9 | 71.3 | 66.5 | 79.3 | 68.5 | 73.8 | 68.6 |
| Ranking-PE (full) pair-level Pareto | 56.0 | 51.1 | 68.2 | 61.4 | 73.3 | 60.0 | 65.8 | 57.5 |
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
| max_pairs | AUROC | Bal. Acc. |
|---|---|---|
| (default) |
| Atelectasis | Cardiomegaly | Consolidation | Avg | |||||
| Method | AUROC | Bal. Acc. | AUROC | Bal. Acc. | AUROC | Bal. Acc. | AUROC | Bal. Acc. |
| Qwen3-VL-8B + SFT (VE-tuned) | ||||||||
| Accuracy-PE | 68.0 | 55.1 | ||||||
| BAcc-Select | 67.6 | 55.9 | ||||||
| Class-Weighted PE | 68.0 | 55.8 | ||||||
| Scalar-AUROC PE | 69.6 | 63.3 | ||||||
| Atelectasis | Cardiomegaly | Consolidation | Avg | |||||
|---|---|---|---|---|---|---|---|---|
| Method | AUROC | Bal. Acc. | AUROC | Bal. Acc. | AUROC | Bal. Acc. | AUROC | Bal. Acc. |
| Accuracy-PE | 68.0 | 55.1 | ||||||
| + AUROC final selection | 61.3 | 57.3 | ||||||
| + pair-level Pareto | 70.7 | 62.7 | ||||||
| + ranking feedback (full) | 73.8 | 68.6 | ||||||
| Ranking-PE (full) pair-level Pareto | 65.8 | 57.5 | ||||||
| Model | SFT setting | Atelectasis | Cardiomegaly | Consolidation | Avg |
|---|---|---|---|---|---|
| Qwen3-VL-8B | None | 68.4 | |||
| VE-frozen | 70.3 | ||||
| VE-tuned | 73.8 | ||||
| Gemma3-4B | None | 51.6 | |||
| MedGemma-4B | None | 69.7 |
| Atelectasis | Cardiomegaly | Consolidation | Avg | |||||
|---|---|---|---|---|---|---|---|---|
| Method | AUROC | Bal. Acc. | AUROC | Bal. Acc. | AUROC | Bal. Acc. | AUROC | Bal. Acc. |
| Self-reflector | 70.7 | 70.9 | 71.3 | 66.5 | 79.3 | 68.5 | 73.8 | 68.6 |
| GPT-5-nano reflector | 59.9 | 55.7 | 68.0 | 61.8 | 72.1 | 57.5 | 66.7 | 58.3 |
| GPT-5-mini reflector | 58.3 | 50.0 | 68.0 | 61.8 | 77.3 | 58.6 | 67.9 | 56.8 |
| Atelectasis | Cardiomegaly | Consolidation | Avg | |||||
| Method | AUROC | Bal. Acc. | AUROC | Bal. Acc. | AUROC | Bal. Acc. | AUROC | Bal. Acc. |
| Qwen3-VL-8B + SFT (VE-tuned) | 63.3 | 45.7 | 68.0 | 59.5 | 72.6 | 62.5 | 68.0 | 55.9 |
| + Accuracy-PE | 63.6 | 45.4 | 68.1 | 58.0 | 72.4 | 61.8 | 68.0 | 55.1 |
| + Ranking-PE (ours) | 70.7 | 70.9 | 71.3 | 66.5 | 79.3 | 68.5 | 73.8 | 68.6 |
| Cross-model transfer: prompt optimized on MedGemma-4B, evaluated on Qwen3-VL-8B + SFT | ||||||||
| + Transferred Ranking-PE prompt | 64.4 | 58.2 | 68.1 | 63.5 | 78.1 | 66.1 | 70.2 | 62.6 |
| Method | Atelectasis | Cardiomegaly | Consolidation | Avg |
|---|---|---|---|---|
| Qwen3-VL-8B | ||||
| Qwen3-VL-8B + SFT (VE-tuned) | ||||
| + Accuracy-PE | ||||
| + BAcc-Select | 86.3 | |||
| + Class-Weighted PE | 87.5 | |||
| + Scalar-AUROC PE | 87.1 |
| String appended to | |
|---|---|
| TRUE POSITIVE --- correctly identified as present. | |
| TRUE NEGATIVE --- correctly ruled out . | |
| FALSE NEGATIVE (clinically dangerous) --- missed . Re-examine the image for subtle evidence. | |
| FALSE POSITIVE --- over-called . Require converging evidence before committing. |
| Feedback level | AUROC | Bal. Acc. |
|---|---|---|
| Basic (correctness only) | 77.6 | 64.6 |
| + clinical error type | 78.2 | 65.5 |
| + confidence | 79.8 | 63.8 |
| + cross-example rank context (full) | 79.3 | 68.5 |