Organizations: School of Informatics, Xiamen University · Institute of Artificial Intelligence, Xiamen University · Zhejiang Expressway Co., Ltd. · School of Transportation Sclence and Engineering, Beihang University · School of Science and Engineering, Chinese University of Hong Kong · Shanghai Artificial Intelligence Laboratory
Medical Anomaly Detection (MedAD) offers a promising direction for medical image analysis with Large Multimodal Models (LMMs). However, progress is limited by fragmented datasets and the tendency of Supervised Fine-Tuning (SFT) to learn superficial image-text correlations rather than verifiable diagnostic reasoning. Consequently, current models often generate fluent explanations that are either insufficiently grounded in the image or inconsistent with their final answers, limiting their reliability in high-stakes medical applications. To address these issues, we introduce MedAD-38K, a large-scale, multimodal, and multicenter benchmark containing structured Visual Question Answering pairs and quality-controlled diagnostic Chain-of-Thought annotations across five core MedAD tasks. Based on this benchmark, we propose a two-stage framework. Cognitive Injection first uses SFT to inject domain-specific medical knowledge and establish a structured think-then-answer format. Consistency Group Relative Policy Optimization (Con-GRPO) then employs an Evidence-Aware Consistency Reward to reinforce reasoning that remains grounded in the image and logically supports the final answer. The resulting MedAD-R1 achieves state-of-the-art performance on MedAD-38K and consistently outperforms all evaluated baselines across five source-disjoint external datasets spanning diverse imaging modalities. Beyond accuracy, it achieves higher reasoning-answer consistency and visual-grounding scores across multiple backbones. With only 0.8B parameters, MedAD-R1 offers practical potential for resource-constrained deployment. These results demonstrate the generality of evidence-aware consistency optimization for interpretable MedAD. Project resources are available at https://github.com/zhtstar/MedAD-R1.
Figures & tables
Figure 1: Comparison of traditional MedAD methods, general-purpose LMMs, and MedAD-R1. General LMMs may hallucinate findings (top) or produce reasoning inconsistent with the final answer (bottom), whereas MedAD-R1 remains visually grounded and answer-consistent.
Figure 2: Examples of the five MedAD VQA tasks in MedAD-38K. Questions are formulated as multiple-choice problems to support standardized evaluation.
Figure 3: Overview of the MedAD-38K construction and MedAD-R1 training pipeline. Top panel: source labels, masks, and metadata determine the VQA answers. MedGemma-27B generates auxiliary image descriptions, Gemini 2.5 Pro drafts the CoT annotations, and human reviewers verify their quality. Bottom panel: Cognitive Injection first adapts the model through SFT, followed by Con-GRPO with an Evidence-Aware Consistency Reward scored by MedGemma-4B along five dimensions: visual grounding (VG), task relevance (TR), answer support (AS), no-hallucination (NH), and no-contradiction (NC).
Figure 4: The equal-area five-region spatial grid used to derive lesion-localization labels in MedAD-38K.
Model
Params
Anat.
Anom.
Loc.
Mod.
Path.
Overall
Consist.
EOverall
EConsist.
Qwen3.5-0.8B
0.8B
94.22
57.88
33.74
93.45
14.93
76.43
92.12
68.12
88.75
InternVL3.5-1B
1B
95.04
50.09
33.53
89.27
15.52
73.10
93.53
69.81
91.53
Qwen3VL-2B
2B
91.20
55.54
37.08
92.57
10.75
74.91
94.84
67.35
94.41
Qwen2.5VL-3B
3B
78.56
50.29
30.63
78.97
11.04
64.90
93.25
51.07
92.01
MiMo-VL-7B
7B
61.71
55.20
30.04
63.71
28.81
56.86
94.57
39.49
93.78
Qwen2.5VL-7B
7B
75.43
56.10
26.90
82.28
15.52
66.30
94.17
50.12
91.84
Table 1: Comparison on MedAD-38K and the external zero-shot test set. Anat.: anatomy identification; Anom.: anomaly detection; Loc.: lesion localization; Mod.: modality classification; Path.: pathology characterization. Overall and Consist. denote accuracy and reasoning–answer consistency on MedAD-38K; EOverall and EConsist. denote the corresponding external-set metrics. Overall values are micro-averaged over applicable VQA instances. ∗ marks medical LMMs and † marks larger models fine-tuned on our training data. Boldface denotes the best result and underlining the second best. Means and standard deviations over five inference passes are reported in supplementary Tables C.1 and C.4 .
Backbone
Setting
Overall ↑
Consist. ↑
EOverall ↑
Qwen3.5-0.8B
Zero-shot
76.43
92.12
68.12
SFT only
90.89
94.35
88.36
Con-GRPO
95.12
97.26
90.09
InternVL3.5-1B
Zero-shot
73.10
93.53
69.81
SFT only
86.40
93.75
81.63
Con-GRPO
92.31
96.73
83.72
Table 2: Cross-backbone effectiveness and ablation. For each backbone we report zero-shot, SFT-only, and full Con-GRPO. Overall and Consist. are measured on MedAD-38K; EOverall is measured on the external set. All metrics are reported as percentages, and higher is better ( ↑ ).
Configuration
Overall ↑
Consist. ↑
RL-only (no SFT)
81.84 ± 2.41
88.90
SFT only
90.89 ± 0.27
94.35
SFT + Acc
92.76 ± 0.44
93.71
SFT + Acc + EAC (Ours)
95.12 ± 0.04
97.26
Table 3: Ablation of the training stages and reward terms on MedAD-38K. Acc: accuracy reward only. EAC: Evidence-Aware Consistency Reward. All values are percentages; higher is better.
Configuration
λfmt
λacc
λeac
Overall ↑
Format-focused
0.8
0.1
0.1
89.34
Accuracy-focused
0.1
0.8
0.1
93.12
Consistency-focused
0.1
0.1
0.8
93.88
Balanced (Ours)
1/3
1/3
1/3
95.12
Table 4: Ablation of Con-GRPO reward weights on MedAD-38K. λeac weights the Evidence-Aware Consistency Reward.
Appendix figures & tables24 assets
Supplementary material from the paper’s appendix.
Appendix
Source
Modality
Anatomy
Normal
Abnormal
Images
VQA Instances
ATLAS
CE-MRI
Liver
2,571
2,173
4,744
16,405
BACH
Microscopy
Breast
100
300
400
1,500
BUSI
Ultrasound
Breast
133
647
780
2,987
Br35H
MRI
Brain
1,500
801
2,301
7,704
BraTS2021
MRI
Brain
8,179
3,109
11,288
36,973
CVC-ClinicDB
Endoscopy
Gastrointestinal
0
612
612
2,448
Appendix
Table A.1: Source-level composition of MedAD-38K. VQA counts differ across sources because lesion localization and pathology characterization are generated only when the required source annotations are available.
Table 10Table 11
Task
Source of Correct Answer
Candidate Pool
Applicability
Modality Classification
Source modality metadata
10 modalities
All images
Anatomy Identification
Source anatomy metadata
10 anatomical regions
All images
Anomaly Detection
Source normality label
Yes / No
All images
Lesion Localization
Source segmentation mask
Five spatial regions
Abnormal images with masks
Pathology Characterization
Source disease label
Predefined pathology set
Images with disease labels
Appendix
Table A.6: Mapping from source annotations to the five VQA tasks.
Figure A.1: Overall token-count distribution of the generated CoT annotations. The histogram (left) marks the mean and median with dashed lines, while the violin and box plots (right) show the density, quartiles, and long-tail outliers. Token counts use the Qwen3.5 tokenizer and include only the text inside the <think>...</think> tags, excluding all tags and final answers. Across the complete training set ( N=78,121 ), the mean is 68.1 tokens, the median is 67.0, the standard deviation is 10.9, and the observed range is 36–266 tokens.
Figure A.2: Token-count distributions of the CoT annotations for the five MedAD tasks. Boxes represent the interquartile ranges, horizontal lines indicate medians, and red markers indicate means. Pathology characterization requires the longest rationales on average, while lesion localization uses the shortest; anomaly detection exhibits the most pronounced upper-tail outliers.
Figure A.3: Word cloud of the generated CoT annotations, with word size proportional to corpus frequency. The dominant terms describe anatomy, imaging modalities, tissue appearance, visual evidence, and spatial relationships.
Final Disposition Category
Traces
Proportion
Accepted without modification (Round 1)
81,172
62.37%
Accepted after limited manual revision
10,997
8.45%
Accepted after iterative regeneration
37,977
29.18%
Total Retained
130,146
100.00%
Appendix
Table A.7: Composition of the final 130,146 retained CoT annotations for MedAD-38K after the iterative refinement process.
Dimension
Agreement
Macro-F1
Cohen’s κ
Visual Grounding
91.20%
0.865
0.768
Task Relevance
95.60%
0.924
0.872
Answer Support
93.80%
0.895
0.821
No Hallucination
92.50%
0.878
0.795
No Contradiction
94.70%
0.912
0.854
Average
93.56%
0.895
0.822
Appendix
Table A.8: Inter-annotator agreement on 1,000 independently reviewed CoT annotations. Agreement is reported as a percentage; Macro-F1 and Cohen’s κ are unitless.
Split
Images
Image Ratio
VQA
VQA Ratio
Test
15,360
40.0%
52,025
40.0%
Train-SFT
16,150
42.0%
54,709
42.0%
Train-RL
6,912
18.0%
23,412
18.0%
Training
23,062
60.0%
78,121
60.0%
Total
38,422
100.0%
130,146
100.0%
Appendix
Table A.9: Image-level split of MedAD-38K. All VQA instances associated with the same image remain in one subset.
Dimension
Initial Prompt
Calibrated Prompt
Agreement
κ
Agreement
κ
Visual Grounding
82.40%
0.615
95.80%
0.892
Task Relevance
86.50%
0.712
97.20%
0.935
Answer Support
84.10%
0.668
96.50%
0.914
No Hallucination
81.20%
0.584
95.40%
0.885
No Contradiction
85.30%
0.695
96.80%
0.921
Appendix
Table B.1: Agreement between MedGemma-4B and human annotations on the held-out calibration set, before and after prompt calibration.
Judge Pair
Agreements
Samples
Rate
GLM-4.6V-Flash vs. InternVL3.5-8B
873
1,000
87.30%
GLM-4.6V-Flash vs. Qwen3-VL-8B-Instruct
909
1,000
90.90%
InternVL3.5-8B vs. Qwen3-VL-8B-Instruct
916
1,000
91.60%
Average
2,698
3,000
89.93%
Appendix
Table B.2: Agreement among the three evaluation judges on 1,000 outputs generated by the four SFT-stage backbone models from randomly sampled MedAD-38K training instances.
Model
Params
Anat.
Anom.
Loc.
Mod.
Path.
Overall
Qwen3.5-0.8B
0.8B
94.22 ± 0.53
57.88 ± 0.84
33.74 ± 0.96
93.45 ± 0.65
14.93 ± 2.70
76.43 ± 0.49
InternVL3.5-1B
1B
95.04 ± 1.94
50.09 ± 3.00
33.53 ± 2.83
89.27 ± 6.81
15.52 ± 8.19
73.10 ± 2.63
Qwen3VL-2B
2B
91.20 ± 6.27
55.54 ± 1.67
37.08 ± 5.74
92.57 ± 2.44
10.75 ± 9.13
74.91 ± 3.34
Qwen2.5VL-3B
3B
78.56 ± 8.33
50.29 ± 6.69
30.63 ± 4.75
78.97 ± 6.00
11.04 ± 12.19
64.90 ± 6.05
MiMo-VL-7B
7B
61.71 ± 8.46
55.20 ± 2.15
30.04 ± 2.94
63.71 ± 7.87
28.81 ± 5.55
56.86 ± 5.47
Qwen2.5VL-7B
7B
75.43 ± 8.82
56.10 ± 1.46
26.90 ± 1.78
82.28 ± 5.38
15.52 ± 10.40
66.30 ± 4.60
Appendix
Table C.1: Detailed MedAD-38K test accuracy over five independent inference passes of each fixed checkpoint, reported as mean ± standard deviation (%). No retraining is performed between passes. Overall is micro-averaged over all applicable VQA instances. ∗ marks medical LMMs and † marks larger models fine-tuned on our training data. Boldface denotes the best result and underlining the second best.
Model
Params
Vis. Ground.
Task Rel.
Ans. Supp.
No Hall.
No Contr.
Prompt A
Consist.
Insuff.
Qwen3.5-0.8B
0.8B
78.56
98.82
81.45
96.83
97.54
4.53
92.12
4.45
InternVL3.5-1B
1B
76.24
97.63
77.82
96.87
96.15
4.45
93.53
0.45
Qwen3-VL-2B
2B
80.56
97.68
82.15
96.82
97.53
4.55
94.84
0.42
Qwen2.5-VL-3B
3B
60.15
96.54
62.52
96.91
97.23
4.13
93.25
2.74
MiMo-VL-7B
7B
50.24
96.51
52.49
94.56
95.55
3.89
94.57
0.47
Qwen2.5-VL-7B
7B
75.42
97.83
76.68
97.16
97.66
4.45
94.17
3.22
Appendix
Table C.2: Main results on the pooled outputs from all five MedAD-38K inference passes. The five dimensions average the three independent judge scores, and Prompt A is their sum on a 0–5 scale; Consist. and Insuff. use the majority-voted Prompt-B verdict. SFT-only rows are the Stage-1 counterparts of the four complete MedAD-R1 models. Vis. Ground.: Visual Grounding; Task Rel.: Task Relevance; Ans. Supp.: Answer Support; No Hall.: No Hallucination; No Contr.: No Contradiction. All metrics except Prompt A are percentages. Higher is better except for Insuff. Boldface marks the best value in each column. ∗ marks a medical LMM and † a larger model fine-tuned on our data.
Dataset
Modality
Anatomy
Images
VQA
ColonDB
Endoscopy
Gastrointestinal
380
1,520
HeadCT
CT
Brain
200
800
EndoTect (Endo)
Endoscopy
Gastrointestinal
200
800
TN3K
Ultrasound
Thyroid
3,493
13,972
NMSC Histopathology
Microscopy
Skin
290
1,160
Total
4,563
18,252
Appendix
Table C.3: Composition of the five source-disjoint external evaluation datasets.
Model
Params
Anatomy ID.
Anomaly Det.
Lesion Loc.
Modality Class.
Pathology Char.
Overall
Qwen3.5-0.8B
0.8B
72.80 ± 0.97
55.80 ± 4.87
43.94 ± 0.98
99.29 ± 0.27
71.20 ± 6.79
68.12 ± 1.25
InternVL3.5-1B
1B
89.26 ± 2.94
55.82 ± 7.12
37.40 ± 5.57
95.73 ± 7.52
81.20 ± 6.11
69.81 ± 4.95
Qwen3-VL-2B
2B
85.88 ± 10.10
42.60 ± 3.01
42.26 ± 9.59
98.21 ± 1.77
62.00 ± 17.03
67.35 ± 5.53
Qwen2.5-VL-3B
3B
45.16 ± 5.73
39.45 ± 7.26
31.40 ± 5.51
87.60 ± 3.75
60.40 ± 15.81
51.07 ± 4.79
MiMo-VL-7B
7B
42.79 ± 11.25
39.61 ± 7.48
26.31 ± 1.23
48.50 ± 7.87
58.00 ± 19.06
39.49 ± 5.86
Qwen2.5-VL-7B
7B
51.31 ± 10.54
34.87 ± 8.69
27.22 ± 2.18
85.77 ± 7.91
82.20 ± 11.44
50.12 ± 5.88
Appendix
Table C.4: External-test accuracy over five independent inference passes of each fixed checkpoint, reported as mean ± standard deviation (%). No retraining is performed between passes. Overall is micro-averaged across all 18,252 applicable VQA instances rather than computed as an unweighted mean of the task columns. Boldface marks the best value in each column. ∗ marks a medical LMM and † a larger model fine-tuned on our data.
Model
Params
Vis. Ground.
Task Rel.
Ans. Supp.
No Hall.
No Contr.
Prompt A
Consist.
Insuff.
Qwen3.5-0.8B
0.8B
72.34
97.72
74.19
97.41
98.20
4.40
88.75
6.76
InternVL3.5-1B
1B
68.96
98.96
70.17
97.09
96.14
4.31
91.53
0.22
Qwen3-VL-2B
2B
78.24
98.93
79.25
97.29
98.65
4.52
94.41
0.18
Qwen2.5-VL-3B
3B
55.17
97.55
56.89
96.91
98.33
4.05
92.01
2.33
MiMo-VL-7B
7B
46.93
97.50
48.58
95.84
97.92
3.87
93.78
0.13
Qwen2.5-VL-7B
7B
69.30
98.79
70.39
98.07
98.72
4.35
91.84
3.79
Appendix
Table C.5: Reasoning-quality results on the pooled outputs from all five external-set inference passes. The five dimensions average the three independent judge scores, and Prompt A is their sum on a 0–5 scale; Consist. and Insuff. use the majority-voted Prompt-B verdict. SFT-only rows are the Stage-1 counterparts of the four complete MedAD-R1 models. The abbreviations and symbols follow Table C.2 ; all metrics except Prompt A are percentages, and lower is better only for Insuff.
Backbone
Setting
Anatomy
Modality
Anomaly
Localization
Pathology
Overall
Δ
Qwen3.5-0.8B
Zero-shot
94.22
93.45
57.88
33.74
14.93
76.43
–
SFT only
95.11
95.07
82.73
90.68
81.09
90.89
–
Con-GRPO
98.96
99.57
88.30
91.49
85.57
95.12
+4.23
InternVL3.5-1B
Zero-shot
95.04
89.27
50.09
33.53
15.52
73.10
–
SFT only
92.58
91.51
76.24
83.61
79.10
86.40
–
Con-GRPO
97.88
98.51
81.40
90.45
80.85
92.31
+5.91
Appendix
Table C.6: Per-task accuracy for the four backbones under zero-shot, SFT-only, and complete Con-GRPO training. Overall is micro-averaged over all applicable test instances rather than computed as an unweighted mean of the task columns. Boldface marks the best value in each metric column.
Setting
Stage 1 (SFT)
Stage 2 (Con-GRPO)
Training subset
Train-SFT
Train-RL
Epochs
2
1
Training GPUs
6
6 policy + 2 reward
Per-device batch size
8
1
Gradient accumulation
1
2
Effective prompt batch
48
12
Appendix
Table C.7: Optimization settings for the two MedAD-R1 training stages.
Model
Params (B)
TTFT (s)
TPOT (ms)
Tokens/s
E2E (s)
Peak Alloc. (GB)
Peak Reserv. (GB)
Lingshu-7B
7.0
0.0932
25.50
39.22
6.625
16.73
16.98
InternVL3.5-8B
8.0
0.0610
26.40
37.88
6.831
17.18
17.34
Qwen3-VL-8B
8.0
0.2910
29.60
33.78
7.864
17.67
17.86
GLM-4.6V-Flash
9.0
0.0922
30.55
32.73
7.921
20.73
20.88
Llama-3.2-Vision
11.0
0.2845
31.20
32.05
8.267
22.30
22.90
Gemma-4-26B
26.0
0.2261
50.83
19.67
13.238
52.00
52.21
Appendix
Table D.1: Inference efficiency under the fixed-length setting (256 output tokens). TTFT: time to first token. TPOT: time per output token. E2E: end-to-end latency. Peak Alloc. and Peak Reserv. are peak allocated and reserved GPU memory. MedAD-R1 variants are grouped at the bottom. Boldface marks the best value in each column.
Model
Params (B)
TTFT (s)
TPOT (ms)
E2E (s)
Out. Tokens
Peak Alloc. (GB)
Peak Reserv. (GB)
Lingshu-7B
7.0
0.0933
25.50
3.618
138.2
16.73
16.98
InternVL3.5-8B
8.0
0.0612
26.40
4.925
184.3
17.22
17.34
Qwen3-VL-8B
8.0
0.2940
29.60
5.490
175.5
17.71
17.86
GLM-4.6V-Flash
9.0
0.0932
30.55
15.734
512.0 ∗
20.73
20.88
Llama-3.2-Vision
11.0
0.2847
31.20
6.166
188.6
22.30
22.90
Gemma-4-26B
26.0
0.2269
50.83
15.275
295.7
52.05
52.21
Appendix
Table D.2: Inference behavior under the natural setting, where each model stops on its own. End-to-end latency here is governed by output length and is not a controlled speed comparison; see Table D.1 for the fixed-length comparison. Boldface marks the lowest numerical value in each column; end-to-end latency and output length remain descriptive rather than controlled comparisons. ∗ GLM-4.6V-Flash reached the maximum generation limit of 512 tokens because it did not emit an explicit end-of-sequence token in this zero-shot setting.
Case
Gold
Ours
InternVL
Gemma
Huatuo ∗
LLaVA ∗
Lingshu ∗
MedVLM ∗
Liver-mass location
B
B
A
A
A
C
C
C
Thyroid finding
B
B
A
A
A
A
A
A
Appendix
Table E.1: Answer-level summary of the two manually reviewed qualitative cases. ∗ denotes a domain-specific medical LMM. All 12 comparison answers are incorrect under the source-derived labels.
Figure E.1: Lesion-localization example (source-derived answer: B, Center-Region). The annotated finding spans a large portion of the liver and is assigned to the center region by the mask-based five-region rule.
Figure E.2: Thyroid-ultrasound anomaly-detection example (source-derived answer: B, Yes). A focal echotextural abnormality is distinguishable from the surrounding thyroid parenchyma.
High-stakes clinical use of large vision-language models (LVLMs) requires reasoning that is grounded in visual evidence and clinical knowledge, not just correct final answers. We introduce OpenMedReason, a large-scale, open multimodal medical reasoning corpus comprising approximately 450K image-question-answer instances whose reasoning traces are primarily derived from curated biomedical, human-authored scientific articles. OpenMedReason provides high-fidelity supervision beyond synthetic chains of thought, covering diverse medical domain vision modalities such as radiological scans, microscopic images, visible light photographs, charts, and others. We complement it with OpenMedReason-Bench, a held-out benchmark that allows fine-grained evaluation of LVLMs along three complementary axes of capability, including perception, medical knowledge, and rationale, enabling diagnostic evaluation beyond final-answer accuracy. OpenMedReason is a rich training resource that exhibits its effectiveness in both supervised fine-tuning (SFT) and reinforcement-based alignment. Training with OpenMedReason yields a 20% average improvement in VQA accuracy over the base model and achieves performance within 4.2% of the strongest comparable-scale medical LVLMs. Fine-grained performance analysis confirms that the gains are not concentrated in any single axis: OpenMedReason improves perception, medical knowledge, and rationale jointly, and its reasoning traces are preferred over those of the base model in 86.1% of pairwise comparisons. We release the code and dataset at huggingface.co/datasets/neginb/OpenMedReason.
Negin Baghbanzadeh, Pritam Sarkar, Michael Colacci +6
York University · Vector Institute · University of British Columbia +5
Chain-of-Thought (CoT) reasoning has proven effective in enhancing large language models by encouraging step-by-step intermediate reasoning, and recent advances have extended this paradigm to Multimodal Large Language Models (MLLMs). In the medical domain, where diagnostic decisions depend on nuanced visual cues and sequential reasoning, CoT aligns naturally with clinical thinking processes. However, current benchmarks for medical image understanding generally focus on the final answer while ignoring the reasoning path. Such opaque reasoning processes lack reliable bases for judgment, making it difficult to assist doctors in diagnosis. To address this gap, we introduce a new M3CoTBench benchmark specifically designed to evaluate the correctness, efficiency, impact, and consistency of CoT reasoning in medical image understanding. M3CoTBench features 1) a diverse, multi-level difficulty dataset covering 24 examination types, 2) 13 varying-difficulty tasks, 3) a suite of CoT-specific evaluation metrics (correctness, efficiency, impact, and consistency) tailored to clinical reasoning, and 4) a performance analysis of multiple MLLMs. M3CoTBench systematically evaluates CoT reasoning across diverse medical imaging tasks, revealing current limitations of MLLMs in generating reliable and clinically interpretable reasoning, and aims to foster the development of transparent, trustworthy, and diagnostically accurate AI systems for healthcare. Project page at https://juntaojianggavin.github.io/projects/M3CoTBench/.
Juntao Jiang, Jiangning Zhang, Yali Bi +7
1Zhejiang University · University of Science and Technology of China · 3East China Normal University +2
A benchmark score credits final answers, but not the route by which an item can be answered. In medical multimodal multiple-choice questions (MCQs), this distinction matters because a correct answer can be supported by the intended image finding or by benchmark-preserved cues in the wording of answers, non-visual clinical text, visible image text, artificial annotations, or device/context artifacts. We call the resulting score-level overinterpretation reasoning inflation. Here, a route is an observable input path that can support answer selection, not a claim about the model's hidden cognition. Across six medical multimodal MCQ datasets, we separate candidate cues from behavioral evidence through prompt- and image-side audits, modality ablations, and matched repairs that preserve the medical target and answer key. In a 13-configuration open-model panel, full-input accuracy is 62.63%, while text-only and options-only settings achieve 53.96% and 29.71%, respectively. Removing length-gap, absolute/conspicuous, and spatial/prepositional cues lowers accuracy by 6.58, 3.50, and 4.77 percentage points. We also construct MedQA-MM, a 1,000-item shortcut-mitigated subset, where text-only and options-only accuracy fall to 5.21% and 12.33%. This does not imply that models never use images; it shows that medical image-reasoning claims require route-level evidence.
Benlu Wang, Yifan Zhang, Jiaqing Yu +7
University of Massachusetts Amherst · 4Yale University · University of Massachusetts Lowell +6