BrainTRACE: Tracing Longitudinal, Multimodal, and Volumetric Evidence in Brain MRI Clinical Reasoning
Organizations: University of Texas Health Science Center at Houston · University of Alabama at Birmingham · University of Pittsburgh · The University of Texas MD Anderson Cancer Center · University of Dublin, Trinity College · Yale University
Abstract
Brain MRI interpretation is a longitudinal clinical reasoning problem: radiologists compare serial studies, integrate information across MRI sequences, localize findings within volumetric anatomy, and translate this evidence into report-grounded assessments. Existing medical VQA and 3D imaging benchmarks capture important parts of this workflow, but often evaluate brain MRI through isolated images, static volumes, or ungrounded report-style answers, thereby obscuring failures in the evidence chain that support clinical validity. We introduce BrainTRACE, a report-grounded benchmark for evaluating whether vision-language models can trace the evidence structure required for longitudinal brain MRI interpretation. BrainTRACE contains 7,273 scored VQA instances derived from 1,778 longitudinal patients, 7,299 MRI studies, and approximately 29k co-registered 3D MRI sequence volumes. The benchmark is organized by five levels of clinical reasoning, from acquisition recognition to case-level synthesis, and by evidence demands covering longitudinal comparison, report-grounded references, multi-sequence integration, and volumetric spatial evidence. BrainTRACE supports rendered inputs compatible with standard VLM interfaces, a 3D-evidence condition, and a decomposed case-reasoning track that audits six steps in a longitudinal evidence chain. Evaluation of 20 VLM configurations shows that current systems can identify isolated visual cues but rarely compose them into grounded longitudinal interpretations. We release the benchmark specification, evaluation lists, scoring implementation, scoring rubrics, and audit-record format to support reproducible progress in brain MRI VLM evaluation.
Figures & tables
| Condition | Evidence retained | Accuracy (%) | Drop (pp) | ||||
|---|---|---|---|---|---|---|---|
| Long. | Multi-seq. | Image | Qwen | GPT | Qwen | GPT | |
| Full evidence | 50.0 | 50.0 | – | – | |||
| Latest timepoint only | – | 12.5 | 26.7 | -37.5 | -23.3 | ||
| T2 sequence only | – | 29.2 | 38.3 | -20.8 | -11.7 | ||
| No image | – | – | – | 15.0 | 9.2 | -35.0 | -40.8 |
Appendix figures & tables19 assets
Supplementary material from the paper’s appendix.
Appendix
| Rank | Model | Pass | QS | Crit. | BLEU4 | ROUGE-L | chrF++ | BERTScore |
|---|---|---|---|---|---|---|---|---|
| 1 | GPT-5.4 | 19.8 | 2.55 | 22.2 | 0.065 | 0.198 | 0.281 | 0.043 |
| 2 | Qwen3-VL-30B | 14.4 | 2.27 | 36.6 | 0.039 | 0.193 | 0.273 | 0.149 |
| 3 | Gemini 2.5 Pro | 14.2 | 2.28 | 32.9 | 0.040 | 0.183 | 0.267 | 0.144 |
| 4 | Qwen3-VL-8B | 13.8 | 2.29 | 29.0 | 0.081 | 0.202 | 0.272 | 0.157 |
| 5 | Qwen3-VL-4B | 11.8 | 2.24 | 28.0 | 0.070 | 0.190 | 0.263 | 0.134 |
| 6 | GPT-5-mini | 10.6 | 2.07 | 40.3 | 0.056 | 0.176 | 0.264 | 0.102 |
| Ranks 1–10 | Ranks 11–20 | ||||||
|---|---|---|---|---|---|---|---|
| Rank | Model | Closed [95% CI] | Pass | Rank | Model | Closed [95% CI] | Pass |
| 1 | GPT-5.4 | 42.2 [39.6, 45.1] | 15.6 | 11 | Qwen2.5-VL-7B | 33.2 [30.5, 36.0] | 3.9 |
| 2 | GPT-5 | 38.3 [35.4, 41.0] | 11.7 | 12 | Gemini 2.5 Pro | 32.0 [29.3, 34.7] | 13.0 |
| 3 | GPT-5-mini | 37.5 [34.7, 40.3] | 10.4 | 13 | M3D-LaMed | 29.9 [27.0, 32.6] | 0.0 |
| 4 | Qwen3-VL-4B | 35.4 [32.6, 37.9] | 5.2 | 14 | HuatuoGPT-Vision-7B | 29.1 [26.4, 31.7] | 0.0 |
| 5 | Lingshu-7B | 35.4 [32.6, 38.1] | 3.9 | 15 | MedGemma-4B-it | 18.9 [16.7, 21.4] | 3.9 |
| Level | Task template | Failure mode | Error type | Images |
|---|---|---|---|---|
| L1 | one_sentence_image_description | Reversed laterality (left right) | wrong_laterality | 1 |
| L2 | morphology_and_boundary | Missed lesion / false negative | missing_finding | 1 |
| L3 | open_interval_change_at_site | Decrease reported as increase | wrong_direction | 6 |
| L4 | open_trajectory_summary | Mislocated + wrong trajectory | wrong_anatomy, wrong_direction | 3 |
| L5 | full_impression_3to5_sentences | Wrong diagnostic category | wrong_diagnosis | 6 |