Knee3DVLM: Dual-Sequence Full-Volume Vision-Language Modeling for Comprehensive Knee MRI Assessment
Authors: Maryam Baizhigitova, Andrew Seohwan Yu, Po-Hao Chen, Naveen Subhas, Sixu Chen, Xinxin Wang, Kunio Nakamura, Richard Lartey, +2 more
Organizations: Cleveland Clinic, Cleveland, OH 44106, USA · Cleveland State University, Cleveland, OH 44115, USA · Case Western Reserve University, Cleveland, OH 44106, USA · Kent State University, Kent, OH 44242, USA
Vision-language models (VLMs) are increasingly being applied to three-dimensional medical imaging, but their application to knee MRI remains limited, particularly for interpreting the complementary sequences used in clinical practice. We introduce Knee3DVLM, a sequence-aware VLM that uses full-volume DESS and fluid-sensitive TSE MRI to predict 57 anatomically resolved binary diagnostic targets derived from the MRI Osteoarthritis Knee Score (MOAKS) for structured reporting. We evaluated DESS-only, TSE-only, and paired DESS-TSE configurations using subject-disjoint Osteoarthritis Initiative partitions. In a held-out cohort of 1,074 examinations, the fused model achieved 72.98% average accuracy, 71.17% balanced accuracy, 78.96% mean ROC-AUC, and 78.74% macro ROC-AUC, the highest values among the three configurations. In a secondary multiclass analysis aligned with the released 3DReasonKnee cohort, Knee3DVLM was numerically higher than the strongest reported 3DReasonKnee configuration across five pathology categories. These findings support dual-sequence full-volume modeling for comprehensive knee MRI assessment.
Figures & tables
Figure 1: Class imbalance in the binary and multiclass evaluation tasks. Panel A summarizes ground-truth normal-versus-abnormal labels across the 57 targets, grouped by pathology. Panel B shows the pooled ground-truth grade distribution across all 109 multiclass MOAKS component attributes on a logarithmic axis, exposing rare grades while retaining the dominant grade 0 class (85.4%).
Table 1: Binary target schema. Parenthetical counts indicate the number of binary targets contributed by each anatomical grouping.
Figure 2: Knee3DVLM architecture. Independent trainable DCFormer encoders map the full DESS and TSE volumes to 768-dimensional representations. Sequence-specific MLP-Mixers form 16 visual tokens of width 512 per sequence. Learned modality embeddings preserve sequence identity before the 32 tokens undergo cross-sequence token mixing and shared channel refinement. A per-token projector maps the fused representation from width 512 to the language-model embedding dimension dLM . The LoRA-adapted language backbone predicts 57 anatomically resolved binary targets, which are converted into a structured report by deterministic mapping. Flame symbols identify trainable modules.
Accuracy
Threshold-dependent classification
Discrimination
Cohort
Model
Average accuracy
Balanced accuracy
Sensitivity
Specificity
Mean ROC-AUC
Macro ROC-AUC
Examinations ( n )
Knee3DVLM-DESS
72.59
70.42
67.04
73.80
76.93
76.89
1,074
Knee3DVLM-TSE
70.36
69.10
65.24
72.96
76.83
76.45
1,074
Knee3DVLM
72.98
71.17
66.88
75.46
78.96
78.74
1,074
Table 2: Primary binary 57-target performance on the matched held-out test cohort. Values are presented in percentages.
Model
BML
Cartilage
Osteophyte
Meniscus
ACL
PCL *
Effusion
Synovitis
Knee3DVLM-TSE
71.34
70.62
69.61
68.40
80.42
50.66
72.86
65.71
Knee3DVLM-DESS
69.55
72.77
71.29
67.18
75.95
32.51
71.64
64.22
Knee3DVLM
71.60
73.48
71.67
69.02
78.37
51.60
72.41
63.60
Table 3: Pathology-wise balanced accuracy for the three Knee3DVLM input configurations. Values are percentages. TSE-only and DESS-only values were recomputed directly from their held-out target-level prediction exports using the unweighted mean of target-level balanced accuracy within each pathology group; Knee3DVLM values were computed identically. Bold indicates the highest value in each column.
Pathology
Targets
Average accuracy
Balanced accuracy
Mean ROC-AUC
Macro ROC-AUC
BML
15
73.05
71.60
78.88
78.86
Cartilage
14
76.01
73.48
80.45
80.38
Meniscus
12
74.31
69.02
77.58
78.03
Osteophyte
12
71.13
71.67
79.16
79.16
ACL
1
70.14
78.37
87.99
87.99
Effusion
1
67.60
72.41
77.80
77.80
Table 4: Knee3DVLM performance by pathology group. Values are percentages, computed directly from the same target-level export used for Tables 2 and 3 . Average accuracy and macro ROC-AUC are unweighted means of per-target values within each pathology group. Mean ROC-AUC weights each target-level ROC-AUC by its number of evaluable examinations within the pathology group.
Model
Method
BML
Cartilage
Osteophyte
Meniscus
Others
Qwen2.5VL-3B †
SFT with CoT
66.70
60.00
34.20
34.00
59.00
Med3DVLM †
SFT with CoT
67.10
57.90
33.30
34.30
60.50
Med3DVLM †
SFT without CoT
67.00
57.80
33.00
34.60
60.00
Qwen2.5VL-3B †
SFT without CoT
70.60
60.20
34.50
38.70
61.60
Our VLM (DESS-only)
Full training without CoT
70.80
65.30
36.10
32.90
64.30
Our VLM (TSE-only)
Full training without CoT
71.91
55.40
34.05
41.60
65.40
Table 5: Category-wise average accuracy comparison with the published 3DReasonKnee multiclass grading results. † Published results transcribed from Sambara et al. [3]. Bold indicates the highest value in each category.
Multi-parametric magnetic resonance imaging (mpMRI) is a cornerstone for brain tumor diagnosis and treatment, yet current AI models face critical limitations: their lack of natural language interaction and interpretability impedes spatial information integration and cross-modal reasoning required clinically. Key challenges arise from significant physical meaning differences across modalities, spatial misalignment due to scan intervals, and the need for complex multi-feature interpretation in tasks like glioma grading. While visual-language models (VLMs) show promise in cross-modal understanding, existing methods focus mainly on 2D image modeling, neglecting direct perception of 3D volumetric space. Although 3D VLMs have been proposed for report generation and feature alignment in 3D CT imaging, mpMRI applications demand collaborative inference across multiple imaging modalities-a requirement unmet by current solutions. To address this, we introduce Mr3D-VL, a dedicated visual-language foundation model for multi-parametric 3D MRI. With 4 billion parameters, it employs an unsupervised pre-trained shared 3D encoder and 4D rotational positional embedding for dual modality-spatial integration. Its cross-modal projection layer uses a multi-resolution feature implantation strategy to enhance feature perception across resolutions. Experimental results show significant improvements over existing 4B/7B/30B domain-specific and general-purpose models in text generation tasks, achieving a BERTScore of 0.856 for report generation, with question-answering accuracy at 0.713 and multiple-choice accuracy at 0.912.
Purpose: To develop and evaluate cross-sequence transfer learning for automatic femoral cartilage segmentation, testing bidirectional transfer between dual-echo steady-state (DESS) and sagittal proton density-weighted 3D fast spin-echo (Cube) sequences. Materials and Methods: We optimized a modified 2D U-Net on 507 DESS images from the Osteoarthritis Initiative (OAI). We then established same-sequence baselines using subject-level cross-validation on a subset of 44 OAI DESS images and 44 Cube images acquired at the Istituto Ortopedico Rizzoli, Bologna, Italy. Each subset included 22 non-lesioned and 22 lesioned subjects. Finally, we performed transfer learning across sequences by fine-tuning the pretrained models on the target sequence with increasing training set sizes to study convergence, while keeping validation and test sets fixed. Segmentations were evaluated using Dice similarity coefficient (DSC) and average surface distance (ASD). Lesion effects were assessed with two-sided Mann-Whitney U tests with Bonferroni correction. Results: Same-sequence training yielded higher accuracy on DESS than Cube (DSC, 0.900 vs 0.830; P<.001). Cube-to-DESS transfer matched DESS performance (DSC, 0.903±0.032 vs 0.900±0.027), reaching a performance plateau at 9 training subjects. DESS-to-Cube yielded a lower combined DSC (0.802±0.049 vs 0.830±0.042), reaching a plateau at 24 training subjects. Lesions did not affect DESS (P≥.39) but reduced Cube accuracy (DSC, 0.805 vs 0.856; P<.001). Conclusion: Transfer learning across sequences can substantially reduce target-sequence annotation requirements for femoral cartilage segmentation, but performance is direction- and sequence-dependent, and the effects of lesions on segmentation may vary across MRI sequences.
Francesco Chiumento, Gianluigi Crimi, Elisa Moretta +7
School of Electronic Engineering, Dublin City University, Dublin, Ireland · Bioengineering and Computing Laboratory, IRCCS Istituto Ortopedico Rizzoli, Bologna, Italy · Dipartimento di Scienze Mediche e Chirurgiche (DIMEC), Alma Mater Studiorum – Università di Bologna, Bologna, Italy +4
Radiology reporting is time-consuming and subject to inter-rater variability, making automated report generation an attractive clinical application for Vision-Language Models (VLMs). We benchmark state-of-the-art VLMs on lumbar spine MRI with a focus on diagnostic accuracy and demonstrate that standard lexical and semantic metrics poorly reflect clinical correctness: fluent, well-structured reports can score highly while containing clinically meaningful diagnostic errors. To address this failure mode, we propose an architecture-agnostic framework that augments VLM inputs with spatially localized, disc-level anomaly heatmaps generated by a semi-supervised U-Net++ model. These heatmaps both improve anatomical sensitivity through explicit visual grounding and provide an independent interpretability output for clinical oversight, moving us closer to diagnostically reliable, visually grounded VLMs for lumbar spine MRI interpretation.
Bruno Palau, Franziska Vogt, Daria Laslo +4
ETH Zurich, Switzerland · 2Swiss Institute of Bioinformatics (SIB), Switzerland