Knee3DVLM: Dual-Sequence Full-Volume Vision-Language Modeling for Comprehensive Knee MRI Assessment
Authors: Maryam Baizhigitova, Andrew Seohwan Yu, Po-Hao Chen, Naveen Subhas, Sixu Chen, Xinxin Wang, Kunio Nakamura, Richard Lartey, +2 more
Organizations: Cleveland Clinic, Cleveland, OH 44106, USA · Cleveland State University, Cleveland, OH 44115, USA · Case Western Reserve University, Cleveland, OH 44106, USA · Kent State University, Kent, OH 44242, USA
Vision-language models (VLMs) are increasingly being applied to three-dimensional medical imaging, but their application to knee MRI remains limited, particularly for interpreting the complementary sequences used in clinical practice. We introduce Knee3DVLM, a sequence-aware VLM that uses full-volume DESS and fluid-sensitive TSE MRI to predict 57 anatomically resolved binary diagnostic targets derived from the MRI Osteoarthritis Knee Score (MOAKS) for structured reporting. We evaluated DESS-only, TSE-only, and paired DESS-TSE configurations using subject-disjoint Osteoarthritis Initiative partitions. In a held-out cohort of 1,074 examinations, the fused model achieved 72.98% average accuracy, 71.17% balanced accuracy, 78.96% mean ROC-AUC, and 78.74% macro ROC-AUC, the highest values among the three configurations. In a secondary multiclass analysis aligned with the released 3DReasonKnee cohort, Knee3DVLM was numerically higher than the strongest reported 3DReasonKnee configuration across five pathology categories. These findings support dual-sequence full-volume modeling for comprehensive knee MRI assessment.
Figures & tables
Figure 1: Class imbalance in the binary and multiclass evaluation tasks. Panel A summarizes ground-truth normal-versus-abnormal labels across the 57 targets, grouped by pathology. Panel B shows the pooled ground-truth grade distribution across all 109 multiclass MOAKS component attributes on a logarithmic axis, exposing rare grades while retaining the dominant grade 0 class (85.4%).
Table 1: Binary target schema. Parenthetical counts indicate the number of binary targets contributed by each anatomical grouping.
Figure 2: Knee3DVLM architecture. Independent trainable DCFormer encoders map the full DESS and TSE volumes to 768-dimensional representations. Sequence-specific MLP-Mixers form 16 visual tokens of width 512 per sequence. Learned modality embeddings preserve sequence identity before the 32 tokens undergo cross-sequence token mixing and shared channel refinement. A per-token projector maps the fused representation from width 512 to the language-model embedding dimension dLM . The LoRA-adapted language backbone predicts 57 anatomically resolved binary targets, which are converted into a structured report by deterministic mapping. Flame symbols identify trainable modules.
Accuracy
Threshold-dependent classification
Discrimination
Cohort
Model
Average accuracy
Balanced accuracy
Sensitivity
Specificity
Mean ROC-AUC
Macro ROC-AUC
Examinations ( n )
Knee3DVLM-DESS
72.59
70.42
67.04
73.80
76.93
76.89
1,074
Knee3DVLM-TSE
70.36
69.10
65.24
72.96
76.83
76.45
1,074
Knee3DVLM
72.98
71.17
66.88
75.46
78.96
78.74
1,074
Table 2: Primary binary 57-target performance on the matched held-out test cohort. Values are presented in percentages.
Model
BML
Cartilage
Osteophyte
Meniscus
ACL
PCL *
Effusion
Synovitis
Knee3DVLM-TSE
71.34
70.62
69.61
68.40
80.42
50.66
72.86
65.71
Knee3DVLM-DESS
69.55
72.77
71.29
67.18
75.95
32.51
71.64
64.22
Knee3DVLM
71.60
73.48
71.67
69.02
78.37
51.60
72.41
63.60
Table 3: Pathology-wise balanced accuracy for the three Knee3DVLM input configurations. Values are percentages. TSE-only and DESS-only values were recomputed directly from their held-out target-level prediction exports using the unweighted mean of target-level balanced accuracy within each pathology group; Knee3DVLM values were computed identically. Bold indicates the highest value in each column.
Pathology
Targets
Average accuracy
Balanced accuracy
Mean ROC-AUC
Macro ROC-AUC
BML
15
73.05
71.60
78.88
78.86
Cartilage
14
76.01
73.48
80.45
80.38
Meniscus
12
74.31
69.02
77.58
78.03
Osteophyte
12
71.13
71.67
79.16
79.16
ACL
1
70.14
78.37
87.99
87.99
Effusion
1
67.60
72.41
77.80
77.80
Table 4: Knee3DVLM performance by pathology group. Values are percentages, computed directly from the same target-level export used for Tables 2 and 3 . Average accuracy and macro ROC-AUC are unweighted means of per-target values within each pathology group. Mean ROC-AUC weights each target-level ROC-AUC by its number of evaluable examinations within the pathology group.
Model
Method
BML
Cartilage
Osteophyte
Meniscus
Others
Qwen2.5VL-3B †
SFT with CoT
66.70
60.00
34.20
34.00
59.00
Med3DVLM †
SFT with CoT
67.10
57.90
33.30
34.30
60.50
Med3DVLM †
SFT without CoT
67.00
57.80
33.00
34.60
60.00
Qwen2.5VL-3B †
SFT without CoT
70.60
60.20
34.50
38.70
61.60
Our VLM (DESS-only)
Full training without CoT
70.80
65.30
36.10
32.90
64.30
Our VLM (TSE-only)
Full training without CoT
71.91
55.40
34.05
41.60
65.40
Table 5: Category-wise average accuracy comparison with the published 3DReasonKnee multiclass grading results. † Published results transcribed from Sambara et al. [3]. Bold indicates the highest value in each category.
School of Electronic Engineering, Dublin City University, Dublin, Ireland · Bioengineering and Computing Laboratory, IRCCS Istituto Ortopedico Rizzoli, Bologna, Italy · Dipartimento di Scienze Mediche e Chirurgiche (DIMEC), Alma Mater Studiorum – Università di Bologna, Bologna, Italy +4