Automated identification of DICOM image series is essential for large-scale medical image analysis, quality control, protocol harmonization, and reliable downstream processing. However, DICOM series classification remains challenging due to heterogeneous slice content, variable series length, and entirely missing, incomplete or inconsistent DICOM metadata. We propose an end-to-end multimodal framework for DICOM series classification that jointly models image content and acquisition metadata while explicitly accounting for all these challenges. (i) Images and metadata are encoded with modality-aware modules and fused using a bi-directional cross-modal attention mechanism. (ii) Metadata is processed by a sparse, missingness-aware encoder based on learnable feature dictionaries and value-conditioned modulation. By design, the approach does not require any form of imputation. (iii) Variability in series length and image data dimensions is handled via a 2.5D visual encoder and attention operating on equidistantly sampled slices. We evaluate the proposed approach on the publicly available Duke Liver MRI dataset and a large multi-institutional in-house cohort, assessing both in-domain performance and out-of-domain generalization. Across all evaluation settings, the proposed method consistently outperforms relevant image only, metadata-only and multimodal 2D/3D baselines. The results demonstrate that explicitly modeling metadata sparsity and cross-modal interactions improves robustness for DICOM series classification.
Figures & tables
Figure 1: Proposed method: pixel data of S DICOM slices is embedded in visual feature pathway. DICOM metadata is embedded by the Sparse Metadata Encoder. Bi-directional cross-modal attention contextualizes all image and metadata embeddings. Final integration to a series-level representation is done by learnable pooling.
Exp.
Modality
Image Enc.
MD Enc.
Imputer
Fusion
(1)
Image
2D DenseNet121
N/A
N/A
N/A
(2)
Image [ 14 ]
3D ResNet18
N/A
N/A
N/A
(3)
Metadata
N/A
XGBoost
No
N/A
(4)
Joint
2.5D DenseNet121
MLP (dense)
Zero
Concat
(5)
Joint
2.5D DenseNet121
MLP (dense)
Learned
Concat
(6)
Joint [ 8 ]
2D DenseNet121
RF
No
No
Table 1: Baselines for in-domain evaluation on the Duke dataset.
Figure 2: In-domain evaluation: five-fold cross-validation per-class F1 scores (%) on the Duke Liver MRI dataset.
Table 3: Out-of-domain evaluation: Model trained on in-house cohort training split. Per-label F1 (%) on the in-house holdout and the external Duke dataset.
S=1
S=3
S=5
S=10
S=20
85.09 ± 1.79
95.46 ± 0.69
96.17 ± 1.00
96.66 ± 1.03
95.88 ± 0.82
Table 4: Ablation on the number of input slices S . Weighted F1 scores (%) on Duke (five-fold cross-validation) for S∈{1,3,5,10,20} with all other settings fixed.
Medical diagnosis requires the effective synthesis of visual manifestations and clinical metadata. However, existing methods often treat metadata as isolated tags, failing to exploit the rich semantic knowledge embedded in clinical descriptions. We propose PRIMA (Pre-training with Risk-integrated Image-Metadata Alignment), a framework that integrates domain-specific knowledge into multi-modal representation learning. We first curate an expert corpus of risk--disease correlations via Retrieval-Augmented Generation (RAG) to refine Clinical ModernBERT, embedding diagnostic priors into the text encoder. To bridge the modality gap, we introduce a dual-encoder pre-training strategy utilizing DINOv3 and our refined Clinical ModernBERT, optimized by a suite of four complementary loss functions. These losses are designed to capture multi-granular semantic alignment and handle the ambiguity of clinical correlations through soft labels. Finally, we leverage Qwen3 to fuse these aligned features for precise disease classification. Extensive experiments demonstrate that PRIMA effectively harmonizes pixel-level features with abstract clinical expertise, consistently outperforming other state-of-the-art methods. Notably, our framework achieves strong performance without requiring massive data collection or exhaustive computational resources. Our code will is available at https://github.com/yqwang01/PRIMA.
Yiqing Wang, Chunming He, Ming-Chen Lu +4
Department of Biomedical Engineering, Duke University, Durham, NC, USA · Department of Ophthalmology and Visual Sciences, University of Michigan, Ann Arbor, MI, USA
Adapting natural-image foundation models like DINOv3 to multi-modal medical imaging is challenging due to the significant domain gap between natural color images and multi-channel medical scans. We present a unified, patch-based framework that processes raw multimodal imaging through training-free registration, automated localization, and mask-filtered patch extraction. This architecture culminates in a hierarchical strategy that aggregates patch-level insights into subject-level diagnostics. Using liver fibrosis staging as a case study, we evaluate four patch-level feature representations: handcrafted Radiomics features, learned ResNet features, pre-trained foundation model SAM-Med2D features, and frozen DINOv3 features. To ensure a controlled comparison, all models utilize the same lightweight MLP head and are evaluated across both rigid and deformable registration settings. Our training protocol focuses on mild fibrosis (S1) and cirrhosis (S4) classes only, enabling a single classifier to address both substantial fibrosis detection and cirrhosis staging. Evaluated via 10 random train (90%)/ test (10%) splits on 360 subjects from the CARE 2025 Liver Track 4 cohort, our DINOv3-based framework significantly outperforms all baselines, achieving the best classification accuracy of 78.4% for S1 and 75.8% for S4.
Boya Wang, Ruizhe Li, Chao Chen +1
School of Computer Science, University of Nottingham, UK · Nottingham Biomedical Research Centre (BRC), School of Medicine, University of Nottingham, UK
Multi-modal medical imaging enables comprehensive diagnostics, yet current foundation models process 2D (e.g. X-ray) and 3D (e.g. CT) data with separate, dimensionality-specific architectures. We present MultiMedVision, a unified framework for joint 2D/3D representation learning built on a Sparse Vision Transformer. Our model uses 3D Rotary Positional Embeddings and variable-length sequence packing to process mixed-modality batches natively within a shared latent space, without modality-specific adapters or treating 3D volumes as 2D slice sequences. Trained with a self-supervised objective on chest X-rays (MIMIC-CXR) and CT scans (CT-RATE), and using a single shared encoder with 5x less data, MultiMedVision achieves competitive performance on both 2D benchmarks (Macro AUROC 0.82 on MIMIC, 0.84 on CheXpert) and 3D tasks (0.85 on CT-RATE). Analysis of the learned representations reveals coexisting modality-specific and shared feature subspaces, demonstrating that unified cross-dimensional representation learning is feasible without sacrificing modality-specific performance.
Frank Li, Bardia Khosravi, Mohammadreza Chavoshi +5