Multi-task learning for the automatic grading of enlarged perivascular space burden using MRI
Authors: Jesse Phitidis, William N. Whiteley, Joanna M. Wardlaw, Miguel O. Bernabeu, Yajun Cheng, Xiaodi Liu, Junfang Zhang, Una Clancy, +6 more
Organizations: Centre for Clinical Brain Sciences, University of Edinburgh, Edinburgh, UK · Canon Medical Research Europe, Edinburgh, UK · UK Dementia Research Institute, Centre at The University of Edinburgh, Edinburgh, UK · Usher Institute, University of Edinburgh, Edinburgh, UK · Department of Radiology, Chongqing General Hospital, Chongqing University, Chongqing, China · Department of Psychiatry and Behavioral Sciences, University of California, San Francisco, USA · Centre for Rural Health, University of Aberdeen, Inverness, UK · Department of Psychology, University of Edinburgh, Edinburgh, UK
Enlarged perivascular spaces (PVS) visible in brain magnetic resonance imaging (MRI) are increasingly thought to be linked to poor brain health. PVS are elongated structures of less than 3 mm in diameter and can be numerous. To reflect the incidence of PVS, radiologists visually score their burden following a clinical grading scale - a task that would benefit from automation to accelerate analyses and overcome the influence of inter-observer differences. We developed and evaluated methods for training machine learning models to score PVS incidence in the basal ganglia (BG) and centrum semiovale (CSO) leveraging the Potters/Wardlaw scale. The novelty in our work lies in the use of imperfect, semi-automatically generated "silver-standard" PVS segmentation masks during training, in addition to PVS radiological scores. We comparatively evaluated a conditional convolutional neural network (CNN) which accepts PVS masks as an extra input channel, a multi-task CNN which performs both PVS segmentation and scoring, and a logistic regression model which utilises features derived from PVS masks to predict PVS scores. Multi-task learning was the most effective method, achieving a mean average precision of 64.08% compared to 60.22% for the conditional CNN, 52.11% for a baseline CNN trained only to predict PVS scores, and 49.32% for the logistic regression model. The multi-task model showed an ability to localise individual PVS not shown by the other CNNs, and behaved in a probabilistically sensible way, predicting with lower confidence on inherently harder classes. Age, sex, hypertension status, white matter hyperintensity volume, and ischaemic stroke lesion status were shown to be associated with the multi-task model's PVS score predictions and the ground truth in a similar way.
Figures & tables
Figure 1 : Example showing PVS . T1-weighted (T1w), T2-weighted (T2w), and fluid-attenuated inversion recovery (FLAIR) (left-right) magnetic resonance imaging (MRI) scans. Visible perivascular spaces (PVS) shown in red box, most noticeable as white speckles on T2w scans.
Dataset
Sequence
Type
Field Strength (T)
Resolution (mm)
MSS1
T1w
2D
1.5
0.94 × 0.94 × 6.50
T2w
2D
1.5
0.94 × 0.94 × 6.50
FLAIR
2D
1.5
0.94 × 0.94 × 6.50
MSS2
T1w
3D
1.5
0.90 × 1.29 × 1.29
T2w
2D
1.5
0.47 × 0.47 × 6.00
FLAIR
2D
1.5
0.47 × 0.47 × 6.00
Table 1 : MRI scan details . Magnetic resonance imaging (MRI) scan details (in RAS + coordinate system).
Figure 2 : Example of generated region of interest masks . T1-weighted (T1w) magnetic resonance imaging (MRI) scan showing basal ganglia (BG) (green) and centrum semiovale (CSO) (yellow) masks.
Figure 3 : Figure showing disagreement between PVS scores and masks . Distribution of maximum perivascular space (PVS) counts (as derived from silver-standard segmentation masks using the Potters/Wardlaw counting criteria) for subjects with different ground truth scores (as rated by a trained observer). Dashed red lines show the count thresholds at which the score on the Potters/Wardlaw scale would be expected to change, with score 0 being on the first line, and score 4 being above the last line. The results show very high disagreement between the counts and the scores for the basal ganglia (BG) and centrum semiovale (CSO). Box plot whiskers extend to 1.5 times the interquartile range.
Dataset
Training
Validation
Test
Masks
Scores
Total
Masks
Scores
Total
Masks
Scores
Total
MSS1
0
67
67
0
8
8
0
18
18
MSS2
177
177
178
17
17
17
61
62
62
MSS3
160
160
160
28
28
28
40
40
40
LBC1936
346
463
463
59
71
71
96
128
128
VALDO
6
0
6
0
0
0
0
0
0
Table 2 : Availability of PVS scores and masks in our dataset . Table showing the number of subjects from each dataset and each data split with available silver-standard perivascular space (PVS) segmentation masks and ground truth PVS scores.
Figure 4 : Histograms of PVS scores in our dataset . Top : the frequency of basal ganglia (BG) and centrum semiovale (CSO) perivascular space (PVS) scores. Bottom : the frequency of our modified BG and CSO PVS scores. The modified scores merge 0 into 1 and 4 into 3.
Figure 5 : Overview of the proposed PVS score prediction methods . Baseline CNN : trained to classify basal ganglia (BG) and centrum semiovale (CSO) perivascular space (PVS) scores. Conditional CNN : trained to classify BG and CSO PVS scores while utilising PVS segmentation predictions from a U-Net that is pre-trained on the subset of data with silver-standard PVS segmentation masks. Multi-task CNN : trained to classify BG and CSO PVS scores and to predict the silver-standard PVS segmentation masks. Logistic regression : Extracts statistical features from the pre-trained U-Net’s PVS segmentation predictions (e.g., volumes/count) and combines these with age/sex to form two sets of predictor variables (one for the BG model and one for CSO model), with feature selection on the validation set. CNN: convolutional neural network.
Figure 6 : Qualitative evaluation of the U-Net’s PVS mask predictions . T2-weighted (T2w) magnetic resonance imaging (MRI) scan from test set and segmentation overlays. Red shows the silver-standard perivascular space (PVS) mask and blue shows the segmentation U-Net’s predicted PVS mask, with purple showing overlap.
Figure 7 : Volumetric analysis of the U-Net’s PVS mask predictions . Bland-Altman plots showing maximum PVS counts within the BG and CSO (according to the Potters/Wardlaw counting criteria) as derived from the silver-standard perivascular space (PVS) masks and the PVS masks predicted by the segmentation U-Net. Data points include all training, validation, and test subjects with silver-standard PVS masks available.
Method
Mean
BG
CSO
mAP
F1
ACC
mAP
F1
ACC
mAP
F1
ACC
Logistic regression
49.32 [45.65, 55.18]
49.58 [44.74, 54.10]
52.62 [48.19, 57.26]
58.19 [51.80, 66.59]
57.06 [49.90, 63.74]
61.69 [55.65, 67.74]
40.45 [36.77, 46.71]
42.09 [35.75, 48.17]
43.55 [37.50, 49.60]
Baseline CNN
52.11 [48.24, 58.27]
51.95 [46.93, 56.64]
53.23 [48.39, 57.86]
58.08 [51.86, 66.05]
57.11 [50.28, 63.33]
58.47 [52.42, 64.52]
46.15 [41.70, 53.55]
46.80 [40.27, 52.93]
47.98 [41.94, 54.03]
Conditional CNN
60.22 [55.91, 66.10]
55.11 [50.12, 59.78]
57.06 [52.42, 61.69]
64.92 [58.14, 73.66]
58.85 [51.95, 65.30]
62.90 [56.85, 68.95]
55.52 [49.89, 62.48]
51.37 [45.05, 57.39]
51.21 [45.16, 57.26]
Multi-task CNN
64.08 [59.74, 69.27]
57.23 [52.49, 61.46]
58.47 [54.03, 62.70]
69.89 [63.34, 76.69]
62.95 [56.38, 69.11]
65.73 [59.68, 71.37]
58.27 [52.35, 65.36]
51.51 [45.00, 57.42]
51.21 [44.76, 57.26]
Table 3 : Comparative results on classification metrics . Comparison of methods for perivascular space (PVS) score prediction on the test set. Reported as mean [95% CIs] with confidence intervals (CIs) calculated over 10,000 bootstrap iterations with best results for each metric in bold. mAP: mean average precision, F1: macro-average F1 score, ACC: accuracy, BG: basal ganglia, CSO: centrum semiovale, CNN: convolutional neural network.
Figure 8 : Confusion matrices . Basal ganglia (BG) and centrum semiovale (CSO) perivascular space (PVS) score predictions, on the test set. CNN: convolutional neural network.
Figure 9 : The multi-task CNN localises PVS . Left : T2-weighted (T2w) magnetic resonance imaging (MRI) scan with a perivascular space (PVS) score of 3 for the basal ganglia (BG) and centrum semiovale (CSO). Middle : Silver-standard PVS mask. Right : Gradient-weighted class activation map (GradCAM) for the first layer of the multi-task CNN (random seed 2), using the mean of the predicted probability of score 3 for BG and CSO as the target for calculation of the gradient. We can see that some of the highest areas of activation are on visible PVS. CNN: convolutional neural network.
Figure 10 : Distribution and confidence of the multi-task CNN’s PVS score predictions . Boxplots showing the distribution of predicted probabilities from the multi-task CNN for each ground truth label (top) and predicted label (bottom). Confidence measured as the maximum possible Shannon’s entropy minus the actual Shannon’s Entropy, given by −log2(31)+∑s=13pslog2(ps) for predicted probability ps of score s . CNN: convolutional neural network, PVS: perivascular space.
Predictor
BG
CSO
Ground truth
Predicted
Ground truth
Predicted
Age (cont.)
-0.0689 (0.6498)
-0.0683 (0.6623)
0.0117 (0.9355)
0.0296 (0.8399)
Female (bin.)
-0.3360 (0.2614)
-0.0921 (0.7618)
0.4463 (0.1070)
0.2618 (0.3416)
Hypertension (bin.)
0.5308 (0.0844)
0.5477 (0.0847)
0.3699 (0.2034)
0.3601 (0.2089)
WMH volume (cont.)
0.7276 (<0.0001)*
1.2834 (<0.0001)*
0.2716 (0.0546)
0.7027 (0.0007)*
ISL (bin.)
0.5916 (0.0729)
0.5336 (0.1182)
-0.0438 (0.8897)
-0.0712 (0.8259)
Table 4 : Association of clinical variables with the multi-task CNN’s PVS score predictions . Normalised beta coefficients and p-values of predictor variables when either the ground truth or multi-task CNN predicted perivascular space (PVS) scores are the target variables of an ordinal logistic regression model. Reported on a subset (N=200) of the test set with available predictor variables. BG: basal ganglia, CSO: centrum semiovale, WMH: white matter hyperintensity, ISL: ischaemic stroke lesion, cont: continuous, bin: binary, CNN: convolutional neural network.
Figure 11 : Inter-rater reliability . Linear-weighted Cohen’s Kappa between two observers and the multi-task CNN model. Reported on a subset (N=62) of the test rated by two independent trained observers. P-values corrected for multiple comparisons (Bonferroni). CNN: convolutional neural network, PVS: perivascular space.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Mask type
Mean
BG
CSO
mAP
F1
ACC
mAP
F1
ACC
mAP
F1
ACC
Original
53.06 [48.39, 59.76]
52.67 [46.99, 57.79]
54.57 [49.24, 59.64]
58.13 [51.32, 67.42]
61.82 [53.79, 68.92]
65.48 [58.88, 72.08]
48.00 [42.71, 55.53]
43.52 [36.56, 50.27]
43.65 [37.06, 50.76]
Pseudo-labels
54.00 [49.46, 60.46]
51.35 [45.50, 56.71]
55.84 [50.76, 61.17]
58.58 [51.69, 67.44]]
56.24 [47.82, 63.90]
64.97 [58.38, 71.57]
49.43 [43.93, 57.23]
46.46 [39.14, 53.49]
46.70 [39.59, 53.81]
Appendix
Table 5 : Ablation of PVS mask sources — logistic regression . Logistic regression model performance when features are derived from the original silver-standard perivascular space (PVS) masks, or the binarised pseudo-labels produced by the segmentation U-Net. Reported as mean [95% CIs] with confidence intervals (CIs) calculated over 10,000 bootstrap iterations with best results for each metric in bold. Reported on a subset (N=197) of the test data where silver-standard PVS masks are available. mAP: mean average precision, F1: F1 score, ACC: accuracy, BG: basal ganglia, CSO: centrum semiovale.
Mask type
Mean
BG
CSO
mAP
F1
ACC
mAP
F1
ACC
mAP
F1
ACC
Original
55.07 [50.49, 61.75]
51.40 [45.63, 56.71]
52.03 [46.70, 57.36]
61.23 [53.86, 70.57]
55.81 [48.16, 62.91]
56.35 [49.24, 63.45]
48.91 [43.22, 57.05]
46.99 [39.56, 53.95]
47.72 [40.61, 54.82]
Pseudo-labels
57.01 [52.28, 64.07]
53.30 [47.67, 58.55]
54.31 [48.98, 59.64]
60.02 [53.55, 69.77]
56.24 [48.84, 63.02]
58.38 [51.78, 64.97]
54.00 [47.70, 62.20]
50.37 [43.14, 57.31]
50.25 [43.15, 57.36]
Appendix
Table 6 : Ablation of PVS mask sources — conditional CNN . Conditional CNN model performance when input masks are the original silver-standard perivascular space (PVS) masks, or the binarised pseudo-labels produced by the segmentation U-Net. Reported as mean [95% CIs] with confidence intervals (CIs) calculated over 10,000 bootstrap iterations with best results for each metric in bold. Reported on a subset (N=197) of the test data where silver-standard PVS masks are available. mAP: mean average precision, F1: F1 score, ACC: accuracy, BG: basal ganglia, CSO: centrum semiovale.
Figure 12 : Baseline CNN’s GradCAM — first layer . Gradient-weighted class activation maps (GradCAM) for the first layer. The target class is perivascular space (PVS) score 3 for the basal ganglia (BG) (top) and centrum semiovale (CSO) (bottom). Results shown for three different training runs (left-right). There are no noticeable PVS-related activations. CNN: convolutional neural network.
Figure 13 : Conditional CNN’s GradCAM — first layer . Gradient-weighted class activation maps (GradCAM) for the first layer. The target class is perivascular space (PVS) score 3 for the basal ganglia (BG) (top) and centrum semiovale (CSO) (bottom). Results shown for three different training runs (left-right). There are no noticeable PVS-related activations. CNN: convolutional neural network.
Figure 14 : Multi-task CNN’s GradCAM — first layer . Gradient-weighted class activation maps (GradCAM) for the first layer. The target class is perivascular space (PVS) score 3 for the basal ganglia (BG) (top) and centrum semiovale (CSO) (bottom). Results shown for three different training runs (left-right). The highest activations in all images are overlapping with PVS. Both BG and CSO PVS are activated for both BG and CSO score 3 targets. This suggests that PVS in both regions might be considered in the prediction of the scores for both regions. CNN: convolutional neural network.
Figure 15 : Baseline CNN’s GradCAM — final layer . Gradient-weighted class activation maps (GradCAM) for the final layer. The target class is perivascular space (PVS) score 3 for the basal ganglia (BG) (top) and centrum semiovale (CSO) (bottom). Results shown for three different training runs (left-right). Middle column shows strong activation in BG for both targets. Overall these maps are not very interpretable. CNN: convolutional neural network.
Figure 16 : Conditional CNN’s GradCAM — final layer . Gradient-weighted class activation maps (GradCAM) for the final layer. The target class is perivascular space (PVS) score 3 for the basal ganglia (BG) (top) and centrum semiovale (CSO) (bottom). Results shown for three different training runs (left-right). All columns show strong activation in BG for the BG target. Columns 1 and 3 show more activation outside the BG for the CSO target than for the BG target. CNN: convolutional neural network.
Figure 17 : Multi-task CNN’s GradCAM — final layer . Gradient-weighted class activation maps (GradCAM) for the final layer. The target class is perivascular space (PVS) score 3 for the basal ganglia (BG) (top) and centrum semiovale (CSO) (bottom). Results shown for three different training runs (left-right). All columns show strong activation in BG for the BG target. All columns show more activation outside the BG for the CSO target than for the BG target. CNN: convolutional neural network.
Background: The lateral ventricle choroid plexus (LVCP) is gaining recognition as a key imaging biomarker for multiple sclerosis (MS) related to physical disability and neuroinflammation. Yet, manual segmentation of the LVCP is highly tedious, restricting its use in broad clinical trials and longitudinal assessments. This research aims to develop a SwinUNETR-driven pipeline that leverages targeted intra- and peri-ventricular small patch sampling to automatically segment the LVCP in MS from both standalone and multi-modal MRI inputs. Methods: We retrospectively assessed 3T MRI scans across three sets of data stemming from two separate MS-dominant cohorts (Dataset 1: n=177; Dataset 2: n=177; expanded test set: n=388). Our method employed a SwinUNETR architecture trained on 32x32x32 voxel patches, benchmarking it against the 3D UXNET model. The primary metric for evaluation was the Dice Similarity Coefficient (DSC), supplemented by computational demand (GFLOPs) and the 95th percentile Hausdorff Distance (HD95). Results: On the extended test set, the SwinUNETR model secured a mean DSC of 0.868 (95% CI: 0.863-0.872) with MPRAGE and FLAIR combined, showing a statistically significant gain over UXNET (DSC: 0.858 [95% CI: 0.853-0.862], p<0.0001). When restricted to standalone FLAIR inputs, the transformer-based approach sustained a high DSC of 0.863, while the spatial localization of UXNET worsened considerably (HD95: 1.86 vs. 3.00 mm). Importantly, the proposed framework lowered computational load by 99% (91.8 vs. 22,080 GFLOPs). By integrating localized patch sampling with a SwinUNETR architecture, this methodology offers an accurate, robust, and statistically superior alternative to current leading models for LVCP segmentation. Its vast reduction in computational cost makes it ideal for widespread implementation in clinical and research environments.
Po-Jui Lu, Alessandro Cagol, Mario Ocampo-Pineda +12
Translational Imaging in Neurology (ThINk) Basel, Department of Biomedical Engineering, Faculty of Medicine, University Hospital Basel and University of Basel, Basel, Switzerland · Department of Neurology, University Hospital Basel, Basel, Switzerland · Research Center for Clinical Neuroimmunology and Neuroscience Basel (RC2NB), University Hospital Basel and University of Basel, Basel, Switzerland +4
Perineural invasion (PNI) is a clinically relevant indicator of tumor aggressiveness and can influence surgical decision-making, motivating interest in reliable preoperative assessment. The subtle MRI features of PNI, however, often resemble nearby anatomy, complicating noninvasive prediction. These fine perineural cues are easily attenuated by routine downsampling or overly global feature aggregation, reducing the effectiveness of conventional volumetric models. We present LoSA-Net, a localized and scale-adaptive architecture for boundary-sensitive PNI prediction in 3D MRI. Talking Neighborhood Attention (TNA) preserves nerve-aligned detail through localized self-attention with head-wise mixing, and Scale-Adaptive Feature Mixing (SAFM) modulates the receptive field using multi-scale depthwise processing. Cross-Scale Refinement and Alignment (CSRA) maintains consistency between semantic context and high-resolution boundaries across stages. In contrast-enhanced MRI scans from 168 patients with cholangiocarcinoma, LoSA-Net achieves an AUC of 0.7567 and outperforms representative convolutional and transformer baselines under matched preprocessing and optimization settings.
Youngung Han, Hyunsu Go, Kyeonghun Kim +9
Seoul National University, Seoul, Republic of Korea · OUTTA, Seoul, Republic of Korea · Chung-Ang University, Seoul, Republic of Korea +3
Automated assessment of degenerative pathology in the lumbar spine on magnetic resonance imaging (MRI) requires access to large-scale datasets of expert-annotated radiological gradings. In contrast, segmentation pseudo-labels can be generated by automated tools at negligible radiologist cost. We examine whether pre-training on segmentation can effectively replace a fraction of the manual grading annotations required for downstream supervision. We pre-train a 3D ResNet encoder to segment the vertebrae, intervertebral discs (IVDs), and the spinal canal, then fine-tune lightweight task-specific grading heads using different proportions of the available training data, ranging from 10% to 100%. On a multicentre dataset of ∼2,000 subjects across 11 pathologies, segmentation pre-training, achieving a Dice score of 0.94 against pseudo-labels, improved the task-averaged (macro) one-vs-rest ROC-AUC at all proportions. With only 20% of grading labels after pre-training, the method achieved near full-supervision performance, with the largest gains observed for either low-prevalence or spatially grounded pathologies.
Monzon Maria, Zisserman Andrew, Jutzeler Catherine R. +1
Biomedical Data Science Lab, Dept. D-HEST, ETH Zurich, Zurich, Switzerland · Swiss Institute of Bioinformatics (SIB), Lausanne, 1015, Switzerland · Visual Geometry Group, Dept. of Engineering Science, University of Oxford, UK