ConPro: Contrast Projection Pretraining for Label-Efficient Vessel Segmentation in DSA Sequences
Authors: Xinge Guo, Yuanhao Wang, Liqi Shu, Yang Liu, Min Xu
Organizations: Duke University, Durham, NC, USA · Carnegie Mellon University, Pittsburgh, PA, USA · University of Pittsburgh Medical Center, Pittsburgh, PA, USA
Dense vessel annotation in digital subtraction angiography (DSA) is labor-intensive, yet every unlabeled sequence records how contrast passes through the vessels. Semi-supervised methods take their targets from the current model, and generic self-supervised pretexts reconstruct static appearance, so this signal goes unused. We propose ConPro, a self-supervised pretraining scheme whose target is a contrast projection, the normalized drop of every pixel below its temporal median over the sequence. On DIAS and DSCA, with 10%, 20% and 50% of the training cases labeled, ConPro improves on training from scratch at every label fraction and is the best of the compared methods on DSCA at 20% and 50% labels. Controlled comparisons show that the gain comes from the target. A temporal-median target with the same input, loss and budget stays at scratch level, and using the projection directly instead of learning it, as an input channel or a pseudo-label, helps little or hurts. ConPro provides pretrained weights without changing the segmentation architecture, so it combines with semi-supervised training, and UniMatch, the strongest baseline, gains 0.5 to 2.0 Dice and 0.9 to 2.3 clDice at every label fraction when started from ConPro weights, reaching 75.4 Dice on DIAS and 81.3 on DSCA.
Figures & tables
Figure 1: Annotation-free pretraining. The network sees a subset of the frames (cyan) and predicts the contrast projection C⋆ of the whole sequence. Grey frames are withheld.
Figure 2: ConPro pipeline. Stage I predicts C⋆ from a subset of frames XΩ , and Stage II transfers the backbone F to vessel segmentation.
DIAS
DSCA
Method
Type
10%
20%
50%
10%
20%
50%
Supervised
Sup.
65.1 / 58.7
71.8 / 65.2
73.5 / 67.2
79.4 / 75.2
80.2 / 76.1
80.2 / 76.1
CPS [ 6 ]
Semi
69.1 / 61.6
71.7 / 64.5
72.3 / 64.9
79.8 / 75.5
80.5 / 76.4
80.6 / 76.5
CorrMatch [ 7 ]
Semi
64.1 / 55.5
69.2 / 61.2
71.2 / 63.2
76.1 / 70.3
77.5 / 72.2
78.2 / 73.5
UniMatch [ 8 ]
Semi
71.2 / 65.0
73.1 / 66.6
74.9 / 68.8
80.0 / 76.0
79.6 / 76.1
79.3 / 75.5
RPST ∗ [ 2 ]
Semi
70.6 / 63.7
72.3 / 65.5
72.1 / 65.2
76.6 / 70.7
76.7 / 71.2
77.0 / 71.8
Table 1: Dice / clDice (%; three-seed means), matched backbone and Stage II update budget. Sup.: labeled sequences only; Semi: labeled and unlabeled sequences; Self: self-supervised pretraining, then supervised fine-tuning. The three C⋆ rows use the contrast projection directly, without pretraining. The last three rows are ours, ConPro fine-tuned with labels only, and RPST ∗ and UniMatch started from ConPro weights. Bold : best; underline : second best.
Figure 3: Qualitative results at 20% labels, one test sequence per dataset. Green marks correct vessel pixels, red marks errors.
DIAS
DSCA
Protocol
Pretraining
Input
Target
Loss
10%
20%
50%
10%
20%
50%
Target
None
65.1 / 58.7
71.8 / 65.2
73.5 / 67.2
79.4 / 75.2
80.2 / 76.1
80.2 / 76.1
Median, all frames
XΩ
med(X)
BCE+Dice
64.1 / 57.6
70.7 / 63.7
71.7 / 64.3
78.8 / 74.3
79.9 / 75.8
80.1 / 76.2
Median, all frames
XΩ
med(X)
L1
61.9 / 58.9
68.7 / 61.5
71.2 / 64.4
78.8 / 74.3
79.6 / 75.6
79.9 / 75.9
Median, withheld frames
XΩ
med(XΩˉ)
L1
61.8 / 57.1
69.3 / 62.9
70.6 / 63.1
78.7 / 74.3
79.7 / 75.8
79.7 / 75.7
Contrast projection
XΩ
C⋆
BCE+Dice
69.3 / 63.5
71.9 / 66.0
74.4 / 68.4
80.0 / 75.8
81.2 / 77.4
81.4 / 77.6
Table 2: Stage I ablations (Dice / clDice %; three-seed means). Target: input fixed to a random half XΩ , target varied. Input: contrast projection kept, frame feeding varied. Protocols trained independently; bold : highest within a protocol.
Figure 4: Input variants of Table 2 on the sequences of Fig. 3 , alone and with UniMatch. Colors as in Fig. 3 .
Learning from minimal human supervision is a long-standing goal in medical image analysis, where dense expert annotations are costly. We study retinal vessel segmentation in an extreme semi-supervised setting with one annotated image and a pool of unlabeled images. We propose ESRVS, which selects a representative reference image for manual annotation and transfers vessel cues using target-domain-adapted DINOv3 features. ESRVS constructs a multi granular vessel prototype, combines prototype-similarity maps with a physics-inspired prior to generate initial pseudo-labels, and refines the transferred supervision through weighted pseudo-label training and adversarial refinement. Across eight public datasets, ESRVS achieves the best Dice and clDice on six datasets, and the best HD95 on all eight datasets among the compared semi-supervised methods, although those methods use 10 to 20% labeled data. With Mask2Former, ESRVS retains on average 93.7% of fully supervised Dice and 95.1% of fully supervised clDice. These results demonstrate the potential of foundation-model label propagation for highly label-efficient retinal vessel segmentation. Code is available at https://github.com/IAANNH/ESRVS.
Mingzhi Xu, Yizhe Zhang
Nanjing University of Science and Technology, Nanjing, China
Accurate segmentation of coronary Digital Subtraction Angiography (DSA) images is essential for diagnosing and treating coronary artery disease (CAD). Despite advances in deep learning, challenges such as high intra-class variance and class imbalance limit precise vessel delineation. Existing approaches for coronary DSA segmentation cannot effectively address these issues. Furthermore, existing segmentation network encoders do not directly generate semantic embeddings, which could enable the decoder to reconstruct segmentation masks more effectively. We propose a Supervised Prototypical Contrastive Loss (SPCL) that combines supervised and prototypical contrastive learning to enhance coronary DSA image segmentation. The supervised contrastive loss enforces semantic embeddings in the encoder, improving feature differentiation. The prototypical contrastive loss enables the model to focus on the foreground class while alleviating high intra-class variance and class imbalance by concentrating only on hard-to-classify background samples. We implement the proposed SPCL within MSA-UNet3+, a Multi-Scale Attention-Enhanced UNet3+ architecture. The architecture integrates a Multi-Scale Attention Encoder (M-encoder), a Multi-Scale Dilated Bottleneck (MSD-Bottleneck) for multi-scale feature extraction, and a Contextual Attention Fusion Module (CAFM) to preserve fine-grained details while improving contextual understanding. Experiments on a private coronary DSA dataset demonstrate that MSA-UNet3+ outperforms state-of-the-art methods, achieving the highest Dice coefficient and F1-score while significantly reducing ASD and ACD. The framework provides precise vessel segmentation for accurate identification of coronary stenosis and supports informed diagnostic and therapeutic decisions. The code will be released at https://github.com/rayanmerghani/MSA-UNet3plus.
Shenzhen Institutes of Advanced Technology (SIAT), Chinese Academy of Sciences (CAS), Shenzhen, 518055, China · University of Chinese Academy of Sciences (UCAS), Beijing, 101408, China · Department of Biomedical Engineering and Systems, Faculty of Engineering, Cairo University, Cairo, 12613, Egypt +2
Accurate segmentation of vascular structures in digital subtraction angiography (DSA) images remains challenging due to the thin, elongated, and branching nature of blood vessels. Pixel-wise deep learning approaches such as U-Net achieve strong general-purpose segmentation performance but often produce fragmented or discontinuous predictions in fine vascular regions, since they do not explicitly enforce structural connectivity. Region growing algorithms preserve spatial context and topological continuity, but are highly sensitive to seed point initialization and can be computationally expensive. We propose UI-VISA (U-Net Initialized Vascular Image Segmentation Architecture), a hybrid pipeline that combines the complementary strengths of both approaches. UI-VISA uses U-Net's foreground predictions as informed seed points for a CNN-guided region growing algorithm, which then iteratively refines the segmentation by enforcing local connectivity and recovering fine vessel details that U-Net alone tends to miss or over-predict. We evaluate UI-VISA against standalone U-Net and a prior region-growing-based method (VISA) using 5-fold cross-validation on 26 DSA images. UI-VISA achieves the highest mean Dice and clDice scores across folds, and a paired Wilcoxon signed-rank test shows the improvement in clDice is statistically significant (p=0.023), consistent with the method's design goal of preserving vascular connectivity, while the improvement in Dice does not reach significance (p=0.104).