Textual descriptions can reduce ambiguity in medical image segmentation by specifying the finding and location to be delineated. Existing text-guided methods mainly improve where image and language features interact but generally retain a single learned update pathway across all image-text pairs. We propose MRSeg, a parameter-efficient framework that uses each image-text pair to route the adaptation of visual and textual features before dense prediction. Frozen ConvNeXt-Tiny and PubMedBERT encoders provide multiscale visual features and clinical text tokens. A joint router uses the deepest visual feature and pooled text to predict a sparse mixture over low-rank adapter bases. The resulting route is shared across separate adapter banks for two visual scales and text, coordinating their adaptation while keeping the feature-specific parameters separate. Region Bridge uses text-derived queries to aggregate dense visual tokens into latent regions, refines these regions through self-attention and text cross-attention, and redistributes the refined information back to the feature maps. Finally, a multiscale decoder combines refined semantic features with shallow image evidence. On QaTa-COV19 and MosMedData+, MRSeg achieves 90.90/83.32 and 81.53/68.82 Dice/mIoU, respectively, with 7.11M trainable parameters and 7.60 GFLOPs. Code: https://github.com/maklachur/MRSeg.
Figures & tables
Figure 1: Overview of MRSeg. Frozen, parameter-efficiently adapted encoders produce multiscale visual features and text tokens. The router uses only GAP(f4) and Mean(t) to derive a sparse route α . Pair Adapter reuses this route across separate f3 , f4 , and text adapter banks. Two Region Bridge modules independently produce f3′′ and f4′′ , which are decoded together with f1 , f2 , and shallow features from the adapted image.
Figure 2: Core components of MRSeg. (a) Routed multimodal adaptation: the image–text route controls separate adapter banks for t , f3 , and f4 . (b) Text-aware region refinement: text-derived queries aggregate dense visual tokens, regional and cross-modal interactions refine them, and dense reprojection returns the correction to the original feature map.
Method
Venue
Text
Params ↓
FLOPs ↓
QaTa-COV19
MosMedData+
(M)
(G)
Dice ↑
mIoU ↑
Dice ↑
mIoU ↑
U-Net [ 20 ]
MICCAI’15
×
14.8
50.3
79.02
69.46
64.60
50.73
nnUNet [ 11 ]
Nature’21
×
19.1
412.7
80.42
70.81
72.59
60.36
Swin-UNet [ 2 ]
ECCV’22
×
82.3
67.3
78.07
68.34
63.29
50.19
LAVT [ 23 ]
CVPR’22
✓
118.6
83.8
79.28
69.89
73.29
60.41
LViT [ 12 ]
IEEE TMI’23
✓
29.7
54.1
83.66
75.11
74.57
61.33
Table 1: Comparison with SOTA methods on QaTa-COV19 and MosMedData+. Best and second-best results are shown in bold and underlined , respectively. ↑ / ↓ indicate higher/lower is better. — indicates values not reported by the original paper.
Figure 3: Qualitative comparison on QaTa-COV19 and MosMedData+. Overlays indicate true positives (yellow), false negatives (red), and false positives (green).
Model Variant
QaTa-COV19
MosMedData+
Dice ↑
mIoU ↑
HD95 ↓
Dice ↑
mIoU ↑
HD95 ↓
Image-only (No Text)
87.56
77.87
27.27
78.94
65.21
19.22
w/o Medical Image Adapter
90.68
82.95
16.45
80.99
68.06
15.18
w/o Encoder LoRA
90.49
82.64
16.90
80.76
67.73
16.11
Independent Image/Text Routers
90.59
82.72
16.68
80.82
67.82
15.94
w/o Region Bridges
89.21
80.52
23.07
79.94
66.58
17.83
Table 2: Ablation study on QaTa-COV19 and MosMedData+. Dice and mIoU are reported in %, while HD95 in pixels. w/o means without. Best values are bolded.
Training Data
QaTa-COV19
MosMedData+
Dice ↑
mIoU ↑
HD95 ↓
Dice ↑
mIoU ↑
HD95 ↓
30%
89.60
81.17
17.07
77.83
63.71
18.53
70%
90.48
82.62
15.57
80.11
66.83
17.94
100%
90.90
83.32
14.34
81.53
68.82
14.73
Table 3: Robustness under reduced supervision. Models are trained on randomly sampled {30%, 70%, 100%} subsets of the training split and evaluated on the unchanged full test set. Dice and mIoU are reported in %, and HD95 in pixels.
a School of Computer Science, The University of Sydney, Sydney, Australia. · b Institute of Translational Medicine, Shanghai Jiao Tong University, Shanghai, China. · c Zhongguancun Academy & Zhongguancun Institute of Artificial Intelligence, Beijing, China. +1