Textual descriptions can reduce ambiguity in medical image segmentation by specifying the finding and location to be delineated. Existing text-guided methods mainly improve where image and language features interact but generally retain a single learned update pathway across all image-text pairs. We propose MRSeg, a parameter-efficient framework that uses each image-text pair to route the adaptation of visual and textual features before dense prediction. Frozen ConvNeXt-Tiny and PubMedBERT encoders provide multiscale visual features and clinical text tokens. A joint router uses the deepest visual feature and pooled text to predict a sparse mixture over low-rank adapter bases. The resulting route is shared across separate adapter banks for two visual scales and text, coordinating their adaptation while keeping the feature-specific parameters separate. Region Bridge uses text-derived queries to aggregate dense visual tokens into latent regions, refines these regions through self-attention and text cross-attention, and redistributes the refined information back to the feature maps. Finally, a multiscale decoder combines refined semantic features with shallow image evidence. On QaTa-COV19 and MosMedData+, MRSeg achieves 90.90/83.32 and 81.53/68.82 Dice/mIoU, respectively, with 7.11M trainable parameters and 7.60 GFLOPs. Code: https://github.com/maklachur/MRSeg.
Figures & tables
Figure 1: Overview of MRSeg. Frozen, parameter-efficiently adapted encoders produce multiscale visual features and text tokens. The router uses only GAP(f4) and Mean(t) to derive a sparse route α . Pair Adapter reuses this route across separate f3 , f4 , and text adapter banks. Two Region Bridge modules independently produce f3′′ and f4′′ , which are decoded together with f1 , f2 , and shallow features from the adapted image.
Figure 2: Core components of MRSeg. (a) Routed multimodal adaptation: the image–text route controls separate adapter banks for t , f3 , and f4 . (b) Text-aware region refinement: text-derived queries aggregate dense visual tokens, regional and cross-modal interactions refine them, and dense reprojection returns the correction to the original feature map.
Method
Venue
Text
Params ↓
FLOPs ↓
QaTa-COV19
MosMedData+
(M)
(G)
Dice ↑
mIoU ↑
Dice ↑
mIoU ↑
U-Net [ 20 ]
MICCAI’15
×
14.8
50.3
79.02
69.46
64.60
50.73
nnUNet [ 11 ]
Nature’21
×
19.1
412.7
80.42
70.81
72.59
60.36
Swin-UNet [ 2 ]
ECCV’22
×
82.3
67.3
78.07
68.34
63.29
50.19
LAVT [ 23 ]
CVPR’22
✓
118.6
83.8
79.28
69.89
73.29
60.41
LViT [ 12 ]
IEEE TMI’23
✓
29.7
54.1
83.66
75.11
74.57
61.33
Table 1: Comparison with SOTA methods on QaTa-COV19 and MosMedData+. Best and second-best results are shown in bold and underlined , respectively. ↑ / ↓ indicate higher/lower is better. — indicates values not reported by the original paper.
Figure 3: Qualitative comparison on QaTa-COV19 and MosMedData+. Overlays indicate true positives (yellow), false negatives (red), and false positives (green).
Model Variant
QaTa-COV19
MosMedData+
Dice ↑
mIoU ↑
HD95 ↓
Dice ↑
mIoU ↑
HD95 ↓
Image-only (No Text)
87.56
77.87
27.27
78.94
65.21
19.22
w/o Medical Image Adapter
90.68
82.95
16.45
80.99
68.06
15.18
w/o Encoder LoRA
90.49
82.64
16.90
80.76
67.73
16.11
Independent Image/Text Routers
90.59
82.72
16.68
80.82
67.82
15.94
w/o Region Bridges
89.21
80.52
23.07
79.94
66.58
17.83
Table 2: Ablation study on QaTa-COV19 and MosMedData+. Dice and mIoU are reported in %, while HD95 in pixels. w/o means without. Best values are bolded.
Training Data
QaTa-COV19
MosMedData+
Dice ↑
mIoU ↑
HD95 ↓
Dice ↑
mIoU ↑
HD95 ↓
30%
89.60
81.17
17.07
77.83
63.71
18.53
70%
90.48
82.62
15.57
80.11
66.83
17.94
100%
90.90
83.32
14.34
81.53
68.82
14.73
Table 3: Robustness under reduced supervision. Models are trained on randomly sampled {30%, 70%, 100%} subsets of the training split and evaluated on the unchanged full test set. Dice and mIoU are reported in %, and HD95 in pixels.
Text-guided medical image segmentation leverages clinical semantics to improve lesion delineation, yet many existing models bind cross-modal fusion, supervision, and decoder design into a task-specific architecture. Such tight coupling makes it difficult to reuse language guidance modules across heterogeneous vision and text backbones, and often requires redesigning the network when the encoder pair changes. This paper presents BTHA, a backbone-transferable hierarchical adapter framework for text-guided medical image segmentation. BTHA is built around a stable feature-level interface: given multi-scale visual features and a text representation, it injects semantic guidance through shape-preserving adapters while maintaining the decoder-side tensor contract. To make this interface effective, we introduce a Hierarchical Coarse-to-Fine Supervision Strategy that decomposes learning into global image-text alignment, multi-scale auxiliary localization, and boundary-aware final mask refinement. We further design a Scale-Adaptive Gated Semantic Guidance (SAGSG) adapter, where resolution-specific gates adaptively control textual injection and channel recalibration suppresses redundant cross-modal responses. Evaluations across diverse vision and text backbones show that the same adapter and supervision design remains effective across convolutional and transformer-based visual encoders as well as different language encoders. Experiments on four public datasets further demonstrate that BTHA improves strong text-guided baselines with modest computational overhead.
Yungeng Liu, Xuanzi Fang, Haijin Zeng +2
Harbin Institute of Technology (Shenzhen) · Shenzhen, China · NingBo No.2 Hospital +1
Clinical text can narrow down what to segment, but recent text-guided designs emphasize spatial alignment while overlooking frequency content that governs texture and boundaries. We propose Dual-Domain Cross-Modal Decoding (DD-CMD) for clinical text-guided pulmonary infection segmentation, integrating two complementary forms of language guidance during decoding. In the spatial domain, Text-Guided Spatial Cross-Attention (TGSA) aligns multi-scale visual tokens with text semantics and updates features through gated residual fusion. In the frequency domain, Spectral-Text Adaptive Modulation (STAM) applies a 2D DCT to compute learnable band-energy statistics and predicts text-conditioned FiLM parameters to recalibrate decoder channels for frequency-aware decoding. DD-CMD embeds TGSA and STAM into a coarse-to-fine decoder (7x7 to 56x56) and restores full-resolution masks using a lightweight two-stage refinement module. Experiments on QaTa-COV19 and MosMedData+ show that DD-CMD achieves 91.46% Dice / 84.26% mIoU and 81.95% Dice / 69.42% mIoU, respectively, with average gains of +1.96 Dice and +2.67 mIoU over the strongest prior baselines. Code: https://github.com/maklachur/DD-CMD.
Md Maklachur Rahman, Tracy Hammond
Texas A&M University, College Station, TX 77843, USA
Medical image segmentation is a fundamental task in numerous medical engineering applications. Recently, language-guided segmentation has shown promise in medical scenarios where textual clinical reports are readily available as semantic guidance. Clinical reports contain diagnostic information provided by clinicians, which can provide auxiliary textual semantics to guide segmentation. However, existing language-guided segmentation methods neglect the inherent pattern gaps between image and text modalities, resulting in sub-optimal visual-language integration. Contrastive learning is a well-recognized approach to align image-text patterns, but it has not been optimized for bridging the pattern gaps in medical language-guided segmentation that relies primarily on medical image details to characterize the underlying disease/targets. Current contrastive alignment techniques typically align high-level global semantics without involving low-level localized target information, and thus cannot deliver fine-grained textual guidance on crucial image details. In this study, we propose a Target-informed Multi-level Contrastive Alignment framework (TMCA) to bridge image-text pattern gaps for medical language-guided segmentation. TMCA enables target-informed image-text alignments and fine-grained textual guidance by introducing: (i) a target-sensitive semantic distance module that utilizes target information for more granular image-text alignment modeling, (ii) a multi-level contrastive alignment strategy that directs fine-grained textual guidance to multi-scale image details, and (iii) a language-guided target enhancement module that reinforces attention to critical image regions based on the aligned image-text patterns. Extensive experiments on four public benchmark datasets demonstrate that TMCA enabled superior performance over state-of-the-art language-guided medical image segmentation methods.
Mingjian Li, Mingyuan Meng, Shuchang Ye +4
a School of Computer Science, The University of Sydney, Sydney, Australia. · b Institute of Translational Medicine, Shanghai Jiao Tong University, Shanghai, China. · c Zhongguancun Academy & Zhongguancun Institute of Artificial Intelligence, Beijing, China. +1