cs.CVSep 24, 2026

Multimodal Routing and Region Refinement for Language-Guided Medical Image Segmentation

Authors: Md Maklachur Rahman, Md Hasan Al Banna, Saraf Anjum, Assame Arnob, Tracy Hammond

Organizations: Texas A&M University, College Station, TX 77843, USA · Independent Researcher

Abstract

Textual descriptions can reduce ambiguity in medical image segmentation by specifying the finding and location to be delineated. Existing text-guided methods mainly improve where image and language features interact but generally retain a single learned update pathway across all image-text pairs. We propose MRSeg, a parameter-efficient framework that uses each image-text pair to route the adaptation of visual and textual features before dense prediction. Frozen ConvNeXt-Tiny and PubMedBERT encoders provide multiscale visual features and clinical text tokens. A joint router uses the deepest visual feature and pooled text to predict a sparse mixture over low-rank adapter bases. The resulting route is shared across separate adapter banks for two visual scales and text, coordinating their adaptation while keeping the feature-specific parameters separate. Region Bridge uses text-derived queries to aggregate dense visual tokens into latent regions, refines these regions through self-attention and text cross-attention, and redistributes the refined information back to the feature maps. Finally, a multiscale decoder combines refined semantic features with shallow image evidence. On QaTa-COV19 and MosMedData+, MRSeg achieves 90.90/83.32 and 81.53/68.82 Dice/mIoU, respectively, with 7.11M trainable parameters and 7.60 GFLOPs. Code: https://github.com/maklachur/MRSeg.

Figures & tables

Explore similar work

CardsList
  1. Decoupling Language Guidance from Backbones for Text-Guided Medical Segmentation

    Jul 10, 2026Yungeng Liu, Xuanzi Fang, Haijin Zeng +2Semi-Supervised Medical Image SegmentationVision Encoders

  2. Dual-Domain Cross-Modal Decoding for Clinical Text-Guided Medical Image Segmentation

    Aug 11, 2026Md Maklachur Rahman, Tracy HammondSemi-Supervised Medical Image SegmentationMultimodal Clinical Data

  3. Language-guided Medical Image Segmentation with Target-informed Multi-level Contrastive Alignments

    Dec 18, 2024Mingjian Li, Mingyuan Meng, Shuchang Ye +4Semi-Supervised Medical Image SegmentationContrastive Alignment