Organizations: School of Artificial Intelligence and Computer Science, Nantong University, Nantong 226019, China · School of Telecommunications Engineering, Xidian University, Xi’an 710126, China · College of Computer Science and Technology, Zhejiang University, Hangzhou 310058, China · School of Biomedical Engineering, Shanghai Jiao Tong University, Shanghai 200030, China · Department of Computer Science, Durham University, Durham, UK
Cross-domain diagnosis remains a major challenge in cervical cell pathology due to pronounced domain shifts across institutions and the subtle visual differences among disease stages, which jointly impair model generalization. To address these issues, this paper proposes a two-stage framework for cross-domain cervical cell detection. In the first stage, we propose the Spatially-Continuous Unpaired Neural Schrödinger Bridge (SC-UNSB), which constructs a synthetic intermediate domain to mitigate cross-domain distribution shifts by modeling image translation as an entropy-regularized optimal transport process. In the second stage, we propose a dual-level feature alignment strategy within a knowledge distillation, which progressively aligns shallow structural features and deep semantic representations to facilitate the transfer of domain-invariant knowledge from the source to the target model. Experimental results demonstrate that the proposed method effectively mitigates domain shift and category ambiguity, improving the cross-domain detection performance.
Deep learning-based computer-aided diagnosis (CAD) systems have shown strong performance in breast cancer diagnosis, particularly for classification tasks in mammography. However, domain shifts across multi-site datasets remain a challenge, especially when models are applied to unseen domains. In this work, we proposed a calcification classification framework to improve malignant versus benign breast disease classification across multi-site mammography datasets. The framework consisted of two components: (1) an unsupervised domain adaptation module based on style transfer models (AdaIN and CycleGAN) to generate vendor-specific and technique-specific training samples without additional annotations, and (2) a supervised classification module using Swin Transformer V2 as the backbone. We evaluated the proposed method on three datasets: cross-validation on OPTIMAM (National Health Service, United Kingdom; n=2994), followed by external validation on EMBED (Emory University; n=125), and Duke Calcification Dataset v1 (n=788). These datasets cover multiple vendors and include both full-field digital mammography and synthetic 2D images derived from digital breast tomosynthesis. The proposed framework improved cross-site performance for both EMBED (AUC 0.68 to 0.72) and the Duke Calcification Dataset (AUC 0.68 to 0.73). These findings indicate that domain adaptation can reduce domain shifts and improve the generalization for calcification classification across multi-site datasets.
Cross-domain cell detection for microscopic images suffers from performance degradation due to distribution shifts across imaging domains. Unsupervised Domain Adaptation (UDA) strategies, attempt to overcome domain sift without requiring annotated data from target. However, requirement of availability of annotated data from the source domain and large-size data from target domain are both challenging limitations for realistic scenarios. This is especially true in medical imaging, where privacy requirements might prevent access to annotated source data, and costly data acquisition restricts extensive sampling of the target domain. To address these challenges, we propose AdaptiveCDM, a modular framework for Source-Free Few-Shot Domain Adaptive Object Detection (SF-FSDAOD) setting, that adapts a pretrained source model using only few labeled target images without accessing source data. AdaptiveCDM combines Resolution-Aware Augmentation (RAug) and Category-Aware Representation Learning (CARL). RAug alleviates the scarcity and class imbalance by augmenting instance balanced training examples, while preserving the scale fidelity and morphological properties of cellular structures. CARL enhances discriminative representation learning by encouraging class-consistent proposals, improving both localization and classification. We also introduce two competitive baselines for proposed setting: Faster-FreeShot and MT-FreeShot. Our approach achieves 40.4/43.4 mAP0.5 on M5 and 67.1/75.5 mAP0.5 on Raabin-WBC under 2-/5-shot adaptation. Despite using only a few labeled target images and no source data, AdaptiveCDM achieves competitive or superior performance compared with SOTA methods under their respective supervision settings. Ablations and qualitative analyses further substantiate the contribution of each component and the effectiveness of AdaptiveCDM in low-data regimes. Code/models will be available.
Multimodal based approaches often outperform single modality approaches in downstream tasks as the different modalities provide complementary information, yet acquiring paired clinical data remains a significant challenge in real world scenarios. While cross-modal knowledge distillation addresses this, existing methods often struggle with large modality gaps and the propagation of noise from uncertain source-domain predictions. To overcome these challenges, we propose UnDA, an anchor-guided framework for unpaired cross-modal distillation. Our approach introduces a backbone-agnostic Alignment Module that extracts semantically structured class tokens via an attention based pooling mechanism. To ensure robust knowledge transfer, we propose Uncertainty-Weighted Optimal Transport (UCT-OT), which dynamically weights feature-level alignment based on prediction confidence, effectively suppressing noisy supervision. Furthermore, a per-class ProtoNCE objective maintains stable prototype memories to enforce global discriminability across unpaired batches. Evaluations on representative segmentation tasks under strictly unpaired settings show consistent improvements in accuracy and boundary precision in the target modality, demonstrating that meaningful structural knowledge can be transferred across heterogeneous data sources without paired datasets.