cs.CVSep 29, 2026

Detail in Context: A Dual-Scale Machine Learning Framework for Mycosis Fungoides Detection

Authors: Mohamed Hazem, Tarek Waleed, Omar Khaled, Nada Omar, Mahmoud Raslan, Marwa Mohamed Fawzy, Aya Fahim, Rania M. Mogawer, +3 more

Organizations: Faculty of Engineering, Cairo University, Giza, Egypt · Dermatology Department, Faculty of Medicine, Cairo University, Cairo, Egypt

Abstract

Mycosis fungoides (MF) is a rare form of cutaneous T-cell lymphoma that is often misdiagnosed in early stages due to its visual similarity to benign inflammatory dermatoses. Early and accurate diagnosis is critical for improving patient outcomes. In this paper, we propose a comprehensive diagnostic framework for automated MF detection that combines dual- scale histopathological image analysis with deep learning. To distinguish MF from other lymphoproliferative skin conditions, the proposed approach leverages a late-fusion ensemble of dual- magnification (10x and 20x) convolutional neural networks (CNNs), complemented by a random forest classifier trained on 16 clinical features. Experimental results on an expanded dataset of 6,267 images (4,306 MF; 1,961 Non-MF) across 463 patients demonstrate that strong detection performance is obtained by prioritizing higher-resolution cytological details (20x) within broader architectural context (10x). The image-based late-fusion model achieves an accuracy of 83.58% and a sensitivity of 89.13%, while the clinical random forest model achieves an accuracy of 96.6% and sensitivity of 93.8%, highlighting the po- tential of this multimodal framework as a robust clinical decision support system in dermatology. This framework addresses two distinct clinical objectives: an image-based dual-scale pipeline optimized for the early diagnostic screening of MF versus non- MF dermatoses, and a complementary clinical metadata model designed for the subsequent staging of confirmed MF cases (patch/plaque versus tumor)

Figures & tables

Explore similar work

Apr 30, 2026cs.CV

JI-ADF: Joint-Individual Learning with Adaptive Decision Fusion for Multimodal Skin Lesion Classification

Skin lesion classification is essential for early dermatological diagnosis, yet many existing computer-aided systems rely primarily on dermoscopic images and underutilize the multimodal evidence routinely available in clinical practice. To address this gap, we propose \textbf{JI-ADF}, a trimodal deep learning framework that integrates dermoscopic images, clinical photographs, and structured patient metadata for clinically grounded skin lesion classification. The proposed architecture combines joint multimodal representation learning with modality-specific auxiliary supervision and an adaptive decision fusion mechanism that dynamically calibrates modality contributions on a per-sample basis. To enhance cross-modal reasoning while preserving modality-specific evidence, we further introduce a multimodal fusion attention (MMFA) module. We evaluate JI-ADF on the large-scale MILK10k benchmark, which reflects real-world clinical acquisition conditions and severe class imbalance. The proposed method demonstrates strong and well-balanced performance across lesion categories, improving sensitivity and Dice score while maintaining high specificity and good calibration. Extensive analyses, including modality ablation, calibration evaluation, and Grad-CAM visualization, further confirm the robustness and clinically meaningful behavior of the model. These results indicate that JI-ADF provides a reliable and practical foundation for multimodal skin lesion classification in real-world clinical settings.
Sep 12, 2026cs.CV

A Comparative Evaluation of Pre-trained Convolutional Neural Networks for Melanoma Detection

Early diagnosis of melanoma is critical for improving patient survival rates. However, accurately distinguishing melanoma from other skin lesions remains a significant clinical challenge due to the high visual similarity among lesion types and variability in image acquisition conditions. Artificial intelligence, particularly machine learning, has emerged as a promising tool to support dermatological diagnosis by automating feature extraction from medical images. Among the available approaches, convolutional neural networks (CNNs) have demonstrated strong performance in image classification tasks, making them well-suited for analyzing both dermatoscopic and histopathological images, given their ability to capture hierarchical visual patterns relevant to lesion characterization. Nevertheless, despite numerous pre-trained CNN architectures having been proposed, selecting the most appropriate one for a given imaging modality remains an open challenge. In this study, we evaluate pre-trained convolutional neural networks (CNNs) for skin lesion classification using dermatoscopic and histopathological image datasets. Experiments were conducted on the HAM10000, ISIC 2018, and CR-AI4SkIN datasets, evaluating the ResNet50, VGG16, VGG19, MobileNet, and InceptionV3 architectures under the same training protocol. The experimental evaluation showed that the models achieved accuracies ranging from 71% (InceptionV3 on ISIC 2018) to 84% (ResNet50 on HAM10000) on dermatoscopic images. For histopathological images, accuracies ranged from 72% (VGG19) to 83% (ResNet50) on the CR-AI4SkIN dataset. The results demonstrate that model performance differs between dermatoscopic and histopathological image modalities, showing that architectures exhibiting similar performance on dermatoscopic images exhibit different performance on histopathological data.
Aug 4, 2026cs.CV

Multimodal Skin Lesion Classification with Swin Transformer and Clinical Metadata Fusion

Skin lesion classification plays an important role in supporting the early diagnosis of skin cancer. However, automated analysis remains challenging due to class imbalance, inter-class similarity, and intra-class variability in dermoscopic images. This paper proposes a multimodal classification framework that combines Swin Transformer-based image features with structured clinical metadata to improve diagnostic performance through integrated visual-context learning. Experiments on a publicly available dataset show that the proposed model achieves a test accuracy of 92.55% and a macro F1-score of 91.33%, with strong performance across minority classes. Temperature scaling is applied as a post-hoc calibration method, resulting in a reduction in expected calibration error and improving prediction reliability, while uncertainty estimation is incorporated to further assess the confidence of model predictions. Qualitative explainability analysis further shows that the model focuses on lesion regions during inference. Therefore, the results demonstrate that multimodal fusion, combined with calibration and interpretability analysis, provides an effective and trustworthy approach for automated skin lesion classification.