cs.CVJun 18, 2026

Geometry-Aware Superpixel Graph Transformer with Metadata for Skin Lesion Classification

Authors: Muhammad AzeemTanveer HussainAmr AhmedArdhendu Behera

Abstract

Automated skin cancer classification from dermoscopic images remains challenging due to heterogeneous lesion structure, strong intra-class variability, and subtle visual differences between benign and malignant cases. Existing CNN/ViT pipelines typically rely on global or patch-level features and often combine patient metadata via late fusion, which limits spatially grounded multimodal reasoning. We present a novel region-based graph learning framework that explicitly models lesions as graphs of spatially coherent superpixel regions represented as frozen CNN features. To capture fine-grained lesion arrangements, we encode inter-regional geometry as edge attributes and introduce a dedicated metadata context node connected to all regions, providing structured integration of demographic/clinical variables within the same relational space. Node representations are updated using our edge-aware graph transformer followed by attention-driven propagation, and a final graph-level embedding for benign-malignant classification. Experiments on four public benchmarks demonstrate that explicit region-level relational modeling and graph-native multimodal fusion yield consistent gains over the state-of-the-art. Consequently, we establish a new graph-centric perspective in which CNN features are modeled as relational nodes and improved through contextual integration, yielding more expressive and robust classifications.

Explore similar work

Aug 4, 2026cs.CV

Multimodal Skin Lesion Classification with Swin Transformer and Clinical Metadata Fusion

Skin lesion classification plays an important role in supporting the early diagnosis of skin cancer. However, automated analysis remains challenging due to class imbalance, inter-class similarity, and intra-class variability in dermoscopic images. This paper proposes a multimodal classification framework that combines Swin Transformer-based image features with structured clinical metadata to improve diagnostic performance through integrated visual-context learning. Experiments on a publicly available dataset show that the proposed model achieves a test accuracy of 92.55% and a macro F1-score of 91.33%, with strong performance across minority classes. Temperature scaling is applied as a post-hoc calibration method, resulting in a reduction in expected calibration error and improving prediction reliability, while uncertainty estimation is incorporated to further assess the confidence of model predictions. Qualitative explainability analysis further shows that the model focuses on lesion regions during inference. Therefore, the results demonstrate that multimodal fusion, combined with calibration and interpretability analysis, provides an effective and trustworthy approach for automated skin lesion classification.
Nethmi Pathirana, Isuru Munasinghe, Dileeka Alwis
Aug 6, 2026cs.AI

Integrating Implicit and Explicit Relational Biases through Graph-Based Multiple Instance Learning: A Case Study in Skin Lesion Diagnosis

Relational inductive biases are essential for capturing structural dependencies among data. This study investigates a dual-level relational framework for image classification, bridging the gap between implicit representation learning and explicit structural modelling. We begin by establishing a baseline using an EfficientNetB3 architecture. To move beyond standard convolutional biases, we adopt a patch-based strategy, employing a convolutional masked autoencoder to learn implicit inter-patch relationships through self-supervised reconstruction. We then extend this approach by incorporating explicit relational modelling, organizing the learned embeddings into various graph topologies, including grid-based, random, and k-nearest neighbour structures. Experimental results on the ISIC-2018 and ISIC-2019 skin lesion diagnosis benchmarks show that combining implicit inter-patch modelling with explicit graph-based message passing yields the best performance. On the ISIC-2018 test set, the baseline model achieves a balanced accuracy of 76.17%, which improves to 77.12% with implicit patch-based relational modelling. The fully integrated grid-structured Graph Attention Network further increases performance to 79.27%. Similarly, on ISIC-2019, the implicit approach reaches 59.84% balanced accuracy, while the combination of implicit and explicit modelling yields 60.67%.
Rafał Buler, Jakub Buler, Maciej Bobowicz +1
Apr 30, 2026cs.CV

JI-ADF: Joint-Individual Learning with Adaptive Decision Fusion for Multimodal Skin Lesion Classification

Skin lesion classification is essential for early dermatological diagnosis, yet many existing computer-aided systems rely primarily on dermoscopic images and underutilize the multimodal evidence routinely available in clinical practice. To address this gap, we propose \textbf{JI-ADF}, a trimodal deep learning framework that integrates dermoscopic images, clinical photographs, and structured patient metadata for clinically grounded skin lesion classification. The proposed architecture combines joint multimodal representation learning with modality-specific auxiliary supervision and an adaptive decision fusion mechanism that dynamically calibrates modality contributions on a per-sample basis. To enhance cross-modal reasoning while preserving modality-specific evidence, we further introduce a multimodal fusion attention (MMFA) module. We evaluate JI-ADF on the large-scale MILK10k benchmark, which reflects real-world clinical acquisition conditions and severe class imbalance. The proposed method demonstrates strong and well-balanced performance across lesion categories, improving sensitivity and Dice score while maintaining high specificity and good calibration. Extensive analyses, including modality ablation, calibration evaluation, and Grad-CAM visualization, further confirm the robustness and clinically meaningful behavior of the model. These results indicate that JI-ADF provides a reliable and practical foundation for multimodal skin lesion classification in real-world clinical settings.
Phan Nguyen, Dat Cao, Hien Kha +4