Deep learning classifiers for dermoscopic skin lesions often reach high in-distribution accuracy while quietly relying on spurious background cues such as skin tone, device vignetting, and embedded rulers, rather than on lesion morphology. This undermines robustness and fairness across skin tones. This work asks whether Explainable AI (XAI), typically used only to audit a finished model, can instead be repurposed as an active training signal that corrects this shortcut without sacrificing diagnostic accuracy. We introduce CAMEO (Class Activation Mapped Equitable Overlay), a framework that improves skin-lesion classification by selecting stable model explanations and using them to separate lesions from their backgrounds. It then replaces the background with realistic synthetic skin while keeping the lesion unchanged. On HAM10000 and dark-skin ISIC images, CAMEO maintained accuracy while reducing background-driven errors by nearly four times. It also made the model's attention more consistent when backgrounds changed. Results across multiple tests show that reducing reliance on background information improves robustness, with Fitzpatrick-based backgrounds providing a realistic and interpretable approach. Results show that XAI-guided augmentation can make dermoscopic classifiers measurably more robust and fair at no cost to accuracy. They also clarify that it is the mechanism and not the specific tone palette that matters, and that the lasting contribution of XAI here lies in stability-screened, annotation-free lesion localisation rather than in the robustness number itself.
Figures & tables
Set
N
Description and purpose
T0
342
Frozen class-balanced HAM10000 test set (in-distribution).
T1
100
External dark-skin ISIC set (real cross-tone shift).
T_swap
342
Synthetic background swap of T0 (controlled robustness).
Table 1: Dataset notation used throughout. T1 (external, real) and T_swap (synthetic, controlled) are different sets and are never compared directly. Two probes use T_swap-style compositing: the decision-flip and attention probe (Section 5.6 ) re-composites 150 lesions onto six tones, whereas the T_swap accuracy metric uses one fixed swap per T0 image ( N=342 ).
Name
Core
Background strategy
Role
Base-Skewed (M1)
Full HAM10000, imbalanced (6,224 and 774)
None (raw images)
Naive baseline; reflects the natural class imbalance.
Base-Balanced (M2)
Balanced subset (774 and 774)
None (raw images)
Primary baseline to beat; class imbalance removed by undersampling.
CAMEO (ours) (M3_AugB_M2, v4)
Balanced subset (774 and 774)
ChromaSwap , symmetric on both classes
Headline model; the primary result reported throughout the paper.
CAMEO-Skewed (M3_AugB_M1)
Full HAM10000, imbalanced core
ChromaSwap , minority class only
Tests whether augmentation alone can substitute for explicit class balancing.
CAMEO-RandomBG
Balanced subset (774 and 774)
Random, non-Fitzpatrick background
Internal control: isolates whether Fitzpatrick-specific tones are necessary.
BiasAdv [ 26 ]
Balanced subset (774 and 774)
Bias-conflicting background against a frozen biased classifier
External debiasing baseline, named as in its source publication.
Table 2: Model and baseline glossary. Every name used in the paper is defined here in one place. “Core” is the training data recipe and “Background strategy” is what replaces the region outside the lesion mask, if anything. Identifiers from earlier pipeline versions are given in parentheses for traceability.
Figure 1: Overview of the CAMEO pipeline. Top row: data preparation, XAI extraction and the SSIM stability screen that decides which explanations may be acted on and how strongly each sample is augmented. Middle row: ruler artifact suppression and composite mask generation. Bottom row: ChromaSwap augmentation, retraining and evaluation. The screen is a gate, not a diagnostic: no attribution map reaches the masking stage without passing through it.
Model
Accuracy
Macro F1
AUC-ROC
Cohen’s κ
Base-Skewed
0.743 [0.696, 0.789]
0.736
0.819 [0.774, 0.862]
0.485
Base-Balanced
0.848 [0.810, 0.886]
0.848
0.912 [0.881, 0.941]
0.696
CAMEO (ours)
0.836 [0.795, 0.874]
0.836
0.913 [0.880, 0.942]
0.673
CAMEO-Skewed
0.795 [0.751, 0.836]
0.791
0.898 [0.862, 0.931]
0.591
Table 3: Clean accuracy and related metrics on the frozen balanced test set T0 ( N=342 ). Because T0 is class-balanced, accuracy equals balanced accuracy, so a single accuracy column is reported. Bracketed values are bootstrap 95% confidence intervals (10,000 resamples) for the adjacent metric.
Model
Train
Val.
Test (T0)
Train − Test
Base-Skewed
0.643
0.735
0.743
− 0.100
Base-Balanced
0.981
0.795
0.848
+ 0.133
CAMEO (ours)
0.995
0.824
0.836
+ 0.159
CAMEO-Skewed
0.992
0.759
0.795
+ 0.196
Table 4: Train, validation and test accuracy (overfitting check). Train accuracy is measured on a random 2,000-image subset with evaluation transforms only. All values are accuracies.
Figure 2: Training and validation loss and accuracy against epoch for the augmented models, showing convergence and the observed gap between training and validation performance.
Statistic
Value
Images analysed
3,000
Mean SSIM
0.758
Median SSIM
0.766
SSIM range (min to max)
0.451 to 0.967
XAI-unstable (SSIM <0.7 )
744 (24.8%)
Background-bias flagged
1,559 (52.0%)
Table 5: XAI-instability statistics over the 3,000-image extraction set. SSIM is the mean structural similarity of the GradCAM map under three Gaussian input perturbations ( σ=0.01 ).
Set
N
Mean R2
Rshift
ID rate
HAM10000 val. (calibration)
1,498
0.966
0.000
95.0%
Frozen T0 (control)
342
0.965
0.131
93.3%
T1 ISIC dark (external)
100
0.929
1.751
63.0%
Table 6: Covariate-shift statistics in the Base-Balanced feature space (tap 4). R2 is the per-image reconstruction R-square in the HAM10000 training PCA basis, Rshift is the mean normalised per-feature shift, and the in-distribution (ID) rate is the fraction of images with R2≥τ=0.9426 .
Figure 3: R2 distributions for the HAM10000 validation split, the frozen T0 test set and the external T1 dark-skin set. The vertical dashed line marks the in-distribution threshold τ=0.9426 . T0 aligns with HAM10000 validation, whereas T1 shows a pronounced low- R2 tail that corresponds to out-of-distribution images.
Model
Accuracy
Accuracy
Δ
AUC
(all)
(ID)
(ID)
Base-Skewed
0.820
0.857
+ 0.037
0.861
Base-Balanced
0.640
0.810
+ 0.170
0.796
CAMEO (ours)
0.660
0.794
+ 0.134
0.676
CAMEO-Skewed
0.630
0.778
+ 0.148
0.726
Table 7: Accuracy on T1 (external dark-skin ISIC set): full set versus in-distribution (ID) subset ( nID=63 , nOOD=37 ). Δ is the accuracy gain obtained by restricting to ID images.
Model
Prob. range
Flip rate
CAM SSIM
Base-Balanced
0.250
19.8%
0.720
CAMEO (ours)
0.123
5.4%
0.842
Table 8: Background-counterfactual robustness. Decision stability: lower probability range and lower flip rate are better (150 images, 6 tones). Attention stability: higher GradCAM++ map SSIM before versus after a background swap is better (120 images).
Figure 4: Background shortcut proof on three lesions. Columns 1 to 3: original image with the ground-truth (GT) class and each model’s melanoma probability. Column 4: the same lesion with only the background swapped to a dark synthetic tone. Columns 5 and 6: predictions after the swap, green for stable or correct and red for a flip or error. Base-Balanced flips when the background darkens; CAMEO does not.
Figure 5: GradCAM++ attention on the same lesion with its original and a dark synthetic background. Base-Balanced (columns 2 and 3) drifts onto the background when the tone changes, giving a low map SSIM, whereas CAMEO (columns 4 and 5) stays on the lesion. Per-image SSIM is printed beneath each pair.
Figure 6: Decision stability under background swap. Left: per-image melanoma-probability range across tones, lower being more invariant. Right: mean probability standard deviation and label-flip rate. Bottom: a lesion whose probability under Base-Balanced swings widely with background tone while CAMEO stays nearly constant.
Group
N
Mean IoU
Mean Dice
Median Dice
All
2,332
0.422
0.570
0.611
Non-melanoma
1,858
0.428
0.576
0.618
Melanoma
474
0.397
0.545
0.577
Table 9: Composite XAI mask versus the HAM10000 ground-truth lesion mask (ISBI 2018; Tschandl et al. [ 39 ] ). IoU is Intersection over Union.
Method
Mean IoU
Otsu (grayscale) [ 36 ]
0.552
Composite XAI mask (ours)
0.427
HSV colour threshold
0.426
SkinSAM without prompt [ 27 ]
≈ 0.16
Table 10: Unsupervised segmentation baselines against ground truth ( N=2,332 ). Otsu achieves higher raw Intersection over Union (IoU) on this light-skinned benchmark, but it relies on an intensity assumption that weakens on darker skin and it does not isolate the model-attributed lesion interior that ChromaSwap augmentation needs.
Figure 7: Qualitative validation of composite XAI masks against HAM10000 ground-truth lesion segmentations. The green outline is the clinical lesion mask and the heatmap overlay is the binary XAI mask used for ChromaSwap augmentation.
Figure 8: Exhaustive ChromaSwap augmentation across Fitzpatrick I to VI (rows), with three examples each (columns). The same lesions are carried through every row, so the effect of tone level reads down each column. The lesion interior is preserved while the background is substituted with a textured tone. Hex codes and labels are shown at the left.
Checkpoint
Pipeline / recipe
T0 accuracy
T_swap accuracy
Legacy-LesionJitter
comparison / unbalanced + LesionJitter
0.798
0.792
Legacy-ChromaSwap
comparison / unbalanced + ChromaSwap
0.827
0.810
Legacy-Combined
comparison / unbalanced + A + B
0.827
0.792
Legacy-ChromaSwap-p70
threshold sweep
0.825
0.807
Legacy-ChromaSwap-p80
threshold sweep
0.807
0.789
Legacy-FlatBG
v2 flat background
0.787
0.564
Table 11: Legacy augmentation ablations (v2, v3 and threshold-sweep checkpoints; these are not the headline CAMEO v4 model, see the text). All checkpoints are evaluated on the frozen T0 and T_swap manifests without retraining. T_swap is the synthetic background-counterfactual set and not the external T1 set.
Model
T0 accuracy
T_swap accuracy
Accuracy drop
Base-Balanced
0.827±0.016
0.745±0.031
0.082±0.035
CAMEO (ours)
0.836±0.008
0.815±0.017
0.021±0.022
Table 12: Multi-seed variance. Mean ± standard deviation of accuracy over five seeds on the frozen T0 and T_swap manifests, where T_swap is the synthetic background-counterfactual set of Table 1 and not the external T1 set. “Accuracy drop” is the T0 to T_swap accuracy loss.
Model
Method
T0 accuracy
T_swap accuracy
p vs. ours
CAMEO (ours)
XAI ChromaSwap (Fitzpatrick)
0.836
0.810
n/a
CAMEO-RandomBG
Random non-Fitzpatrick ChromaSwap
0.830
0.816
0.87
BiasAdv
BiasAdv-style conflicting background [ 26 ]
0.830
0.798
0.62
AdvDebias-GRL
Adversarial tone debiasing by gradient reversal [ 25 ]
0.833
0.734
< 0.01
Table 13: Head-to-head debiasing baselines, all evaluated on the same T_swap manifest. “CAMEO (ours)” is the v4 headline checkpoint of Table 3 ; the three baselines are single-seed (42) retrains on the same balanced core with identical hyper-parameters. The last column is the two-sided McNemar p value against ours.
Figure 9: Qualitative pipeline results on three representative examples (rows). From left to right: original image, ruler-corrected GradCAM++ heatmap, cleaned composite mask (GradCAM++ plus Integrated Gradients, p60 ), and the ChromaSwap augmented image with a Fitzpatrick-scale background.
Model
Clean acc.
ε=.001
ε=.01
ε=.03
ε=.05
ε=.10
Base-Skewed
0.743
0.345
0.158
0.213
0.447
0.500
Base-Balanced
0.848
0.213
0.085
0.105
0.377
0.500
CAMEO (ours)
0.836
0.325
0.193
0.219
0.415
0.500
CAMEO-Skewed
0.795
0.409
0.272
0.289
0.386
0.503
Table 14: FGSM adversarial accuracy on the frozen T0 test set ( N=342 ). Epsilon values are given in normalised input space. All values are accuracies.
Skin cancer diagnosis from dermoscopic images remains challenging due to high intra-class variability, inter-class similarity, class imbalance, and the limited interpretability of deep learning models. This paper proposes an uncertainty-aware and explainable deep learning framework for multi-class skin lesion classification. The framework combines a vision transformer model (MaxViT-Tiny) with CNN-based models (ConvNeXt-Tiny and EfficientNetV2-B0) through deep ensemble learning. Monte Carlo (MC) Dropout estimates predictive uncertainty and identifies unreliable predictions, while Grad-CAM++, an explainable AI (XAI) technique, provides visual explanations by highlighting lesion regions that influence model decisions. Evaluated on the HAM10000 dataset, the framework achieves 96% accuracy and 99% ROC-AUC under uncertainty-aware filtering (entropy < 1.0, confidence >= 0.7), with macro-average precision, recall, and F1-score of 94%, 95%, and 95%, respectively, and 96% weighted-average scores across all three metrics. The results demonstrate accurate, interpretable, and uncertainty-aware skin lesion classification for trustworthy computer-aided diagnosis.
Rofiqul Islam, Lilatul Ferdouse
Department of Computer Science and Physics, Wilfrid Laurier University, Waterloo, ON, Canada
This study proposes a domain-specific LLM-based Visual Explanation Evaluation Framework for assessing Grad-CAM explanations in facial skin disease diagnosis models. While previous studies have primarily focused on improving classification performance through data augmentation techniques, relatively few studies have systematically examined whether model explanations are grounded in clinically relevant lesion regions. In this study, geometric augmentation, color-based augmentation, and mixed augmentation strategies were applied to facial skin disease classification models based on EfficientNet-B0, MobileNetV3, and ResNet18. Grad-CAM was employed to generate visual explanations representing the models' decision-making processes. Furthermore, an LLM-as-a-Judge evaluation framework was designed using GPT-5.5, Gemini 3.5 Flash, and Claude Sonnet 4.6 to assess Grad-CAM explanations from the perspectives of lesion localization and explanation trustworthiness. To improve evaluation consistency and clinical grounding, a progressive prompt engineering strategy was introduced, incorporating evaluation rubrics, clinical knowledge, penalty rules, and structured output formats.
Gyuyeon Na
AI and Business Analytics, Ewha Womans University, Seoul, Republic of Korea
Automated segmentation of skin lesions using deep learning models for dermoscopic images can be very helpful in finding melanomas earlier than they would normally be detected. However, most deep learning methods available do not perform well. The aim of this paper is to present a parameter-efficient fine-tuning method called PEFT-MedSAM for adapting the Medical Segment Anything Model (MedSAM) to automatically segment dermoscopic skin lesions. The PEFT-MedSAM method uses only the lightweight mask decoder for training the model while keeping the pre-trained image encoder and prompt encoder frozen. The experiments performed on the ISIC 2018 benchmark dataset shows that PEFT-MedSAM obtains a dice coefficient of .9411 and an intersection over union value of .8918 when compared to both a fully trained U-Net baseline (.8715 dice coefficient) and zero-shot MedSAM inference (.8997 dice coefficient). The external validation of the model using PH2 dataset shows .9467 dice coefficient with +/- .0310 standard deviation. Supportive evidence for these claims include a p-value less than .0001 for Wilcoxon signed rank tests comparing the two datasets and bootstrap-estimated 95% confidence intervals of [.9364,.9447] that represent the estimated range of possible values for the average dice coefficient obtained by repeating the test. To increase clinical trustworthiness, we used Grad-CAM explainability along with a pointing game based evaluation methodology to evaluate the CNN baseline model on the validation set. The results showed that we had an accuracy rate of 98.27% on the validation set of 519 images and confirmed that the model classified regions containing skin lesions.
Asad Channa, Abdullah Khan, Asghar Ali Chandio +4
Department of Computer Science, Quaid-e-Awam University of Engineering, Sciences & Technology · Department of Artificial Intelligence, Quaid-e-Awam University of Engineering, Sciences & Technology, Pakistan · Department of Computer Science, Sindh Madressatul Islam University, City Campus, Karachi +1