Early detection of oral cancer via photographic imaging presents a promising avenue for large-scale oral cavity screening. However, the development of robust deep learning models is frequently hampered by the scarcity of high-quality, annotated datasets. To address this limitation, a novel and well-curated resource, the Photographic Multi-purpose Oral Cancer Imaging (PhotoMOCI) dataset, is introduced for developing models across multiple diagnostic tasks in oral oncology. Then, a comprehensive benchmark study was conducted to investigate how various data augmentation strategies influence the performance of image classifiers. Our analysis spans different generative AI frameworks, evaluating the efficacy of traditional methods against advanced generative approaches, including Generative Adversarial Networks (GANs) and Diffusion Models (DMs). Additionally, we propose the Synthetic Image Filter (SIF), a mechanism to select specific samples based on two auxiliary models: Synthetic Proxy Classifier to ensure samples are representative of the target class and Synthetic Image Detector to verify they appear realistic, thereby selecting only the high-utility images that contribute to improving downstream performance. Across the evaluated datasets and classifiers, the best SIF-filtered setup improves accuracy over traditional augmentation in all cases, with gains of +1.73% and +2.35% on PhotoMOCI and +2.38% and +2.08% on KOCD for ResNet50 and ViT, respectively. Our findings reveal that while the direct application of generative data augmentation may yield performance drops, the integration of SIF, considering (i) how synthetic data looks real and (ii) how it reflects the discriminative features of the belonging class, provides a simple yet effective mechanism to filter out synthetic samples that confuse the classifier during training.
Figures & tables
Name
Year
#Imgs (w/ lesions)
Downstream tasks
Kaggle OC Lips and Tongue Shivam and Prakrut (2020)
Table 1 : Existing photographic oral cancer datasets for computer vision. The column #Imgs (w/ lesions) reports total images and lesion‑positive ones (within round brackets).
Figure 1 : PhotoMOCI metadata distribution by sex in (a) and age group in (b), stratified by pathology class.
Figure 2 : Overview of the proposed methodology for oral lesion classification. In the ochre dotted box, the Synthetic Image Generation based on generative models. In the green dotted box, the Synthetic Image Filter combining the Synthetic Proxy Classifier and the Synthetic Image Detector to select informative synthetic samples. In the light-blue dotted boxes, three training setups: using the original small dataset (top), the augmented synthetic dataset (middle), and the augmented synthetic filtered dataset (bottom), the latter aiming to improve generalization.
Figure 3: Two examples of diffusion processes using SD(txt+img) with conditioning mechanisms based on text and image for neoplastic and traumatic lesions in (a) and (b) from PhotoMOCI, respectively. In each example, ten denoising processes with image conditioning performed at different steps. The red -bordered images indicate where the conditioning image is introduced. In the last row, with water-green borders, the original conditioning image injected at the 100th step, overriding earlier ones.
Synthetic img. gen. setup
FID ↓
KID ↓
Prec. ↑
Recall ↑
Cov. ↑
Tr. aug.
230.30
0.126
0.317
0.183
0.058
SD(txt)
223.92
0.079
0.149
0.660
0.191
SD(txt+img)
179.24
0.072
0.543
0.319
0.521
SG3(lbl)
195.77
0.099
0.404
0.234
0.287
AC-SG3
290.42
0.230
0.426
0.011
0.032
Table 2 : Synthetic Image Generation results on PhotoMOCI.
Synthetic img. gen. setup
FID ↓
KID ↓
Prec. ↑
Recall ↑
Cov. ↑
Tr. aug.
262.91
0.126
0.284
0.632
0.242
SD(txt)
225.38
0.074
0.273
0.762
0.231
SD(txt+img)
143.87
0.034
0.867
0.573
0.888
SG3(lbl)
169.57
0.062
0.888
0.112
0.748
AC-SG3
319.78
0.245
0.035
0.035
0.042
Table 3 : Synthetic Image Generation results on KOCD.
Figure 4 : Synthetic image examples generated from the PhotoMOCI dataset using four synthetic image generation setups: Stable Diffusion conditioned on text (top-left), Stable Diffusion conditioned on text and image (top-right), StyleGAN3 conditioned on text (bottom-left), and AC-SG3 (bottom-right). Within each block, the first, second, and third rows correspond to traumatic, aphthous, and neoplastic lesions, respectively.
Figure 5 : Synthetic image examples generated from the KOCD using four synthetic image generation setups: Stable Diffusion conditioned on text (top-left), Stable Diffusion conditioned on text and image (top-right), StyleGAN3 conditioned on text (bottom-left), and AC-SG3 (bottom-right). Within each block, the first row corresponds to the cancer class, and the second row corresponds to the non-cancer class.
Synthetic img. gen. setup
Single image time (s)
PhotoMOCI 2100 images
KOCD 3000 images
SD(txt)
1.86 ± 0.05
1 h 05 min 06 s
1 h 33 min 00 s
SD(txt+img)
0.82 ± 0.06
28 min 42 s
41 min 00 s
SG3(lbl)
0.045 ± 0.003
1 min 35 s
2 min 15 s
AC-SG3
0.048 ± 0.004
1 min 41 s
2 min 24 s
Table 4 : Computational time for generating synthetic images on an NVIDIA A100 GPU.
Dataset
Synthetic img. gen. setup
Accuracy
SPC ↑
SID ↓
PhotoMOCI (ours)
SD(txt)
0.368 ± 0.045
0.788 ± 0.032
SD(txt+img)
0.771 ± 0.037
0.643 ± 0.040
SG3(lbl)
0.740 ± 0.029
0.596 ± 0.038
AC-SG3
0.893 ± 0.036
0.921 ± 0.036
KOCD Rashid (2024)
SD(txt)
0.617 ± 0.025
0.709 ± 0.030
Table 5 : SIF modules performance: classification results of SPC and SID on PhotoMOCI and KOCD considering the four synthetic data generation.
Synthetic img. gen. setup
PhotoMOCI
KOCD
Total
Traumatic
Aphthous
Neoplastic
Total
Cancer
Non-cancer
SD(txt)
139 (6.6%)
30 (4.3%)
83 (11.9%)
26 (3.7%)
308 (10.3%)
91 (6.1%)
217 (14.5%)
SD(txt+img)
470 (22.4%)
157 (22.4%)
119 (17.0%)
194 (27.7%)
653 (21.8%)
362 (24.1%)
291 (19.4%)
SG3(lbl)
367 (17.5%)
141 (20.1%)
137 (19.6%)
89 (12.7%)
776 (25.9%)
415 (27.7%)
361 (24.1%)
AC-SG3
70 (3.3%)
12 (1.7%)
19 (2.7%)
39 (5.6%)
152 (5.1%)
24 (1.6%)
128 (8.5%)
Table 6 : Number and percentage of synthetic images retained by SIF on PhotoMOCI and KOCD, reported in aggregate and by class.
Model
Conditioning
SIF
Acc
Prec
Recall
F1-score
ResNet He et al. (2016)
No aug.
✗
0.81050 ± 0.00532
0.81538 ± 0.00553
0.81076 ± 0.00581
0.81414 ± 0.00601
Tr.aug.(base)
✗
0.82519 ± 0.00600
0.83130 ± 0.00572
0.82316 ± 0.00603
0.82873 ± 0.00612
SD(txt)
0.74153 ± 0.00590
0.74797 ± 0.00619
0.74225 ± 0.00528
0.74550 ± 0.00493
✓
0.80076 ± 0.00544
0.82653 ± 0.00485
0.80567 ± 0.00521
0.81237 ± 0.00495
SD(txt+img)
0.83655 ± 0.00520
0.83704 ± 0.00446
0.83679 ± 0.00499
0.83698 ± 0.00524
✓
0.84253 ± 0.00498
0.85137 ± 0.00507
0.84294 ± 0.00481
0.85095 ± 0.00511
Table 7 : Classification performance on PhotoMOCI using ResNet and ViT under different data augmentation strategies. The evaluated setups include training without augmentation (No aug.) and with traditional augmentation (Tr. aug.), marked with ✗ since SIF is not applicable, as well as four generative-AI–based setups evaluated both without and with SIF, distinguished by ✓. Highlighted in gray , the setups where SIF improves performance over non-SIF counterparts and surpasses traditional augmentation. Highlighted in bold , the best-performing augmentation setup for each classifier.
Model
Conditioning
SIF
Acc
Prec
Recall
F1-score
ResNet He et al. (2016)
No aug.
✗
0.90935 ± 0.00453
0.90028 ± 0.00452
0.90074 ± 0.00449
0.90026 ± 0.00454
Tr.aug.(base)
✗
0.93109 ± 0.00554
0.93103 ± 0.00547
0.93120 ± 0.00563
0.93095 ± 0.00576
SD(txt)
0.88564 ± 0.00637
0.89160 ± 0.00656
0.88713 ± 0.00625
0.89121 ± 0.00700
✓
0.91616 ± 0.00410
0.91937 ± 0.00474
0.91668 ± 0.00422
0.91871 ± 0.00390
SD(txt+img)
0.95011 ± 0.00498
0.95225 ± 0.00501
0.94917 ± 0.00533
0.95075 ± 0.00511
✓
0.95492 ± 0.00467
0.95719 ± 0.00501
0.95266 ± 0.00461
0.95543 ± 0.00521
Table 8 : Classification performance on KOCD using ResNet and ViT under different data augmentation strategies. The evaluated setups include training without augmentation (No aug.) and with traditional augmentation (Tr. aug.), marked with ✗ since SIF is not applicable, as well as four generative-AI–based setups evaluated both without and with SIF, distinguished by ✓. Highlighted in gray , the setups where SIF improves performance over non-SIF counterparts and surpasses traditional augmentation. Highlighted in bold , the best-performing augmentation setup for each classifier.
Figure 6 : Confusion matrices for PhotoMOCI comparing traditional augmentation and SD(txt+img)+SIF for ResNet50 and ViT.
Dataset
Model
χ2
p
Kendall’s W
KOCD
ResNet
54.938
<0.0001
0.610
KOCD
ViT
53.520
<0.0001
0.595
PhotoMOCI
ResNet
56.749
<0.0001
0.631
PhotoMOCI
ViT
50.787
<0.0001
0.564
Table 9 : Overall statistical significance analysis using Friedman tests on accuracy.
KOCD
PhotoMOCI
Synthetic img. gen. setup
ResNet
ViT
ResNet
ViT
Δ Acc.
p
Δ Acc.
p
Δ Acc.
p
Δ Acc.
p
SD(txt)+SIF
-1.49
0.0645
-0.59
0.3750
-2.44
0.0371
-0.84
0.4316
SD(txt+img)+SIF
+2.38
0.0371
+2.08
0.0059
+1.73
0.0273
+2.35
0.0195
SG3(lbl)+SIF
+1.72
0.0488
+1.17
0.0934
+1.08
0.0754
+1.79
0.0515
AC-SG3+SIF
+0.44
0.6953
+0.21
0.6250
-1.53
0.0195
-0.17
0.6250
Table 10 : Planned Wilcoxon comparisons between synthetic+SIF setups and traditional augmentation on accuracy. Values report accuracy difference in percentage points and p -value.
KOCD
PhotoMOCI
Synthetic img. gen. setup
ResNet
ViT
ResNet
ViT
Δ Acc.
p
Δ Acc.
p
Δ Acc.
p
Δ Acc.
p
SD(txt)
+3.05
0.0039
+3.58
0.0020
+5.92
0.0020
+6.31
0.0020
SD(txt+img)
+0.48
0.0566
+0.71
0.0316
+0.60
0.0566
+0.80
0.0316
SG3(lbl)
+0.81
0.0754
+0.63
0.0750
+0.32
0.0250
+0.57
0.0695
AC-SG3
+0.74
0.0602
+0.91
0.0223
+1.84
0.0195
+1.33
0.0324
Table 11 : Planned Wilcoxon comparisons between SIF-filtered and unfiltered synthetic augmentation on accuracy. Values report accuracy difference in percentage points and p -value.
Figure 7 : Ablation study evaluating the impact of SPC and SID within the SIF design on (a) PhotoMOCI and (b) KOCD using ViT. The study compares training without augmentation (blue bar) to training with integrated synthetic data: unfiltered (ochre bar), filtered using SPC only (water-green bar), SID only (orange bar), and SIF (pink bar).
Figure 8 : SID threshold calibration on (a) PhotoMOCI and (b) KOCD. Each plot reports accuracy as a function of the SID threshold τ for ResNet50 (left) and ViT (right), with shaded 95% confidence intervals over 10 runs and a marker indicating the best threshold.