Early detection of oral cancer via photographic imaging presents a promising avenue for large-scale oral cavity screening. However, the development of robust deep learning models is frequently hampered by the scarcity of high-quality, annotated datasets. To address this limitation, a novel and well-curated resource, the Photographic Multi-purpose Oral Cancer Imaging (PhotoMOCI) dataset, is introduced for developing models across multiple diagnostic tasks in oral oncology. Then, a comprehensive benchmark study was conducted to investigate how various data augmentation strategies influence the performance of image classifiers. Our analysis spans different generative AI frameworks, evaluating the efficacy of traditional methods against advanced generative approaches, including Generative Adversarial Networks (GANs) and Diffusion Models (DMs). Additionally, we propose the Synthetic Image Filter (SIF), a mechanism to select specific samples based on two auxiliary models: Synthetic Proxy Classifier to ensure samples are representative of the target class and Synthetic Image Detector to verify they appear realistic, thereby selecting only the high-utility images that contribute to improving downstream performance. Across the evaluated datasets and classifiers, the best SIF-filtered setup improves accuracy over traditional augmentation in all cases, with gains of +1.73% and +2.35% on PhotoMOCI and +2.38% and +2.08% on KOCD for ResNet50 and ViT, respectively. Our findings reveal that while the direct application of generative data augmentation may yield performance drops, the integration of SIF, considering (i) how synthetic data looks real and (ii) how it reflects the discriminative features of the belonging class, provides a simple yet effective mechanism to filter out synthetic samples that confuse the classifier during training.
Figures & tables
Name
Year
#Imgs (w/ lesions)
Downstream tasks
Kaggle OC Lips and Tongue Shivam and Prakrut (2020)
Table 1 : Existing photographic oral cancer datasets for computer vision. The column #Imgs (w/ lesions) reports total images and lesion‑positive ones (within round brackets).
Figure 1 : PhotoMOCI metadata distribution by sex in (a) and age group in (b), stratified by pathology class.
Figure 2 : Overview of the proposed methodology for oral lesion classification. In the ochre dotted box, the Synthetic Image Generation based on generative models. In the green dotted box, the Synthetic Image Filter combining the Synthetic Proxy Classifier and the Synthetic Image Detector to select informative synthetic samples. In the light-blue dotted boxes, three training setups: using the original small dataset (top), the augmented synthetic dataset (middle), and the augmented synthetic filtered dataset (bottom), the latter aiming to improve generalization.
Figure 3: Two examples of diffusion processes using SD(txt+img) with conditioning mechanisms based on text and image for neoplastic and traumatic lesions in (a) and (b) from PhotoMOCI, respectively. In each example, ten denoising processes with image conditioning performed at different steps. The red -bordered images indicate where the conditioning image is introduced. In the last row, with water-green borders, the original conditioning image injected at the 100th step, overriding earlier ones.
Synthetic img. gen. setup
FID ↓
KID ↓
Prec. ↑
Recall ↑
Cov. ↑
Tr. aug.
230.30
0.126
0.317
0.183
0.058
SD(txt)
223.92
0.079
0.149
0.660
0.191
SD(txt+img)
179.24
0.072
0.543
0.319
0.521
SG3(lbl)
195.77
0.099
0.404
0.234
0.287
AC-SG3
290.42
0.230
0.426
0.011
0.032
Table 2 : Synthetic Image Generation results on PhotoMOCI.
Synthetic img. gen. setup
FID ↓
KID ↓
Prec. ↑
Recall ↑
Cov. ↑
Tr. aug.
262.91
0.126
0.284
0.632
0.242
SD(txt)
225.38
0.074
0.273
0.762
0.231
SD(txt+img)
143.87
0.034
0.867
0.573
0.888
SG3(lbl)
169.57
0.062
0.888
0.112
0.748
AC-SG3
319.78
0.245
0.035
0.035
0.042
Table 3 : Synthetic Image Generation results on KOCD.
Figure 4 : Synthetic image examples generated from the PhotoMOCI dataset using four synthetic image generation setups: Stable Diffusion conditioned on text (top-left), Stable Diffusion conditioned on text and image (top-right), StyleGAN3 conditioned on text (bottom-left), and AC-SG3 (bottom-right). Within each block, the first, second, and third rows correspond to traumatic, aphthous, and neoplastic lesions, respectively.
Figure 5 : Synthetic image examples generated from the KOCD using four synthetic image generation setups: Stable Diffusion conditioned on text (top-left), Stable Diffusion conditioned on text and image (top-right), StyleGAN3 conditioned on text (bottom-left), and AC-SG3 (bottom-right). Within each block, the first row corresponds to the cancer class, and the second row corresponds to the non-cancer class.
Synthetic img. gen. setup
Single image time (s)
PhotoMOCI 2100 images
KOCD 3000 images
SD(txt)
1.86 ± 0.05
1 h 05 min 06 s
1 h 33 min 00 s
SD(txt+img)
0.82 ± 0.06
28 min 42 s
41 min 00 s
SG3(lbl)
0.045 ± 0.003
1 min 35 s
2 min 15 s
AC-SG3
0.048 ± 0.004
1 min 41 s
2 min 24 s
Table 4 : Computational time for generating synthetic images on an NVIDIA A100 GPU.
Dataset
Synthetic img. gen. setup
Accuracy
SPC ↑
SID ↓
PhotoMOCI (ours)
SD(txt)
0.368 ± 0.045
0.788 ± 0.032
SD(txt+img)
0.771 ± 0.037
0.643 ± 0.040
SG3(lbl)
0.740 ± 0.029
0.596 ± 0.038
AC-SG3
0.893 ± 0.036
0.921 ± 0.036
KOCD Rashid (2024)
SD(txt)
0.617 ± 0.025
0.709 ± 0.030
Table 5 : SIF modules performance: classification results of SPC and SID on PhotoMOCI and KOCD considering the four synthetic data generation.
Synthetic img. gen. setup
PhotoMOCI
KOCD
Total
Traumatic
Aphthous
Neoplastic
Total
Cancer
Non-cancer
SD(txt)
139 (6.6%)
30 (4.3%)
83 (11.9%)
26 (3.7%)
308 (10.3%)
91 (6.1%)
217 (14.5%)
SD(txt+img)
470 (22.4%)
157 (22.4%)
119 (17.0%)
194 (27.7%)
653 (21.8%)
362 (24.1%)
291 (19.4%)
SG3(lbl)
367 (17.5%)
141 (20.1%)
137 (19.6%)
89 (12.7%)
776 (25.9%)
415 (27.7%)
361 (24.1%)
AC-SG3
70 (3.3%)
12 (1.7%)
19 (2.7%)
39 (5.6%)
152 (5.1%)
24 (1.6%)
128 (8.5%)
Table 6 : Number and percentage of synthetic images retained by SIF on PhotoMOCI and KOCD, reported in aggregate and by class.
Model
Conditioning
SIF
Acc
Prec
Recall
F1-score
ResNet He et al. (2016)
No aug.
✗
0.81050 ± 0.00532
0.81538 ± 0.00553
0.81076 ± 0.00581
0.81414 ± 0.00601
Tr.aug.(base)
✗
0.82519 ± 0.00600
0.83130 ± 0.00572
0.82316 ± 0.00603
0.82873 ± 0.00612
SD(txt)
0.74153 ± 0.00590
0.74797 ± 0.00619
0.74225 ± 0.00528
0.74550 ± 0.00493
✓
0.80076 ± 0.00544
0.82653 ± 0.00485
0.80567 ± 0.00521
0.81237 ± 0.00495
SD(txt+img)
0.83655 ± 0.00520
0.83704 ± 0.00446
0.83679 ± 0.00499
0.83698 ± 0.00524
✓
0.84253 ± 0.00498
0.85137 ± 0.00507
0.84294 ± 0.00481
0.85095 ± 0.00511
Table 7 : Classification performance on PhotoMOCI using ResNet and ViT under different data augmentation strategies. The evaluated setups include training without augmentation (No aug.) and with traditional augmentation (Tr. aug.), marked with ✗ since SIF is not applicable, as well as four generative-AI–based setups evaluated both without and with SIF, distinguished by ✓. Highlighted in gray , the setups where SIF improves performance over non-SIF counterparts and surpasses traditional augmentation. Highlighted in bold , the best-performing augmentation setup for each classifier.
Model
Conditioning
SIF
Acc
Prec
Recall
F1-score
ResNet He et al. (2016)
No aug.
✗
0.90935 ± 0.00453
0.90028 ± 0.00452
0.90074 ± 0.00449
0.90026 ± 0.00454
Tr.aug.(base)
✗
0.93109 ± 0.00554
0.93103 ± 0.00547
0.93120 ± 0.00563
0.93095 ± 0.00576
SD(txt)
0.88564 ± 0.00637
0.89160 ± 0.00656
0.88713 ± 0.00625
0.89121 ± 0.00700
✓
0.91616 ± 0.00410
0.91937 ± 0.00474
0.91668 ± 0.00422
0.91871 ± 0.00390
SD(txt+img)
0.95011 ± 0.00498
0.95225 ± 0.00501
0.94917 ± 0.00533
0.95075 ± 0.00511
✓
0.95492 ± 0.00467
0.95719 ± 0.00501
0.95266 ± 0.00461
0.95543 ± 0.00521
Table 8 : Classification performance on KOCD using ResNet and ViT under different data augmentation strategies. The evaluated setups include training without augmentation (No aug.) and with traditional augmentation (Tr. aug.), marked with ✗ since SIF is not applicable, as well as four generative-AI–based setups evaluated both without and with SIF, distinguished by ✓. Highlighted in gray , the setups where SIF improves performance over non-SIF counterparts and surpasses traditional augmentation. Highlighted in bold , the best-performing augmentation setup for each classifier.
Figure 6 : Confusion matrices for PhotoMOCI comparing traditional augmentation and SD(txt+img)+SIF for ResNet50 and ViT.
Dataset
Model
χ2
p
Kendall’s W
KOCD
ResNet
54.938
<0.0001
0.610
KOCD
ViT
53.520
<0.0001
0.595
PhotoMOCI
ResNet
56.749
<0.0001
0.631
PhotoMOCI
ViT
50.787
<0.0001
0.564
Table 9 : Overall statistical significance analysis using Friedman tests on accuracy.
KOCD
PhotoMOCI
Synthetic img. gen. setup
ResNet
ViT
ResNet
ViT
Δ Acc.
p
Δ Acc.
p
Δ Acc.
p
Δ Acc.
p
SD(txt)+SIF
-1.49
0.0645
-0.59
0.3750
-2.44
0.0371
-0.84
0.4316
SD(txt+img)+SIF
+2.38
0.0371
+2.08
0.0059
+1.73
0.0273
+2.35
0.0195
SG3(lbl)+SIF
+1.72
0.0488
+1.17
0.0934
+1.08
0.0754
+1.79
0.0515
AC-SG3+SIF
+0.44
0.6953
+0.21
0.6250
-1.53
0.0195
-0.17
0.6250
Table 10 : Planned Wilcoxon comparisons between synthetic+SIF setups and traditional augmentation on accuracy. Values report accuracy difference in percentage points and p -value.
KOCD
PhotoMOCI
Synthetic img. gen. setup
ResNet
ViT
ResNet
ViT
Δ Acc.
p
Δ Acc.
p
Δ Acc.
p
Δ Acc.
p
SD(txt)
+3.05
0.0039
+3.58
0.0020
+5.92
0.0020
+6.31
0.0020
SD(txt+img)
+0.48
0.0566
+0.71
0.0316
+0.60
0.0566
+0.80
0.0316
SG3(lbl)
+0.81
0.0754
+0.63
0.0750
+0.32
0.0250
+0.57
0.0695
AC-SG3
+0.74
0.0602
+0.91
0.0223
+1.84
0.0195
+1.33
0.0324
Table 11 : Planned Wilcoxon comparisons between SIF-filtered and unfiltered synthetic augmentation on accuracy. Values report accuracy difference in percentage points and p -value.
Figure 7 : Ablation study evaluating the impact of SPC and SID within the SIF design on (a) PhotoMOCI and (b) KOCD using ViT. The study compares training without augmentation (blue bar) to training with integrated synthetic data: unfiltered (ochre bar), filtered using SPC only (water-green bar), SID only (orange bar), and SIF (pink bar).
Figure 8 : SID threshold calibration on (a) PhotoMOCI and (b) KOCD. Each plot reports accuracy as a function of the SID threshold τ for ResNet50 (left) and ViT (right), with shaded 95% confidence intervals over 10 runs and a marker indicating the best threshold.
Generative Adversarial Networks (GANs) can help overcome data scarcity in computer vision tasks by generating additional training samples. In this work, we explore generative data augmentation in two low-resource domains: Bangla handwritten character recognition and chest X-ray image analysis. We use DCGAN-based models trained on 64x64 images to generate synthetic samples and evaluate their quality using Inception Score (IS), Fréchet Inception Distance (FID), and visualization methods such as t-SNE and UMAP. To measure practical usefulness, we train image classifiers using real data and a combination of real and synthetic data. Experimental results show that synthetic augmentation improves data diversity and consistently increases classification performance in limited-data settings. We also investigate training stability techniques, including gradient penalty and spectral normalization, and perform ablation studies on synthetic-to-real data ratios and sample filtering strategies. In addition, we discuss challenges related to medical image evaluation, dataset licensing, and privacy concerns of synthetic data. Our approach is simple, reproducible, and provides a strong baseline for generative augmentation in resource-constrained imaging applications.
Md. Sohanuzzaman Soad, Mahady Al Hady, S M Rafiuddin Rifat +1
Department of Computer Science and Engineering University of Asia Pacific Dhaka, Bangladesh
Generative augmentation is often proposed as a remedy for small medical-image datasets, but synthetic images are only useful when they improve downstream task performance. "Augmentation" here means synthetic supplementation: GAN-generated samples added to the real training pool, not geometric or photometric transforms of existing images. Twelve class-plane StyleGAN2-ADA generators were trained on constrained BRISC 2025 partitions to test whether their output, with or without InceptionV3 feature-space filtering, improves held-out tumour classification across three classifier families: a random forest (RF) on InceptionV3 features, a compact two-headed convolutional neural network (CNN), and MobileViTV2, a mobile hybrid convolutional-transformer. Each was evaluated at 1:1 and 1:2 real-to-synthetic ratios. An independent GPT-5.5 blind test placed gated real-versus-synthetic discrimination at 57.73% (95% CI: 54.48--60.92%) on the model-legible subset -- modestly above chance. The RF classifier did not benefit from the synthetic MRIs. The CNN showed consistent mean gains that did not survive Holm correction. MobileViTV2 showed the clearest benefit: filtered 1:1 augmentation improved tumour classification accuracy by 1.02% absolute (95% CI: 0.54--1.54%; Holm-corrected p = 0.0104). A secondary efficiency analysis found that every augmented CNN condition selected its checkpoint 42--64% earlier than baseline, while compute-matched MobileViTV2 runs reached selection after 50--67% fewer real-data epochs. Overall, augmentation utility was found to be architecture- and ratio-dependent, not guaranteed by visual fidelity alone.
Background/Objectives: Dermoscopic skin-lesion classifiers lose accuracy when images arrive from a new clinic or a new device. We asked which data augmentations reduce that loss, and measured the effect under a protocol that keeps policy selection separate from policy evaluation. Methods: A ConvNeXt-Large binary malignant-versus-non-malignant classifier was trained on six dermoscopic sources (25,903 images); HAM10000 and ISIC 2016-2020 were held out of training entirely. Single augmentations, photometric combinations and eleven composite policies were ranked on a development split of 1511 held-out images. The winning policy was then evaluated on a confirmation set of 8073 held-out images that took no part in that ranking and from which we removed every image sharing a lesion identifier with the training data and every image contributed by an institution represented in training. Both policies were retrained with four random seeds each and compared with an exact permutation test. Results: The mix policy raised confirmation-set ROC-AUC from 0.787 to 0.826 (+0.039; per-seed ranges 0.772-0.797 and 0.815-0.840, non-overlapping; exact permutation p=0.029), with the same direction on each contributing source. At matched sensitivity the gain is larger in clinical terms: specificity rose from 0.612 to 0.713 at a sensitivity of 0.80, and from 0.284 to 0.397 at a sensitivity of 0.95. In-domain ROC-AUC was preserved (0.938 to 0.941). On an independent clinical cohort acquired with a different device at a different institution (472 images, 22 malignant), performance was maintained (0.934 versus 0.930). Conclusions: Augmentations that model the physical causes of domain shift improve cross-source transfer at no cost to in-domain accuracy, and the improvement survives a selection-disjoint, contamination-free evaluation.
Alexander Kozachok, Ilya Latyshev, Evgeny Karpulevich +3
Trusted AI Research Center, Russian Academy of Sciences, 109004 Moscow, Russia