cs.CVSep 28, 2026

Generative AI-Based Data Augmentation for Oral Lesion Classification: The PhotoMOCI Dataset and Benchmark

Authors: Marco Parola, Mario G. C. A. Cimino, Sabrina Senatore

Organizations: Aalborg University, Denmark. · University of Pisa, Italy. · University of Salerno, Italy.

Abstract

Early detection of oral cancer via photographic imaging presents a promising avenue for large-scale oral cavity screening. However, the development of robust deep learning models is frequently hampered by the scarcity of high-quality, annotated datasets. To address this limitation, a novel and well-curated resource, the Photographic Multi-purpose Oral Cancer Imaging (PhotoMOCI) dataset, is introduced for developing models across multiple diagnostic tasks in oral oncology. Then, a comprehensive benchmark study was conducted to investigate how various data augmentation strategies influence the performance of image classifiers. Our analysis spans different generative AI frameworks, evaluating the efficacy of traditional methods against advanced generative approaches, including Generative Adversarial Networks (GANs) and Diffusion Models (DMs). Additionally, we propose the Synthetic Image Filter (SIF), a mechanism to select specific samples based on two auxiliary models: Synthetic Proxy Classifier to ensure samples are representative of the target class and Synthetic Image Detector to verify they appear realistic, thereby selecting only the high-utility images that contribute to improving downstream performance. Across the evaluated datasets and classifiers, the best SIF-filtered setup improves accuracy over traditional augmentation in all cases, with gains of +1.73% and +2.35% on PhotoMOCI and +2.38% and +2.08% on KOCD for ResNet50 and ViT, respectively. Our findings reveal that while the direct application of generative data augmentation may yield performance drops, the integration of SIF, considering (i) how synthetic data looks real and (ii) how it reflects the discriminative features of the belonging class, provides a simple yet effective mechanism to filter out synthetic samples that confuse the classifier during training.

Figures & tables

Explore similar work

CardsList
  1. Cross-Domain Adversarial Augmentation: Stabilizing GANs for Medical and Handwriting Data Scarcity

    May 3, 2026Md. Sohanuzzaman Soad, Mahady Al Hady, S M Rafiuddin Rifat +1Data AugmentationGenerative Adversarial Network

  2. Searching for Robust Augmentations to Improve Out-of-Domain Generalization in Dermoscopic Skin Cancer Classification

    Jul 29, 2026Alexander Kozachok, Ilya Latyshev, Evgeny Karpulevich +3Skin Lesion ClassificationData Augmentation