Visual Anomaly Synthesis for Model Selection in Data Scarcity
Authors: Daniel Pröll, Thomas Kraxner, Tobias Schaefer, Sebastian Hegenbart
Organizations: Digital Factory Vorarlberg GmbH, Dornbirn, Austria · University of Salzburg, Salzburg, Austria · illwerke vkw AG, Bregenz, Austria · Vorarlberg University of Applied Sciences, Dornbirn, Austria
Defect detection systems for industrial condition monitoring can only be relied upon if they are validated, yet defective samples are rare and, for a specific asset, often nonexistent. We present a framework that synthesizes severity-graded defects on real non-defective images without any defect references for the target asset, that can be used for model selection and validation. A defect taxonomy for common failure modes is distilled from literature into prescriptive prompts at varying defect severities. Regions of interest are cropped from in defect-free images and edited with a pre-trained image generation model ("FLUX.2 [klein]"). Color-matching and blending are employed to improve structural coherence with the original image. Generations are filtered out by a scorer and by estimated detection difficulty. Model selection experiments on MVTecAD show image AUROC choice regret over model selection can be nearly halved compared to the best fixed model chosen with access to test data. Experiments show the need for severity-graded anomaly synthesis. A case study investigates the proposed method for in-situ monitoring of Pelton turbine runners in hydropower, where real defect images are rare and expensive to collect. A PatchCorebased anomaly detection model is fit on Pelton turbine images and selected and validated using synthetic images, showing strong detection performance (94 % correct detection at optimal threshold and AUROC 0.97). The model reliably detects moderate and advanced defects, while early-stage defects remain challenging, indicating the synthetic data meaningfully stresses detector sensitivity.
Figures & tables
Figure 1: Overview of the three steps of the proposed method: prompt and taxonomy creation (once per object), image editing and post-processing (once per image) and downstream use in evaluation and model selection. Pre-trained models shown in green.
Figure 2: Examples of generated images with increasing severity for two MVTec AD categories and the Pelton use-case.
I-AUC
Fix
Ours
O
1
0.918
0.927
0.952
2
0.926
0.945
0.961
4
0.938
0.951
0.965
8
0.956
0.969
0.977
mean
0.934
0.948
0.964
PRO
Fix
Ours
O
Table 1: Left: real anomaly detection scores of the model each rule selected, by training-set size n , for the best single model (Fix), from synthetic data (Ours) and by the oracle (O). Mean over categories and seeds; higher is better. Right: image AUROC choice regret, on a reduced model set, of the model selected with each generator’s synthetic data, for the best single model (Fix), our pipeline (Ours), DRAEM (DRA) and MIRAGE (MIR). Mean over shots and seeds; lower is better. Bold = best within each side; oracle not ranked.
Set
med.
mean
max
Fix
0.026
0.029
0.105
1
0.014
0.016
0.043
2
0.019
0.025
0.084
3
0.020
0.024
0.069
4
0.018
0.025
0.095
All
0.019
0.017
0.050
Table 2: Left: MVTec AD I-AUC choice regret for model selection based on different subsets, for the best fixed model (Fix) and synthetic anomalies of different severities from least (1) to most severe (4) and all severities (all). Median, mean and max over the per-category means (over shots and seeds); lower is better, bold = best in each column. Right: I-AUC of the synthetic Pelton defects of each mode (abr.: abrasive erosion, cav.: cavitation, crack: fatigue crack, imp.: stone impact) and severity, with synthetic negatives (defects vs. real crops + synthetic defect-free images), mean over four models, higher is better. Below: AUROC of synthetic negatives against real negatives. Near 0.5 is best.
Synthetic anomaly generation helps expand industrial anomaly datasets when real defects are scarce or unavailable. Existing approaches lie at two extremes: procedural approaches are fast but struggle to represent complex anomalies, while generative approaches produce diverse defects but require costly per-sample generation. We present FLASH, a framework that decouples defect generation from anomaly synthesis under a ``generate once, synthesize many'' paradigm. Given only normal images, FLASH uses Vision-Language Model (VLM) guidance and an image-generation model to produce a small set of defect images, from which it extracts, validates, and banks reusable defect patches. For synthesis of anomalous images, Object Boundary Suppression (OBS) first identifies the probable foreground object-aware region of the host image, while Multi-Resolution Spectral Pyramid (MRSP) noise generates diverse, size-controllable masks that determine the defect location and spatial extent. It then composes a large and diverse synthetic anomalous image set by localizing the defect region, sampling size-controllable placement masks and seamlessly blending retrieved defects onto new defect-free images without further need for image generation. Experiments on the MVTec AD 2 dataset show that FLASH-generated anomalies nearly close the calibration gap on real defects, reaching 78.1% image-level F1 against an 83.6% real-anomaly upper bound and providing the most consistent calibration transfer across detectors among procedural and generative alternatives. Moreover, FLASH synthesizes anomalies more than 11.95x faster than per-sample generative approaches.
The bottleneck in learning-based industrial defect detection is often limited not by model capacity, but by the scarcity of labeled defect data: defects are rare, annotations are expensive, and collecting balanced training sets is slow. We present an end-to-end pipeline for synthetic defect generation and annotation, combining Vision-Language-Model-based prompts, LoRA-adapted diffusion, mask-guided inpainting, and sample filtering with automatic label derivation, and demonstrates the potential of real data with realistic synthetic samples to overcome data scarcity. The evaluation is conducted on, a challenging dataset of pitting defects on ball screw drives, and then on a subset of the Mobile phone screen surface defect segmentation dataset (MSD) dataset to test cross-domain transfer. Beyond downstream detector performance, we analyze key stages of the pipeline, including prompt construction, LoRA selection, and sample filtering with DreamSim and CLIPScore, to understand which synthetic samples are both realistic and useful. Experiments with YOLOv26, YOLOX, and LW-DETR show that synthetic-only training does not replace real data. When combined with real data, synthetic defects can preserve performance and yield modest gains in selected BSData training regimes. The MSD transfer study shows that the overall pipeline structure carries over to a second industrial inspection domain, while also highlighting the importance of domain-specific adaptation and annotation-quality control. Overall, the paper provides an end-to-end assessment of diffusion-based industrial defect synthesis and shows that its strongest value lies in strengthening scarce real datasets rather than substituting for them.
Paul Julius Kühn, Mika Pommeranz, Arjan Kuijper +1
Industrial anomaly detection and localization are limited by scarce real anomalies and pixel-level annotations, a bottleneck that synthetic image-mask pairs can alleviate. However, existing few-shot mask-guided generation may over-follow mask geometry, produce weak anomalies, or use condition masks incompatible with the current object instance. We propose OSAGEN, which combines object-aware mask priors with multistage decoupled diffusion. Its three-stage adaptation sequentially learns normal appearance, defect appearance under coarse conditions, and fine-grained mask calibration, improving defect realization and local control. QBG injects object structure from a matched normal image into mask diffusion to produce object-aware priors, while ISC restricts anomaly propagation and preserves normal content during sampling. A lightweight materialization step recovers pixel-level labels aligned with the realized defects. On MVTec AD and VisA, OSAGEN achieves AP-P/F1-P scores of 88.1/82.2 and 68.5/66.1, respectively, under a unified downstream localization protocol. The code will be released upon acceptance.
Jinyi Xu, Peng Chen, Yunkang Cao +3
School of Cyber Science and Technology, Shenzhen Campus of Sun Yat-sen University · School of Artificial Intelligence and Robotics, Hunan University · Department of Computer and Information Science, University of Macau +1