Synthetic anomaly generation helps expand industrial anomaly datasets when real defects are scarce or unavailable. Existing approaches lie at two extremes: procedural approaches are fast but struggle to represent complex anomalies, while generative approaches produce diverse defects but require costly per-sample generation. We present FLASH, a framework that decouples defect generation from anomaly synthesis under a ``generate once, synthesize many'' paradigm. Given only normal images, FLASH uses Vision-Language Model (VLM) guidance and an image-generation model to produce a small set of defect images, from which it extracts, validates, and banks reusable defect patches. For synthesis of anomalous images, Object Boundary Suppression (OBS) first identifies the probable foreground object-aware region of the host image, while Multi-Resolution Spectral Pyramid (MRSP) noise generates diverse, size-controllable masks that determine the defect location and spatial extent. It then composes a large and diverse synthetic anomalous image set by localizing the defect region, sampling size-controllable placement masks and seamlessly blending retrieved defects onto new defect-free images without further need for image generation. Experiments on the MVTec AD 2 dataset show that FLASH-generated anomalies nearly close the calibration gap on real defects, reaching 78.1% image-level F1 against an 83.6% real-anomaly upper bound and providing the most consistent calibration transfer across detectors among procedural and generative alternatives. Moreover, FLASH synthesizes anomalies more than 11.95x faster than per-sample generative approaches.
Figures & tables
Figure 1 : Five-stage FLASH architecture. Stages 1–3 generate, extract, validate, and store reusable defects, while Stages 4–5 localize and synthesize retrieved defects onto separate normal host images.
Figure 2 : DiffMask-based defect extraction from registered normal and generated anomaly images.
Figure 3 : Qualitative results across all eight MVTec AD 2 categories, illustrating the end-to-end FLASH pipeline.
Model
Real
Perlin
AnoStyler
FLASH (Ours)
PaDiM
80.37
69.41
62.37
79.11
PatchCore
82.46
67.50
43.02
74.64
AnomalyDINO
81.47
62.19
49.85
77.05
Dinomaly
81.81
60.42
55.37
77.51
SuperADD
83.64
79.12
64.36
78.13
Table 1 : (a) Image-level F1 and (b) pixel-level F1 scores (%) averaged across categories of the MVTec AD 2.
Category
Real
Perlin
AnoStyler
FLASH (Ours)
Can
0.02
0.01
0.00
0.00
Fabric
78.36
18.57
44.30
69.06
Fruit Jelly
56.07
40.02
55.93
55.57
Rice
58.75
6.21
25.36
51.90
Sheet Metal
35.87
2.53
7.79
7.39
Vial
57.05
56.46
46.63
39.41
Table 2 : Per-category pixel-level F1 scores (%) using SuperADD.
Method
Ia Gen. (s/image)
Isyn Gen. (ms/image)
Total (s/category)
Perlin
–
14.4
1.3
AnoStyler
–
14,982.1
1,348.1
FLASH (Ours)
15
586.2
112.8
Table 3 : Estimated generation time for Ia and Isyn . Total (s/category) indicates the total time taken to generate the synthetic images for a category of MVTec AD 2.
Hyperparameter
Value
Purpose
Synthesis resolution
1024 px
Provides sufficient spatial support for localized defects and blending.
DiffMask context expansion
1.8
Retains surrounding substrate context around the recovered defect.
Minimum defect region
500 px
Suppresses small reconstruction artifacts during defect extraction.
Defect contrast threshold
ΔE=6.0
Filters weak defect evidence using perceptual color difference.
MRSP noise scale
14
Controls the characteristic spatial scale of the placement field.
MRSP octaves
6
Provides multi-resolution spatial structure for defect placement.
Table 4: Principal hyperparameters governing defect extraction, object-aware placement, adaptive defect sizing, and hybrid compositing. The configuration is fixed across categories after development.
Component
Configuration
Anomalib
2.5.2
Python
3.13
PyTorch
2.13.0 (CUDA 13.0)
PyTorch Lightning
2.6.5
TorchMetrics
1.9.0
Torchvision
0.28.0
Table 5 : Software and hardware environment used for all experiments.
Figure 4 : Joint comparison of noise primitives in terms of generation time, EMD, and calibrated F1 . (a) Calibrated F1 on the real test set versus EMD to the real anomaly-score distribution; lower EMD and higher F1 are preferred, with the horizontal axis reversed accordingly. (b) Noise-generation time on a logarithmic scale (measured on images at 224×224 resolution), where lower values are preferred. Marker shapes distinguish single-octave and six-octave fBm configurations, while MRSP is highlighted as the proposed primitive. MRSP provides a favorable trade-off among computational cost, similarity to real anomaly-score distributions, and calibration transfer.
Configuration
Value
Purpose
Random seeds
{0,1,2}
Provides repeated paired evaluations under controlled stochastic variation.
VLM
Qwen2.5-VL-7B-Instruct
Used for semantic anomaly specification and defect validation.
Synthetic anomalies
90 / category / seed
Fixed synthetic calibration budget.
SuperADD backbone
DINOv3 ViT-H/16+
Frozen representation used across all evaluation settings.
Patch-feature cap
8000
Fixed memory-bank capacity across evaluation settings.
Table 6: Fixed corpus and downstream evaluation configuration used in the reported experiments.
Figure 5 : MRSP ablation at the MVTec AD 2 operating point. Placement masks are evaluated across spectral exponent α and pyramid levels L at fixed coverage c=0.018 . The selected configuration, α=1.5 and L=6 , provides a coherent placement region while avoiding excessive fragmentation or over-smoothing.
Figure 6 : Construction of the object-aware placement mask Mf for Can, Fabric, Fruit Jelly, and Rice. Each row shows the MRSP field, its restriction to the object region Ω , thresholded response, retained largest connected component, and seed-dependent placement regions.
Figure 7 : Construction of the object-aware placement mask Mf for Sheet Metal, Vial, Wallplugs, and Walnuts. Each row shows the MRSP field, its restriction to the object region Ω , thresholded response, retained largest connected component, and seed-dependent placement regions.
Stage
Time (ms)
Share (%)
OBS foreground extraction
154
26.3
MRSP noise generation
12
2.0
Placement-mask computation
37
6.3
Defect placement
5
0.9
CIELAB color harmonization
139
23.7
Poisson blending
212
36.2
Table 7 : Per-image timing breakdown of FLASH synthesis at 1024 × 1024 resolution on a single NVIDIA RTX 3090. Stage times are means over fresh, non-cached images; the sub-stage sum (559 ms) is slightly below the measured total (586 ms) because stages overlap and cache.
Category
Real
Perlin
AnoStyler
FLASH (Ours)
PaDiM
Can
71.49
44.60
64.67
68.93
Fabric
73.89
73.17
58.37
73.17
Fruit Jelly
88.27
69.37
87.24
85.71
Rice
81.31
81.08
6.27
81.08
Sheet Metal
89.64
88.24
65.13
88.24
Table 8 : Per-category image-level F1 (%) across all five detectors on MVTec AD 2. The Mean row averages over the eight categories; bold marks the best synthetic source per detector (the Real oracle is excluded).
Category
Real
Perlin
AnoStyler
FLASH (Ours)
PaDiM
Can
0.16
0.05
0.05
0.05
Fabric
3.36
1.44
0.93
1.01
Fruit Jelly
12.41
4.46
5.13
2.96
Rice
6.26
2.03
6.21
1.74
Sheet Metal
11.95
3.70
11.27
9.08
Table 9 : Per-category pixel-level F1 (%) across all five detectors on MVTec AD 2. The Mean row averages over the eight categories; bold marks the best synthetic source per detector (the Real oracle is excluded).
Figure 8 : Visual comparison of synthetic anomaly generation across the eight MVTec AD 2 categories. Each row uses the same normal host image across methods, with columns showing the clean host followed by DRAEM, NSA, GLASS, AnoStyler, and FLASH (ours).
Figure 9 : Per-category synthesis results for Can, Fabric, Fruit Jelly, and Rice. Each row traces a synthesized instance through the host image, object region Ω , placement mask Mf , retrieved defect, ground-truth mask Mgt , the alpha and Poisson composites, their zoomed views, and the absolute difference of each composite against the host. Rows correspond to three seeds of the same category.
Figure 10 : Per-category synthesis results for Sheet Metal, Vial, Wallplugs, and Walnuts; columns as in Figure 9 . These categories carry the constrained supports: a thin specular strip, a transparent vial, and two multi-object scenes in which the placement mask selects which instance receives the defect.
School of Cyber Science and Technology, Shenzhen Campus of Sun Yat-sen University · School of Artificial Intelligence and Robotics, Hunan University · Department of Computer and Information Science, University of Macau +1