Adaptive Adversarial Augmentation for Controllable Face Synthesis
Authors: Saransh Suri, Shivang Agarwal, Mayank Vatsa, Richa Singh
Organizations: Manipal University Jaipur, India · Indian Institute of Technology Jodhpur, India · Birla Institute of Technology and Science Pilani Dubai Campus, Dubai, United Arab Emirates
Synthetic data provides a scalable alternative to real-world datasets for training face recognition models, particularly under challenging conditions such as low resolution, occlusion, and masks. Yet, most approaches lack diversity and fail to generalize effectively. We propose Ensemble Feedback Controllable Synthesis (EFCS), a guided framework that generates diverse and challenging samples while preserving visual realism. EFCS expands distributional variability, often reflected in higher FID and KID scores compared to single-feedback and random synthesis, while maintaining high precision. Recognition models trained on EFCS data consistently outperform baselines across multiple benchmarks, showing improved generalization to real-world scenarios. Furthermore, we introduce an analytically motivated formulation linking perturbation-induced difficulty, sample utility, and performance degradation, offering principled insights into balancing synthetic data complexity for optimal training. Together, these contributions establish EFCS as an effective and analytically grounded approach for bridging the gap between synthetic and real datasets.
Figures & tables
Figure 1: Illustration highlighting the necessity of adaptive synthesis. Controlled perturbations (hard samples) introduce targeted complexity to synthetic data, enhancing model robustness beyond traditional synthesis methods.
Figure 2: Qualitative comparison of synthetic samples: (I) constant degradation, (II) single feedback synthesis [ 17 ] , and (III) proposed EFCS method.
Figure 3: Overview of the proposed EFCS framework showing synthesis, perturbation, and ensemble feedback.
Dataset
Identities
Images/ID
Total Images
CASIA-WebFace (original)
10,575
∼ 47 avg
494,414
Real-only subset
10,575
20
211,500
EFCS Synthetic
10,575
20
211,500
Mixed (Real + EFCS)
10,575
40
423,000
Table 1: Dataset statistics used for EFCS training. Synthetic images are generated per identity using the EFCS synthesis model.
Sample Type
Avg. SER-FIQ
Original
0.7776
Synthesized
0.7338
Single Feedback [ 17 ]
0.7264
EFCS (ours)
0.7277
Table 2: Average SER-FIQ comparison across synthesis strategies (higher is better).
Table 5: Core evaluation on custom test sets (R1/R5 and TAR@{1e-5,1e-4,1e-3}, %).
Figure 4: Effect of synthesis noise and difficulty on training utility and recognition performance. Results averaged across ArcFace, AdaFace, and ElasticFace on CASIA-WebFace. Moderate perturbation improves utility, while excessive noise causes degradation.
Model
Training strategy
M1 (EFCS)
Scratch on EFCS
M2 (Mix-50)
Scratch on EFCS+Real (50%)
M3 (Mix-33)
Scratch on EFCS+Real (33%)
M4 (Mix-25)
Scratch on EFCS+Real (25%)
M5 (Real)
Scratch on Real
M6 (Pretr.+EFCS)
Pretrained → finetune EFCS
Table 6: Training settings of evaluated models.
LFW
CFP-FP
AgeDB
CALFW
CPLFW
CFP-FF
Model
Acc
1e-4
1e-3
Acc
1e-4
1e-3
Acc
1e-4
1e-3
Acc
1e-4
1e-3
Acc
1e-4
1e-3
Acc
1e-4
1e-3
M1
98.03
29.81
89.52
87.10
0.07
0.73
87.95
5.76
27.22
89.90
14.90
50.80
81.97
22.50
26.20
97.49
1.90
19.02
M2
98.68
29.15
92.33
90.57
0.18
1.80
90.70
9.56
33.13
91.28
6.02
51.70
84.65
5.65
37.70
98.51
3.25
32.53
M3
98.65
27.98
90.78
90.27
0.23
2.30
90.73
10.22
38.37
91.40
14.34
57.87
84.55
7.64
31.30
98.41
3.30
33.28
M4
98.83
30.39
94.97
90.21
0.52
5.20
91.00
6.04
43.03
92.03
15.51
54.53
84.20
6.32
29.18
98.29
3.40
33.95
M5
98.88
30.72
94.60
91.33
0.38
3.80
90.98
13.70
48.63
92.08
15.60
62.00
85.03
6.35
39.10
98.41
3.60
36.02
Table 7: Verification performance (%) on six public benchmarks. Metrics: accuracy (Acc) and TAR@FAR.
The shortage of legally compliant data for face recognition training has sparked growing interest in using synthetic data as an alternative. While recent diffusion-based methods enable the generation of photorealistic face images with strong identity adherence and data diversity, their downstream recognition performance still exhibits a significant synthetic-real gap. This paper identifies visual tendency as a previously underexplored limitation, whereby synthetic data exhibit an unrealistic prevalence of visual attributes and thus deviate from the real-data distribution. Visual tendency can be attributed to the generator's conditioning on identity embeddings, through which co-occurring residual visual cues are unintentionally absorbed into learned identity semantics. To discourage the generator from exploiting such visual cues, this paper proposes SteerFace, a simple and efficient training framework that perturbs identity embeddings by steering them toward random orthogonal directions on the embedding hypersphere. The perturbation serves as an identity-preserving regularizer that penalizes the generator's reliance on non-identity components, as supported by theoretical analysis. This paper further introduces an adaptive strategy that learns perturbation strengths with both sample-wise preference and favorable overall statistics. Extensive experiments show that SteerFace effectively mitigates visual tendency, outperforms prior methods in downstream face recognition, and generalizes well across different training datasets and generation pipelines.
Yuxi Mi, Qiuyang Yuan, Jianqing Xu +5
Fudan University Shanghai, China · Youtu Lab, Tencent Shanghai, China · WeChat Pay Lab33, Tencent Shenzhen, China
Synthetic data generation is increasingly used in machine learning for training and data augmentation. Yet, current strategies often rely on external foundation models or datasets, whose usage is restricted in many scenarios due to policy or legal constraints. We propose ScoreMix, a self-contained synthetic generation method to produce hard synthetic samples for recognition tasks by leveraging the score compositionality of diffusion models. The approach mixes class-conditioned scores along reverse diffusion trajectories, yielding domain-specific data augmentation without external resources. We systematically study class-selection strategies and find that mixing classes distant in the discriminator's embedding space yields larger gains, providing up to 3% additional average improvement, compared to selection based on proximity. Interestingly, we observe that condition and embedding spaces are largely uncorrelated under standard alignment metrics, and the generator's condition space has a negligible effect on downstream performance. Across 8 public face recognition benchmarks, ScoreMix improves accuracy by up to 7 percentage points, without hyperparameter search, highlighting both robustness and practicality. Our method provides a simple yet effective way to maximize discriminator performance using only the available dataset, without reliance on third-party resources. Paper website: https://parsa-ra.github.io/scoremix/.
Face Recognition (FR) systems in surveillance settings often encounter Low Resolution (LR) faces, those whose face region falls below the standard 112 × 112 input size. While labelled High Resolution (HR) training data is abundant, labelled native-LR data, and above all paired native LR/HR data, is scarce. One workaround is to synthesize LR data from the available HR faces, but how much synthesis effort is repaid in recognition accuracy remains unclear. We present a study of simple synthetic generation strategies for a compact, edge device-oriented face recognition system, spanning interpolation-based degradation, knowledge distillation, a Prepended Domain Transformer (PDT), Real ESRGAN-style degradation, and a learned Super Resolution (SR) front-end with an identity-aware loss. We evaluate these strategies on synthetic cross-resolution face benchmarks (LFW, CFP-FP, AgeDB-30) and on TinyFace, a real-world native LR dataset, and expose a synthetic-real gap: the degradation setting that is optimal on synthetic benchmarks is not the one that is optimal on real LR. We find that more synthesis effort does not help monotonically: the learned SR front-end does not surpass a direct feed of the aligned LR image into a strong backbone, while simple interpolation augmentation of a compact backbone is the only synthesis that improves over its own baseline. We conclude that generative methods for LR face recognition must be validated on real LR and against a direct-feed baseline, and release our pipeline at https://idiap.ch/paper/synth-lrfr
Luis S. Luevano, Ünsal Öztürk, Hatef Otroshi Shahreza +2