Organizations: Innopolis University, Innopolis, Russia · Laboratory of Innovative Technologies for Processing Video Content, Innopolis University, Innopolis, Russia
Fusing a learned camera-noise fingerprint with an RGB appearance backbone is an appealing route to generator-independent deepfake detection, because the noise residual is grounded in image-formation physics rather than in the texture statistics of a particular generator. We test, on FaceForensics++, whether a Noiseprint++ residual channel carries information \emph{complementary} to an RGB Xception backbone for face-swap detection. A three-model ablation (RGB-only, residual-only, late-fusion) shows that fusion does not improve over RGB alone and that the residual branch alone is near chance. A seven-level bottleneck diagnostic localizes the cause: the noise maps do carry a discriminative signal, but it is statistical---carried by the per-sample first and second moments (mean, variance, energy) of the residual---and the per-sample \texttt{InstanceNorm} layer placed at the noise-branch input, following the TruFor template, standardizes exactly those moments away (five-fold cross-validated AUC drops from 0.747 to 0.554). A context-crop control rules out cropping geometry, and two fixed-fusion variants that remove the bottleneck recover the statistical signal yet still fail to beat RGB on every dataset. We conclude that, on this manipulation distribution, the noise residual is redundant with RGB rather than complementary, and we give concrete guidance for practitioners adopting noise-residual fusion for face-swap detection.
Figures & tables
Model
FF++
Celeb-DF
DFDC
DFF
RGB-Only
0.928
0.888
0.717
0.531
Late-Fusion
0.924
0.881
0.712
0.475
Residual-Only
0.579
0.547
0.578
0.212
SBI (ref.)
0.865
0.907
0.884
0.614
TABLE I: In-distribution and cross-dataset video-level AUC (frame-level for DFF). RGB-only ≈ Late-Fusion (Wilcoxon p=0.977 ); Residual-only is near chance. SBI is an off-the-shelf reference, shown for context.
Classifier
Input
AUC
L1: LogReg
mean ∣ noise ∣
0.653±0.018
L2: LogReg
6 scalar stats
0.664±0.018
L3: MLP (ceiling)
6 scalar stats
0.747±0.023
L4: LogReg
InstanceNorm-pooled noise
0.554±0.014
L5: LogReg
raw-pooled noise
0.559±0.026
L6: ConvNet (2-layer)
raw noise
0.602±0.021
TABLE II: Seven-level bottleneck diagnostic on Noiseprint++ maps. Five-fold cross-validated AUC on 3,000 FF++ test samples. L3 sets the statistical ceiling ( 0.747 ); L4 localizes the InstanceNorm collapse ( 0.554 ).
Fig. 1: Seven-level bottleneck AUC (L1–L7) with five-fold standard-deviation bars. The L3 → L4 drop localizes the loss to standardize-then-pool; L7 recovers part of the statistical ceiling.
Model
FF++ AUC
Celeb-DF AUC
RGB-Only (ref.)
0.928
0.888
Late-Fusion (original)
0.924
0.881
Residual-Only (original)
0.579
0.547
StatNoise-Fusion (fixed)
0.845±0.009
0.786±0.015
ResAware-Fusion (fixed)
0.856±0.009
0.805±0.015
TABLE III: Fixed-fusion variants (mean ± std over five seeds). Removing the InstanceNorm bottleneck recovers the statistical signal but still underperforms RGB-only on both datasets.
Fig. 2: Paired per-dataset difference in frame-level AUC relative to RGB-Only, by perturbation family (JPEG, Gaussian blur, resize, gamma), across FF++, Celeb-DF, and DFDC; shaded bands are 95% CIs ( n=3 datasets, paired within dataset). Late-Fusion − RGB-Only (blue) hugs the zero line at every operating point—never significantly positive, its only CI-excluding-zero point being negative (blur σ=2 )—so the noise branch adds nothing over RGB. SBI − RGB-Only (orange), included as a reference for what a genuinely different model looks like, is positive throughout and largest under heavy JPEG. DFF is excluded (image-only and near chance for the RGB and fusion models).
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
Item
Setting
Data
Source corpus
FF++ + DeepFakeDetection extension
Split granularity
Video level (no leakage across splits)
Train / val / test
19,571 / 3,573 / 4,675 crops
Cross-dataset eval
Celeb-DF, DFDC, DFF
Class balance
Class-weighted cross-entropy
Appendix
TABLE IV: Data protocol and training configuration. Splits are at the video level so no source video crosses partitions.
Current face video forgery detectors use wide or dual-stream backbones. We show that a single, lightweight fusion of two handcrafted cues can achieve higher accuracy with a much smaller model. Based on the Xception baseline model (21.9 million parameters), we build two detectors: LFWS, which adds a 1x1 convolution to combine a low-frequency Wavelet-Denoised Feature (WDF) with a phase-spectrum channel derived from Spatial-Phase Shallow Learning (SPSL), and LFWL, which merges WDF with Local Binary Patterns (LBP) in the same way. This extra module adds only 292 parameters, keeping the total at 21.9 million, smaller than F3Net (22.5 million) and less than half the size of SRM (55.3 million). Even with this minimal overhead, the fused models increase the average area under the curve (AUC) from 74.8% to 78.6% on FaceForensics++ and from 70.5% to 74.9% on DFDC-Preview, gains of 3.8% and 4.4% over the Xception baseline. They also consistently outperform F3Net, SRM, and SPSL in eight public benchmarks, without extra data or test-time augmentation. These results show that carefully paired, handcrafted features, combined through the lightweight fusion block, can provide competitive robustness at a significantly lower cost than comparable frequency-based detectors. Our findings suggest a need to reevaluate scale-driven design choices in face video forgery detection.
Recent deepfake detection methods demonstrate improved cross-dataset generalization, yet the underlying mechanisms remain underexplored. We introduce the Alpha Blending Hypothesis, positing that state-of-the-art frame-based detectors primarily function as alpha blending searchers; rather than learning semantic anomalies or specific generative neural fingerprints, they localize low-level compositing artifacts introduced during the integration of manipulated faces into target frames. We experimentally validate the hypothesis, demonstrating that deepfake detectors exhibit high sensitivity to the so-called self-blended images (SBI) and non-generative manipulations. We propose the method BlenD that leverages a large-scale, diverse dataset of real-only facial images augmented with SBI. This approach achieves the best average cross-dataset generalization on 15 compositional deepfake datasets released between 2019 and 2025 without utilizing explicitly generated deepfakes during training. Furthermore, we show that predictions from explicit blending searchers and models resilient to blending shortcuts are highly complementary, yielding a state-of-the-art AUROC of 94.0% in an ensemble configuration. The code with experiments and the trained model will be publicly released.
Andrii Yermakov, Jan Cech, Mario Fritz +1
Czech Technical University in Prague · CISPA Helmholtz Center for Information Security
We introduce IRIS-GAN, a specialist forensic detector for synthetic face images under cross-generator shift. Rather than addressing universal synthetic-image detection, we focus on faces generated by generative adversarial networks (GANs), which are state-of-the-art in deepfake content, and train the detector through staged exposure to increasingly demanding GAN families while retaining earlier generators. The final model reaches fake-detection rates above 99% across the GAN families considered and classifies an external real-face dataset with 98.9% accuracy. Grad-CAM analysis further reveals measurable generator-dependent spatial response patterns, which remain informative for a secondary heatmap-only classifier. Out-of-family tests on diffusion-generated faces confirm that IRIS-GAN is a specialist detector, with some capability to reach non-GAN deepfakes. These results establish staged training as an effective strategy for robust GAN-face forensics.
Jaume M. Trenchs, Veronica Sanz
aDepartamento de Física Teórica, Universitat de València, Burjassot, Spain · bInstituto de Física Corpuscular (IFIC), CSIC–Universitat de València, Valencia, Spain