Organizations: Innopolis University, Innopolis, Russia · Laboratory of Innovative Technologies for Processing Video Content, Innopolis University, Innopolis, Russia
Fusing a learned camera-noise fingerprint with an RGB appearance backbone is an appealing route to generator-independent deepfake detection, because the noise residual is grounded in image-formation physics rather than in the texture statistics of a particular generator. We test, on FaceForensics++, whether a Noiseprint++ residual channel carries information \emph{complementary} to an RGB Xception backbone for face-swap detection. A three-model ablation (RGB-only, residual-only, late-fusion) shows that fusion does not improve over RGB alone and that the residual branch alone is near chance. A seven-level bottleneck diagnostic localizes the cause: the noise maps do carry a discriminative signal, but it is statistical---carried by the per-sample first and second moments (mean, variance, energy) of the residual---and the per-sample \texttt{InstanceNorm} layer placed at the noise-branch input, following the TruFor template, standardizes exactly those moments away (five-fold cross-validated AUC drops from 0.747 to 0.554). A context-crop control rules out cropping geometry, and two fixed-fusion variants that remove the bottleneck recover the statistical signal yet still fail to beat RGB on every dataset. We conclude that, on this manipulation distribution, the noise residual is redundant with RGB rather than complementary, and we give concrete guidance for practitioners adopting noise-residual fusion for face-swap detection.
Figures & tables
Model
FF++
Celeb-DF
DFDC
DFF
RGB-Only
0.928
0.888
0.717
0.531
Late-Fusion
0.924
0.881
0.712
0.475
Residual-Only
0.579
0.547
0.578
0.212
SBI (ref.)
0.865
0.907
0.884
0.614
TABLE I: In-distribution and cross-dataset video-level AUC (frame-level for DFF). RGB-only ≈ Late-Fusion (Wilcoxon p=0.977 ); Residual-only is near chance. SBI is an off-the-shelf reference, shown for context.
Classifier
Input
AUC
L1: LogReg
mean ∣ noise ∣
0.653±0.018
L2: LogReg
6 scalar stats
0.664±0.018
L3: MLP (ceiling)
6 scalar stats
0.747±0.023
L4: LogReg
InstanceNorm-pooled noise
0.554±0.014
L5: LogReg
raw-pooled noise
0.559±0.026
L6: ConvNet (2-layer)
raw noise
0.602±0.021
TABLE II: Seven-level bottleneck diagnostic on Noiseprint++ maps. Five-fold cross-validated AUC on 3,000 FF++ test samples. L3 sets the statistical ceiling ( 0.747 ); L4 localizes the InstanceNorm collapse ( 0.554 ).
Fig. 1: Seven-level bottleneck AUC (L1–L7) with five-fold standard-deviation bars. The L3 → L4 drop localizes the loss to standardize-then-pool; L7 recovers part of the statistical ceiling.
Model
FF++ AUC
Celeb-DF AUC
RGB-Only (ref.)
0.928
0.888
Late-Fusion (original)
0.924
0.881
Residual-Only (original)
0.579
0.547
StatNoise-Fusion (fixed)
0.845±0.009
0.786±0.015
ResAware-Fusion (fixed)
0.856±0.009
0.805±0.015
TABLE III: Fixed-fusion variants (mean ± std over five seeds). Removing the InstanceNorm bottleneck recovers the statistical signal but still underperforms RGB-only on both datasets.
Fig. 2: Paired per-dataset difference in frame-level AUC relative to RGB-Only, by perturbation family (JPEG, Gaussian blur, resize, gamma), across FF++, Celeb-DF, and DFDC; shaded bands are 95% CIs ( n=3 datasets, paired within dataset). Late-Fusion − RGB-Only (blue) hugs the zero line at every operating point—never significantly positive, its only CI-excluding-zero point being negative (blur σ=2 )—so the noise branch adds nothing over RGB. SBI − RGB-Only (orange), included as a reference for what a genuinely different model looks like, is positive throughout and largest under heavy JPEG. DFF is excluded (image-only and near chance for the RGB and fusion models).
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
Item
Setting
Data
Source corpus
FF++ + DeepFakeDetection extension
Split granularity
Video level (no leakage across splits)
Train / val / test
19,571 / 3,573 / 4,675 crops
Cross-dataset eval
Celeb-DF, DFDC, DFF
Class balance
Class-weighted cross-entropy
Appendix
TABLE IV: Data protocol and training configuration. Splits are at the video level so no source video crosses partitions.
aDepartamento de Física Teórica, Universitat de València, Burjassot, Spain · bInstituto de Física Corpuscular (IFIC), CSIC–Universitat de València, Valencia, Spain