In this paper, we show that standard evaluations of high-resolution Model Inversion Attacks (MIAs) significantly underestimate training-data privacy leakage. State-of-the-art privacy defenses, standard training techniques such as MixUp and Adversarial Training, and undefended models all leak training images at rates 1.16 to 6.59 times higher on FaceScrub under simple adaptive changes to the attack, with the largest increases among defenses reporting the strongest privacy. We further show that measured leakage depends on the feature basis of the external classifier used to evaluate reconstructions: for the same reconstructed images, an adversarially trained Inception evaluator identifies the targeted identity at different rates than the standard Inception evaluator. Our results suggest that standard MIA evaluation can mistake optimization and measurement failures for privacy. These underestimated leakage rates also concealed a broader relationship between privacy and adversarial robustness. Once we adapt the attack and vary the evaluator, reconstruction leakage closely tracks adversarial robustness across recent defenses and standard training regimes, suggesting that robustness provides an attack-agnostic proxy for reconstruction vulnerability that applies far more broadly than previously theorized. This raises an open question: can a practical defense reduce training-data reconstruction without paying a corresponding cost in adversarial robustness?
Figures & tables
Figure 1: The privacy–robustness relationship in Torp et al. (2026) does not generalize under standard MIA evaluation. Opaque points were included in the original Torp et al. (2026) regression.
Figure 2: Leakage (AttAcc@1 & L2-Face) vs. test accuracy. Shapes denote attacks. Vertical lines measure vanilla PPA vs. the adaptive attacks, and opaque markers highlight the worst-case leakage.
Figure 3: The privacy–robustness relationship generalizes substantially better under adaptive attacks (bottom row) and a Robust Inception-v3 evaluator (right column). Opaque points were used for fitting (the same set of defenses that were used to fit in the Torp et al. (2026) regression).
Figure 4: Heatmaps of inversion-time loss gradients ∇xLcls(x) , where x=G(z) is a GAN-generated reconstruction at various PPA iterations on ResNet-152. Gradient magnitudes are normalized , i.e., incomparable across target models. Rows 1 & 2 correspond to FaceScrub class 12 at attack iterations 9 and 49, respectively. Rows 3 & 4 correspond to the same iterations of class 16.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: Regression predicting leakage from the alignment score. The fitted alignment coefficient is positive (6.16), consistent with greater alignment being associated with greater leakage, although held-out predictive fit is weak.
Figure 6: Regressions fit on all data points using the worst-case adaptive attack.