cs.LGOct 6, 2026

Adaptive Model Inversion Attacks Generalize a Privacy-Robustness Tradeoff

Authors: Shailen Smith, Rasmus Torp, Adam Breuer

Organizations: Department of Computer Science, Dartmouth College · Department of Government, Dartmouth College

Abstract

In this paper, we show that standard evaluations of high-resolution Model Inversion Attacks (MIAs) significantly underestimate training-data privacy leakage. State-of-the-art privacy defenses, standard training techniques such as MixUp and Adversarial Training, and undefended models all leak training images at rates 1.16 to 6.59 times higher on FaceScrub under simple adaptive changes to the attack, with the largest increases among defenses reporting the strongest privacy. We further show that measured leakage depends on the feature basis of the external classifier used to evaluate reconstructions: for the same reconstructed images, an adversarially trained Inception evaluator identifies the targeted identity at different rates than the standard Inception evaluator. Our results suggest that standard MIA evaluation can mistake optimization and measurement failures for privacy. These underestimated leakage rates also concealed a broader relationship between privacy and adversarial robustness. Once we adapt the attack and vary the evaluator, reconstruction leakage closely tracks adversarial robustness across recent defenses and standard training regimes, suggesting that robustness provides an attack-agnostic proxy for reconstruction vulnerability that applies far more broadly than previously theorized. This raises an open question: can a practical defense reduce training-data reconstruction without paying a corresponding cost in adversarial robustness?

Figures & tables

Appendix figures & tables2 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

Jul 14, 2026cs.LG

Reducing information dependency does not cause training data privacy. Adversarially non-robust features do

In this paper, we challenge the prevailing view that information dependency (including rote memorization) drives training data exposure to image reconstruction attacks. We show that extensive exposure can persist without rote memorization and is instead caused by a tunable connection to adversarial robustness. We begin by presenting three surprising results: (1) recent defenses that inhibit reconstruction by Model Inversion Attacks (MIAs), which evaluate leakage under an idealized attacker, do not reduce standard measures of information dependency (HSIC); (2) models that maximally memorize their training datasets remain robust to MIA reconstruction; and (3) models trained without seeing 97% of the training pixels, where recent information-theoretic bounds give arbitrarily strong privacy guarantees under standard assumptions, can still be devastatingly reconstructed by MIA. To explain these findings, we provide causal evidence that privacy under MIA arises from what the adversarial examples literature calls ``non-robust'' features (generalizable but imperceptible and unstable features). We further show that recent MIA defenses obtain their privacy improvements by unintentionally shifting models toward such features. To establish this causal relationship, we introduce Anti Adversarial Training (AT-AT), a training regime that intentionally learns non-robust features to obtain both superior reconstruction defense and higher accuracy than state-of-the-art defenses. Our results revise the prevailing understanding of training data exposure and reveal a new privacy-robustness tradeoff.
Sep 30, 2026cs.CR

On the Relationship between Model Quantization and Model Inversion Attacks

Model quantization reduces the numerical precision of neural network weights and activations to lower storage and computational costs. Model inversion attacks recover or reconstruct sensitive training data or inference inputs from model outputs or intermediate features, so quantization may also alter their effectiveness. However, two questions remain unresolved: How does model quantization affect model inversion? How do data characteristics influence this relationship? To address the first, we bound quantization-induced changes in mutual information between inputs and a categorical variable defined by prediction probabilities, distinguishing informational effects from attack optimization obstacles. To address the second, we identify data-dependent changes in feature distributions and inversion outcomes, with pronounced quantization sensitivity differences at 4 bits. These insights guide a privacy-aware post-training quantization method that improves inversion resistance while recovering utility. It uses a Fisher-type task-sensitivity proxy for budget-aware bit allocation, calibrates activation ranges, and jointly optimizes weight and activation scales and weight-rounding decisions with task-recovery and geometry-retention objectives and scale and rounding regularization. Experiments cover multiple metrics, neural network architectures, and face, palmprint, and iris recognition tasks. On ResNet-50, Palm at 4 bits reduces RL-MIA's strict success from 54% to 26%, while accuracy decreases from 99.01% to 96.55% relative to FP32. Our method also supports output-level defenses: adding Stealthy Shield Defense (SSD, epsilon = 0.1) to Iris at 4.5 bits reduces BREP-MI's strict success from 63.33% to 37.33%, while accuracy decreases from 92.8% to 87.6% relative to quantization alone.
Jul 14, 2026cs.CR

VanillaBench: The Hidden Accuracy Cost of Adversarial Robustness

Adversarial robustness research has produced hundreds of defended models over the past decade, yet the literature almost universally reports robustness results in isolation: standard (clean) accuracy and adversarial accuracy of the robust model are shown, but the gap to the corresponding vanilla model is rarely quantified. We introduce VanillaBench, a systematic benchmark that makes this gap explicit. For every adversarially-trained model catalogued by RobustBench across four threat models, we compute the accuracy difference against multiple vanilla references from Papers with Code, computed over both all entries and no-extra-data entries, the best vanilla model as of the robust model's publication year, and an architecture-matched baseline. Across all 186 robust models, the mean delta clean relative to the best vanilla model ranges from -7.7 to -29.5 percentage points, and even the single most robust model per track still trails its temporal vanilla counterpart by 4.0-21.0 points. The architecture-matched comparison, which isolates the effect of adversarial training from architectural differences, reveals a mean gap of -3.5 to -17.5 points. Restricting this architecture-matched comparison to models whose vanilla accuracy is known for the exact same architecture, rather than approximated from a related one, narrows the gap to -4.0 to -14.0 points. These results demonstrate that the robustness-accuracy trade-off is substantially larger than what is typically conveyed by individual papers. This information is critical for practitioners and decision-makers. When deploying models in real-world settings, the accuracy cost of robustness directly affects business outcomes, yet current publications do not provide the vanilla baseline needed to assess it. We argue that future robustness evaluations should report vanilla-referenced accuracy gaps as a standard component.