Model quantization reduces the numerical precision of neural network weights and activations to lower storage and computational costs. Model inversion attacks recover or reconstruct sensitive training data or inference inputs from model outputs or intermediate features, so quantization may also alter their effectiveness. However, two questions remain unresolved: How does model quantization affect model inversion? How do data characteristics influence this relationship? To address the first, we bound quantization-induced changes in mutual information between inputs and a categorical variable defined by prediction probabilities, distinguishing informational effects from attack optimization obstacles. To address the second, we identify data-dependent changes in feature distributions and inversion outcomes, with pronounced quantization sensitivity differences at 4 bits. These insights guide a privacy-aware post-training quantization method that improves inversion resistance while recovering utility. It uses a Fisher-type task-sensitivity proxy for budget-aware bit allocation, calibrates activation ranges, and jointly optimizes weight and activation scales and weight-rounding decisions with task-recovery and geometry-retention objectives and scale and rounding regularization. Experiments cover multiple metrics, neural network architectures, and face, palmprint, and iris recognition tasks. On ResNet-50, Palm at 4 bits reduces RL-MIA's strict success from 54% to 26%, while accuracy decreases from 99.01% to 96.55% relative to FP32. Our method also supports output-level defenses: adding Stealthy Shield Defense (SSD, epsilon = 0.1) to Iris at 4.5 bits reduces BREP-MI's strict success from 63.33% to 37.33%, while accuracy decreases from 92.8% to 87.6% relative to quantization alone.
Figures & tables
Fig. 1: Overview of model quantization and model inversion attacks in biometric recognition. Our research aims to understand how quantization and data characteristics affect inversion risk and to improve the trade-off between predictive utility and inversion risk.
Fig. 2: Blocked inversion gradients in the original PLG-MI attack [ 52 ] on INT8-quantized Palm classifiers. The checkpoints were produced using PyTorch’s official static PTQ APIs [ 39 ] . Across 20 Stage-1 generator updates, the inversion-loss gradient with respect to the generated image was zero for (a) ViT and unavailable for (b) ResNet-50; unavailable gradients are plotted as zero. Separate Stage-2 first-step checks likewise found a zero latent gradient for ViT and no autograd path for ResNet-50.
Fig. 3: Illustrative bit-width dependence of information-related bounds. Both panels use Δb=2/(2b−1) for an idealized grid containing both endpoints of [−1,1] , with clipping omitted and propagation factors normalized. This grid differs from ( 1 ). (a) The categorical information bound ( 21 ) with M=50 and KΔ=1 . (b) The per-dimension form of ( 33 ), −21log[1−(2Δb+Δb2)] , with unit feature-spread and regularized covariance-floor parameters and representation error bounded by Δb . The required condition 2Δb+Δb2<1 holds throughout. Full derivations appear in Appendices D and G.
Fig. 4: Data characteristics and model responses to quantization. (a) Input variability in a shared 32-dimensional PCA space [ 40 ] : the horizontal axis shows an absolute Gaussian within-class entropy proxy, and the vertical axis shows relative conditional volume. Error bars denote 95% class-bootstrap confidence intervals. (b) Within-class variability of FP32 representations across layers. Inner-circle area relative to the outer circle represents the geometric mean of per-coordinate within-class-to-total variance ratios; color indicates the corresponding normalized Gaussian entropy proxy. Conditioning in (a) and (b) uses the ground-truth class label Y . (c) Dataset means of the global high-frequency energy ratio (HFER) and the per-image patch-level 95th-percentile HFER, with radial cutoff rc=0.25 , based on spectral texture analysis [ 18 , 44 ] . (d) FP32-minus-quantized accuracy under 4-bit weight-and-activation quantization across datasets and architectures. Panels use the analysis populations specified in Appendices H and I; comparisons across panels are at the domain level.
Fig. 5: Overview of privacy-aware mixed-precision PTQ. (1) Estimate task sensitivity at candidate bit widths. (2) Allocate weight and activation precision under separate budgets and calibrate activation ranges. (3) Refine quantization scales and weight-rounding decisions with fixed bit widths and network weights. The FP32 teacher guides utility recovery, while the initial quantized model provides the reference for limiting inter-class expansion.
(a) Target-data split and auxiliary-pool size
Dataset
IDs
Train pool
Test
Auxiliary
FaceScrub
530
33,756 (80.00%)
8,440 (20.00%)
195,721 ( 5.80× )
Palm
50
617 (75.24%)
203 (24.76%)
4,682 ( 7.59× )
Iris-Thousand
150
2,250 (75.00%)
750 (25.00%)
17,000 ( 7.56× )
TABLE I: Experimental setup for quantization and model inversion.
Fig. 6: Utility and model inversion results on FaceScrub and Palm with ResNet-50 targets. Open circles and filled markers show accuracy before and after quantizer refinement; open squares (P) denote independent calibration-only PTQ references, and dashed lines indicate FP32 accuracy. Heatmap entries are changes from the FP32 baseline for each attack and target; absolute baselines appear in brackets. Changes in Strict success and Evaluator-KNN@1 are in percentage points; LPIPS-Alex and FID retain their original units. Blue indicates lower identity-hit rates or larger distances, and orange indicates the reverse; each heatmap has its own zero-centered color scale. 0 , + , and − denote refinement with η=0 , 0.1 , and −0.001 , respectively. Superscripts 1 and 2 distinguish mixed-precision candidate sets {6,5,4} and {6,5,4,3} . Palm additionally includes CDM-MIA.
Fig. 7: Utility and model inversion results on Iris-Thousand with ResNet-50 and Swin targets. Plotting conventions follow Fig. 6 .
Fig. 8: Qualitative identity recovery under quantization. Each panel shows three real images of one target identity and final attack candidates from FP32, calibration-only 4-bit PTQ, and our 4-, 3.5-, and 3-bit models ( η=0 ). The real images are identity references, not paired reconstruction targets. Both 4-bit columns use uniform W4A4; candidate-set superscripts follow Fig. 6 . Percentages are target-test accuracies. Check marks denote Strict success (both the target and independent evaluator predict the target identity); crosses denote failure. All shown identities satisfy FP32 Strict success for BREP-MI and RL-MIA. The Iris examples additionally require Strict failure in at least two of the three displayed Ours configurations for each of these attacks; LOKT and CDM-MIA outcomes do not constrain selection. The Iris panels use different identities.
Fig. 9: Composition with SSD [ 56 ] on Iris-Thousand with a ResNet-50 target. Open and filled circles show test accuracy without (Base) and with SSD for each matched checkpoint. SSD uses T=0.1 and the displayed per-model ϵ , frozen under model-specific validation budgets. Accuracy uses one response per test image (750 images). Heatmaps use each attack’s original FP32 baseline without SSD. Superscript definitions and other heatmap conventions follow Fig. 6 . The 4-bit models use uniform W4A4. Results include all 150 final candidates per non-adaptive attack, retaining failures.
In this paper, we show that standard evaluations of high-resolution Model Inversion Attacks (MIAs) significantly underestimate training-data privacy leakage. State-of-the-art privacy defenses, standard training techniques such as MixUp and Adversarial Training, and undefended models all leak training images at rates 1.16 to 6.59 times higher on FaceScrub under simple adaptive changes to the attack, with the largest increases among defenses reporting the strongest privacy. We further show that measured leakage depends on the feature basis of the external classifier used to evaluate reconstructions: for the same reconstructed images, an adversarially trained Inception evaluator identifies the targeted identity at different rates than the standard Inception evaluator. Our results suggest that standard MIA evaluation can mistake optimization and measurement failures for privacy. These underestimated leakage rates also concealed a broader relationship between privacy and adversarial robustness. Once we adapt the attack and vary the evaluator, reconstruction leakage closely tracks adversarial robustness across recent defenses and standard training regimes, suggesting that robustness provides an attack-agnostic proxy for reconstruction vulnerability that applies far more broadly than previously theorized. This raises an open question: can a practical defense reduce training-data reconstruction without paying a corresponding cost in adversarial robustness?
Shailen Smith, Rasmus Torp, Adam Breuer
Department of Computer Science, Dartmouth College · Department of Government, Dartmouth College
Post-training quantization (PTQ) converts a trained full-precision model into low-bit weights without task-level retraining, while quantization-aware training (QAT) incorporates quantization into the training loop. Although PTQ is efficient and often accurate at moderate bitwidths, it can fail sharply at aggressive bitwidths; QAT is more expensive but can often recover the lost accuracy. We propose a unified geometric framework that explains both PTQ failure and QAT recovery. We model full-precision training as following a low-loss \emph{river} inside a wider \emph{valley}: a normal neighborhood of the river forms a nearly flat \emph{basin}, while leaving this basin incurs a sharp loss increase. When the quantization grid is comparable to the basin width, local PTQ objectives, including rounding and Hessian-based second-order reconstruction, can select a high-loss deployed quantized point outside the basin even when nearby low-loss quantized points exist. In this regime, straight-through-estimator-based QAT has a useful bias: it evaluates gradients at the deployed quantized weights while updating latent full-precision weights, causing the gradient to sense the valley wall and acquire an inward component that steers subsequent quantized iterates back into the basin. We formalize this mechanism through a local landscape model, construct a geometric PTQ failure mode, and prove finite-time QAT recovery under local quantizer-compatibility assumptions. Experiments across vision and language models under multiple neural-network quantization schemes corroborate the predicted basin-crossing failure of PTQ and the corresponding recovery mechanism of QAT.
Hanyang Li, Jianhao Ma, Ying Cui
Department of IEOR, University of California, Berkeley · Department of Statistics and Data Science, University of Pennsylvania
LLM quantization has become essential for memory-efficient deployment. Recent work has shown that quantization schemes can pose critical security risks: an adversary may release a model that appears benign in full precision but exhibits malicious behavior once quantized by users. However, existing quantization-conditioned attacks have been limited to relatively simple quantization methods, where the attacker can estimate weight regions that remain invariant under the target quantization. Notably, prior attacks have consistently failed to compromise more popular and sophisticated schemes, limiting their practical impact. In this work, we introduce the first quantization-conditioned attack that consistently induces malicious behavior that can be triggered by a broad range of advanced quantization techniques, including AWQ, GPTQ, and GGUF I-quants. Our attack exploits a simple property shared by many modern quantization methods: large outliers can cause other weights to be rounded to zero. Consequently, by injecting outliers into specific weight blocks, an adversary can induce a targeted, predictable weight collapse in the model. This effect can be used to craft seemingly benign full-precision models that exhibit a wide range of malicious behaviors after quantization. Through extensive evaluation across three attack scenarios and LLMs, we show that our attack achieves high success rates against a broad range of quantization methods on which prior attacks fail. Our results demonstrate, for the first time, that the security risks of quantization are not restricted to simpler schemes but are broadly relevant across complex, widely-used quantization methods.