cs.LGOct 22, 2025

Estimating Model-Level Membership Inference Vulnerability Without Reference Models

Authors: Euodia Dodd, Nataša Krčo, Igor Shilov, Matthew Wicker, Yves-Alexandre de Montjoye

Organizations: Imperial College London

Abstract

Membership inference attacks (MIAs) have emerged as the standard tool for evaluating the privacy risks of AI models. However, state-of-the-art attacks require training numerous, often computationally expensive, reference models, limiting their practicality. We present a novel approach for estimating model-level vulnerability to the Likelihood Ratio Attack (LiRA), the strongest available attack, directly from the train and test loss distributions of the target model and without training any reference models. We show that LiRA's per-sample signal decomposes into a variance-ratio term and a residual mean-shift term, with the relative contribution of each determined by how much training collapses model uncertainty at the trained sample. This places models on a continuum, with different regimes calling for different reference-free loss-based statistics as proxies for LiRA TPR. The shapes of the loss distributions themselves indicate which proxy applies. We instantiate the framework with two natural proxies. At the heavy-tailed end, the LOSS attack TNR predicts LiRA TPR@FPR=10−310^{-3} with RMSE 0.036 across 10 image classification architectures and 4 datasets, outperforming low-cost reference-model attacks such as RMIA. At the symmetric end, the LOSS attack AUC predicts LiRA TPR with RMSE 0.018 across five GPT-2 sizes from 10M to 1B parameters.

Explore similar work

Jun 16, 2026cs.LG

CheckMIABench: Firm Foundations For Membership Inference Attacks on Language Models

Membership inference attacks (MIAs) are a canonical way to assess a machine learning model's privacy properties. Although several attempts have been made to evaluate MIAs on language models, the extant literature has suffered numerous difficulties in constructing clean evaluations to test new techniques. In particular, subtle distribution shifts between member and non-member sets can undermine the statistical validity of MIAs; recent work has underscored this by showing that "blind" methods with no access to the underlying model can perform far better than published methods on the same benchmarks. This paper constructs a benchmark for principled evaluation of MIAs against LLMs, by leveraging the insight that training data before and after a fixed point during training are drawn from the same distribution. Therefore, all open-source models with intermediate checkpoints and public training data can be converted into MIA testbeds. We apply our framework to a half-dozen published attacks on the Pythia and OLMo family of models, from 70M to 7B parameters. To facilitate further privacy research, we open-source a modular library for designing and implementing attacks in this setting: https://github.com/safr-ai-lab/pandora_llm.
May 25, 2026cs.LG

On Reliability of Membership Inference Vulnerability Evaluation

Membership inference attacks (MIAs) are popular methods for empirically assessing the leakage of sensitive information in the training data through models or statistics learned from the data. The MI vulnerability is often evaluated through a binary classifier that tries to predict whether a particular sample was in the training data. In order to evaluate the effectiveness of MIAs multiple \textit{shadow models} are trained using random partitions of a larger dataset. After training the shadow models the MI vulnerability can be evaluated for all the samples for which we obtained shadow models. In order to evaluate the MI vulnerability reliably one needs a lot of shadow models which can be computationally infeasible. Therefore instead of reporting the actual sample level vulnerabilities aggregates over multiple samples are often reported in practice. We demonstrate two key weaknesses in typical MIA evaluation pipeline. First, we show that sampling the shadow datasets from a fixed superset leads to finite sample bias inflating the vulnerability estimates. Second, we show that evaluating the true positive rate (TPR) by concatenating MIA scores across multiple individuals, commonly used in the very low false positive rate (FPR) regime, is not calibrated across the per-sample FPRs. For both weaknesses we propose fixes that in the most simple approximate form do not incur any additional computation cost. We show that with additional computation one can further improve the reliability of the vulnerability estimation.
Sep 14, 2026stat.ML

Membership Inference via Pairwise Likelihood Ratios

Membership inference attacks (MIAs) are the standard tool for auditing the privacy risks of machine learning models. Given a query point, an MIA aims to determine whether that point was used to train the target model. In practice, such inference must rely on the statistical signals exposed by the model's outputs, such as confidence scores, logits, and intermediate feature representations. However, existing methods often fail to efficiently summarize and combine these statistical signals. To address this limitation, we propose Pairwise Likelihood MIA (PL-MIA), a unified method that combines a Gaussian likelihood-ratio (GLR) statistic with population calibration and the Cauchy combination test. We characterize theoretically how the GLR retains variance-contraction signals and establish conditions under which population calibration and Cauchy combination improve attack power. We obtain pp-values from pairwise comparisons between the query point and reference points not used for training, and aggregate these continuous signals using the Cauchy combination test. This preserves the evidence strength that is discarded when each pairwise comparison is reduced to a binary vote. Extensive experiments demonstrate that PL-MIA outperforms strong baselines, improving the true positive rate (TPR) by over 25% in the critical low-false-positive regime, corroborating our theoretical findings. These results demonstrate how statistical principles can turn noisy model outputs into more powerful, calibrated, and reproducible evidence for membership privacy auditing.