Membership inference attacks (MIAs) test whether a candidate example was used to train a language model. Existing attacks on fine-tuned discrete diffusion language models (dLLMs) often aggregate reconstruction signals across many mask configurations, requiring repeated model evaluations. We propose Joint Uncertainty Guided Mask Probing (JUMP), an efficient MIA that exploits the ability of dLLMs to predict masked tokens in parallel. Using the pre-fine-tuning checkpoint as a reference, JUMP selects low-confidence positions, masks them jointly, and aggregates clipped target-reference reconstruction gaps. This focuses the attack on positions that reveal stronger membership signals from fine-tuning. After mask selection, all selected tokens are evaluated with one scoring query per model. Across six MIMIR domains, JUMP improves mean ROC-AUC over a prior multi-mask attack from 0.819 to 0.902 on LLaDA and from 0.851 to 0.942 on Dream. Including mask selection, it requires only three forward passes per example, compared with 32 for the baseline. We further extend JUMP to the target-only setting by replacing target-reference scoring with relative token preference, which compares the observed token with alternative predictions at the same masked position. Target-Only JUMP achieves mean ROC-AUCs of 0.609 and 0.638 on LLaDA and Dream.
Figures & tables
Figure 1: Overview of JUMP . The attack first selects informative positions under the reference model, then masks the selected positions jointly and scores them with the target and reference dLLMs. The final MIA statistic is computed from clipped token-level target/reference reconstruction gaps. In the visualization, general tokens are indicated by solid-line boxes, while confidence scores are denoted by dashed-line boxes.
LLaDA-8B-Base
ArXiv
GitHub
HackerNews
Method
Attack quality ( ↑ )
Cost ( ↓ )
Attack quality ( ↑ )
Cost ( ↓ )
Attack quality ( ↑ )
Cost ( ↓ )
AUC
T@10
T@1
T@0.1
NFE
AUC
T@10
T@1
T@0.1
NFE
AUC
T@10
T@1
T@0.1
NFE
Loss
0.52
0.12
0.01
0.00
16
0.59
0.20
0.04
0.01
16
0.51
0.11
0.01
0.01
16
ZLIB
0.52
0.13
0.01
0.00
16
0.61
0.23
0.07
0.02
16
0.51
0.10
0.01
0.01
16
SAMA
0.82
0.54
0.28
0.07
32
0.77
0.43
0.14
0.07
32
0.71
0.34
0.05
0.01
32
Table 1: Per-domain reference-assisted results on LLaDA-8B-Base and Dream-v0-Base-7B. We report ROC-AUC and TPR at fixed FPR ( ↑ ) with total NFE ( ↓ ). SAMA uses T=16 (32 forwards) and JUMP uses 3; best quality values are in bold .
Table 4: Checkpoint-selection robustness on GitHub and HackerNews. “Selected” denotes the checkpoint used in the main evaluation ( ∼ 3 epochs), while “Epoch 4” evaluates checkpoint 80.
Domain
AUC
T@10
T@1
T@0.1
ArXiv
0.94 [.93,.95]
0.84 [.80,.86]
0.55 [.45,.60]
0.30 [.10,.46]
GitHub
0.86 [.84,.87]
0.66 [.62,.70]
0.32 [.24,.44]
0.18 [.00,.25]
HackerNews
0.76 [.74,.78]
0.39 [.35,.45]
0.10 [.07,.16]
0.02 [.00,.08]
PubMed Central
0.92 [.90,.93]
0.75 [.71,.80]
0.43 [.31,.52]
0.12 [.04,.33]
Wikipedia (en)
0.98 [.97,.99]
0.95 [.94,.97]
0.80 [.76,.85]
0.67 [.61,.78]
Pile-CC
0.96 [.95,.97]
0.90 [.88,.93]
0.60 [.52,.67]
0.10 [.01,.53]
Appendix
Table 5: Per-domain JUMP results on fine-tuned LLaDA-8B-Base, with 95% bootstrap CIs in brackets.
Domain
AUC
T@10
T@1
T@0.1
ArXiv
0.93 [.92,.94]
0.81 [.78,.84]
0.52 [.39,.58]
0.32 [.22,.40]
GitHub
0.91 [.89,.92]
0.75 [.69,.80]
0.36 [.28,.48]
0.01 [.00,.29]
HackerNews
0.93 [.92,.94]
0.81 [.77,.83]
0.49 [.28,.58]
0.23 [.15,.30]
PubMed Central
0.95 [.94,.96]
0.86 [.83,.89]
0.53 [.49,.66]
0.35 [.28,.51]
Wikipedia (en)
0.96 [.95,.97]
0.89 [.87,.92]
0.63 [.51,.73]
0.29 [.22,.52]
Pile-CC
0.97 [.97,.98]
0.94 [.92,.95]
0.77 [.66,.81]
0.38 [.13,.68]
Appendix
Table 6: Per-domain JUMP results on fine-tuned Dream-v0-Base-7B, with 95% bootstrap CIs in brackets.
Domain
AUC
T@10
T@1
T@0.1
ArXiv
0.62 [.59,.64]
0.21 [.17,.26]
0.03 [.01,.06]
0.01 [.00,.02]
GitHub
0.60 [.57,.62]
0.21 [.17,.24]
0.03 [.02,.05]
0.01 [.00,.02]
HackerNews
0.57 [.54,.59]
0.15 [.12,.19]
0.02 [.01,.03]
0.00 [.00,.02]
PubMed Central
0.62 [.60,.64]
0.21 [.17,.26]
0.03 [.02,.06]
0.01 [.00,.02]
Wikipedia (en)
0.64 [.62,.66]
0.22 [.18,.26]
0.03 [.02,.05]
0.01 [.00,.02]
Pile-CC
0.62 [.59,.64]
0.20 [.16,.24]
0.02 [.01,.04]
0.01 [.00,.01]
Appendix
Table 7: Per-domain Target-Only JUMP results on LLaDA-8B-Base, with 95% bootstrap CIs in brackets.
Domain
AUC
T@10
T@1
T@0.1
ArXiv
0.66 [.63,.68]
0.23 [.18,.28]
0.04 [.02,.08]
0.02 [.01,.03]
GitHub
0.58 [.56,.61]
0.13 [.10,.17]
0.01 [.00,.02]
0.00 [.00,.01]
HackerNews
0.66 [.63,.68]
0.22 [.19,.25]
0.04 [.01,.07]
0.00 [.00,.02]
PubMed Central
0.64 [.61,.66]
0.20 [.16,.25]
0.02 [.00,.04]
0.00 [.00,.01]
Wikipedia (en)
0.65 [.63,.68]
0.19 [.15,.25]
0.02 [.00,.04]
0.00 [.00,.01]
Pile-CC
0.64 [.62,.66]
0.23 [.18,.27]
0.04 [.02,.06]
0.01 [.00,.02]
Appendix
Table 8: Per-domain Target-Only JUMP results on Dream-v0-Base-7B, with 95% bootstrap CIs in brackets.
Exposure
Loss AUC
ZLIB AUC
SAMA AUC
JUMP AUC
JUMP TPR@1%
Epoch 1
0.5020
0.5055
0.5137
0.5263
0.0127
Epoch 2
0.5056
0.5091
0.6600
0.7316
0.0942
Selected ≤4
0.5273
0.5311
0.8193
0.9024
0.4663
Appendix
Table 9: Attack performance as fine-tuning exposure increases. Values are six-domain means.
Exposure
Member loss
Non-member loss
Gap
Held-out gain
Epoch 1
2.4271
2.4432
0.0161
0.0020
Epoch 2
2.3878
2.4179
0.0301
0.0273
Selected ≤4
2.2604
2.3453
0.0850
0.0998
Appendix
Table 10: Member/non-member reconstruction losses for the same exposure controls. “Held-out gain” is the reduction in non-member loss relative to the base model.
Method
NFE
AUC
TPR@10%
TPR@1%
TPR@0.1%
SAMA ( T=2 )
4
0.6283
0.2482
0.0515
0.0197
SAMA ( T=16 )
32
0.7116
0.3660
0.1265
0.0510
JUMP
3
0.7466
0.4192
0.1450
0.0572
Appendix
Table 11: Mixed six-domain target. The primary result is the equal-weight mean over six within-domain ROCs.
Learned selector
AUC
TPR@10%
TPR@1%
TPR@0.1%
LLaDA-8B-Base
0.9024
0.7500
0.4663
0.2320
Distilled 5.36B
0.8966
0.7245
0.4090
0.1885
LLaDA-1.5 (8B-class)
0.9035
0.7470
0.4548
0.3152
LLaDA-8B-Instruct
0.9040
0.7470
0.4655
0.2550
Appendix
Table 12: Selector-backbone transfer with the target and exact-base scoring pair fixed.
Scoring reference
AUC
TPR@10%
TPR@1%
TPR@0.1%
Exact pre-FT LLaDA-8B
0.9024
0.7500
0.4663
0.2320
Distilled 5.36B
0.6052
0.2137
0.0233
0.0067
LLaDA-1.5
0.6949
0.3422
0.0666
0.0182
LLaDA-8B-Instruct
0.6953
0.3448
0.0670
0.0190
Appendix
Table 13: Diagnostic results with non-exact scoring references.
Family
Method
ROC-AUC
TPR@10%
TPR@1%
TPR@0.1%
LLaDA
Random Joint
0.8173
0.5297
0.2073
0.0890
Entropy Top
0.8867
0.7043
0.3650
0.1453
RareToken (C4-freq)
0.8235
0.5489
0.2078
0.0394
SAMA ( T=16 )
0.8193
0.5393
0.2625
0.1220
JUMP
0.9024
0.7500
0.4663
0.2320
Dream
Random Joint
0.8700
0.6337
0.2927
0.1365
Appendix
Table 14: Selector and joint-probing controls across model families (six-domain mean). Joint variants use K=64 and τ=log1.5 ; RareToken uses C4 frequencies.
Selector corpus
Mean ROC-AUC
C4
0.9024
SlimPajama
0.8983
FineWeb
0.8972
Appendix
Table 15: Six-domain mean AUC for matched selector-training corpora.
Masked Diffusion Language Models MDLMs replace autoregressive generation with iterative demasking and their privacy properties are largely unstudied. We study membership inference attacks MIA on fine tuned MDLMs and show they are significantly more vulnerable than current grey box baselines suggest. We extract a 46 dimensional feature vector from the models reconstruction loss at four masking ratios and train XGBoost and MLP classifiers on top. On the MIMIR benchmark across six text domains XGBoost achieves mean AUC 0.878 peaking at 0.930 on Pile CC and beats the SAMA grey box baseline by 0.062 AUC on average. A leave one signal out ablation shows that the ELBO trajectory alone drives most of this with a mean drop of 0.130 when removed while attention features add almost nothing below 0.003. We also design a shadow model transfer attack where K equals 3 surrogate MDLMs trained on data from unrelated domains generate classifier labels with no access to the target domain. This achieves 0.858 mean AUC within 0.020 of the white box oracle and establishes shadow model transfer as a practical and near equally effective attack path.
Membership inference attacks (MIAs) are a canonical way to assess a machine learning model's privacy properties. Although several attempts have been made to evaluate MIAs on language models, the extant literature has suffered numerous difficulties in constructing clean evaluations to test new techniques. In particular, subtle distribution shifts between member and non-member sets can undermine the statistical validity of MIAs; recent work has underscored this by showing that "blind" methods with no access to the underlying model can perform far better than published methods on the same benchmarks. This paper constructs a benchmark for principled evaluation of MIAs against LLMs, by leveraging the insight that training data before and after a fixed point during training are drawn from the same distribution. Therefore, all open-source models with intermediate checkpoints and public training data can be converted into MIA testbeds. We apply our framework to a half-dozen published attacks on the Pythia and OLMo family of models, from 70M to 7B parameters. To facilitate further privacy research, we open-source a modular library for designing and implementing attacks in this setting: https://github.com/safr-ai-lab/pandora_llm.
Diffusion large language models (dLLMs) generate text by iteratively denoising partially masked sequences under bidirectional context, exposing a safety surface distinct from autoregressive LLMs. Because mask tokens are native inputs and tokens are committed by confidence rather than position, harmful content can be induced through infilling and outside the monitored prefix. Existing jailbreaks either miss this native infill capability or rely on low-diversity mask-bearing templates applied uniformly across goals, with little structural adaptation or accumulated attack experience. We propose MaskForge, a fully black-box adaptive attack that casts dLLM red-teaming as optimized search over a growing library of structural patterns. MaskForge abstracts successful attempts into reusable schemas, selects goal-compatible patterns with a UCB bandit, and invokes a scorer-guided fallback when the current library fails. Successful attempts are distilled back into the pattern library, enabling experience to accumulate across goals. Across five public dLLMs and three benchmarks, MaskForge achieves an average attack success rate of 79.3%, a 17.6% relative improvement over the strongest competing dLLM baseline. The matured pattern library further transfers to AdvBench without any updates, achieving a 88.2% attack success rate and a 67% relative improvement over the strongest competing baseline.
Yingzi Ma, Zhengyue Zhao, Xiaogeng Liu +3
University of Wisconsin-Madison · Johns Hopkins University · Responsible AI Research (RAIR) Centre, The University of Adelaide +1