Masked diffusion language models (dLLMs) generate text by iteratively denoising masked positions, re-predicting each token multiple times before it is committed. An autoregressive decoder exposes an answer's distribution once, at the step that commits it; a dLLM exposes it at every denoising step before commitment, and we show that an adversary can exploit this. Since an answer remains open to revision over many denoising steps, an adversary with access to internal activations can watch how likely the model is to produce a chosen answer and adjust the intervention accordingly. Building on this observation, we study targeted bias injection, an attack that steers a frozen dLLM toward a demographic answer selected by the adversary. The attack uses a simple proportional-integral (PI) controller that tracks the target-answer probability during denoising and adapts the strength of a steering vector on the fly. On ambiguous BBQ questions where the correct answer is abstention, our attack raises LLaDA-8B-Instruct's preference for the targeted group from 1.8 to 16.7 percentage points, more than three times the strongest fixed-strength steering baseline, and on SocialStigmaQA it raises the selection of stigmatizing answers from 17.6% to 58.1%. Fitted to other demographic targets, the same attack shifts answers by up to 37 percentage points, and each attack takes about 40 minutes on one GPU. On the primary target, feedback is what makes the attack work: constant steering at the same average strength over the token-committing steps produces a far smaller shift while corrupting nearly three times as many outputs, and a constant strength set separately for each example still falls well short. Our findings identify the denoising trajectory as a new control channel in dLLMs and call for bias audits that examine the serving stack rather than the frozen model alone.
Figures & tables
Figure 1: Feedback assigns each example its own steering strength. Primary Black-target PI run on LLaDA-8B-Instruct, 1,200 sequences, split by strict outcome. (a) Each point is one sequence: its target probability after the first controlled step against its mean command over steps 1–32; the dashed line is the effort-matched constant. (b,c) Mean target probability and command per step with 95% bands. Shading marks steps after the last token commits.
Figure 2: Feedback during denoising steers a frozen model toward an unsupported demographic answer. Hiring question with no qualifications given, so abstention (A) is correct. (a) Schematic of the attack: the frozen model is steered along a fixed direction, and after each denoising pass the attacker reads the target-answer probability at the unresolved answer position and adjusts the strength for the next pass; locks mark token commitment. (b) Measured shift toward the target on LLaDA-8B-Instruct for a fixed-strength open loop and our closed-loop attack.
Figure 3: Closed-loop bias injection on LLaDA-8B-Instruct. (a) Held-out pairs complete the same question with comparator and target option texts; their mean answer-token residual difference at block 14 defines the fixed unit direction v . (b) At forward t , PI control uses the preceding target-answer probability pt−1 to compute a clipped strength αt . The offset αtv is added at every position in all 32 frozen blocks; the sensor reads pt at the first generated position, underlined in blue. Commitment fixes the token but does not stop the controller. (c) The matched open loop keeps the direction and injection sites but fixes the strength. Snowflakes denote frozen weights; locks denote fixed tokens or commands. Token states and activation clusters are schematic.
Method
Tgt. (%)
Cmp. (%)
Abst. (%)
Inv. (%) ↓
Gap (pp) ↑
Δg (pp)
Unsteered base
12.8
11.0
76.3
0.0
1.8
—
CAA
13.7
10.3
76.0
0.0
3.3
+1.6
ActAdd
14.1
10.0
75.9
0.0
4.1
+2.3
Mean-AcT
21.9
18.4
59.4
0.3
3.5
+1.7
Local Gaussian AcT
25.6
23.0
51.3
0.1
2.6
+0.8
ITI-C
18.8
16.7
64.6
0.0
2.1
+0.3
Table 1: Black-target BBQ: 400 items, three answer rotations. Tgt., Cmp., and Abst. are target, comparator, and abstention rates; Inv. is the invalid-output rate, retained in all denominators. Bold marks the best gap. Baseline operating points: Figure 5 ; 95% CIs: Figure 6 and Figure 4 .
Method
Tgt. (%)
Cmp. (%)
Abst. (%)
Inv. (%) ↓
Δg (pp) ↑
Δgsem (pp) ↑
Unsteered base
17.6
61.9
20.5
0.0
—
—
CAA
21.0
56.8
22.3
0.0
+8.6
+8.6
ActAdd
20.7
56.4
22.9
0.0
+8.6
+8.6
Mean-AcT
32.5
32.9
34.6
0.0
+44.0
+44.0
Local Gaussian AcT
0.0
0.0
0.0
100.0
+44.3
+47.4
ITI-C
42.5
42.9
14.5
0.0
+44.0
+44.0
Table 2: SocialStigmaQA: 518 scored items, three answer rotations. Columns as in Table 1 ; Δg and Δgsem use the strict and semantic parsers. The unsteered gap is −44.3 pp, so all-invalid rows (Inv. =100 ) score Δg=44.3 by construction and are failures, not attacks. Bold marks the best Δg per column. Operating points, CIs, and audits: Appendix C.8 .
Setting
Method
Abst. (%)
Inv. (%) ↓
Gap (pp) ↑
Δg (pp)
Arab, LLaDA-8B
Unsteered base
69.8
0.0
−0.2
—
Open loop ( α=4 )
44.1
0.0
12.8
+12.9
Decode PI (ours)
16.0
9.7
22.3
+22.5
Old, LLaDA-8B
Unsteered base
50.2
0.0
−5.5
—
Open loop ( α=4 )
6.0
19.0
25.8
+31.3
Decode PI (ours)
9.9
9.3
31.8
+37.2
Table 3: Attack generality: other demographic targets and models. Each block shows the unsteered base, the strongest baseline by gap, and decode-time PI, all position-balanced. Bold marks each block’s best gap; full method grids and CIs are in Appendix C .
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
Quantity
Value (pp)
95% CI (pp)
Black, calibration items removed
PI Δg
14.1
[10.0,18.4]
Black, calibration items removed
PI − ActAdd (item-clustered)
12.1
[8.1,16.3]
White
Open loop ( α=1 ) Δg
2.5
[0.7,4.3]
Asian
Open loop ( α=2 ) Δg
1.0
[−1.0,3.0]
Black, Dream-7B
ITI-C Δg
−0.8
[−6.6,4.9]
Black, LLaDA-MoE
PI (block 8) − CAA
−2.3
[−3.6,−0.9]
Appendix
Table 4: Point estimates and 95% intervals for comparisons quoted in the appendix text. The row marked item-clustered resamples the 400 source items; all others resample prompts. Differences are PI minus the named method.
Figure 4: Item-clustered 95% intervals for the paired comparisons stated in the body. Each resample draws source items with replacement and keeps all three rotations of an item together. Filled blue: PI ahead with the interval above zero. Hollow grey: the interval contains zero. Orange: PI behind. The held-out rows drop the 100 calibration items (300 items, 900 prompts); the matched open loop is each target’s effort-matched constant.
Figure 5: Black-target outcomes on LLaDA-8B-Instruct (Table 1 ). Left: shares of target, comparator, abstention, and invalid outputs over 1,200 prompts under the strict parser. Right: target–comparator gap with item-clustered 95% intervals. Operating points are CAA and ActAdd at α=16 , Mean-AcT (unit, s=2 ), Local Gaussian AcT ( s=1 ), ITI-C ( K=48 , α=8 ), AurA-inspired amplification ( γ=4 ), and layer-wise PI ( α=2 ). The static arms hold one α per item for all steps (Appendix C.7 ).
Figure 6: Black-target comparisons and cross-axis effects. Points show the gap g with 95% prompt-level bootstrap intervals; the axis is in fractional units, so 0.1 equals 10 pp. Left: decode-time control against every baseline. Right: PI and base gaps on Black, Arab, and Old.
Figure 7: Decode-time PI and the largest-shift comparison arm in each setting. Each row places PI (blue diamond) beside the tested arm with the largest change from the unsteered gap (grey circle, labelled), selected from the prior-work baselines, fixed open loop, and effort-matched constant. The connecting line is blue where PI is ahead and orange where it is behind. Hollow markers flag strict invalidity above 15%, where the measured gap is not treated as a valid attack. On LLaDA-MoE, CAA at α=2 is identical to the matched constant. Paired intervals are in Figure 4 .
Figure 8: SocialStigmaQA parser sensitivity across methods. Left and center: change from the unsteered biased–safe gap under strict, semantic (answer-text), and strict-valid scoring, with 95% template-clustered bootstrap intervals. Right: the corresponding invalid-output rates. The answer-text parser accepts an initial exact option text and rejects responses that re-list multiple options. Headline answer rates are in Table 2 .
Diffusion large language models (dLLMs) generate text through iterative denoising rather than left-to-right decoding. This generation paradigm introduces two axes that can influence safety alignment: when tokens are generated during denoising and where they appear in the response. In this paper, we measure dLLM safety behavior under harmful prompts by tracing intermediate token distributions and commitment decisions throughout denoising. Our analysis shows that refusal signals are concentrated in early denoising steps and leading response positions, and the tokens committed early can strongly shape the final safety outcome. Our measurements further show that the denoising step and persistence of refusal-token commitment are important for understanding dLLM safety. Based on these findings, we propose Refusal-Aware Early Commitment (RAEC), a simple training-free decoding method that commits persistent refusal signals from early steps. Experiments on LLaDA and Dream show that RAEC reduces attack success rates while largely preserving utility. The code is available at https://github.com/Glresearch1/RAEC.
Guoli Wang, Haonan Shi, Tu Ouyang +1
Case Western Reserve University 10900 Euclid Avenue, Cleveland, OH 44106, USA
Discrete Masked diffusion language models generate text by iterative parallel decoding, but few-step decoding suffers from a tradeoff between length and quality: with a fixed step budget, standard methods can generate a short, high-quality output, or they can produce long but repetitive text. Continuous denoising can sidestep this tradeoff by evolving all positions jointly in embedding space, but building such a model from scratch at scale remains an open problem. We show that a pretrained masked DLM can instead be lightly adapted to support continuous embedding-space denoising. Starting from LLaDA-8B-Instruct, we continue-pretrain for only 1,000 steps with Discrete Stochastic Localization (DSL), replacing binary masking with continuous per-token Gaussian noise as a soft mask. The adapted model supports continuous inference that evolves all positions jointly in embedding space and defers hard token commitment to the final step. On zero-shot summarization at low step budgets (<=16 forward passes), DSL-LLaDA-SDE achieves the best ROUGE-1 on all four benchmarks and largely avoids the premature-termination / repetition tradeoff of iterative unmasking. The same adaptation also yields selective noisy-state robustness: the model corrects corrupted tokens while preserving clean ones. Control experiments using standard masked diffusion training with the same compute demonstrate neither behavior.
Longxuan Yu, Yunshu Wu, Yu Fu +5
University of California, Riverside · Georgia Institute of Technology · Microsoft
Diffusion large language models (dLLMs) offer an efficient alternative to autoregressive models through parallel decoding, yet existing post-training methods largely rely on random masking strategies that overlook intrinsic token dependencies. In this work, we present an empirical analysis of attention in dLLMs and show that tokens attending more strongly to unmasked context exhibit greater generation stability and play a critical role in reasoning. Motivated by these findings, we propose AGDO, an attention-guided denoising and optimization framework that aligns both training and optimization with attention-derived dependencies. AGDO determines the denoising order based on attention structure and emphasizes attention-critical tokens during supervised fine-tuning and reinforcement learning. Experiments on mathematical and coding benchmarks demonstrate that AGDO consistently improves reasoning performance, outperforming state-of-the-art post-training methods for dLLMs.
Jia Deng, Junyi Li, Wayne Xin Zhao +3
Gaoling School of Artificial Intelligence, Renmin University of China · Department of Data Science, City University of Hong Kong · Beijing Key Laboratory of Research on Large Models and Intelligent Governance +2