This work presents Diffusion Layer Integrated Gradients (DLIG), a token attribution method for diffusion language models (DLMs) that extends Integrated Gradients (IG~\cite{sundararajan2017axiomatic}) to arbitrary layers and denoising steps. DLIG attributes a DLM's progressive commitment to a self-generated or fixed completion for an input prompt. We establish direct correspondences between DLIG and the IG axioms of completeness, implementation invariance, linearity, and symmetry preservation. As a lightweight complement to interventional analysis, DLIG provides an inexpensive first check of mechanistic hypotheses across the denoising trajectory. We demonstrate this on word-sense disambiguation, multi-hop graph reasoning, and sentence infilling, revealing how DLMs draw on inputs across positions, layers, and denoising steps.
Figures & tables
Figure 1 : Layer-wise DLIG attribution for a single WiC example (target word channel ; gold label same senses , model correctly answers Yes ). Bars are the per-token DLIG score averaged over all T=64 denoising steps. Blue (positive) marks tokens driving the model toward its committed answer, red (negative) those suppressing it. Takeaway: Attribution concentrates on the two occurrences of the pivot word channel and the surrounding content tokens that fix its shared sense, largest in the shallow layers and thinning with depth, while the prompt scaffolding stays near zero.
Figure 2 : Yes- vs. No-responses along the attribution-depth and commitment-timing axes, grouped by predicted label. Takeaway: The two response types read out from essentially the same depth (a) but commit at different times (b), with Yes confident from the first step and No only barely clearing the threshold.
Figure 3 : Attribution on the multi-hop reasoning graph (ProsQA, successes). Takeaway: Positive contrastive mass concentrates on the gold reasoning path well above a random-edge null (a), localizes to the mid-to-upper layers and early denoising steps (b), and rises nearly monotonically along the chain toward the answer edge (c).
Figure 4 : Contrastive magnitude versus attribution location, by outcome. Left: mean per-token ∣d∣ over fact tokens (how much support the facts carry). Right: gold-path precision gap (where that support lands). Takeaway: Attribution localization is preserved across outcomes: failures localize the gold path as well as successes, while showing only a non-significant trend toward higher attribution magnitude.
Figure 5 : Context reliance and attribution shape for sentence infilling (self-generated target). Takeaway: The reliance–quality association is absent early and strengthens across denoising ( 5(a) ); the attribution field is bidirectional and left-skewed, peaking just outside each span boundary, with better infills carrying slightly more mass throughout ( 5(b) ).
Appendix figures & tables19 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6 : Relative completeness error ∣∑DLIG−ΔF∣/∣ΔF∣ at m=1000 over the full grid (layers × denoising steps). Every cell satisfies the combined tolerance εabs+εrel∣ΔF∣ ; darker cells indicate smaller error.
Figure 7 : Partial-forward vs. full-path DLIG, per layer: relative difference against the 10−2 pass/fail threshold. All layers agree exactly.
Task
Outcome
n
% [95% CI]
WiC ( n=500 )
Overall accuracy
318
63.6[59.2,67.8]
Same-sense (Yes) accuracy
227
90.8[87.2,94.4]
Different-sense (No) accuracy
91
36.4[30.4,42.0]
ProsQA ( n=500 )
Success (correct ∪ concept-only)
275
55.0[51.0,59.4]
Fail (wrong option)
64
12.8[10.0,15.8]
Off-manifold (wrong subject / invalid concept)
161
32.2[28.2,36.2]
Appendix
Table 1: Task-level performance of DiffuGPT-M on WiC, ProsQA, and ROCStories infilling. WiC: accuracy on all 500 test examples, split by gold label into same-sense (Yes, n=250 ) and different-sense (No, n=250 ). ProsQA: outcomes over n=500 test examples, bucketed into success (exact-match correct ∪ conceptually correct), fail (wrong option among the two offered), and off-manifold (wrong subject or an invalid/unlisted concept). Infill: ROUGE-1/2/L (self-generated span vs. gold sentence) over n=1000 ROCStories cases.
Bucket
Model generation ( gen )
Gold
correct
Bob is a yerpus. Every yerpus is a zumpusus. ### Bob is a shumpus.
Bob is a shumpus.
concept_only
Bob is a gorpus. Every gorpus is a grimpus. Every grus is a a sterpus. Every sterpus is a numpus. ### ### Bob is a numpus.
Bob is a numpus.
wrong_valid (fail)
Rex is a jelp. Every jelpus is a tumpus. Every tumpus is a wump. Every wumpus is a lempus. ### Rex is a real lempus.
Rex is a rompus.
subj_wrong
Sally is a boompus. Every boompus is a remusus. Every remusus is a sterpus. Every Sally is a sterpus.
Sally is a sterpus.
concept_invalid
Davis is a lorpus. Every lorpus is a numpus. ### Davis is a numpus. Every nump is a worpus. ### Davis is a worpus.
Davis is a worpus.
Appendix
Table 2: One example per ProsQA bucket. correct : exact gold match. concept_only : right subject and option, wrong wording. wrong_valid : right subject, valid option ( lempus ), just not gold. subj_wrong : final sentence names the wrong subject ( Every , a parsing artifact, not a real entity). concept_invalid : right subject, but the reasoning chain lands on a concept ( numpus ) outside the offered pair {worpus, chorpus}.
Figure 8 : Pooled over steps: total context mass vs. infill ROUGE-1 under the self-generated target (per-token normalized); the binned mean is roughly flat. The weak pooled correlation hides the trajectory that strengthens across denoising (main text, Figure 5(a) ).
Analysis
Statistic
Value
Total mass, high- vs. low-ROUGE (pooled; pseudo-replicated, see per-step)
Table 3: Context mass vs. infill quality, pooled statistics (self-generated target, median ROUGE split, nhigh=525 , nlow=475 stories). Per-step AUC in Table 4 .
t
AUC [95% CI]
p
1
0.501 [0.401, 0.606]
0.986
3
0.500 [0.441, 0.560]
0.989
5
0.485 [0.435, 0.536]
0.565
7
0.498 [0.456, 0.545]
0.938
9
0.504 [0.462, 0.547]
0.849
11
0.490 [0.450, 0.530]
0.620
Appendix
Table 4: Mann–Whitney AUC that a high-ROUGE story carries more total context mass than a low-ROUGE one, per denoising step (self-generated target, median ROUGE split, nhigh=525 , nlow=475 stories; 0.5= no difference). 95% bootstrap CIs over stories. No step reaches significance ( p>0.07 throughout).
Figure 9 : Normalization robustness: per-step Pearson r (mass, ROUGE-1) for the self-generated target without per-token normalization ( 95% bootstrap CI). The near-zero-early, strengthening-over-denoising pattern of Fig. 5(a) survives; per-token normalization only sharpens the late-step peak.
Figure 10 : Fixed-target robustness: context reliance vs. infill quality under the fixed gold target (per-token normalized). The correlation is an order of magnitude larger than the self-generated value and, unlike it, does not rise over denoising.
Figure 11 : Fixed-target robustness: mean normalized DLIG magnitude by denoising step and signed distance from the span (fixed gold target). The bidirectional, left-skewed shape matches the self-generated profile (Fig. 5(b) ); only the overall magnitude differs.
Analysis
Statistic
Value
Overall task accuracy
Overall accuracy ( n=500 )
63.6%
Same-sense (Yes) accuracy ( n=250 )
90.8%
Different-sense (No) accuracy ( n=250 )
36.4%
Behavioral divergence: predicted-Yes vs. predicted-No
Attribution depth (centroid): mean pred-Yes / pred-No
7.017 / 7.014
Attribution depth: AUC / p
0.503 / 0.916 (n.s.)
Commitment timing: mean pred-Yes / pred-No
0.073 / 0.271
Appendix
Table 6: Full statistical results for word-sense discrimination on WiC (§ 5 ).
t
zero-mass %
t
zero-mass %
1
93.0
33
7.8
3
84.2
35
7.4
5
75.6
37
6.2
7
67.1
39
5.4
9
59.7
41
4.4
11
51.1
43
3.6
Appendix
Table 7: Fraction of WiC test examples with zero prompt-token attribution mass ( Mt≤10−9 ) at each denoising step (DiffuGPT-M, n=499 parseable predictions). Zero-mass steps occur before the Yes/No span commits, when the self-generated target carries no attributable signal, and are excluded from the attribution-depth centroid. Aggregate over all 32 steps: 22.3% .
Analysis
Statistic
Value
Attribution concentrates on the gold path (successes, pooled)
Precision − chance
+0.150
Wilcoxon p
<10−15∗∗∗
Random-edge null: mean
−0.020
Random-edge null: 95th percentile
+0.075
Faithfulness on fails: mass tracks the model’s cited chain
Cited chain vs. other facts: per-token gap ( n=66 )
+4.77
Cited chain vs. other facts: Wilcoxon p
3.5×10−6∗∗∗
Appendix
Table 8: Full statistical results for multi-hop graph reasoning (§ 6 ).
Analysis
Statistic
Value
ROUGE-1 distribution (self-generated span vs. gold)
Mean / median
0.212 / 0.200
Range
[0.000,0.833]
Self-target attribution tracks output quality when pooled
Ratio(right/left) vs. ROUGE: Pearson r ( n=28514 )
0.027
Mean attribution ratio (right/left)
0.984
Total-mass AUC (pooled, high vs. low ROUGE)
0.510[0.503,0.517]
Total-mass p
3.4×10−3∗∗
Appendix
Table 9: Full statistical results for the ROCStories infilling attribution analysis (§ 7 ), self-generated target mode.
t
n
mean mass
t
n
mean mass
1
124
101.96
33
995
60.93
3
361
99.76
35
995
60.11
5
506
96.04
37
998
59.67
7
622
91.08
39
1000
59.36
9
719
87.21
41
1000
59.00
11
792
83.71
43
1000
58.96
Appendix
Table 10: Mean total context attribution mass per denoising step, ROCStories infilling (self-generated target).
Section
Parameter
Value
Data (WiC)
Train / val / test
4,928 / 638 / 500
Test split
label-stratified hold-out from train (official test labels not public)
Prompt format
“Sentence 1: … Sentence 2: … Does the word “<w>” have the same meaning in both sentences?”, empty system prompt
Target
binary ### Yes. / ### No. ( <eos> -terminated)
Fine-tuning
Objective
diffusion SFT ( ddm-sft ), full fine-tuning
Answer-token loss weight
10× on the Yes / No token (counteracts Yes -bias)
Appendix
Table 11: Data and fine-tuning configuration for the WiC word-sense disambiguation task.
Diffusion Language Models (DLMs) generate text by iteratively denoising a masked sequence, independently predicting multiple tokens at each step. This conditional independence discards inter-token dependencies and degrades coherence-an issue that parallels the multi-modality problem in Non-Autoregressive Translation (NAT). Drawing on the Directed Acyclic Transformer (DAT), which tackles this problem in NAT via a Directed Acyclic Graph (DAG), we propose DA-DLM, a model that adapts DAG-based dependency modeling to DLMs' iterative setting through a position-oriented DAG design. The position-oriented DAG binds node groups to fixed output positions so that tokens fixed in earlier steps anchor neighboring predictions via learned transitions, and evolves with denoising to focus on remaining uncertainty as anchors accumulate. On language modeling, open-ended generation, and summarization, DA-DLM consistently outperforms Block Diffusion, especially under fewer denoising steps, and matches autoregressive models while preserving the parallel generation advantage. Our code is publicly available at https://github.com/jipy0222/DA-DLM.
Pengyu Ji, Zichen Zhang, Xiang Hu +1
‡School of Information Science and Technology, ShanghaiTech University · §Shanghai Engineering Research Center of Intelligent Vision and Imaging · ¶Tencent
Diffusion large language models (dLLMs) offer an efficient alternative to autoregressive models through parallel decoding, yet existing post-training methods largely rely on random masking strategies that overlook intrinsic token dependencies. In this work, we present an empirical analysis of attention in dLLMs and show that tokens attending more strongly to unmasked context exhibit greater generation stability and play a critical role in reasoning. Motivated by these findings, we propose AGDO, an attention-guided denoising and optimization framework that aligns both training and optimization with attention-derived dependencies. AGDO determines the denoising order based on attention structure and emphasizes attention-critical tokens during supervised fine-tuning and reinforcement learning. Experiments on mathematical and coding benchmarks demonstrate that AGDO consistently improves reasoning performance, outperforming state-of-the-art post-training methods for dLLMs.
Jia Deng, Junyi Li, Wayne Xin Zhao +3
Gaoling School of Artificial Intelligence, Renmin University of China · Department of Data Science, City University of Hong Kong · Beijing Key Laboratory of Research on Large Models and Intelligent Governance +2
Diffusion language models expose an explicit denoising trajectory, making it possible to ask when different kinds of information become measurable during generation. We study three independent 32-step runs of LLaDA-8B-Base on masked WikiText-103 text, each with 1{,}000 probe-training sequences and 200 held-out evaluation sequences. From saved trajectories, we derive four temporal measurements: token commitment; linear recoverability of part-of-speech (POS), coarse semantic category, and token identity; confidence and entropy dynamics; and sensitivity under mid-trajectory re-masking. Across seeds, the same ordering recurs: content categories stabilize earlier than function-heavy categories, POS and coarse semantic labels remain substantially more linearly recoverable than exact lexical identity under our probe setup, uncertainty remains higher for tokens that ultimately resolve incorrectly even though late confidence becomes less calibrated, and perturbation sensitivity peaks in the middle of the trajectory. A direct/collateral decomposition shows that this peak is overwhelmingly local to the perturbed positions themselves. In this LLaDA+WikiText setting, denoising time is therefore a useful analysis axis: under our measurements, coarse labels are recovered earlier and more robustly than lexical identity, trajectory-level uncertainty tracks eventual correctness, and mid-trajectory states are the most intervention-sensitive.
Harry Lu
School of Mathematics, University of Minnesota, Minneapolis MN, United States.