This work presents Diffusion Layer Integrated Gradients (DLIG), a token attribution method for diffusion language models (DLMs) that extends Integrated Gradients (IG~\cite{sundararajan2017axiomatic}) to arbitrary layers and denoising steps. DLIG attributes a DLM's progressive commitment to a self-generated or fixed completion for an input prompt. We establish direct correspondences between DLIG and the IG axioms of completeness, implementation invariance, linearity, and symmetry preservation. As a lightweight complement to interventional analysis, DLIG provides an inexpensive first check of mechanistic hypotheses across the denoising trajectory. We demonstrate this on word-sense disambiguation, multi-hop graph reasoning, and sentence infilling, revealing how DLMs draw on inputs across positions, layers, and denoising steps.
Figures & tables
Figure 1 : Layer-wise DLIG attribution for a single WiC example (target word channel ; gold label same senses , model correctly answers Yes ). Bars are the per-token DLIG score averaged over all T=64 denoising steps. Blue (positive) marks tokens driving the model toward its committed answer, red (negative) those suppressing it. Takeaway: Attribution concentrates on the two occurrences of the pivot word channel and the surrounding content tokens that fix its shared sense, largest in the shallow layers and thinning with depth, while the prompt scaffolding stays near zero.
Figure 2 : Yes- vs. No-responses along the attribution-depth and commitment-timing axes, grouped by predicted label. Takeaway: The two response types read out from essentially the same depth (a) but commit at different times (b), with Yes confident from the first step and No only barely clearing the threshold.
Figure 3 : Attribution on the multi-hop reasoning graph (ProsQA, successes). Takeaway: Positive contrastive mass concentrates on the gold reasoning path well above a random-edge null (a), localizes to the mid-to-upper layers and early denoising steps (b), and rises nearly monotonically along the chain toward the answer edge (c).
Figure 4 : Contrastive magnitude versus attribution location, by outcome. Left: mean per-token ∣d∣ over fact tokens (how much support the facts carry). Right: gold-path precision gap (where that support lands). Takeaway: Attribution localization is preserved across outcomes: failures localize the gold path as well as successes, while showing only a non-significant trend toward higher attribution magnitude.
Figure 5 : Context reliance and attribution shape for sentence infilling (self-generated target). Takeaway: The reliance–quality association is absent early and strengthens across denoising ( 5(a) ); the attribution field is bidirectional and left-skewed, peaking just outside each span boundary, with better infills carrying slightly more mass throughout ( 5(b) ).
Appendix figures & tables19 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6 : Relative completeness error ∣∑DLIG−ΔF∣/∣ΔF∣ at m=1000 over the full grid (layers × denoising steps). Every cell satisfies the combined tolerance εabs+εrel∣ΔF∣ ; darker cells indicate smaller error.
Figure 7 : Partial-forward vs. full-path DLIG, per layer: relative difference against the 10−2 pass/fail threshold. All layers agree exactly.
Task
Outcome
n
% [95% CI]
WiC ( n=500 )
Overall accuracy
318
63.6[59.2,67.8]
Same-sense (Yes) accuracy
227
90.8[87.2,94.4]
Different-sense (No) accuracy
91
36.4[30.4,42.0]
ProsQA ( n=500 )
Success (correct ∪ concept-only)
275
55.0[51.0,59.4]
Fail (wrong option)
64
12.8[10.0,15.8]
Off-manifold (wrong subject / invalid concept)
161
32.2[28.2,36.2]
Appendix
Table 1: Task-level performance of DiffuGPT-M on WiC, ProsQA, and ROCStories infilling. WiC: accuracy on all 500 test examples, split by gold label into same-sense (Yes, n=250 ) and different-sense (No, n=250 ). ProsQA: outcomes over n=500 test examples, bucketed into success (exact-match correct ∪ conceptually correct), fail (wrong option among the two offered), and off-manifold (wrong subject or an invalid/unlisted concept). Infill: ROUGE-1/2/L (self-generated span vs. gold sentence) over n=1000 ROCStories cases.
Bucket
Model generation ( gen )
Gold
correct
Bob is a yerpus. Every yerpus is a zumpusus. ### Bob is a shumpus.
Bob is a shumpus.
concept_only
Bob is a gorpus. Every gorpus is a grimpus. Every grus is a a sterpus. Every sterpus is a numpus. ### ### Bob is a numpus.
Bob is a numpus.
wrong_valid (fail)
Rex is a jelp. Every jelpus is a tumpus. Every tumpus is a wump. Every wumpus is a lempus. ### Rex is a real lempus.
Rex is a rompus.
subj_wrong
Sally is a boompus. Every boompus is a remusus. Every remusus is a sterpus. Every Sally is a sterpus.
Sally is a sterpus.
concept_invalid
Davis is a lorpus. Every lorpus is a numpus. ### Davis is a numpus. Every nump is a worpus. ### Davis is a worpus.
Davis is a worpus.
Appendix
Table 2: One example per ProsQA bucket. correct : exact gold match. concept_only : right subject and option, wrong wording. wrong_valid : right subject, valid option ( lempus ), just not gold. subj_wrong : final sentence names the wrong subject ( Every , a parsing artifact, not a real entity). concept_invalid : right subject, but the reasoning chain lands on a concept ( numpus ) outside the offered pair {worpus, chorpus}.
Figure 8 : Pooled over steps: total context mass vs. infill ROUGE-1 under the self-generated target (per-token normalized); the binned mean is roughly flat. The weak pooled correlation hides the trajectory that strengthens across denoising (main text, Figure 5(a) ).
Analysis
Statistic
Value
Total mass, high- vs. low-ROUGE (pooled; pseudo-replicated, see per-step)
Table 3: Context mass vs. infill quality, pooled statistics (self-generated target, median ROUGE split, nhigh=525 , nlow=475 stories). Per-step AUC in Table 4 .
t
AUC [95% CI]
p
1
0.501 [0.401, 0.606]
0.986
3
0.500 [0.441, 0.560]
0.989
5
0.485 [0.435, 0.536]
0.565
7
0.498 [0.456, 0.545]
0.938
9
0.504 [0.462, 0.547]
0.849
11
0.490 [0.450, 0.530]
0.620
Appendix
Table 4: Mann–Whitney AUC that a high-ROUGE story carries more total context mass than a low-ROUGE one, per denoising step (self-generated target, median ROUGE split, nhigh=525 , nlow=475 stories; 0.5= no difference). 95% bootstrap CIs over stories. No step reaches significance ( p>0.07 throughout).
Figure 9 : Normalization robustness: per-step Pearson r (mass, ROUGE-1) for the self-generated target without per-token normalization ( 95% bootstrap CI). The near-zero-early, strengthening-over-denoising pattern of Fig. 5(a) survives; per-token normalization only sharpens the late-step peak.
Figure 10 : Fixed-target robustness: context reliance vs. infill quality under the fixed gold target (per-token normalized). The correlation is an order of magnitude larger than the self-generated value and, unlike it, does not rise over denoising.
Figure 11 : Fixed-target robustness: mean normalized DLIG magnitude by denoising step and signed distance from the span (fixed gold target). The bidirectional, left-skewed shape matches the self-generated profile (Fig. 5(b) ); only the overall magnitude differs.
Analysis
Statistic
Value
Overall task accuracy
Overall accuracy ( n=500 )
63.6%
Same-sense (Yes) accuracy ( n=250 )
90.8%
Different-sense (No) accuracy ( n=250 )
36.4%
Behavioral divergence: predicted-Yes vs. predicted-No
Attribution depth (centroid): mean pred-Yes / pred-No
7.017 / 7.014
Attribution depth: AUC / p
0.503 / 0.916 (n.s.)
Commitment timing: mean pred-Yes / pred-No
0.073 / 0.271
Appendix
Table 6: Full statistical results for word-sense discrimination on WiC (§ 5 ).
t
zero-mass %
t
zero-mass %
1
93.0
33
7.8
3
84.2
35
7.4
5
75.6
37
6.2
7
67.1
39
5.4
9
59.7
41
4.4
11
51.1
43
3.6
Appendix
Table 7: Fraction of WiC test examples with zero prompt-token attribution mass ( Mt≤10−9 ) at each denoising step (DiffuGPT-M, n=499 parseable predictions). Zero-mass steps occur before the Yes/No span commits, when the self-generated target carries no attributable signal, and are excluded from the attribution-depth centroid. Aggregate over all 32 steps: 22.3% .
Analysis
Statistic
Value
Attribution concentrates on the gold path (successes, pooled)
Precision − chance
+0.150
Wilcoxon p
<10−15∗∗∗
Random-edge null: mean
−0.020
Random-edge null: 95th percentile
+0.075
Faithfulness on fails: mass tracks the model’s cited chain
Cited chain vs. other facts: per-token gap ( n=66 )
+4.77
Cited chain vs. other facts: Wilcoxon p
3.5×10−6∗∗∗
Appendix
Table 8: Full statistical results for multi-hop graph reasoning (§ 6 ).
Analysis
Statistic
Value
ROUGE-1 distribution (self-generated span vs. gold)
Mean / median
0.212 / 0.200
Range
[0.000,0.833]
Self-target attribution tracks output quality when pooled
Ratio(right/left) vs. ROUGE: Pearson r ( n=28514 )
0.027
Mean attribution ratio (right/left)
0.984
Total-mass AUC (pooled, high vs. low ROUGE)
0.510[0.503,0.517]
Total-mass p
3.4×10−3∗∗
Appendix
Table 9: Full statistical results for the ROCStories infilling attribution analysis (§ 7 ), self-generated target mode.
t
n
mean mass
t
n
mean mass
1
124
101.96
33
995
60.93
3
361
99.76
35
995
60.11
5
506
96.04
37
998
59.67
7
622
91.08
39
1000
59.36
9
719
87.21
41
1000
59.00
11
792
83.71
43
1000
58.96
Appendix
Table 10: Mean total context attribution mass per denoising step, ROCStories infilling (self-generated target).
Section
Parameter
Value
Data (WiC)
Train / val / test
4,928 / 638 / 500
Test split
label-stratified hold-out from train (official test labels not public)
Prompt format
“Sentence 1: … Sentence 2: … Does the word “<w>” have the same meaning in both sentences?”, empty system prompt
Target
binary ### Yes. / ### No. ( <eos> -terminated)
Fine-tuning
Objective
diffusion SFT ( ddm-sft ), full fine-tuning
Answer-token loss weight
10× on the Yes / No token (counteracts Yes -bias)
Appendix
Table 11: Data and fine-tuning configuration for the WiC word-sense disambiguation task.
‡School of Information Science and Technology, ShanghaiTech University · §Shanghai Engineering Research Center of Intelligent Vision and Imaging · ¶Tencent
Gaoling School of Artificial Intelligence, Renmin University of China · Department of Data Science, City University of Hong Kong · Beijing Key Laboratory of Research on Large Models and Intelligent Governance +2