Uniform-state diffusion models (USDMs) can revise any token at any denoising step, which lets them correct their own mistakes, a key advantage over masked diffusion. Self-correction, however, requires both revising incorrect tokens and retaining correct ones, and we show that current USDMs lack the latter. Even under greedy-tail decoding, state-of-the-art USDMs (DUO, UDLM, and uniform-noise SEDD) keep revising 173--270 of 512 positions at every step, and these large, uncoordinated edits collapse sample diversity. A random-token corruption experiment traces this deficit to the models themselves: they reconstruct clean and corrupted tokens with nearly identical accuracy, even though clean tokens are easier targets. A decomposition of the validation NELBO shows that training barely rewards retention: incorrect predictions are heavily penalized at corrupted positions but almost free at clean ones. We propose Correct-Token Retention Regularization (CTR-Reg), a simple but effective auxiliary loss that trains the model to retain tokens left unperturbed by the forward process and requires no change to the sampler. CTR-Reg improves clean-token accuracy by 26.5 percentage points on average across six benchmarks, while leaving corrupted-token accuracy virtually unchanged, and its per-step revisions converge to only 3--11 positions. With just five greedy-tail steps, generative perplexity more than halves under CTR-Reg for all three models while diversity is preserved, and these gains hold across sampling budgets. Our results identify correct-token retention as a key missing ingredient for self-correcting diffusion language models, and demonstrate an effective fix.
Figures & tables
Figure 1: Effect of CTR-Reg on corrupted- vs. clean-token reconstruction accuracy under random-token perturbation. For SEDD, UDLM, and DUO, we compare base NELBO training (orange) to NELBO + Correct-Token Retention Regularization (CTR-Reg; blue), measuring reconstruction accuracy on corrupted tokens (left) and clean tokens (right), averaged over corruption rates p∈{10,20,30,40}% . Each panel shows OpenWebText and arXiv results, separated by a dotted line. CTR-Reg leaves corrupted-token accuracy essentially unchanged while improving clean-token accuracy by 21.5–29.3 percentage points, across all models and both corpora. All models are trained on 20B OpenWebText tokens (169M parameters, sequence length 512); evaluation uses the OpenWebText validation set ( ∼ 110M tokens) and Scientific Papers (arXiv) validation set ( ∼ 55M tokens).
Figure 2
Figure 4: PPL–entropy frontier vs. greedy-tail steps. Unconditional generation with DUO, UDLM, and SEDD (one per panel) at a fixed budget of 512 steps, sweeping the number of terminal greedy-tail steps g from 0 (fully ancestral) to 20 (annotated at each point); PPL is measured with GPT2-Large over 1024 samples. The dashed line marks real OpenWebText entropy (4.83), below which lower PPL reflects degenerate repetition rather than genuine quality. Greedy-tail steps collapse diversity in NELBO-only models, which fail to preserve correct tokens, whereas CTR-Reg shifts the frontier toward lower PPL at higher entropy and reaches high-quality generations within a few steps.
Figure 5: The NELBO barely penalizes failures of retention. Mean validation NELBO per token of NELBO-only models on OpenWebText (dark) and arXiv (light) versus the noise level 1−αt (fraction of corrupted tokens), for clean (blue) and corrupted (orange) positions where the argmax prediction x^θ equals the clean token x (solid) or not (dashed). The solid–dashed gap is the likelihood penalty for an incorrect prediction: large at corrupted positions, near zero at clean positions except at the highest noise. Magnitudes are comparable within, not across, panels.
Table 5
Appendix figures & tables22 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: Generation quality across sampling budgets (16–1024 steps). PPL ( ↓ ) vs. total sampling steps for DUO, UDLM, and SEDD trained with NELBO or NELBO + CTR-Reg, under default (DUO: +1 greedy step) and greedy-tail sampling (final three steps greedy); each point is annotated with its generation entropy. Under the default sampler, CTR-Reg matches or outperforms NELBO-only training; under greedy-tail sampling, CTR-Reg substantially lowers PPL while preserving diversity, whereas NELBO-only models lower PPL only at the cost of diversity. These gains hold across the full step-budget range, including the ultra-low-step regime.
αt
stay
jump-to-prediction
jump-elsewhere
0.01
0.835
0.002
0.163
0.50
0.992
0.004
0.004
0.99
0.803
0.195
0.002
Appendix
Table 2: Single-step transition probabilities at a disagreeing position under the log-linear schedule ( αt≈1−t ) with T=512 ( h≈1.95×10−3 ).
In-dist.
Zero-shot
Model
Loss
OWT
LM1B
AG News
arXiv
PubMed
WT-103
Avg.
DUO
NELBO
36.59
93.21
114.21
62.53
74.81
51.31
72.11
+ CTR-Reg
36.74
93.66
114.47
64.21
78.09
51.78
73.16
Δ
+0.4%
+0.5%
+0.2%
+2.7%
+4.4%
+0.9%
+1.5%
UDLM
NELBO
40.31
128.92
139.37
77.24
91.03
62.80
89.95
+ CTR-Reg
40.65
129.84
143.68
77.37
91.07
63.69
91.05
Appendix
Table 3: Validation perplexity ( ↓ ) of NELBO-only and CTR-Reg models. Perplexity upper bound computed from the validation NELBO alone; the CTR-Reg term is excluded at evaluation. All models have 169M parameters and are trained on 20B OpenWebText (OWT) tokens with sequence length 512. Δ : relative change of NELBO + CTR-Reg with respect to NELBO-only. CTR-Reg changes perplexity by at most a few percent, and lowers it on average for SEDD.
Figure 8: The NELBO barely penalizes failures of retention on the remaining corpora. Mean validation NELBO per token of NELBO-only DUO, UDLM, and SEDD on LM1B and AG News (top) and PubMed and WikiText-103 (bottom), with the first corpus of each pair dark and the second light, versus the noise level 1−αt (fraction of corrupted tokens). Lines show clean (blue) and corrupted (orange) positions where the argmax prediction x^θ equals the clean token x (solid) or not (dashed), so the solid–dashed gap is the likelihood penalty for an incorrect prediction. As on OpenWebText and arXiv (Figure 5 ), this penalty is large at corrupted positions and near zero at clean ones, except at the highest noise. Magnitudes are comparable within, not across, panels.
Mean NELBO per token (nats)
Number of tokens
Clean
Corrupted
Clean
Corrupted
Corpus
1−αt
x^θ=x †
x^θ=x
x^θ=x
x^θ=x
x^θ=x
x^θ=x
x^θ=x
x^θ=x
OpenWebText
[0.0,0.1)
0.019
0.1587
4.7935
84.0081
7,062,434
3,572,467
378,832
186,779
[0.1,0.2)
0.027
0.1111
1.7759
29.8928
6,139,616
3,545,775
1,074,039
632,058
[0.2,0.3)
0.040
0.1255
1.2435
18.8710
5,059,097
3,420,197
1,627,315
1,195,791
[0.3,0.4)
0.063
0.1586
1.0497
14.1688
4,087,763
3,298,748
2,052,332
1,915,781
Appendix
Table 4: Exact values underlying Figures 5 and 8 for DUO. Mean validation NELBO per token (nats) of the NELBO-only model and the number of tokens behind each mean, by corpus and noise-level bin 1−αt (fraction of corrupted tokens), for clean and corrupted positions where the argmax prediction x^θ equals the clean token x or not. † In units of 10−3 nats. Gray : fewer than 1,000 tokens (noisier estimate).
Mean NELBO per token (nats)
Number of tokens
Clean
Corrupted
Clean
Corrupted
Corpus
1−αt
x^θ=x †
x^θ=x
x^θ=x
x^θ=x
x^θ=x
x^θ=x
x^θ=x
x^θ=x
OpenWebText
[0.0,0.1)
0.018
0.2514
4.8082
86.1054
7,101,661
3,543,893
374,566
191,656
[0.1,0.2)
0.025
0.1482
1.8055
30.4774
6,149,687
3,543,007
1,056,380
651,118
[0.2,0.3)
0.038
0.1566
1.2567
19.2184
5,063,872
3,423,035
1,601,837
1,223,896
[0.3,0.4)
0.060
0.1894
1.0628
14.4402
4,078,135
3,314,353
2,011,194
1,960,158
Appendix
Table 5: Exact values underlying Figures 5 and 8 for UDLM. Mean validation NELBO per token (nats) of the NELBO-only model and the number of tokens behind each mean, by corpus and noise-level bin 1−αt (fraction of corrupted tokens), for clean and corrupted positions where the argmax prediction x^θ equals the clean token x or not. † In units of 10−3 nats. Gray : fewer than 1,000 tokens (noisier estimate).
Mean NELBO per token (nats)
Number of tokens
Clean
Corrupted
Clean
Corrupted
Corpus
1−αt
x^θ=x †
x^θ=x
x^θ=x
x^θ=x
x^θ=x
x^θ=x
x^θ=x
x^θ=x
OpenWebText
[0.0,0.1)
0.006
0.1122
6.9664
60.0927
44,975,769
18,586,155
570,010
382,626
[0.1,0.2)
0.094
0.4139
5.6180
57.0556
3,958,369
2,032,886
559,789
454,652
[0.2,0.3)
0.224
0.6183
5.8430
55.4429
2,015,786
1,270,128
539,883
536,443
[0.3,0.4)
0.441
0.8860
6.0644
53.4810
1,204,759
953,960
513,454
640,979
Appendix
Table 6: Exact values underlying Figures 5 and 8 for SEDD. Mean validation NELBO per token (nats) of the NELBO-only model and the number of tokens behind each mean, by corpus and noise-level bin 1−αt (fraction of corrupted tokens), for clean and corrupted positions where the argmax prediction x^θ equals the clean token x or not. † In units of 10−3 nats. Gray : fewer than 1,000 tokens (noisier estimate).
Accuracy (%) by Evaluation Dataset
Evaluated On
Setting
OWT
AG News
arXiv
LM1B
PubMed
WT-103
Avg
Corrupted-Tokens
DUO (20B)
54.5
41.5
48.4
35.2
46.3
45.5
45.2
+ CTR-Reg (20B)
55.0
41.8
49.7
36.5
47.7
47.2
46.3
DUO (524B)
60.7
48.8
54.4
39.9
53.6
52.3
51.6
Clean-Tokens
DUO (20B)
53.9
40.7
52.7
36.3
50.2
49.1
47.2
+ CTR-Reg (20B)
83.2
72.5
79.6
69.9
78.4
80.8
77.4
Appendix
Table 7: Scaling training data alone does not close the retention gap. Corrupted-token, clean-token, and overall reconstruction accuracy (%, averaged over corruption rates p∈{10,20,30,40}% ) on six evaluation corpora, comparing our from-scratch DUO trained on 20B tokens (NELBO-only, black), the same 20B-token model with NELBO + CTR-Reg (purple), and the officially released DUO checkpoint trained on 524B tokens (blue, ∼ 26 × more data, native sequence length 1024). The bolded value in each column is the highest of the three settings. Scaling to 524B tokens substantially improves corrupted-token accuracy but leaves clean-token accuracy nearly as low as the 20B baseline, so the clean-minus-corrupted gap does not close with scale; CTR-Reg opens a large gap using 26× less data. See Table 18 for the per-corruption-rate breakdown underlying the 524B-token row.
Greedy-Tail Steps ( g )
0
1
2
3
4
5
6
8
10
12
14
16
18
20
PPL ( ↓ )
77.77
71.14
71.91
68.67
57.55
47.61
36.01
22.76
16.19
12.56
11.22
10.44
9.97
9.90
Entropy
5.54
5.21
4.89
4.58
4.26
3.93
3.63
3.07
2.83
2.61
2.49
2.34
2.21
2.14
Appendix
Table 8: Greedy-tail sampling collapses diversity at scale, too. PPL (GPT2-Large, 1024 generated samples) and generation entropy for the officially released DUO checkpoint (524B training tokens, sequence length 1024, total step budget 512) across greedy-tail step counts g . Entropy is colored green when it meets or exceeds the real-OpenWebText reference entropy at sequence length 1024 ( 5.40 ) and red when it falls short, following the same convention as Table 10 (whose reference entropy, 4.83 , is computed at sequence length 512 and is therefore not the value used here). A red entropy means the corresponding PPL was achieved via degenerate, low-diversity generation rather than genuine quality. The bolded PPL is the lowest across all g . Because this checkpoint’s native sequence length (1024) differs from our 20B-token models (512), PPL and entropy values here are not directly comparable in magnitude to Table 10 ; only the qualitative trend, and each model’s standing relative to its own reference entropy, should be compared.
Accuracy (%) by Evaluation Dataset
Evaluated On
Model
Loss
OWT
AG News
arXiv
LM1B
PubMed
WT-103
Avg
Corrupted-Tokens
DUO
NELBO
54.5
41.5
48.4
35.2
46.3
45.5
45.2
NELBO + CTR-Reg
55.0
41.8
49.7
36.5
47.7
47.2
46.3
SEDD
NELBO
47.2
33.6
43.1
28.9
39.7
39.0
38.6
NELBO + CTR-Reg
47.1
32.6
42.6
28.9
38.5
38.1
38.0
UDLM
NELBO
55.3
42.8
47.4
38.8
45.4
47.9
46.3
Appendix
Table 9: Generalization of CTR-Reg across evaluation corpora. For SEDD, UDLM, and DUO, we compare NELBO-only training (black) to NELBO + CTR-Reg (purple) on six corpora: OWT, AG News, arXiv, LM1B, PubMed, and WT-103, reporting corrupted-token, clean-token, and overall reconstruction accuracy averaged over p∈{10,20,30,40}% ; Avg is the mean across all six corpora, and the bolded value in each pair is higher. CTR-Reg leaves corrupted-token accuracy essentially unchanged (0.7 points per cell on average) while substantially improving clean-token and overall accuracy across every dataset and model. All models are trained on 20B OpenWebText tokens (169M parameters, sequence length 512); evaluation uses each corpus’s validation split.
NELBO
NELBO + CTR-Reg
Model
Greedy-Tail Steps ( g )
PPL ( ↓ )
Entropy
PPL ( ↓ )
Entropy
DUO
0
105.37
5.18
105.86
5.19
1
110.88
4.96
64.67
5.09
2
118.86
4.87
56.93
5.08
3
112.91
4.72
53.32
5.06
4
102.94
4.57
51.16
5.06
Appendix
Table 10: Exact PPL and entropy values underlying Figure 4 . For each of three uniform-state diffusion language models and each greedy-tail step count g∈{0,1,2,3,4,5,6,8,10,12,14,16,18,20} (total step budget fixed at 512), we report perplexity (PPL, computed with GPT2-Large over 1024 generated samples; bolded value is the lower, i.e. better, of the two) and generation entropy for the NELBO-only baseline and for NELBO + CTR-Reg . Entropy values are colored green when they meet or exceed the real-OpenWebText reference entropy ( 4.83 ) and red when they fall short, matching Figure 4 ; a red entropy means the corresponding PPL was achieved via degenerate, low-diversity generation rather than genuine quality.
Standard Ancestral Sampling
Greedy-Tail Sampling (final 3 steps)
NELBO
NELBO + CTR-Reg
NELBO
NELBO + CTR-Reg
Model
Steps
PPL
Ent.
PPL
Ent.
PPL
Ent.
PPL
Ent.
DUO
16
147.50
4.94
80.53
5.08
126.75
4.56
63.63
5.04
32
124.84
4.97
71.23
5.09
122.84
4.69
56.90
5.06
64
117.41
4.97
67.88
5.10
119.18
4.71
55.10
5.06
128
111.73
4.97
66.25
5.09
118.33
4.73
54.03
5.07
Appendix
Table 11: Exact PPL and entropy values underlying Figure 7 . For each of three uniform-state diffusion language models and each total sampling-step budget from 16 to 1024, we report perplexity (PPL, computed with GPT2-Large; bolded value is the lower, i.e. better, of the two) and generation entropy for the NELBO-only baseline and for NELBO + CTR-Reg , under standard ancestral sampling (DUO’s default sampler includes one trailing greedy step) and under greedy-tail sampling (final three steps fixed to be greedy). Entropy values are colored green when they meet or exceed the real-OpenWebText reference entropy ( 4.83 ) and red when they fall short. Under greedy-tail sampling, the NELBO-only baseline’s lower PPL is consistently paired with red (sub-threshold) entropy, indicating the gain comes from reduced diversity rather than genuine quality; CTR-Reg keeps entropy green while matching or improving PPL.
Model
Loss
Corruption
Corrupted-Token
Clean-Token
Overall
Rate ( p )
Acc. (%)
Acc. (%)
Acc. (%)
DUO
NELBO
10%
52.6
49.1
49.5
20%
45.6
44.1
44.4
30%
38.0
38.2
38.2
40%
29.6
31.4
30.7
Avg
41.5
40.7
40.7
Appendix
Table 12: Per-corruption-rate results on AG News. Full breakdown of Table 9 by corruption rate p∈{10,20,30,40}% on AG News, for each of three uniform-state diffusion language models, comparing NELBO alone (black) against NELBO + CTR-Reg (purple). The bolded value in each pair is the higher of the two; Avg rows report the corruption-rate-averaged values used in Table 9 .
Model
Loss
Corruption
Corrupted-Token
Clean-Token
Overall
Rate ( p )
Acc. (%)
Acc. (%)
Acc. (%)
DUO
NELBO
10%
57.2
59.8
59.5
20%
52.0
55.8
55.1
30%
45.9
50.8
49.3
40%
38.2
44.4
41.9
Avg
48.4
52.7
51.5
Appendix
Table 13: Per-corruption-rate results on arXiv. Full breakdown of Table 9 by corruption rate p∈{10,20,30,40}% on arXiv, for each of three uniform-state diffusion language models, comparing NELBO alone (black) against NELBO + CTR-Reg (purple). The bolded value in each pair is the higher of the two; Avg rows report the corruption-rate-averaged values used in Table 9 .
Model
Loss
Corruption
Corrupted-Token
Clean-Token
Overall
Rate ( p )
Acc. (%)
Acc. (%)
Acc. (%)
DUO
NELBO
10%
44.9
43.8
43.9
20%
38.7
39.3
39.1
30%
32.0
34.1
33.5
40%
25.0
28.2
26.9
Avg
35.2
36.3
35.8
Appendix
Table 14: Per-corruption-rate results on LM1B. Full breakdown of Table 9 by corruption rate p∈{10,20,30,40}% on LM1B, for each of three uniform-state diffusion language models, comparing NELBO alone (black) against NELBO + CTR-Reg (purple). The bolded value in each pair is the higher of the two; Avg rows report the corruption-rate-averaged values used in Table 9 .
Model
Loss
Corruption
Corrupted-Token
Clean-Token
Overall
Rate ( p )
Acc. (%)
Acc. (%)
Acc. (%)
DUO
NELBO
10%
56.3
58.2
58.0
20%
50.3
53.6
53.0
30%
43.4
48.0
46.6
40%
35.0
40.9
38.5
Avg
46.3
50.2
49.0
Appendix
Table 15: Per-corruption-rate results on PubMed. Full breakdown of Table 9 by corruption rate p∈{10,20,30,40}% on PubMed, for each of three uniform-state diffusion language models, comparing NELBO alone (black) against NELBO + CTR-Reg (purple). The bolded value in each pair is the higher of the two; Avg rows report the corruption-rate-averaged values used in Table 9 .
Model
Loss
Corruption
Corrupted-Token
Clean-Token
Overall
Rate ( p )
Acc. (%)
Acc. (%)
Acc. (%)
DUO
NELBO
10%
56.5
57.5
57.4
20%
49.8
52.9
52.3
30%
42.2
46.9
45.5
40%
33.6
39.3
37.0
Avg
45.5
49.1
48.0
Appendix
Table 16: Per-corruption-rate results on WT-103. Full breakdown of Table 9 by corruption rate p∈{10,20,30,40}% on WT-103, for each of three uniform-state diffusion language models, comparing NELBO alone (black) against NELBO + CTR-Reg (purple). The bolded value in each pair is the higher of the two; Avg rows report the corruption-rate-averaged values used in Table 9 .
NELBO
NELBO + CTR-Reg
Model
g
PPL
Ent.
MAUVE
PPL
Ent.
MAUVE
DUO
0
104.96
5.187
0.701
105.57
5.188
0.713
1
108.47
4.950
0.414
64.15
5.087
0.760
3
113.02
4.720
0.180
52.70
5.057
0.804
6
77.57
4.250
0.142
48.96
5.054
0.805
9
48.56
3.816
0.166
47.69
5.049
0.807
Appendix
Table 17: Exact PPL, entropy, and MAUVE values underlying Figures 3 . For each of three uniform-state diffusion language models and each greedy-tail step count g∈{0,1,3,6,9} (applied after 500 ancestral sampling steps), we report perplexity (PPL, computed with GPT2-Large; bolded value is the lower, i.e. better, of the two), generation entropy, and MAUVE score ( bolded value is the higher, i.e. better, of the two) for the NELBO-only baseline and for NELBO + CTR-Reg . Entropy values are colored green when they meet or exceed the real-OpenWebText reference entropy ( 4.83 ) and red when they fall short. As g increases, the NELBO-only baseline’s entropy drops below threshold and MAUVE collapses even as its PPL appears to improve, indicating reduced diversity rather than genuine quality; CTR-Reg keeps entropy green and MAUVE high while matching or improving PPL.
Dataset
Corruption Rate ( p )
Corrupted-Token Acc. (%)
Clean-Token Acc. (%)
Overall Acc. (%)
OWT
10%
71.0
69.6
69.7
20%
65.5
65.2
65.3
30%
58.3
59.7
59.3
40%
48.1
51.4
50.1
Avg
60.7
61.5
61.1
AG News
10%
61.8
58.7
59.0
Appendix
Table 18: Per-corruption-rate results for the official DUO checkpoint (524B tokens). Full breakdown of the 524B-token DUO row of Table 7 by corruption rate p∈{10,20,30,40}% , across all six evaluation corpora. In each row, the bolded value is the higher of corrupted-token vs. clean-token accuracy; Avg rows report the corruption-rate-averaged values used in Table 7 .
λ=0 (NELBO)
λ=0.02
λ=0.1
Model
Greedy-Tail Steps ( g )
PPL ( ↓ )
Entropy
PPL ( ↓ )
Entropy
PPL ( ↓ )
Entropy
DUO
0
105.37
5.18
105.86
5.19
109.03
5.19
1
110.88
4.96
64.67
5.09
73.24
5.13
2
118.86
4.87
56.93
5.08
67.82
5.11
3
112.91
4.72
53.32
5.06
64.56
5.11
4
102.94
4.57
51.16
5.06
63.52
5.10
Appendix
Table 19: Exact PPL and entropy values underlying Figure 6 . For DUO and UDLM and each greedy-tail step count g∈{0,1,2,3,4,5,6,8,10,12,14,16,18,20} (total step budget fixed at 512), we report perplexity (PPL, GPT2-Large, 1024 generated samples; bolded value is the lowest of the three) and generation entropy for λ=0 (NELBO-only), λ=0.02 ( our default ), and λ=0.1 ( blue ). Entropy values are colored green when they meet or exceed the real-OpenWebText reference entropy ( 4.83 ) and red when they fall short, matching Figure 6 .
Dataset
Model
Corruption Rate ( p )
Corrupted-Token Acc. (%)
Clean-Token Acc. (%)
Overall Acc. (%)
AG News
DUO
10%
52.6
89.7
86.0
20%
45.9
86.8
78.6
30%
38.3
82.9
69.5
40%
29.9
78.0
58.7
Avg
41.7
84.3
73.2
UDLM
10%
53.2
88.2
84.7
Appendix
Table 20: Per-corruption-rate results at λ=0.1 ( blue ). Full breakdown of the λ=0.1 row of Table 1 by corruption rate p∈{10,20,30,40}% , across all six evaluation corpora, for DUO and UDLM. Avg rows report the corruption-rate-averaged values used there.
DUO
UDLM
Dataset
p
0
0.02
0.1
0
0.02
0.1
AG News
10%
0.004
0.03
0.15
0.007
0.03
0.08
20%
0.004
0.04
0.18
0.010
0.03
0.09
30%
0.004
0.05
0.23
0.005
0.03
0.09
40%
0.004
0.06
0.30
0.005
0.03
0.09
Avg
0.004
0.05
0.22
0.007
0.03
0.09
Appendix
Table 21: Per-corruption-rate copy rate. Full breakdown of the copy-rate rows of Table 1 by corruption rate p , for DUO and UDLM at λ=0 (NELBO-only), 0.02 ( our default ), and 0.1 ( blue ), in %. Avg rows report the corruption-rate-averaged values used there.
Masked diffusion language models (MDLMs) re-predict every position at each denoising step, but standard samplers commit tokens once revealed, leaving this revision capability unused. Existing approaches either add heuristic or learned mechanisms to revise committed tokens, or remask them back to [MASK] before re-predicting; a principled sampler that directly revises visible tokens without auxiliary modules remains underexplored. We introduce D3IM, a parameter-free sampler derived as a corrector-style reverse update that permits direct visible-to-visible revision without additional modules or auxiliary passes. D3IM also reveals a model-side obstacle we term preservation bias: the model tends to reproduce its own wrong committed tokens rather than correct them. We address this with SCOPE (Self-Conditioned On Prediction Errors), a lightweight post-training procedure that simulates D3IM's sampling process. On LLaDA-8B at 64 denoising steps, SCOPE+D3IM improves over the original LLaDA-8B with standard unmasking by +13.0 on GSM8K (68.3%), +4.8 on MATH-500 (23.6%), +15.3 on HumanEval (29.3%), and +10.4 on MBPP (30.8%), with gains that increase as more denoising steps are used on math and HumanEval.
Diffusion large language models (dLLMs) gain speed by committing multiple tokens in parallel at each denoising step, but any erroneous commitment persists as conditioning context and biases every subsequent prediction. LLaDA2.1 repairs such errors with Token-to-Token (T2T) editing, which re-examines previously unmasked tokens and overwrites them when an alternative becomes sufficiently confident. We argue that this replacement action is itself the limiting factor: under polluted context, a confident replacement can propagate the error, while under a multimodal posterior no alternative may be confident enough to trigger an edit. We propose Token-to-Mask (T2M) remasking, a training-free rule that revokes suspicious commitments by resetting them to [M] and lets the subsequent mask-filling steps re-predict them from a cleaner context. T2M improves accuracy by +13.33 points on AIME 2025 and +8.56 points on CMATH. These results suggest that, for parallel discrete generators, remasking suspect tokens rather than overwriting them is a more reliable self-correction primitive.
Lin Yao
School of Computer Science, Shanghai Jiao Tong University, Shanghai, 200240, China · Zhongguancun Academy, Beijing, 100097, China
Is the uniform-state diffusion framework a more powerful paradigm for discrete diffusion? Recent studies indicate that this may be the case. In combination with predictor-corrector samplers, uniform-state diffusion models (USDMs) produce samples of higher-quality than masked diffusion models (MDMs), and USDMs equal or outperform MDMs in downstream tasks, even though they exhibit greater perplexity. Two issues remain unresolved. First, existing work compares uniform and masked diffusion with un-informed correctors that re-inject noise at random positions, rather than targeting tokens most likely to be wrong. Second, prior work compares full-sequence diffusion models, so we do not know whether the same conclusion holds when tokens are generated block by block. To address these issues, we introduce BlockGen, a blockwise sequence model that we instantiate with both masked and uniform diffusion. BlockGen trains on a mixture of block sizes and its likelihood interpolates between AR and pure diffusion more finely than models with a fixed block size. BlockGen enables AR-informed predictor-corrector sampling (ARPC), which combines AR and diffusion predictions to re-generate unlikely tokens without an auxiliary verifier. Under ancestral sampling, uniform outperforms masked in the block-by-block setting, especially in the few-step regime. Under ARPC, the gap closes and reverses at high NFE. With block size 16 on GSM8K, MDMs reach slightly higher accuracy than USDMs, and we observe a similar trend in Generative Perplexity on OpenWebText. Find our code at https://github.com/jdeschena/blockgen.
Justin Deschenaux, Caglar Gulcehre
EPFL Lausanne, Switzerland · EPFL, Lausanne, Switzerland · Microsoft AI