Masked diffusion language models (MDMs) admit flexible generation orders, making the unmasking strategy an inference decision. Existing methods vary in how they prioritize positions, control parallelism, restrict selection regions, revise predictions, or plan future denoising, yet it remains unclear when these choices should change during generation. We study this question through strategy reversals, where an alternative action becomes preferable to a fixed choice. We organize MDM inference into five axes--score, cardinality, region, commitment, and planning--and define adaptation opportunity as the one-step utility advantage of the best candidate action over a validation-selected fixed action. This view shows that adaptation value depends on both the frequency and magnitude of such reversals. Across three MDMs and ten tasks, adaptation opportunities are highly heterogeneous, with some regimes exhibiting concentrated and predictable one-step gains. This motivates selective adaptation: lightweight detectors calibrated on validation prompts identify high-opportunity states, capturing, for example, 56.9 percent of the candidate-set oracle opportunity by adapting only the top 10 percent of states on LLaDA-8B constrained JSON filling. Our transition-level results suggest that state adaptation is most useful when applied selectively rather than uniformly.
Figures & tables
Figure 1: Selective state adaptation. An axis-wise diagnostic proposes an alternative action, while a validation-calibrated opportunity detector decides whether to deviate from the fixed action.
Method
Score
Card.
Region
Commit.
Planning
LLaDA ( Nie et al., 2025b )
Conf.
Scheduled
Full
Irrev.
–
Fast-dLLM ( Wu et al., 2026 )
Conf.
Thresh. adapt.
Block
Irrev.
–
KLASS ( Kim et al., 2025b )
KL stab. + conf.
Multi-token
Full
Irrev.
–
DUS ( Luxembourg et al., 2026 )
Joint entropy
Grouped
Dilated
Irrev.
–
ReMDM ( Wang et al., 2025 )
Conf. / remask
Sched. dep.
Full
Remask.
–
WINO ( Hong et al., 2026c )
Draft verif.
Parallel draft
Active
Remask.
Draft–verify
Table 1: Representative methods under our five-axis taxonomy. Entries are abbreviated; full terms and an extended mapping are provided in Appendix A . “–” denotes no explicit mechanism beyond the underlying sampler.
Task
Axis
Positive states
Bidir. mass
CSV Missing Cells
region
0.0%
0.0025
Multi-Span Cloze
region
9.2%
0.0401
Constrained JSON Fill
region
13.6%
0.0225
TD07 Sparse Mask
region
4.5%
0.0238
JSON Mode Eval
region
7.0%
0.0418
Carry RTL
region
11.4%
0.0200
Table 2: Opportunity frequency and bidirectional reversal mass on LLaDA-1.5. Positive states are held-out states with positive candidate-set opportunity relative to the validation-selected fixed action. Bidirectional reversal mass is utility-weighted and is not a percentage.
Model
Task
Axis
Pos.
C5⋆
C10⋆
C20⋆
C5det
C10det
C20det
LLaDA-8B
Constrained JSON Fill
region
8.0%
62.7%
100.0%
100.0%
35.3%
56.9%
84.3%
Dream
Carry RTL
cardinality
4.2%
100.0%
100.0%
100.0%
42.9%
64.3%
71.4%
Dream
HTML Close Tags
region
4.7%
100.0%
100.0%
100.0%
16.7%
36.7%
66.7%
LLaDA-1.5
Carry RTL
region
11.4%
53.9%
89.9%
100.0%
4.5%
9.0%
31.5%
Table 3: Opportunity concentration versus detectability. Cc⋆ denotes candidate-set oracle opportunity captured by top- c states, while Ccdet denotes opportunity captured by detector-ranked states.
Model
Task
Axis
AUROC
Spearman
Pos. states
C10
L10
LLaDA-8B
Constrained JSON Fill
region
.854
.381
8.0%
56.9%
+.0453
Dream
Carry RTL
cardinality
.802
.368
4.2%
64.3%
+.0089
Dream
HTML Close Tags
region
.804
.261
4.7%
36.7%
+.0094
LLaDA-1.5
Carry RTL
region
.689
.242
11.4%
9.0%
+.0036
Dream
JSON Mode Eval
score
.891
.136
0.5%
0.0%
+.0000
LLaDA-1.5
JSON Mode Eval
region
.542
.051
7.0%
14.1%
+.0000
Table 4: Held-out opportunity detection. C10 is detector-captured candidate-set oracle opportunity at 10% coverage; L10 is realized one-step lift over the validation-selected fixed action. Prompt-cluster bootstrap 95% confidence intervals are reported in Appendix I .
Coverage
LLaDA-8B: Constrained JSON / region
Dream: Carry / cardinality
Dream: HTML / region
5%
35.3% / +.0281
42.9% / +.0057
16.7% / +.0016
10%
56.9% / +.0453
64.3% / +.0089
36.7% / +.0094
20%
84.3% / +.0594
71.4% / +.0099
66.7% / +.0188
50%
88.2% / +.0594
82.1% / +.0099
90.0% / +.0297
100%
100% / +.0594
100% / +.0099
100% / +.0297
Table 5: Coverage-dependent selective utility for representative successful regimes. Each entry reports the fraction of total candidate-set oracle opportunity captured / realized average one-step utility lift over the validation-selected fixed action.
Model
Task
Axis
Random C10
Detector C10
Oracle-gate C10
Detector lift
Oracle-gate diagnostic lift
LLaDA-8B
Constrained JSON Fill
region
10.0%
56.9%
100.0%
+0.0453
+0.0625
Dream
Carry RTL
cardinality
10.0%
64.3%
100.0%
+0.0089
+0.0099
Dream
HTML Close Tags
region
10.1%
36.7%
100.0%
+0.0094
+0.0375
Table 6: Coverage controls at 10% coverage. Random C10 averages opportunity capture over 10,000 uniform state selections. Oracle-gate C10 ranks states by held-out candidate-set opportunity, while oracle-gate lift applies the same diagnostic action only to those oracle-selected states.
Model
Task
Axis
Ref.
Fixed
Stage
Always
Selective
Δ
Coverage
LLaDA-8B
Unique List Commit
region
0.260
0.365
0.353
0.160
0.570
+0.205
54.2%
LLaDA-1.5
JSON Mode Eval
region
0.423
0.423
0.435
0.384
0.577
+0.154
19.6%
Dream
Carry RTL
commitment
0.433
0.467
0.467
0.367
0.567
+0.100
56.3%
Table 7: End-to-end selective decoding on illustrative settings. Terminal task utility under five decoding policies. Δ denotes Selective − Fixed, and Coverage is the fraction of states where selective adaptation is activated. Full ten-task results are in Appendix Table 18 .
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Score
Cardinality
Region
Commitment
Planning
LLaDA ( Nie et al., 2025b )
Confidence
Scheduled
Full
Irreversible
–
Fast-dLLM ( Wu et al., 2026 )
Confidence
Threshold-adaptive
Block
Irreversible
–
KLASS ( Kim et al., 2025b )
KL stability + confidence
Adaptive multi-token
Full
Irreversible
–
EB-Sampler ( Ben-Hamu et al., 2025 )
Entropy bound
Adaptive
Full
Irreversible
–
DOS ( Zhou et al., 2026 )
Attention dependency
Base sampler
Full
Irreversible
–
DEMASK † ( Ringel et al., 2026 )
Learned pairwise dependency
Dependency-bounded
Full
Irreversible
–
Appendix
Table 8: Extended mapping of masked diffusion inference methods under our five-axis taxonomy. Entries reflect our interpretation under the proposed taxonomy rather than terminology necessarily used by the original authors. “–” denotes no explicit additional mechanism, and † denotes methods requiring an additional learned component.
Task
Horizon
Source
Output
Utility
CSV Missing Cells
16
constructed
one cell
binary
HumanEval
256
official
Python completion
test pass
Multi-Span Cloze
64
constructed
indexed spans
partial
GSM8K
256
official
numeric answer
binary
Carry RTL
32
constructed
three carry bits
partial
TD07 Sparse Mask
48
constructed
three spans
partial
Appendix
Table 9: Task-suite details. “Partial” indicates component- or field-level credit in [0,1] ; “binary” indicates exact task success. The generation horizon is fixed within each task across candidate actions.
Model
Task
Axis
Gap
Bidir. mass
LLaDA-8B
CSV Missing Cells
Score
0.0000 (0.0000, 0.0000)
0.0000
LLaDA-8B
CSV Missing Cells
Region
0.0025 (0.0000, 0.0075)
0.0025
LLaDA-8B
CSV Missing Cells
Cardinality
0.0012 (0.0000, 0.0025)
0.0013
LLaDA-8B
CSV Missing Cells
Commitment
0.0000 (0.0000, 0.0000)
0.0000
LLaDA-8B
HumanEval
Score
0.0000 (0.0000, 0.0000)
0.0000
LLaDA-8B
HumanEval
Region
0.0048 (0.0000, 0.0141)
0.0048
Appendix
Table 10: Complete axis-wise adaptation-opportunity results. Gap is the held-out cross-fitted candidate-set oracle gap, with prompt-cluster bootstrap 95% confidence intervals shown in parentheses. Bidir. mass denotes axis-wise bidirectional reversal mass.
Model
Task
Axis
Corr.
Sign
LLaDA-8B
CSV Missing Cells
Score
–
1.000
LLaDA-8B
CSV Missing Cells
Region
–
0.778
LLaDA-8B
CSV Missing Cells
Cardinality
–
0.998
LLaDA-8B
CSV Missing Cells
Commitment
–
1.000
LLaDA-8B
HumanEval
Score
–
–
LLaDA-8B
HumanEval
Region
–
–
Appendix
Table 11: Complete diagnostic statistics. Corr. and Sign summarize diagnostic behavior for each model–task–axis setting. Dashes denote unavailable or undefined quantities.
Model
Task
Axis
FReg
SReg
DReg
Rec.
Screen
LLaDA-8B
CSV Missing Cells
Score
0.0000
0.0000
0.0000
–
no
LLaDA-8B
CSV Missing Cells
Region
0.0000
0.0000
0.0000
–
no
LLaDA-8B
CSV Missing Cells
Cardinality
0.0016
0.0016
0.0016
0.000
no
LLaDA-8B
CSV Missing Cells
Commitment
0.0000
0.0000
0.0000
–
no
LLaDA-8B
HumanEval
Score
–
–
–
–
–
LLaDA-8B
HumanEval
Region
–
–
–
–
–
Appendix
Table 12: Complete diagnostic action-selection utility. FReg, SReg, and DReg denote regret relative to the cross-fitted candidate-set oracle for the validation-selected fixed, stage-only, and diagnostic actions, respectively. Rec. denotes candidate-set oracle-gap recovery, and Screen indicates whether the exploratory held-out viability criteria are satisfied.
Model
Task
Axis
Fixed
Stage
Diagnostic
Oracle
D–F
D–S
LLaDA-8B
CSV Missing Cells
Score
0.0000
0.0000
0.0000
0.0000
+0.0000
+0.0000
LLaDA-8B
CSV Missing Cells
Region
0.0000
0.0000
0.0000
0.0000
+0.0000
+0.0000
LLaDA-8B
CSV Missing Cells
Cardinality
0.0000
0.0000
0.0000
0.0016
+0.0000
+0.0000
LLaDA-8B
CSV Missing Cells
Commitment
0.0016
0.0016
0.0016
0.0016
+0.0000
+0.0000
LLaDA-8B
HumanEval
Score
–
–
–
–
–
–
LLaDA-8B
HumanEval
Region
–
–
–
–
–
–
Appendix
Table 13: Complete transition-level utility results. Fixed, Stage, Diagnostic, and Oracle denote one-step utility under the validation-selected fixed action, stage-only action, diagnostic action, and cross-fitted candidate-set oracle, respectively. D–F and D–S denote diagnostic lift over the fixed and stage-only actions.
Model
Task
Axis
AUROC
ρ
Pos.
LLaDA-8B
CSV Missing Cells
Score
–
–
0.0%
LLaDA-8B
CSV Missing Cells
Region
–
–
0.0%
LLaDA-8B
CSV Missing Cells
Cardinality
0.500
–
0.2%
LLaDA-8B
CSV Missing Cells
Commitment
–
–
0.0%
LLaDA-8B
HumanEval
Score
–
–
–
LLaDA-8B
HumanEval
Region
–
–
–
Appendix
Table 14: Complete held-out opportunity-detection results. AUROC measures discrimination of positive-opportunity states, ρ denotes Spearman rank correlation with continuous candidate-set opportunity, and Pos. is the fraction of held-out states with positive opportunity. Dashes denote unsupported or undefined quantities.
Model
Task
Axis
5%
10%
20%
50%
100%
LLaDA-8B
CSV Missing Cells
Score
–
–
–
–
–
LLaDA-8B
CSV Missing Cells
Region
–
–
–
–
–
LLaDA-8B
CSV Missing Cells
Cardinality
0.0% (+0.0000)
0.0% (+0.0000)
0.0% (+0.0000)
0.0% (+0.0000)
100.0% (+0.0000)
LLaDA-8B
CSV Missing Cells
Commitment
–
–
–
–
–
LLaDA-8B
HumanEval
Score
–
–
–
–
–
LLaDA-8B
HumanEval
Region
–
–
–
–
–
Appendix
Table 15: Complete coverage-dependent selective-utility results. At each coverage level, each cell reports the fraction of total candidate-set oracle opportunity captured by detector-ranked states, with realized average one-step utility lift over the validation-selected fixed action shown in parentheses. Dashes denote unsupported or undefined quantities.
Model
Task
Axis
AUROC
ρ
C10det
L10det
Dream
Carry RTL
Cardinality
0.802 [0.699, 0.911]
0.368 [0.242, 0.486]
64.3% [43.9, 87.0]
+0.0089 [+0.0052, +0.0130]
Dream
Carry RTL
Commitment
0.498 [0.438, 0.556]
0.001 [-0.077, 0.081]
13.4% [6.9, 20.8]
-0.0005 [-0.0078, +0.0073]
Dream
Carry RTL
Region
0.540 [0.368, 0.862]
0.020 [-0.054, 0.088]
16.7% [0.0, 33.3]
+0.0000 [+0.0000, +0.0000]
Dream
Carry RTL
Score
–
–
–
+0.0000 [+0.0000, +0.0000]
Dream
Constrained JSON Fill
Cardinality
0.500 [0.500, 0.500]
–
4.9% [0.0, 12.1]
+0.0000 [+0.0000, +0.0000]
Dream
Constrained JSON Fill
Commitment
0.402 [0.321, 0.495]
-0.069 [-0.125, -0.003]
0.0% [0.0, 10.0]
+0.0000 [+0.0000, +0.0000]
Appendix
Table 16: Bootstrap robustness of opportunity detection and selective utility. C10det is the fraction of total candidate-set oracle opportunity captured by the detector-selected top 10% of states, and L10det is the corresponding realized average one-step utility lift over the validation-selected fixed action. Each cell reports the point estimate with its 95% confidence interval.
Model
Task
Axis
C10⋆
Oracle-gate lift
Random C10
Dream
Carry RTL
Cardinality
100.0% [100.0, 100.0]
+0.0099 [+0.0062, +0.0141]
10.0% [0.0, 21.4]
Dream
Carry RTL
Commitment
67.9% [56.4, 82.3]
+0.0312 [+0.0208, +0.0422]
10.0% [4.5, 15.7]
Dream
Carry RTL
Region
100.0% [100.0, 100.0]
+0.0000 [+0.0000, +0.0000]
10.1% [0.0, 33.3]
Dream
Carry RTL
Score
–
+0.0000 [+0.0000, +0.0000]
–
Dream
Constrained JSON Fill
Cardinality
100.0% [100.0, 100.0]
+0.0000 [+0.0000, +0.0000]
10.0% [2.4, 19.5]
Dream
Constrained JSON Fill
Commitment
100.0% [100.0, 100.0]
+0.0000 [+0.0000, +0.0000]
9.9% [0.0, 25.0]
Appendix
Table 17: Opportunity-concentration and coverage controls at 10% coverage. C10⋆ is the maximum fraction of total candidate-set oracle opportunity captured by the oracle-ranked top 10% of held-out states. Oracle-gate lift applies the same diagnostic action only to those oracle-selected states, and Random C10 is the opportunity fraction captured by uniformly selecting 10% of held-out states. Each cell reports the point estimate with its 95% confidence interval.
Model
Task
Axis
Ref.
Fixed
Stage
Always
Selective
Δ
Coverage
LLaDA-8B
Carry RTL
region
0.027
0.647
0.647
0.693
0.693
+0.046
45.8%
LLaDA-8B
Constrained JSON Fill
region
0.320
0.000
0.000
0.000
0.000
0.000
62.9%
LLaDA-8B
CSV Missing Cells
region
0.040
0.080
0.080
0.080
0.080
0.000
49.6%
LLaDA-8B
GSM8K
region
0.660
0.760
0.760
0.760
0.760
0.000
30.1%
LLaDA-8B
HTML Close Tags
commitment
0.160
0.160
0.160
0.160
0.160
0.000
50.4%
LLaDA-8B
HumanEval
commitment
0.141
0.141
0.141
0.141
0.141
0.000
0.0%
Appendix
Table 18: Full end-to-end selective-decoding results. Terminal task utility under the reference decoder, validation-selected best-fixed action, stage-only schedule, always-diagnostic adaptation, and selective adaptation. Δ denotes Selective − Fixed, and Coverage is the fraction of states at which selective adaptation is activated. Macro rows report the unweighted average across the ten tasks for each model.
Masked diffusion language models can reduce inference steps by revealing multiple tokens per denoising iteration, but this parallelism is fragile: positions that are individually confident may be unsafe to commit together when their predictions are coupled. Existing training-free samplers such as Top-k, Fast-dLLM, and EB-Sampler mainly control how many tokens to reveal, while often ranking candidates by token-wise scores that ignore interactions within the selected set. We propose ADAS, a training-free reranking rule that leaves the base sampler's stopping rule unchanged and greedily discounts each token-wise confidence score according to its attention to already selected positions, weighted by their prediction uncertainty. Across LLaDA-8B-Base and Dream-7B-Base on the reasoning benchmarks GSM8K and MATH500 and the code benchmarks HumanEval and MBPP, plugging ADAS into all three samplers improves low-NFE performance at matched denoiser evaluations by 9.11 and 10.46 percentage points on average, respectively, with 3.1% per-forward runtime overhead. Code is available at https://github.com/yusufsahin99/ADAS.
Yusuf Sahin, Ahmed Rockey Saikia, Volkan Cevher +1
University of Bern, Bern, Switzerland · EPFL, Lausanne, Switzerland
Reinforcement learning has proven effective for improving reasoning in large language models, but extending it to Masked Diffusion Language Models (MDLMs) remains challenging due to the intractability of the log-likelihood estimation. Existing approaches approximate this log-likelihood by modeling only the token predictions, ignoring the order in which positions are unmasked during generation. We observe that MDLM generation involves two decisions at each step: what tokens to place at each masked position and which positions to remask. We formalize this as a two-stage action MDP, showing that the policy gradient naturally decomposes into a token term and a masking term. Combining optimization of both terms leads to state-of-the-art outcomes on mathematical reasoning and coding benchmarks, with scores of 87.1% on GSM8K and 53.4% on MBPP.
Haran Raajesh, Kulin Shah, Adam Klivans +1
Department of Computer Science The University of Texas at Austin
Masked diffusion language models (MDLMs) re-predict every position at each denoising step, but standard samplers commit tokens once revealed, leaving this revision capability unused. Existing approaches either add heuristic or learned mechanisms to revise committed tokens, or remask them back to [MASK] before re-predicting; a principled sampler that directly revises visible tokens without auxiliary modules remains underexplored. We introduce D3IM, a parameter-free sampler derived as a corrector-style reverse update that permits direct visible-to-visible revision without additional modules or auxiliary passes. D3IM also reveals a model-side obstacle we term preservation bias: the model tends to reproduce its own wrong committed tokens rather than correct them. We address this with SCOPE (Self-Conditioned On Prediction Errors), a lightweight post-training procedure that simulates D3IM's sampling process. On LLaDA-8B at 64 denoising steps, SCOPE+D3IM improves over the original LLaDA-8B with standard unmasking by +13.0 on GSM8K (68.3%), +4.8 on MATH-500 (23.6%), +15.3 on HumanEval (29.3%), and +10.4 on MBPP (30.8%), with gains that increase as more denoising steps are used on math and HumanEval.