Masked diffusion language models (MDLMs) can generate text efficiently by predicting multiple masked tokens in parallel, but predictions from the same forward pass are not necessarily reliable when committed together. We study when parallel commitment is reliable. Our diagnostics show that confidence alone does not determine a reliable commitment order: confident predictions near the end of the sequence can fix an answer before its supporting computations are established, and downstream predictions become less reliable as the uncertainty of their upstream context grows. At the same time, a single forward pass can already resolve several masked tokens, and predictions that remain stable across the final layers are more likely to be correct. Based on these findings, we propose Reliable Parallel Decoding (RPD), a training-free method that selects candidates by layerwise prediction stability and final confidence, and commits them under a cumulative entropy budget over their preceding masked positions. RPD defers predictions with uncertain upstream context while committing the remaining candidates in parallel, without relying on a fixed block schedule. Across mathematical reasoning and code generation benchmarks on LLaDA and Dream, RPD achieves the highest decoding throughput among the evaluated methods while maintaining or improving accuracy.
Figures & tables
Figure 1: Left: Confidence-based decoding commits the answer early and fixes an incorrect value (1330),whereasdelayingtheansweryieldsthecorrect1430. (a) Accuracy improves under stricter ordering constraints. (b) Initial confidence peaks near the sequence boundaries. Right: (i) Downstream tokens should wait until their upstream values are resolved. (ii) Tokens resolved within one forward pass can be committed together.
Figure 2: Diagnostics on Dream. (a) Answer error rates for low- and high-upstream-entropy masking patterns within the same problem and mask count. (b) Token error rates under increasing prediction persistence K . (c) Token error rates by prediction persistence K and confidence drop ri for predictions with final confidence in [0.6,0.9) .
Figure 3: Overview of RPD. (a) Layerwise stability and confidence identify candidate predictions. Prediction persistence favors consistent predictions, while confidence drop penalizes predictions that lose probability. (b) A cumulative entropy budget determines which candidates can be committed together. A and B are committed in parallel, while C waits because the cumulative entropy of its preceding masked positions exceeds the budget.
Model
Benchmark
Metric
Default
Fast-dLLM
EB-Sampler
LoPA
DAPD
RPD-block
RPD
Dream-7B Instruct
HumanEval
pass@1 ↑
57.32
60.37
56.71
55.49
55.49
59.76
60.98
NFE ↓
256.00
64.54
78.35
35.79
77.74
61.72
63.09
TPS ↑
7.35
31.69
25.57
19.71
20.93
31.71
30.07
MBPP
pass@1 ↑
56.80
58.20
56.80
51.20
56.40
58.00
58.40
NFE ↓
256.00
57.63
62.41
36.95
65.05
55.49
55.89
TPS ↑
4.09
16.82
15.30
11.78
13.82
17.34
15.14
Table 1: Main results across mathematical reasoning and code generation benchmarks. Bold indicates the best result among all methods, while underline indicates the better result between our two variants when neither is best overall.
Figure 4: Accuracy–throughput trade-offs on LLaDA (top) and Dream (bottom). Fast-dLLM and EB-Sampler are swept over their confidence threshold and entropy budget, respectively, with each point labeled by its value. Other methods are shown at their default configurations.
LLaDA-8B-Instruct
Dream-7B-Instruct
Variant
GSM8K
HumanEval
GSM8K
HumanEval
Acc. ↑
NFE ↓
TPS ↑
Acc. ↑
NFE ↓
TPS ↑
Acc. ↑
NFE ↓
TPS ↑
Acc. ↑
NFE ↓
TPS ↑
Confidence
54.59
88.02
28.65
27.44
48.93
31.03
35.56
105.89
1.69
21.95
90.07
3.57
+ block32
76.35
77.09
35.63
43.29
44.48
43.13
81.73
76.04
30.03
60.37
64.54
31.69
+ PP
74.00
44.80
57.50
31.10
28.88
64.57
71.87
48.41
43.53
48.78
47.97
40.06
+ CE
78.92
80.84
33.23
43.29
45.23
42.46
81.88
88.03
24.53
60.98
65.62
29.75
Table 2: Component ablation of RPD. Starting from confidence-based decoding, we add prediction persistence (PP), confidence drop (CD), and cumulative entropy (CE). Confidence (full canvas) commits all predictions with confidence of at least 0.9 without block structure. We report accuracy (Acc.), number of function evaluations (NFE), and throughput (TPS). On Dream, confidence decoding commits EOS tokens early and truncates the response, so few tokens count toward TPS.
Figure 5: (a, b) Commitment order of Full-canvas and RPD on LLaDA GSM8K. Each cell shows the percentage of samples in which an output position is committed at a given normalized decoding progress. (c) Runtime of each method relative to Fast-dLLM, decomposed into backbone forward, layer projection, scoring and gating, and other operations.
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6: GSM8K instruction.
Figure 7: MATH-500 instruction.
Figure 8: MBPP instruction. {test_list} is filled with the first three public tests.
Figure 9: HumanEval instruction.
Figure 10: Chat wrapping for LLaDA-8B-Instruct.
Figure 11: Chat wrapping for Dream-v0-Instruct-7B.
Model
Masks
n
Low entropy
High entropy
Δ
LLaDA
4
40
22.5
51.0
28.5
8
81
14.6
38.8
24.2
16
182
12.7
37.6
24.8
32
813
10.5
39.0
28.6
Dream
4
128
15.8
33.1
17.3
8
322
9.3
27.2
17.9
Appendix
Table 3: Complete upstream-entropy results. n is the number of retained mixed-outcome questions. Error rates are percentages; Δ is the paired high-minus-low difference in percentage points, computed before rounding.
Figure 12: LLaDA commitment order with and without block structure, one commit per forward pass. Each box is one generated token, shaded by the step at which it was committed.
Figure 13: Dream commitment order under its default sampler and under greedy decoding, one commit per forward pass and lowest-entropy-first selection in both panels. Each box is one generated token, shaded by the step at which it was committed; ~ marks a terminal token.
Symbol
Description
Value
θh
High-confidence threshold (bypasses the stability test)
0.9
θc
Minimum confidence for stability-based admission
0.6
θs
Stability threshold
3.5 (LLaDA) / 2.5 (Dream)
Kmax
Cap on prediction persistence
6
w
Weight of the confidence-drop penalty
15 (LLaDA) / 20 (Dream)
β
Cumulative entropy budget
4 nats
Appendix
Table 4: Hyperparameters of RPD.
Figure 14: Final confidence of tokens admitted by the stability test.
Figure 15: Confidence drop ri of tokens admitted by the stability test. The upper limit in each panel is the selection ceiling (Kmax−θs)/w .
Figure 16: Commit step of each output position under three decoders, LLaDA on GSM8K (first document). Lower is earlier.
Masked diffusion language models (MDLMs) enable parallel decoding by predicting all masked positions at each denoising step, yet existing training-free samplers usually decide which positions to commit at token-level granularity. We revisit this granularity and observe that reliable predictions often emerge as contiguous high-confidence spans, suggesting that the unit of parallel commitment can be larger than a single token. We first group adjacent high-confidence candidates into confidence-induced clusters (CICs) as span-level update units. We then use self-attention maps from the same forward pass to estimate inter-cluster dependencies, enabling conflict-aware selection of mutually compatible CICs for parallel commitment. This yields CLAD (Cluster-Level Attention-Guided Decoding), a training-free cluster-level decoder for MDLMs. Experiments on LLaDA and Dream model families across four reasoning and code-generation benchmarks show that CLAD achieves 1.77x--8.47x speedups over Vanilla decoding while maintaining broadly comparable task accuracy in most settings.
Heqiang Qi, Wei Huang, Mingyuan Bai +1
Zhejiang University · RIKEN Center for Advanced Intelligence Project · The Institute of Statistical Mathematics +1
Masked diffusion language models (MDLMs) generate text by iteratively unmasking tokens, but their standard decoder reduces each step to a binary action: a position is either committed to a single token or left fully masked, with no representation of partial belief in between. This all-or-nothing regime discards rich predictive information and forces premature, irrevocable commitments, leading to poor performance under a limited decoding budget. In this paper, we reinterpret mask prediction as clean-state prediction (x-prediction) and show that it can be used to induce a continuous flow in input embedding space. Building on this view, we propose a continuous decoding framework for MDLMs where tokens can accumulate partial progress at each diffusion step and remain revisable. To match the uneven contextual constraints across positions in language, we replace the globally synchronous schedule in image diffusion with a confidence-based asynchronous update in which the diffusion progress is token-wise accumulated. Additionally, we introduce a lightweight policy network and formulate its training as a reinforcement learning problem. Applied to pretrained LLaDA, our continuous decoder reaches 97% of its performance on the HumanEval dataset with 25% of decoding budget.
Weitian Wang, Lianlei Shan, Shubham Rai +2
Robert Bosch GmbH, Germany · University of the Chinese Academy of Sciences, China · Ruhr University Bochum, Germany
Discrete diffusion language models enable parallel token generation, offering a pathway to low-latency decoding. However, selecting tokens independently by marginal confidence limits effective parallelism: tokens that appear reliable in isolation can form incompatible configurations when several positions are updated at once. We introduce a training-free decoding framework that coordinates these parallel updates. At each forward pass, the method assigns a commit score to each masked position and refines these scores using pairwise interactions derived from the model's predictive distributions. A variational relaxation yields a simple fixed-point update that suppresses conflicting simultaneous commitments within a single forward pass. This mechanism allows the decoder to commit more tokens in parallel while maintaining competitive generation quality. The method is lightweight, requires no auxiliary model or retraining, and drops into existing diffusion decoding pipelines without modification. Experiments on reasoning and code-generation benchmarks show consistent improvements in the quality-latency trade-off.
Tamim Zoabi, Ameen Ali, Liran Ringel +1
School of Electrical & Computer Engineering Tel Aviv University, Israel · School of Computer Science and AI Tel Aviv University, Israel · Department of Computer Science Technion, Israel Institute of Technology, Israel