Masked diffusion language models (MDLMs) can generate text efficiently by predicting multiple masked tokens in parallel, but predictions from the same forward pass are not necessarily reliable when committed together. We study when parallel commitment is reliable. Our diagnostics show that confidence alone does not determine a reliable commitment order: confident predictions near the end of the sequence can fix an answer before its supporting computations are established, and downstream predictions become less reliable as the uncertainty of their upstream context grows. At the same time, a single forward pass can already resolve several masked tokens, and predictions that remain stable across the final layers are more likely to be correct. Based on these findings, we propose Reliable Parallel Decoding (RPD), a training-free method that selects candidates by layerwise prediction stability and final confidence, and commits them under a cumulative entropy budget over their preceding masked positions. RPD defers predictions with uncertain upstream context while committing the remaining candidates in parallel, without relying on a fixed block schedule. Across mathematical reasoning and code generation benchmarks on LLaDA and Dream, RPD achieves the highest decoding throughput among the evaluated methods while maintaining or improving accuracy.
Figures & tables
Figure 1: Left: Confidence-based decoding commits the answer early and fixes an incorrect value (1330),whereasdelayingtheansweryieldsthecorrect1430. (a) Accuracy improves under stricter ordering constraints. (b) Initial confidence peaks near the sequence boundaries. Right: (i) Downstream tokens should wait until their upstream values are resolved. (ii) Tokens resolved within one forward pass can be committed together.
Figure 2: Diagnostics on Dream. (a) Answer error rates for low- and high-upstream-entropy masking patterns within the same problem and mask count. (b) Token error rates under increasing prediction persistence K . (c) Token error rates by prediction persistence K and confidence drop ri for predictions with final confidence in [0.6,0.9) .
Figure 3: Overview of RPD. (a) Layerwise stability and confidence identify candidate predictions. Prediction persistence favors consistent predictions, while confidence drop penalizes predictions that lose probability. (b) A cumulative entropy budget determines which candidates can be committed together. A and B are committed in parallel, while C waits because the cumulative entropy of its preceding masked positions exceeds the budget.
Model
Benchmark
Metric
Default
Fast-dLLM
EB-Sampler
LoPA
DAPD
RPD-block
RPD
Dream-7B Instruct
HumanEval
pass@1 ↑
57.32
60.37
56.71
55.49
55.49
59.76
60.98
NFE ↓
256.00
64.54
78.35
35.79
77.74
61.72
63.09
TPS ↑
7.35
31.69
25.57
19.71
20.93
31.71
30.07
MBPP
pass@1 ↑
56.80
58.20
56.80
51.20
56.40
58.00
58.40
NFE ↓
256.00
57.63
62.41
36.95
65.05
55.49
55.89
TPS ↑
4.09
16.82
15.30
11.78
13.82
17.34
15.14
Table 1: Main results across mathematical reasoning and code generation benchmarks. Bold indicates the best result among all methods, while underline indicates the better result between our two variants when neither is best overall.
Figure 4: Accuracy–throughput trade-offs on LLaDA (top) and Dream (bottom). Fast-dLLM and EB-Sampler are swept over their confidence threshold and entropy budget, respectively, with each point labeled by its value. Other methods are shown at their default configurations.
LLaDA-8B-Instruct
Dream-7B-Instruct
Variant
GSM8K
HumanEval
GSM8K
HumanEval
Acc. ↑
NFE ↓
TPS ↑
Acc. ↑
NFE ↓
TPS ↑
Acc. ↑
NFE ↓
TPS ↑
Acc. ↑
NFE ↓
TPS ↑
Confidence
54.59
88.02
28.65
27.44
48.93
31.03
35.56
105.89
1.69
21.95
90.07
3.57
+ block32
76.35
77.09
35.63
43.29
44.48
43.13
81.73
76.04
30.03
60.37
64.54
31.69
+ PP
74.00
44.80
57.50
31.10
28.88
64.57
71.87
48.41
43.53
48.78
47.97
40.06
+ CE
78.92
80.84
33.23
43.29
45.23
42.46
81.88
88.03
24.53
60.98
65.62
29.75
Table 2: Component ablation of RPD. Starting from confidence-based decoding, we add prediction persistence (PP), confidence drop (CD), and cumulative entropy (CE). Confidence (full canvas) commits all predictions with confidence of at least 0.9 without block structure. We report accuracy (Acc.), number of function evaluations (NFE), and throughput (TPS). On Dream, confidence decoding commits EOS tokens early and truncates the response, so few tokens count toward TPS.
Figure 5: (a, b) Commitment order of Full-canvas and RPD on LLaDA GSM8K. Each cell shows the percentage of samples in which an output position is committed at a given normalized decoding progress. (c) Runtime of each method relative to Fast-dLLM, decomposed into backbone forward, layer projection, scoring and gating, and other operations.
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6: GSM8K instruction.
Figure 7: MATH-500 instruction.
Figure 8: MBPP instruction. {test_list} is filled with the first three public tests.
Figure 9: HumanEval instruction.
Figure 10: Chat wrapping for LLaDA-8B-Instruct.
Figure 11: Chat wrapping for Dream-v0-Instruct-7B.
Model
Masks
n
Low entropy
High entropy
Δ
LLaDA
4
40
22.5
51.0
28.5
8
81
14.6
38.8
24.2
16
182
12.7
37.6
24.8
32
813
10.5
39.0
28.6
Dream
4
128
15.8
33.1
17.3
8
322
9.3
27.2
17.9
Appendix
Table 3: Complete upstream-entropy results. n is the number of retained mixed-outcome questions. Error rates are percentages; Δ is the paired high-minus-low difference in percentage points, computed before rounding.
Figure 12: LLaDA commitment order with and without block structure, one commit per forward pass. Each box is one generated token, shaded by the step at which it was committed.
Figure 13: Dream commitment order under its default sampler and under greedy decoding, one commit per forward pass and lowest-entropy-first selection in both panels. Each box is one generated token, shaded by the step at which it was committed; ~ marks a terminal token.
Symbol
Description
Value
θh
High-confidence threshold (bypasses the stability test)
0.9
θc
Minimum confidence for stability-based admission
0.6
θs
Stability threshold
3.5 (LLaDA) / 2.5 (Dream)
Kmax
Cap on prediction persistence
6
w
Weight of the confidence-drop penalty
15 (LLaDA) / 20 (Dream)
β
Cumulative entropy budget
4 nats
Appendix
Table 4: Hyperparameters of RPD.
Figure 14: Final confidence of tokens admitted by the stability test.
Figure 15: Confidence drop ri of tokens admitted by the stability test. The upper limit in each panel is the selection ceiling (Kmax−θs)/w .
Figure 16: Commit step of each output position under three decoders, LLaDA on GSM8K (first document). Lower is earlier.
School of Electrical & Computer Engineering Tel Aviv University, Israel · School of Computer Science and AI Tel Aviv University, Israel · Department of Computer Science Technion, Israel Institute of Technology, Israel