Masked Diffusion Language Models (MDLMs) generate sequences by iteratively replacing masked tokens with model predictions. At each denoising step, the decoder chooses which positions are sufficiently confident to commit. Existing decoding methods typically rely on softmax confidence, which can be miscalibrated. We introduce BayesER (BAYESian Entropy-based Reordering), a post-hoc Bayesian decoding framework that uses predictive uncertainty to guide token commitment. In BayesER, we construct a lightweight approximate posterior centered at the pretrained checkpoint, similar to Laplace-LoRA but without training LoRA adapters. We average predictions over posterior samples and use predictive entropy to prioritize reliable positions. We examine how posterior predictions affect position ordering and token selection across benchmarks spanning code generation, mathematical reasoning, planning, and molecular generation. We show that BayesER reduces sequence-level calibration error while preserving or improving accuracy relative to common decoding schemes, including confidence-threshold decoding. Additionally, a posterior fitted on one code-generation dataset reduces calibration error on another without refitting, suggesting that Bayesian uncertainty may provide a transferable signal for more reliable MDLM decoding.
Figures & tables
Dataset
Method
Accuracy (%) ↑
ECE gen ( ×100 ) ↓
ECE ans ( ×100 ) ↓
Brier gen ( ×100 ) ↓
MBPP
Threshold (Fast-dllm)
57.60
29.18
31.24
31.85
Base Entropy
57.00(−1.04)
28.99(+0.65)
31.53(−0.93)
31.23(+1.97)
Info-gain
54.001.71(−6.25)
31.601.71(−8.29)
33.731.71(−7.97)
33.151.23(−4.07)
BayesER (best accuracy)
57.530.12(−0.12)
23.540.11(+19.33)
25.700.09(+17.73)
28.400.02(+10.84)
BayesER (best calibration)
56.731.01(−1.51)
16.321.21(+44.07)
18.171.02(+41.84)
25.570.36(+19.73)
HumanEval
Threshold (Fast-dllm)
54.88
30.90
30.59
33.27
Table 1: Evaluation on mathematical reasoning and programming tasks. Accuracy is reported in percent; ECE and Brier values and their standard deviations are multiplied by 100. Subscripts are the standard deviations of three seeds on the corresponding column scale. Numbers in parentheses are relative percentage improvements over confidence-thresholding decoding. Bold values indicate the best result and underlined values indicate the second-best result per metric within each benchmark.
Dataset
Method
Accuracy (%) ↑
ECE gen ( ×100 ) ↓
ECE ans ( ×100 ) ↓
Brier gen ( ×100 ) ↓
Sudoku
Threshold (Fast-dllm)
97.60
3.13
3.13
2.24
Base Entropy
97.60(+0.00)
2.98(+4.79)
2.98(+4.79)
2.28(−1.54)
Info-gain
96.070.50(−1.57)
2.540.08(+18.85)
2.540.08(+18.85)
3.390.37(−51.35)
BayesER (best accuracy)
97.530.31(−0.07)
3.960.03(−26.52)
3.960.03(−26.52)
2.400.17(−7.00)
BayesER (best calibration)
97.400.20(−0.20)
2.990.27(+4.47)
2.990.27(+4.47)
2.390.16(−6.56)
Countdown
Threshold (Fast-dllm)
18.90
66.46
61.54
58.56
Table 2: Evaluation on planning tasks.
Method
Quality Yield (%) ↑
Frag. Valid (%) ↑
QED ↑
SA ↓
ECE ↓
Native GenMol
11.78±0.69
99.33±0.33
0.6369±0.0104
4.3932±0.0362
45.08±0.29
Base Entropy
14.89±0.69
99.89±0.19
0.6599±0.0011
4.2809±0.0106
35.40±0.59
Info-gain
3.00±0.67
100.00±0.00
0.6704±0.0206
4.1528±0.0595
24.56±1.71
DPRM ( Bu et al., 2026 )
12.44±0.69
100.00±0.00
0.6619±0.0117
4.2682±0.0742
30.62±1.33
BayesER (Ours)
15.33±0.33
99.67±0.00
0.6602±0.0006
4.2900±0.0062
35.36±0.73
Table 3: Evaluation on GenMol V2 Scaffold Decoration. Quality Yield measures the percentage of generated attempts satisfying both drug-likeness ( QED≥0.6 ) and synthetic accessibility ( SA≤4.0 ).
Figure 1: Accuracy–calibration trade-off for MBPP with different hyperparameter choices.
Figure 2: Dream HumanEval cost versus uncertain depth.
Table 4: Main experimental configuration for the six-benchmark Dream study. HumanEval uses the MBPP-trained posterior without refitting; all other posteriors are fitted on task-specific or synthetic training data.
Task
Target
Ls
Rank
Alpha
Prior
τB
MBPP
Accuracy
1
8
16
300
0.85
MBPP
Calibration
1
8
16
200
0.90
HumanEval
Accuracy
1
8
16
300
0.85
HumanEval
Calibration
1
8
16
200
0.90
GSM8K
Accuracy
4
8
16
2,000
0.90
GSM8K
Calibration
1
8
16
500
0.85
Appendix
Table 5: Hyperparameters for the evaluation in the main results. HumanEval reuses the MBPP-trained posterior; the remaining rows are task-fitted.
Figure 3: HumanEval accuracy versus calibration.
Figure 4: GSM8K accuracy versus calibration.
Figure 5: MATH500 accuracy versus calibration.
Figure 6: Countdown accuracy versus calibration.
Figure 7: Measured Dream GSM8K cost versus uncertain depth. Full test set ( n=1,319 ), GSM8K-fitted rank-8 q/v Laplace-LoRA with S=10 , on H100 GPUs. Color denotes prior precision λ . Dashed lines show the confidence-threshold run. Wall time includes loading and scoring; the loose prior increases decoding steps, especially at four layers.
Method
Accuracy (%) ↑
ECE gen ( ×100 ) ↓
ECE ans ( ×100 ) ↓
Brier gen ( ×100 ) ↓
MBPP (500)
Threshold (Fast-dllm)
39.40
47.63
50.01
46.14
Base Entropy
41.20(+4.57)
44.51(+6.53)
46.84(+6.33)
43.48(+5.76)
Info-gain
41.270.61(+4.74)
44.960.50(+5.59)
48.010.52(+4.00)
43.520.26(+5.67)
BayesER
41.330.58(+4.91)
43.570.67(+8.52)
45.890.66(+8.23)
42.700.45(+7.46)
HumanEval (164)
Appendix
Table 6: LLaDA full-test results. Accuracy is reported in percent; ECE and generated-content Brier values and their standard deviations are multiplied by 100. Subscripts are the standard deviations of three seeds on the corresponding column scale. Numbers in parentheses are relative percentage improvements over confidence-thresholding decoding. Bold values indicate the best result and underlined values indicate the second-best result per metric within each benchmark, including ties.
MIRank
BayesER , τB=0.85
BayesER , τB=0.90
Ls
λ
Acc.
ECE g
ECE a
Steps
Acc.
ECE g
ECE a
Steps
Acc.
ECE g
ECE a
Steps
1
100
28.91
10.89
10.77
25.74
34.38
8.40
6.61
100.70
40.63
12.62
10.44
93.91
1
200
48.44
19.63
21.55
38.02
57.03
16.42
18.14
59.38
55.47
19.27
20.04
67.27
1
500
54.69
27.78
30.58
39.11
57.81
25.77
27.89
43.26
57.81
27.02
29.19
45.13
1
1000
53.13
30.62
33.09
43.25
57.81
27.18
29.16
37.36
58.59
27.30
29.60
42.54
1
2000
57.81
26.34
29.27
40.44
57.81
27.56
29.50
37.62
57.03
29.37
31.60
42.23
Appendix
Table 7: Comparison between MIRank and BayesER on 128 examples of MBPP. Accuracy is in percent; both ECE columns are multiplied by 100.
MIRank
BayesER , τB=0.85
BayesER , τB=0.90
Ls
λ
Acc.
ECE g
ECE a
Steps
Acc.
ECE g
ECE a
Steps
Acc.
ECE g
ECE a
Steps
1
100
59.38
31.56
12.84
30.39
58.59
31.50
16.28
102.63
63.28
38.08
15.94
118.50
1
200
64.84
10.14
19.34
41.83
77.34
5.56
15.76
70.57
75.78
6.65
16.95
77.88
1
500
78.91
4.61
18.37
49.34
79.69
8.37
13.51
56.29
75.78
9.10
17.81
63.16
1
1000
77.34
8.21
19.78
52.70
76.56
10.08
18.80
53.72
78.13
7.45
16.38
60.42
1
2000
77.34
9.39
17.38
51.38
76.56
9.09
18.56
53.38
78.13
7.63
17.32
61.18
Appendix
Table 8: Comparison between MIRank and BayesER on 128 examples of GSM8K. Accuracy is in percent; both ECE columns are multiplied by 100.
Masked diffusion language models (MDLMs) generate text by iteratively unmasking tokens, but their standard decoder reduces each step to a binary action: a position is either committed to a single token or left fully masked, with no representation of partial belief in between. This all-or-nothing regime discards rich predictive information and forces premature, irrevocable commitments, leading to poor performance under a limited decoding budget. In this paper, we reinterpret mask prediction as clean-state prediction (x-prediction) and show that it can be used to induce a continuous flow in input embedding space. Building on this view, we propose a continuous decoding framework for MDLMs where tokens can accumulate partial progress at each diffusion step and remain revisable. To match the uneven contextual constraints across positions in language, we replace the globally synchronous schedule in image diffusion with a confidence-based asynchronous update in which the diffusion progress is token-wise accumulated. Additionally, we introduce a lightweight policy network and formulate its training as a reinforcement learning problem. Applied to pretrained LLaDA, our continuous decoder reaches 97% of its performance on the HumanEval dataset with 25% of decoding budget.
Weitian Wang, Lianlei Shan, Shubham Rai +2
Robert Bosch GmbH, Germany · University of the Chinese Academy of Sciences, China · Ruhr University Bochum, Germany
Masked diffusion language models (MDLMs) can generate text efficiently by predicting multiple masked tokens in parallel, but predictions from the same forward pass are not necessarily reliable when committed together. We study when parallel commitment is reliable. Our diagnostics show that confidence alone does not determine a reliable commitment order: confident predictions near the end of the sequence can fix an answer before its supporting computations are established, and downstream predictions become less reliable as the uncertainty of their upstream context grows. At the same time, a single forward pass can already resolve several masked tokens, and predictions that remain stable across the final layers are more likely to be correct. Based on these findings, we propose Reliable Parallel Decoding (RPD), a training-free method that selects candidates by layerwise prediction stability and final confidence, and commits them under a cumulative entropy budget over their preceding masked positions. RPD defers predictions with uncertain upstream context while committing the remaining candidates in parallel, without relying on a fixed block schedule. Across mathematical reasoning and code generation benchmarks on LLaDA and Dream, RPD achieves the highest decoding throughput among the evaluated methods while maintaining or improving accuracy.
Zhenghao He, Bohan Liu, Guangzhi Xiong +1
Department of Computer Science, University of Virginia
Masked Diffusion Language Models (MDLMs) have emerged as a distinct paradigm for sequence generation. As MDLMs become diverse in capabilities and knowledge coverage, an important question is how to combine their knowledge. Toward this, we first investigate the unique decoding dynamics of MDLMs. We find that successful generations exhibit stable confidence dynamics over answer-relevant positions, while unreliable trajectories can often be corrected by injecting promising intermediate states from other models. Guided by this observation, we propose TIE (Trajectory-based Iterative Ensembling), a knowledge fusion framework in which MDLMs iteratively identify reliable decoding trajectories and relay them across models. TIE tracks confidence dynamics over answer-relevant positions to determine which model currently follows a more reliable trajectory and selectively transfers partially denoised sequences across models. As the model on the more promising trajectory often changes across denoising steps, TIE allows different models to contribute complementary strengths at different stages of generation. Strong performance across diverse reasoning tasks, along with our analyses, suggests that TIE offers a practical approach to the underexplored problem of MDLM ensembling.