Masked Diffusion Language Models (MDLMs) generate sequences by iteratively replacing masked tokens with model predictions. At each denoising step, the decoder chooses which positions are sufficiently confident to commit. Existing decoding methods typically rely on softmax confidence, which can be miscalibrated. We introduce BayesER (BAYESian Entropy-based Reordering), a post-hoc Bayesian decoding framework that uses predictive uncertainty to guide token commitment. In BayesER, we construct a lightweight approximate posterior centered at the pretrained checkpoint, similar to Laplace-LoRA but without training LoRA adapters. We average predictions over posterior samples and use predictive entropy to prioritize reliable positions. We examine how posterior predictions affect position ordering and token selection across benchmarks spanning code generation, mathematical reasoning, planning, and molecular generation. We show that BayesER reduces sequence-level calibration error while preserving or improving accuracy relative to common decoding schemes, including confidence-threshold decoding. Additionally, a posterior fitted on one code-generation dataset reduces calibration error on another without refitting, suggesting that Bayesian uncertainty may provide a transferable signal for more reliable MDLM decoding.
Figures & tables
Dataset
Method
Accuracy (%) ↑
ECE gen ( ×100 ) ↓
ECE ans ( ×100 ) ↓
Brier gen ( ×100 ) ↓
MBPP
Threshold (Fast-dllm)
57.60
29.18
31.24
31.85
Base Entropy
57.00(−1.04)
28.99(+0.65)
31.53(−0.93)
31.23(+1.97)
Info-gain
54.001.71(−6.25)
31.601.71(−8.29)
33.731.71(−7.97)
33.151.23(−4.07)
BayesER (best accuracy)
57.530.12(−0.12)
23.540.11(+19.33)
25.700.09(+17.73)
28.400.02(+10.84)
BayesER (best calibration)
56.731.01(−1.51)
16.321.21(+44.07)
18.171.02(+41.84)
25.570.36(+19.73)
HumanEval
Threshold (Fast-dllm)
54.88
30.90
30.59
33.27
Table 1: Evaluation on mathematical reasoning and programming tasks. Accuracy is reported in percent; ECE and Brier values and their standard deviations are multiplied by 100. Subscripts are the standard deviations of three seeds on the corresponding column scale. Numbers in parentheses are relative percentage improvements over confidence-thresholding decoding. Bold values indicate the best result and underlined values indicate the second-best result per metric within each benchmark.
Dataset
Method
Accuracy (%) ↑
ECE gen ( ×100 ) ↓
ECE ans ( ×100 ) ↓
Brier gen ( ×100 ) ↓
Sudoku
Threshold (Fast-dllm)
97.60
3.13
3.13
2.24
Base Entropy
97.60(+0.00)
2.98(+4.79)
2.98(+4.79)
2.28(−1.54)
Info-gain
96.070.50(−1.57)
2.540.08(+18.85)
2.540.08(+18.85)
3.390.37(−51.35)
BayesER (best accuracy)
97.530.31(−0.07)
3.960.03(−26.52)
3.960.03(−26.52)
2.400.17(−7.00)
BayesER (best calibration)
97.400.20(−0.20)
2.990.27(+4.47)
2.990.27(+4.47)
2.390.16(−6.56)
Countdown
Threshold (Fast-dllm)
18.90
66.46
61.54
58.56
Table 2: Evaluation on planning tasks.
Method
Quality Yield (%) ↑
Frag. Valid (%) ↑
QED ↑
SA ↓
ECE ↓
Native GenMol
11.78±0.69
99.33±0.33
0.6369±0.0104
4.3932±0.0362
45.08±0.29
Base Entropy
14.89±0.69
99.89±0.19
0.6599±0.0011
4.2809±0.0106
35.40±0.59
Info-gain
3.00±0.67
100.00±0.00
0.6704±0.0206
4.1528±0.0595
24.56±1.71
DPRM ( Bu et al., 2026 )
12.44±0.69
100.00±0.00
0.6619±0.0117
4.2682±0.0742
30.62±1.33
BayesER (Ours)
15.33±0.33
99.67±0.00
0.6602±0.0006
4.2900±0.0062
35.36±0.73
Table 3: Evaluation on GenMol V2 Scaffold Decoration. Quality Yield measures the percentage of generated attempts satisfying both drug-likeness ( QED≥0.6 ) and synthetic accessibility ( SA≤4.0 ).
Figure 1: Accuracy–calibration trade-off for MBPP with different hyperparameter choices.
Figure 2: Dream HumanEval cost versus uncertain depth.
Table 4: Main experimental configuration for the six-benchmark Dream study. HumanEval uses the MBPP-trained posterior without refitting; all other posteriors are fitted on task-specific or synthetic training data.
Task
Target
Ls
Rank
Alpha
Prior
τB
MBPP
Accuracy
1
8
16
300
0.85
MBPP
Calibration
1
8
16
200
0.90
HumanEval
Accuracy
1
8
16
300
0.85
HumanEval
Calibration
1
8
16
200
0.90
GSM8K
Accuracy
4
8
16
2,000
0.90
GSM8K
Calibration
1
8
16
500
0.85
Appendix
Table 5: Hyperparameters for the evaluation in the main results. HumanEval reuses the MBPP-trained posterior; the remaining rows are task-fitted.
Figure 3: HumanEval accuracy versus calibration.
Figure 4: GSM8K accuracy versus calibration.
Figure 5: MATH500 accuracy versus calibration.
Figure 6: Countdown accuracy versus calibration.
Figure 7: Measured Dream GSM8K cost versus uncertain depth. Full test set ( n=1,319 ), GSM8K-fitted rank-8 q/v Laplace-LoRA with S=10 , on H100 GPUs. Color denotes prior precision λ . Dashed lines show the confidence-threshold run. Wall time includes loading and scoring; the loose prior increases decoding steps, especially at four layers.
Method
Accuracy (%) ↑
ECE gen ( ×100 ) ↓
ECE ans ( ×100 ) ↓
Brier gen ( ×100 ) ↓
MBPP (500)
Threshold (Fast-dllm)
39.40
47.63
50.01
46.14
Base Entropy
41.20(+4.57)
44.51(+6.53)
46.84(+6.33)
43.48(+5.76)
Info-gain
41.270.61(+4.74)
44.960.50(+5.59)
48.010.52(+4.00)
43.520.26(+5.67)
BayesER
41.330.58(+4.91)
43.570.67(+8.52)
45.890.66(+8.23)
42.700.45(+7.46)
HumanEval (164)
Appendix
Table 6: LLaDA full-test results. Accuracy is reported in percent; ECE and generated-content Brier values and their standard deviations are multiplied by 100. Subscripts are the standard deviations of three seeds on the corresponding column scale. Numbers in parentheses are relative percentage improvements over confidence-thresholding decoding. Bold values indicate the best result and underlined values indicate the second-best result per metric within each benchmark, including ties.
MIRank
BayesER , τB=0.85
BayesER , τB=0.90
Ls
λ
Acc.
ECE g
ECE a
Steps
Acc.
ECE g
ECE a
Steps
Acc.
ECE g
ECE a
Steps
1
100
28.91
10.89
10.77
25.74
34.38
8.40
6.61
100.70
40.63
12.62
10.44
93.91
1
200
48.44
19.63
21.55
38.02
57.03
16.42
18.14
59.38
55.47
19.27
20.04
67.27
1
500
54.69
27.78
30.58
39.11
57.81
25.77
27.89
43.26
57.81
27.02
29.19
45.13
1
1000
53.13
30.62
33.09
43.25
57.81
27.18
29.16
37.36
58.59
27.30
29.60
42.54
1
2000
57.81
26.34
29.27
40.44
57.81
27.56
29.50
37.62
57.03
29.37
31.60
42.23
Appendix
Table 7: Comparison between MIRank and BayesER on 128 examples of MBPP. Accuracy is in percent; both ECE columns are multiplied by 100.
MIRank
BayesER , τB=0.85
BayesER , τB=0.90
Ls
λ
Acc.
ECE g
ECE a
Steps
Acc.
ECE g
ECE a
Steps
Acc.
ECE g
ECE a
Steps
1
100
59.38
31.56
12.84
30.39
58.59
31.50
16.28
102.63
63.28
38.08
15.94
118.50
1
200
64.84
10.14
19.34
41.83
77.34
5.56
15.76
70.57
75.78
6.65
16.95
77.88
1
500
78.91
4.61
18.37
49.34
79.69
8.37
13.51
56.29
75.78
9.10
17.81
63.16
1
1000
77.34
8.21
19.78
52.70
76.56
10.08
18.80
53.72
78.13
7.45
16.38
60.42
1
2000
77.34
9.39
17.38
51.38
76.56
9.09
18.56
53.38
78.13
7.63
17.32
61.18
Appendix
Table 8: Comparison between MIRank and BayesER on 128 examples of GSM8K. Accuracy is in percent; both ECE columns are multiplied by 100.