Supervised fine-tuning of discrete diffusion language models masks some response tokens and trains the model to recover their original values from the visible context. The masking pattern therefore determines both the context available to the model and the tokens it learns to predict. Uniform random masking does not explicitly account for the interaction between these choices. We introduce GoldiMask, which selects tokens to reveal as context by approximately maximizing a submodular objective. This objective uses model signals to balance the benefit of revealing tokens against their value as prediction targets. GoldiMask then weights the remaining targets according to how they benefit from the selected context and their remaining learning potential. Across three backbones and three training datasets, GoldiMask achieves the highest average accuracy in most evaluated settings, demonstrating gains on both reasoning and code generation. Component ablations show that both context selection and target weighting contribute to the gains. GoldiMask also reduces decoding iterations on GSM8K and MATH-500 under confidence-threshold parallel decoding, while maintaining comparable accuracy at higher confidence thresholds.
Figures & tables
Figure 1: Overview of GoldiMask. A first forward pass on the over-masked response provides ground-truth token probabilities and attention scores for selecting a reveal set. A second forward pass evaluates the remaining targets with the selected context. Their positive context gains and post-reveal priorities determine loss weights through water-filling for the training update. The supervision-priority function λ and support response ϕ remain fixed throughout fine-tuning.
Method
Venue
GSM8K
MATH
Countdown
Sudoku
Average
LLaDA-8B-Instruct
LLaDA-8B
–
78.1
36.1
19.6
11.2
36.3
Vanilla ( Sahoo et al., 2024 )
NeurIPS’24
79.2 ± 0.2
36.2 ± 0.9
17.2 ± 2.1
14.5 ± 1.3
36.8 ± 0.3
GIFT ( Xu et al., 2026 )
ICLRw’26
81.0 ± 0.1
37.3 ± 0.4
21.5 ± 2.8
15.8 ± 0.4
38.9 ± 0.5
CART ( Ye et al., 2025 )
arXiv’25
80.4 ± 1.0
35.9 ± 1.6
24.0 ± 1.4
14.4 ± 1.6
38.7 ± 0.6
AGDO † ( Deng et al., 2026 )
ACL’26
78.3 ± 0.8
37.7 ± 0.1
16.9 ± 1.3
12.6 ± 0.7
36.4 ± 0.3
Table 1: Main results. Fine-tuning on s1K across three diffusion language models. Values are means with standard deviations in smaller ± entries. Baselines are retrained in our harness with matched hyperparameters. ∗ LIFT is the stronger of its H=2 and H=3 variants per setting, by four-benchmark average. Relative changes are against the strongest baseline in each block. † denotes our code implementation; ‘w’ denotes a workshop.
Name
GSM8K
MATH
Countdown
Sudoku
Average
LLaDA-8B
78.1
36.1
19.6
11.2
36.3
Vanilla
82.9 ± 1.0
34.6 ± 0.4
19.9 ± 0.3
9.3 ± 0.7
36.7 ± 0.3
GIFT
79.8 ± 0.3
36.1 ± 0.1
23.6 ± 4.7
8.8 ± 1.2
37.1 ± 1.0
CART
80.2 ± 0.9
36.4 ± 2.8
22.4 ± 3.5
15.6 ± 1.2
38.6 ± 1.8
LIFT ∗
81.3 ± 0.7
36.6 ± 0.5
23.6 ± 1.5
11.3 ± 0.3
38.2 ± 0.4
AGDO †
82.3 ± 0.5
34.8 ± 0.4
19.4 ± 1.7
10.3 ± 0.9
36.7 ± 0.3
Table 2: A second corpus. LLaDA-8B-Instruct fine-tuned on LIFT-SFT-12K at an equal step budget. Conventions and markers follow Table 1 .
LLaDA-8B-Instruct
Dream-7B
Name
HE
MBPP
Avg.
HE
MBPP
Avg.
Base model
34.1
41.3
37.7
39.6
52.5
46.1
Vanilla SFT
30.8 ± 1.3
41.6 ± 1.7
36.2 ± 0.2
45.4 ± 5.6
42.4 ± 1.7
43.9 ± 3.6
GIFT
35.4 ± 1.7
40.5 ± 2.8
38.0 ± 0.5
44.5 ± 1.7
46.5 ± 0.8
45.5 ± 0.4
CART
33.5 ± 1.7
39.9 ± 1.4
36.7 ± 0.2
47.5 ± 4.7
50.4 ± 0.3
49.0 ± 2.2
AGDO †
32.9 ± 0.9
40.1 ± 2.2
36.5 ± 0.7
41.8 ± 3.0
53.7 ± 5.5
47.7 ± 4.3
Table 3: Code generation. HE denotes HumanEval; MBPP uses the sanitized test split. Relative changes use each backbone’s strongest fine-tuned baseline; other conventions follow Table 1 .
Figure 2: Accuracy versus parallel-decoding steps. GoldiMask-D/R/P (solid) and SFT baselines (dashed) on GSM8K and MATH-500. Markers show τ=0.70,0.80,0.90,0.95 from left to right; lines connect measured settings. Upper-left is better. L denotes generation length.
Reveal selection
Target weighting
GSM8K
MATH
Countdown
Sudoku
Avg.
None (vanilla SFT)
Uniform
79.2
36.2
17.2
14.5
36.8
Random
Water-filling
80.1
37.5
24.6
14.5
39.2
Static initial-marginal
Water-filling
80.2
37.8
24.6
14.9
39.4
GoldiMask-D
Uniform
79.4
38.1
26.8
15.3
39.9
GoldiMask-D
Water-filling
80.4
39.5
27.7
16.9
41.1
Table 4: Contributions of selection and weighting. Accuracy on LLaDA-8B-Instruct fine-tuned on s1K. All reveal-based configurations use two passes. Bold marks the best result in each column.
Priority
GSM8K
MATH
Countdown
Sudoku
Avg.
λ∝1−p (peak at 0 )
79.2
38.6
20.7
15.9
38.6
λ∝p(1−p) (peak at 1/2 )
80.9
38.0
27.0
16.1
40.5
λ∝p(1−p)2 (peak at 1/3 )
81.0
38.4
26.3
16.9
40.7
λ∝p(1−p)3 (ours, peak at 1/4 )
80.4
39.5
27.7
16.9
41.1
Table 5: Supervision-priority ablation. Accuracy on LLaDA-8B-Instruct fine-tuned on s1K, averaged over three seeds. Avg. is the mean across the four benchmarks. Bold marks the highest average.
Weighting
GSM8K
MATH
Countdown
Sudoku
Avg.
Uniform
79.4
38.1
26.8
15.3
39.9
DRO-tilted hardness
80.6
39.4
24.6
17.9
40.6
Uncapped utility weighting
78.4
37.3
27.1
14.7
39.4
Water-filling ( c=5 )
80.1
38.9
27.5
17.1
40.9
Water-filling ( c=15 )
80.1
38.7
27.9
15.8
40.6
Water-filling ( c=10 , ours)
80.4
39.5
27.7
16.9
41.1
Table 6: Loss-weighting comparison. Accuracy on LLaDA-8B-Instruct fine-tuned on s1K, averaged over three seeds. Bold marks the highest average.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Maximizer
Deterministic
Reference
Greedy (D)
yes
Nemhauser et al. (1978) , Math. Programming
Randomized greedy (R)
no
Buchbinder et al. (2014) , SODA
TwinGreedy (T)
yes
Han et al. (2020) , NeurIPS
Practical 0.385 (P)
no
Tukan et al. (2024) , NeurIPS
Appendix
Table 7: The four maximizers used for reveal-set selection. The selection procedures below describe our implementations.
Name
GSM8K
MATH
Countdown
Sudoku
Average
LLaDA-8B-Instruct / s1K
LIFT ∗
79.1
36.6
28.8
13.5
39.5
GoldiMask-D
80.4
39.5
27.7
16.9
41.1
vs. LIFT
↑ 1.6%
↑ 7.9%
↓ 3.8%
↑ 25.2%
↑ 4.1%
GoldiMask-R
80.4
38.3
29.7
14.4
40.7
vs. LIFT
↑ 1.6%
↑ 4.6%
↑ 3.1%
↑ 6.7%
↑ 3.0%
Appendix
Table 8: All maximizers on all settings. Same benchmarks, protocol and conventions as Tables 1 and 2 . The baseline row is the strongest baseline of each setting by four-benchmark average; GoldiMask-D, GoldiMask-R and baseline rows are those of the main text. GoldiMask-T and -P are two-seed means using validation-selected checkpoints. The relative change beneath each GoldiMask row is measured against that baseline. Bold marks the best entry per column within a setting.
LLaDA-8B / s1K
Method
GSM8K
MATH
Countdown
Sudoku
Average
GIFT
81.0 ± 0.1
37.3 ± 0.4
21.5 ± 2.8
15.8 ± 0.4
38.9 ± 0.5
CART
80.4 ± 1.0
35.6 ± 2.0
23.8 ± 1.7
14.4 ± 1.6
38.6 ± 0.6
AGDO
78.3 ± 0.8
37.2 ± 0.7
15.5 ± 1.3
12.4 ± 0.6
35.8 ± 0.5
LIFT
79.1 ± 0.8
35.9 ± 1.0
24.9 ± 3.0
13.5 ± 1.0
38.4 ± 0.7
GoldiMask-D
80.4 ± 1.2
39.5 ± 0.8
21.9 ± 2.0
16.3 ± 1.0
39.5 ± 0.8
Appendix
Table 9: Fixed task-specific generation lengths. Accuracy (%) at GSM8K/MATH-500/Countdown/Sudoku lengths of 512/512/128/128. Bold marks the highest displayed mean in each column within a panel.
Method
LLaDA-8B
LLaDA-1.5
Dream-7B
Vanilla
6.7 ± 0.0
7.8 ± 1.9
6.7 ± 3.3
GIFT
8.9 ± 3.8
8.9 ± 1.9
8.9 ± 5.1
CART
12.2 ± 3.8
14.4 ± 5.1
11.1 ± 5.1
LIFT
12.2 ± 5.1
12.2 ± 1.9
8.9 ± 3.8
GoldiMask-D
10.0 ± 3.3
12.2 ± 1.9
13.3 ± 6.7
Appendix
Table 10: AIME24 pass@16 after fine-tuning on s1K. Values are mean accuracy (%) ± sample standard deviation across three seeds. Bold marks the highest mean in each column.
Figure 3: Accuracy versus decoding cost on Countdown and Sudoku. Same conventions as Figure 2 : three backbones, thresholds τ∈{0.70,0.80,0.90,0.95} from left to right, GoldiMask-D, -R and -P solid against the SFT baselines dashed. Countdown uses L=128 on all backbones.
Method
Training time relative to Vanilla SFT
Vanilla SFT
1.00×
CART
1.01×
GIFT
1.17×
AGDO
1.17×
LIFT
1.24×
GoldiMask-D
1.21×
Appendix
Table 11: Relative training time. LLaDA-8B-Instruct on four B200 GPUs. Values are averaged across s1K and LIFT-SFT-12K, with Vanilla SFT as the 1.00× reference.
LLaDA-8B
LLaDA-1.5
Supervision rule
ΔNLL↓
Δp↑
ΔNLL↓
Δp↑
Uniform random
−0.013
+0.013
−0.015
+0.011
1−p
−0.017
−0.013
−0.019
−0.014
p(1−p)
−0.001
+0.052
−0.003
+0.051
Our priority p(1−p)3
−0.016
+0.032
−0.018
+0.030
Appendix
Table 12: Learning under a fixed supervision budget. Changes after 300 training steps, with 32 supervised positions per training example. ΔNLL is the change in held-out negative log-likelihood; Δp is the mean correct-token probability change for held-out targets with initial p<0.9 . Arrows indicate the preferred direction. Bold values are best within each column; shading identifies our priority.
Model / corpus
Fitted peak [95% CI]
Held-out R2
LLaDA-8B / s1K
0.159[0.148,0.170]
0.958
LLaDA-1.5 / s1K
0.157[0.147,0.168]
0.952
Dream-7B / s1K
0.164[0.154,0.177]
0.968
LLaDA-8B / 12K
0.163[0.144,0.179]
0.973
Appendix
Table 13: Study 2: absolute probability gain indexed by pre-reveal confidence. Peak intervals use an example-level bootstrap. The power-family fit is trained on half the examples and evaluated on the other half.
Measurement
Result
Held-out R2 , masking interval 0.50 – 0.71
0.033→0.041
Held-out R2 , masking interval 0.71 – 0.84
0.082→0.098
Fitted half-saturation scale s0 across noise bins
0.05 – 0.2
Appendix
Table 14: Study 3: contextual support and diminishing returns. Support fits on LLaDA-8B. Arrows compare held-out within-target R2 for the linear and saturating responses, with s0=0.05 for the latter.
Figure 4: GoldiMask’s functional forms (analytical, not empirical). Left: priorities normalized to unit maximum. Right: support responses for s0=0.05 (default), 0.1 , and 0.2 . Empirical fits appear in Table 14 .
Reinforcement learning has proven effective for improving reasoning in large language models, but extending it to Masked Diffusion Language Models (MDLMs) remains challenging due to the intractability of the log-likelihood estimation. Existing approaches approximate this log-likelihood by modeling only the token predictions, ignoring the order in which positions are unmasked during generation. We observe that MDLM generation involves two decisions at each step: what tokens to place at each masked position and which positions to remask. We formalize this as a two-stage action MDP, showing that the policy gradient naturally decomposes into a token term and a masking term. Combining optimization of both terms leads to state-of-the-art outcomes on mathematical reasoning and coding benchmarks, with scores of 87.1% on GSM8K and 53.4% on MBPP.
Haran Raajesh, Kulin Shah, Adam Klivans +1
Department of Computer Science The University of Texas at Austin
Masked diffusion language models (MDLMs) generate text by iteratively unmasking tokens, but their standard decoder reduces each step to a binary action: a position is either committed to a single token or left fully masked, with no representation of partial belief in between. This all-or-nothing regime discards rich predictive information and forces premature, irrevocable commitments, leading to poor performance under a limited decoding budget. In this paper, we reinterpret mask prediction as clean-state prediction (x-prediction) and show that it can be used to induce a continuous flow in input embedding space. Building on this view, we propose a continuous decoding framework for MDLMs where tokens can accumulate partial progress at each diffusion step and remain revisable. To match the uneven contextual constraints across positions in language, we replace the globally synchronous schedule in image diffusion with a confidence-based asynchronous update in which the diffusion progress is token-wise accumulated. Additionally, we introduce a lightweight policy network and formulate its training as a reinforcement learning problem. Applied to pretrained LLaDA, our continuous decoder reaches 97% of its performance on the HumanEval dataset with 25% of decoding budget.
Weitian Wang, Lianlei Shan, Shubham Rai +2
Robert Bosch GmbH, Germany · University of the Chinese Academy of Sciences, China · Ruhr University Bochum, Germany
Masked diffusion language models (MDLMs) can generate text efficiently by predicting multiple masked tokens in parallel, but predictions from the same forward pass are not necessarily reliable when committed together. We study when parallel commitment is reliable. Our diagnostics show that confidence alone does not determine a reliable commitment order: confident predictions near the end of the sequence can fix an answer before its supporting computations are established, and downstream predictions become less reliable as the uncertainty of their upstream context grows. At the same time, a single forward pass can already resolve several masked tokens, and predictions that remain stable across the final layers are more likely to be correct. Based on these findings, we propose Reliable Parallel Decoding (RPD), a training-free method that selects candidates by layerwise prediction stability and final confidence, and commits them under a cumulative entropy budget over their preceding masked positions. RPD defers predictions with uncertain upstream context while committing the remaining candidates in parallel, without relying on a fixed block schedule. Across mathematical reasoning and code generation benchmarks on LLaDA and Dream, RPD achieves the highest decoding throughput among the evaluated methods while maintaining or improving accuracy.
Zhenghao He, Bohan Liu, Guangzhi Xiong +1
Department of Computer Science, University of Virginia