Masked diffusion language models (dLLMs) have shown strong potential for faster inference through parallel token generation when combined with confidence-based samplers. However, recent work has shown that such methods can defer unmasking high-entropy fork positions at which multiple plausible continuations exist. This results in reduced generation diversity, as shown by worse pass@k scaling, and limits gains obtainable from RL post-training. To avoid this flexibility trap, prior work advocated for autoregressive (AR) sampling. Here, we show that discarding confidence-based sampling is unnecessary and, once inference cost is taken into account, wasteful. We first propose Fork-dLLM, a simple hybrid sampler that uses AR-style ordering only at uncertain fallback steps while retaining parallel generation otherwise. We then extend the same principle to post-training with ForkGRPO, which uses Fork-dLLM rollouts and applies the GRPO objective only at fallback steps, preserving exact policy-likelihood ratios while substantially reducing rollout and optimization cost. In our experiments, Fork-dLLM matches the strong pass@k scaling of AR sampling while being 2-3x more efficient, and ForkGRPO achieves downstream performance comparable to or better than AR-based GRPO baselines at a substantially lower training cost.
Figures & tables
Figure 1: GSM8K accuracy across post-training methods, evaluated with Fast-dLLM at λ=0.95 and temperature τ=0 . The dashed gray line denotes the performance of the base LLaDA-8B-Instruct model. ForkGRPO and JustGRPO-Fast are trained by us; for JustGRPO, we evaluate the released checkpoint from Ni et al. (2026) . Under the same training budget, ForkGRPO substantially outperforms JustGRPO-Fast and reaches accuracy comparable to JustGRPO despite using 8× fewer H100-hours for post-training.
Figure 2
Figure 2: Token frequencies and entropies from Fast-dLLM sampling on HumanEval using LLaDA-8B-Instruct.
Figure 3: Pass@ k (top) and pass@NFE (bottom) curves on LLaDA-8B-Instruct across mathematics and code benchmarks. Both rows show the same generations; the bottom row instead accounts for the number of NFEs required to obtain the corresponding samples. Fork-dLLM, together with the other two confidence-based samplers (Fast-dLLM and TCT), retains the strong pass@ k scaling of AR sampling while requiring substantially fewer NFEs to reach a given performance level. Insets zoom into the high-accuracy regime.
Figure 4: The 30 most frequent tokens in fallback steps (a) and parallel steps (b) from Fork-dLLM on MATH-500 using LLaDA-8B-Instruct. Word font size is scaled to frequency.
Figure 5: Pass@NFE curves of Fork-dLLM and its ablations on HumanEval using LLaDA-8B-Instruct.
Figure 6: Reward, ratio of fallback tokens, and entropy of the tokens committed in parallel and fallback steps over GRPO training on GSM8K. FastGRPO is shown with the tuned Fast-dLLM values, (λ,τ)=(0.7,1.2) , and with those of ForkGRPO, (0.8,0.8) . In an attempt to improve stability, this experiment adds a KL penalty of β=0.02 ; we observe similar behaviour at β=0 .
Figure 7: Training reward (top) and final-checkpoint accuracy versus NFEs (bottom) for ForkGRPO and JustGRPO-Fast under the same training budget ( 4× H100, 12 h). Dashed vertical lines in the top row mark the final GRPO step reached by each method: owing to cheaper rollouts and policy optimization, ForkGRPO completes more than twice as many training steps. The bottom row evaluates the resulting checkpoints with Fast-dLLM decoding ( τ=0 ), sweeping the confidence threshold λ . ForkGRPO yields consistently larger improvements over the base model than JustGRPO-Fast.
ForkGRPO
JustGRPO-Fast (relative)
Benchmark
Step
Acc.
Bwd. passes
Acc. ( Δ )
Bwd. passes ( × )
GSM8K
50
82.9
3.2 k
−0.6
1.72×
100
83.8
5.3 k
+0.1
1.92×
MATH-500
25
35.8
4.2 k
−0.2
1.60×
50
37.0
8.0 k
−1.6
1.63×
Table 1: Step-matched accuracy and backward passes, evaluated at λ=0.95 . JustGRPO-Fast is reported relative to ForkGRPO, with accuracy differences in percentage points.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Method
λ
τ
AR
–
0.6
Fork-dLLM
0.8
0.8
Fast-dLLM
0.7
1.2
TCT
0.7
0.8
Appendix
Table 2: Selected hyperparameter values for each sampling method.
Figure 8: Pass@ k and pass@NFE curves for Fork-dLLM. AR is shown as a baseline.
Figure 9: Pass@ k and pass@NFE curves for Fast-dLLM. AR and Fork-dLLM are shown as baselines.
Figure 10: Pass@ k and pass@NFE curves for TCT. AR and Fork-dLLM are shown as baselines.
Group
Hyperparameter
Value
Objective
Group Size ( G )
16
Prompts per Step
16
Groups with σG=0
Skipped
KL Penalty ( β )
0
Policy Update Steps ( μ )
1
Top-Entropy Ratio ∗
20%
Appendix
Table 3: GRPO hyperparameters. The row marked ∗ applies only to JustGRPO-Fast. With a single policy update step per rollout, the clipping range ϵ never binds and is omitted.
Figure 11: Pass@ k (top) and pass@NFE (bottom) curves on Dream-v0-Instruct-7B.
Method
GSM8K
MATH-500
HumanEval
MBPP
Pass@1
Pass@64
NFE
Pass@1
Pass@64
NFE
Pass@1
Pass@64
NFE
Pass@1
Pass@64
NFE
Fork-dLLM
80.9
98.3
67.8
31.7
76.6
92.3
31.7
78.6
108.4
33.8
74.1
75.8
Fast-dLLM
80.1
97.3
58.3
32.1
73.4
76.0
35.7
75.0
87.9
36.7
73.3
60.5
TCT
75.1
98.3
51.5
28.5
73.4
67.3
31.4
73.2
74.1
33.7
72.1
54.0
AR
80.9
99.0
227.9
31.4
74.4
243.3
34.0
78.8
242.3
34.8
74.7
192.3
Appendix
Table 4: Pass@ 1 and pass@ 64 results on LLaDA-8B-Instruct. Best per column in bold.
Method
GSM8K
MATH-500
HumanEval
MBPP
Pass@1
Pass@64
NFE
Pass@1
Pass@64
NFE
Pass@1
Pass@64
NFE
Pass@1
Pass@64
NFE
Fork-dLLM
81.7
99.3
78.0
38.0
83.4
97.3
46.8
92.8
63.9
50.6
85.6
43.3
Fast-dLLM
81.2
99.3
67.4
38.4
82.8
82.6
45.9
90.2
51.2
51.3
83.4
34.7
TCT
74.2
99.3
59.9
33.7
80.0
73.1
39.9
89.6
47.6
45.6
84.0
35.6
AR
82.4
99.0
198.5
37.9
83.0
223.1
51.1
92.7
131.2
53.0
88.4
71.9
Appendix
Table 5: Pass@ 1 and pass@ 64 results on Dream-v0-Instruct-7B. Best per column in bold.
Figure 12: Pass@ k curves of Fork-dLLM and its ablations on HumanEval using LLaDA-8B-Instruct.
Figure 13: Accuracy versus NFEs of the final ForkGRPO and JustGRPO checkpoints on GSM8K (left) and MATH-500 (right), evaluated with Fast-dLLM (top) and EB (bottom) across thresholds. The JustGRPO checkpoint is taken from Ni et al. (2026) .
A popular selling point of diffusion large language models (dLLMs) is their capacity for parallelism: the ability to generate sequences of text far more efficiently than autoregressive models, which require one forward pass per token. Yet among the many competing paradigms for dLLMs, from masked to uniform to Gaussian diffusion, principled understanding of how these different proposals compare in parallelism remains limited. In this work, we initiate a fine-grained comparison of the capacity for parallelism among these three leading approaches and prove the following: - Uniform and Gaussian diffusion can sample in a number of forward passes which scales with the dual total correlation of the underlying distribution, a measure of intrinsic complexity which can be much smaller than the context length. Previously, it was only known how to achieve this using masked diffusion. - For a certain family of random empirical measures, we show that Θ(d) forward passes are necessary and sufficient to sample using uniform or Gaussian diffusion, yet there exist approximate score oracles for which Ω(d) forward passes are needed for masked diffusion. This establishes the first provable separation in parallelism between the three prevailing dLLM paradigms. Contrary to popular intuition that masked diffusions are harder to parallelize because they must commit to token values, the latter separation instead comes from the fact that the critical windows in masked diffusion sampling are asymptotically narrower than those in uniform and Gaussian diffusion sampling.
Masked diffusion language models can reduce inference steps by revealing multiple tokens per denoising iteration, but this parallelism is fragile: positions that are individually confident may be unsafe to commit together when their predictions are coupled. Existing training-free samplers such as Top-k, Fast-dLLM, and EB-Sampler mainly control how many tokens to reveal, while often ranking candidates by token-wise scores that ignore interactions within the selected set. We propose ADAS, a training-free reranking rule that leaves the base sampler's stopping rule unchanged and greedily discounts each token-wise confidence score according to its attention to already selected positions, weighted by their prediction uncertainty. Across LLaDA-8B-Base and Dream-7B-Base on the reasoning benchmarks GSM8K and MATH500 and the code benchmarks HumanEval and MBPP, plugging ADAS into all three samplers improves low-NFE performance at matched denoiser evaluations by 9.11 and 10.46 percentage points on average, respectively, with 3.1% per-forward runtime overhead. Code is available at https://github.com/yusufsahin99/ADAS.
Yusuf Sahin, Ahmed Rockey Saikia, Volkan Cevher +1
University of Bern, Bern, Switzerland · EPFL, Lausanne, Switzerland
Diffusion large language models (dLLMs) offer an efficient alternative to autoregressive models through parallel decoding, yet existing post-training methods largely rely on random masking strategies that overlook intrinsic token dependencies. In this work, we present an empirical analysis of attention in dLLMs and show that tokens attending more strongly to unmasked context exhibit greater generation stability and play a critical role in reasoning. Motivated by these findings, we propose AGDO, an attention-guided denoising and optimization framework that aligns both training and optimization with attention-derived dependencies. AGDO determines the denoising order based on attention structure and emphasizes attention-critical tokens during supervised fine-tuning and reinforcement learning. Experiments on mathematical and coding benchmarks demonstrate that AGDO consistently improves reasoning performance, outperforming state-of-the-art post-training methods for dLLMs.
Jia Deng, Junyi Li, Wayne Xin Zhao +3
Gaoling School of Artificial Intelligence, Renmin University of China · Department of Data Science, City University of Hong Kong · Beijing Key Laboratory of Research on Large Models and Intelligent Governance +2