Diffusion-based large language models (dLLMs) promise to break the sequential latency bottleneck of autoregressive agents through parallel decoding, but recent evaluations show this efficiency does not transfer to embodied agentic competence: dLLM-backed agents repeatedly fall into retry loops, re-issuing an action long after it has failed. We give a mechanistic account of this failure and a training-free remedy. We trace the retry loop to the adaptivity of masked decoding: the sampler commits the positions it is most confident about and defers the uncertain ones, and at a failure state the context already offers a confident fill for the deferred decision, i.e. the failed action itself, so the retry is committed without the failure feedback ever being confronted. We model the resulting distortion of the action distribution as a task-blind corruption: contextually salient actions (e.g., the action just taken) receive inflated probability by a factor that depends on the state and the action but not on the task. Under this model, we analyse an invariance proposition: the task-blind factor cancels exactly from the reverse conditional, i.e. the likelihood of the task given the state and a candidate action, which coincides with the task posterior of an idealized uncorrupted model. Masked dLLMs evaluate the reverse conditional natively, unlike autoregressive models, by masking the task tokens and denoising, at the cost of a few parallel passes per candidate. We instantiate the rule as Reflect Reverse and evaluate it on four multi-turn embodied benchmarks, where it improves task success and progression rates over forward-scoring baselines.
Figures & tables
Figure 1: Task-blind corruption analyzed over 208 failed states in ALFWorld.
Figure 2: Success rate and progress rate for LLaDA-8B-Instruct and iLLaDA-8B-Instruct.
Model
Benchmark
Success Rate (SR)
Progress Rate (PR)
Reverse
Forward
Δ (%)
Reverse
Forward
Δ (%)
LLaDA
ALFWorld
14.18±2.99
8.21±1.33
+72.72
35.26±2.29
30.59±1.79
+15.27
ScienceWorld
6.85±1.78
4.63±1.78
+47.95
22.82±2.06
21.15±2.43
+7.90
BabyAI
17.56±2.57
17.56±3.32
0.00
26.83±1.35
29.40±2.35
−8.74
Jericho
1.67±4.08
1.67±2.58
0.00
23.63±3.99
19.23±2.61
+22.88
iLLaDA
ALFWorld
10.07±1.40
0.75±0.94
+1242.67
28.20±2.13
16.08±1.56
+75.37
Table 1: Comparison of success rate (SR) and progress rate (PR) for Reflect Reverse and Reflect Forward , averaged across the two sampling temperatures. Reverse and Forward results are reported as mean ± standard deviation (%). The relative change is computed as Δ=(Reverse−Forward)/Forward×100% . Positive relative changes are shown in bold. When the Forward value is zero, the relative change is undefined and reported as N/A.
Figure 3: Retry state percentage of LLaDA-8B-Instruct and iLLaDA-8B-Instruct.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Hyperparameters
LLaDA
iLLaDA
Generation length (tokens)
128
128
Requested diffusion steps
128
128
Block size (tokens)
32
32
Threshold
0.9
0.9
Dual cache
On
Off
Configured context length (tokens)
4,000
8,192
Appendix
Table 2: Evaluation hyperparameters for LLaDA and iLLaDA on ALFWorld, ScienceWorld, BabyAI, and Jericho. The upper table reports model settings and the lower table reports forward and reverse likelihood-scoring parameters
Benchmark
Method
Temperature
SR (%)
PR (%)
ALFWorld
One Pass
0.8
6.47±2.40
26.87±1.46
0.9
8.71±1.55
29.33±0.79
Mean
7.59±2.18
28.10±1.71
Reflect Forward
0.8
7.46±1.49
29.73±2.18
0.9
8.96±0.75
31.45±1.02
Mean
8.21±1.33
30.59±1.79
Appendix
Table 3: Success rate (SR) and progress rate (PR) evaluated by LLaDA-8B-Instruct. Entries are mean ± sample standard deviation. Boldface marks the best method within each benchmark and temperature category (including Mean), with ties included.
Benchmark
Method
Temperature
SR (%)
PR (%)
ALFWorld
One Pass
0.9
5.47±0.43
19.57±1.76
1.0
4.73±0.43
21.06±1.81
Mean
5.10±0.56
20.32±1.79
Reflect Forward
0.9
0.75±0.75
15.80±1.39
1.0
0.75±1.29
16.36±1.97
Mean
0.75±0.94
16.08±1.56
Appendix
Table 4: Success rate (SR) and progress rate (PR) evaluated by iLLaDA-8B-Instruct. Entries are mean ± sample standard deviation. Boldface marks the best method within each benchmark and temperature category (including Mean), with ties included.
Look at pencil under the desklamp.
A: go to desk 1 O: On the desk 1, you see a bowl 2, a creditcard 2, a desklamp 1, a laptop 1, a mug 1, and a pen 2.
A: examine desk 1 O: On the desk 1, you see a bowl 2, a creditcard 2, a desklamp 1, a laptop 1, a mug 1, and a pen 2.
A: go to desk 2 O: On the desk 2, you see a alarmclock 1, a bowl 3, and a pencil 1.
Appendix
Table 5: Single-action example. A and O denote the action and its resulting observation. The repetition loop is marked in red including the initial action of the loop.
Put a clean egg in microwave.
A: go to fridge 1 O: The fridge 1 is closed.
A: open fridge 1 O: You open the fridge 1. The fridge 1 is open. In it, you see a apple 2, a egg 1, a lettuce 1, a pan 2, a plate 1, and a tomato 1.
A: take egg 1 from fridge 1 O: You pick up the egg 1 from the fridge 1.
Appendix
Table 6: Alternating-action example. A and O denote the action and its resulting observation. The repetition loop is marked in red including the pair of the loop.
Diffusion Large Language Models (dLLMs) have emerged as an efficient alternative to autoregressive models, yet aligning them via Reinforcement Learning (RL) requires likelihood surrogates estimated from masked reconstruction subproblems under a small Monte Carlo budget per rollout. Existing methods construct these subproblems by uniform random masking, leaving open the question of which subproblems to prioritize. We identify a systematic upstream/downstream structure in dLLM rollouts. Some tokens, when revealed, trigger large confidence changes in nearby undecoded positions; we call them upstream. Others induce only small local changes and are therefore downstream. We find masking downstream tokens yields substantially better-posed subproblems than masking upstream tokens, a phenomenon we term subproblem difficulty asymmetry. Based on the observation, we propose Informed Masking (IM), which derives a per-token priority score from the denoising trajectory at zero extra inference cost and biases mask sampling toward downstream tokens. IM is plug-and-play: when plugged into three state-of-the-art dLLM RL methods on LLaDA-8B-Instruct, it delivers up to 2.01%, 8.68%, and 5.77% relative average gains on math and planning benchmarks with improved training stability.
Xiaoyi Yu, Enver Sangineto, Pei Fu +6
Gaoling School of Artificial Intelligence, Renmin University of China, Beijing, China · MiLM Plus, Xiaomi Inc., Beijing, China · University of Modena and Reggio Emilia, Italy +1
Diffusion large language models (dLLMs) offer an efficient alternative to autoregressive models through parallel decoding, yet existing post-training methods largely rely on random masking strategies that overlook intrinsic token dependencies. In this work, we present an empirical analysis of attention in dLLMs and show that tokens attending more strongly to unmasked context exhibit greater generation stability and play a critical role in reasoning. Motivated by these findings, we propose AGDO, an attention-guided denoising and optimization framework that aligns both training and optimization with attention-derived dependencies. AGDO determines the denoising order based on attention structure and emphasizes attention-critical tokens during supervised fine-tuning and reinforcement learning. Experiments on mathematical and coding benchmarks demonstrate that AGDO consistently improves reasoning performance, outperforming state-of-the-art post-training methods for dLLMs.
Jia Deng, Junyi Li, Wayne Xin Zhao +3
Gaoling School of Artificial Intelligence, Renmin University of China · Department of Data Science, City University of Hong Kong · Beijing Key Laboratory of Research on Large Models and Intelligent Governance +2
Masked diffusion language models (dLLMs) have shown strong potential for faster inference through parallel token generation when combined with confidence-based samplers. However, recent work has shown that such methods can defer unmasking high-entropy fork positions at which multiple plausible continuations exist. This results in reduced generation diversity, as shown by worse pass@k scaling, and limits gains obtainable from RL post-training. To avoid this flexibility trap, prior work advocated for autoregressive (AR) sampling. Here, we show that discarding confidence-based sampling is unnecessary and, once inference cost is taken into account, wasteful. We first propose Fork-dLLM, a simple hybrid sampler that uses AR-style ordering only at uncertain fallback steps while retaining parallel generation otherwise. We then extend the same principle to post-training with ForkGRPO, which uses Fork-dLLM rollouts and applies the GRPO objective only at fallback steps, preserving exact policy-likelihood ratios while substantially reducing rollout and optimization cost. In our experiments, Fork-dLLM matches the strong pass@k scaling of AR sampling while being 2-3x more efficient, and ForkGRPO achieves downstream performance comparable to or better than AR-based GRPO baselines at a substantially lower training cost.
Stipe Frković, Metod Jazbec, Christian A. Naesseth
University of Amsterdam · UvA-Bosch Delta Lab, University of Amsterdam