Diffusion Large Language Models (dLLMs) have shown strong reasoning capabilities, yet further improving them typically requires costly post-training with additional data and supervision. We ask whether a post-trained dLLM can improve itself at inference time without additional training, data, or reward models. This requires a guidance signal from the checkpoints alone that is well-defined on the partially masked intermediate states of dLLMs, which existing methods fail to provide. Here we propose Reward-Free Guidance (RFG), a training-free framework for inference-time self-improvement of dLLMs that derives such a signal from the checkpoints themselves. We theoretically demonstrate that a reward signal for a partially masked state can be parameterized by the log-likelihood ratio between a post-trained policy dLLM and its reference model. In practice, this signal can be obtained using off-the-shelf checkpoints alone. Extensive experiments show that RFG consistently improves state-of-the-art post-trained dLLMs by up to 16.1%. Furthermore, despite being entirely training-free, RFG delivers gains that rival or even surpass resource-intensive reinforcement learning techniques.
Figures & tables
Figure 1: RFG consistently enhances performance over the original post-trained baseline across diverse benchmarks and model variants. Results are averaged across generation sequence lengths.
GSM8K
MATH-500
HumanEval
MBPP
Method
Train GPU-h
Avg. Acc.
Rel. Gain
Avg. Acc.
Rel. Gain
Avg. Acc.
Rel. Gain
Avg. Acc.
Rel. Gain
LLaDA-Instruct
—
79.9
—
37.8
—
39.0
—
51.9
—
+RL (LLaDA-1.5)
405
80.9
+1.3%
37.5
−0.8%
39.4
+1.0%
52.3
+0.8%
+RL (d1-LLaDA)
110+108
81.8
+2.4%
38.5
+1.9%
—
—
—
—
+RFG (ours)
0
81.0
+1.4%
39.3
+4.0%
40.7
+4.4%
54.2
+4.4%
Table 1: Comparison of RFG (training-free) with RL. Accuracy is averaged across generation lengths. Training time is measured in H100 GPU hours. d1-LLaDA trains a separate checkpoint for each benchmark, so its two numbers are the training time for GSM8K and MATH-500, respectively. Best is in bold and second best is underlined .
Figure 2: Illustration of RFG sampling. By linearly combining logits from the policy and reference models at each denoising step, RFG steers the generation trajectory toward accurate reasoning traces, resulting in improved performance without external reward models.
Seq Len = 128
Seq Len = 256
Seq Len = 512
Model
Original
CFG
RFG
Rel. Gain
Original
CFG
RFG
Rel. Gain
Original
CFG
RFG
Rel. Gain
Instruction Fine-tuning
LLaDA 8B Instruct
76.3
72.6
77.2
+1.2%
79.8
80.7
81.3
+1.9%
83.7
81.5
84.6
+1.1%
Dream 7B Instruct
67.9
63.6
74.2
+9.3%
80.9
72.2
82.1
+1.5%
84.8
80.1
85.1
+0.4%
Reinforcement Learning
d1-LLaDA
79.3
77.4
79.8
+0.6%
82.5
81.3
84.7
+2.7%
83.6
82.2
84.6
+1.2%
Table 2: Performance on GSM8K across different sequence lengths. Best in bold . Relative gain of RFG over the original post-trained model in green .
GSM8K
MATH-500
Model
Original
CFG
RFG
Rel. Gain
Original
CFG
RFG
Rel. Gain
Instruction Fine-tuning
LLaDA 8B Instruct
78.6
77.9
80.7
+2.7%
34.2
28.6
38.0
+11.1%
Dream 7B Instruct
68.5
57.2
73.7
+7.6%
35.4
19.4
37.6
+6.2%
Reinforcement Learning
d1-LLaDA
79.8
78.3
82.0
+2.8%
36.6
30.0
39.2
+7.1%
Table 6: Performance on mathematical reasoning tasks with a sequence length of 256. This setting doubles the number of tokens generated per step compared to Table 5 and Table 5 .
Figure 3: Accuracy under varying guidance strength w . GSM8K and MATH-500 use LLaDA-1.5 with sequence length 128, HumanEval and MBPP use Dream-Instruct with sequence length 512. RFG consistently improves performance over a broad range of guidance strength.
LLaDA 8B Instruct
Dream 7B Instruct
Method
P-strict
P-loose
I-strict
I-loose
P-strict
P-loose
I-strict
I-loose
Original
55.6
60.6
65.7
69.8
55.6
59.1
67.1
70.0
RFG
60.1
64.1
68.3
71.5
57.7
60.4
68.3
70.6
Table 8: IFEval instruction-following accuracy (%). P: prompt-level, I: instruction-level, each with the official strict and loose variants. Best in bold .
GSM8K
MATH-500
Model
Original
Sharpen
RFG
Original
Sharpen
RFG
LLaDA-Instruct
83.7
82.9
84.6
42.6
42.2
44.0
Dream-Instruct
84.8
83.0
85.1
50.4
49.4
52.6
d1-LLaDA
83.6
83.9
84.6
42.6
41.8
43.8
LLaDA-1.5
84.8
84.4
85.3
42.0
42.8
44.0
Table 9: Policy-only sharpening control at sequence length 512. Sharpen applies (1+w)logπθ with the same w as RFG . Best in bold .
Figure 4: Qualitative examples for mathematical reasoning task.
Figure 5: Qualitative examples for code generation task.