Diffusion large language models (dLLMs) generate text by denoising a sequence or successive blocks, allowing several tokens to be revealed in parallel. Reinforcement learning with verifiable rewards (RLVR) reuses terminal feedback across these decisions, even as their conditioning context changes. We propose stepwise risk-sensitive GRPO (StepRS-GRPO), which varies the risk coefficient of the group-advantage transformation across denoising states while retaining the underlying trainer. For binary rewards, we show that this transformation is exactly a prompt- and state-dependent rescaling of centered outcome advantages. A capability-based calibration suggests a coefficient scale, while endpoint and interpolation ablations guide schedule selection. Across multiple dLLM backbones and mathematical reasoning benchmarks, StepRS-GRPO improves both pass@1 accuracy and pass@k coverage over centered GRPO, while increasing answer diversity. In our ablation studies, mass-matched controls support the contributions of state allocation and schedule direction, and the gains persist after matching the root mean square (RMS) of the advantages to that of centered GRPO. Reasoning-trace diagnostics further show that the diversity gains from StepRS-GRPO extend beyond final-answer strings.
Figures & tables
Figure 1: Pass@ k on SDAR-1.7B-Chat and SDAR-4B-Chat at the 160 -update checkpoints. Markers show means and capped error bars show one sample standard deviation across inference seeds 42 , 43 , and 44 , with 64 responses per problem per seed. The selected StepRS schedules are [1.0→0.5] and [1.5→0.75] , respectively. Answer diversity is reported in Figure 12 .
Figure 2: SDAR-4B mechanism ablations. Top: StepRS versus prompt-only mass matching (C1) and a mass-normalized reversed schedule (C2). Bottom: GRPO, constant-risk RS-GRPO ( β=1.5 ), StepRS, and advantage-RMS-matched StepRS (C3). Markers show three-inference-seed means; capped error bars show one sample standard deviation. Controls are defined in the text, with implementation details in Appendix D.1 .
Figure 3: Answer-removed trace diversity on SDAR-4B: (a) Dngram ; (b) Demb . Points and capped bars show means and sample standard deviations across three inference seeds, with 64 responses per question per seed. Higher scores indicate greater variation; panel scales differ. Definitions are in Appendix C.3 .
Figure 4: Schedule-shape ablation on SDAR-4B with fixed endpoints [1.5→0.75] . Linear, cosine, and two-stage schedules are compared at all seven sampling budgets. Markers show means and capped error bars show one sample standard deviation across three inference seeds.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Backbone
Default StepRS schedule
Constant-risk sweep
Qwen3-0.6B-MDLM
[0.5→0.25]
{0.25,0.5,1.0}
SDAR-1.7B-Chat
[1.0→0.5]
{0.25,0.5,1.0}
SDAR-4B-Chat
[1.5→0.75]
{0.5,1.0,1.5,2.0}
Appendix
Table 1: Selected schedules and evaluated constant-risk coefficients.
Figure 5: Qwen3-0.6B-MDLM evaluation curves. The selected [0.5→0.25] schedule improves pass@ k over GRPO across GSM8K, MATH500, and AMC23. All curves show means and capped error bars show one sample standard deviation over three seeds.
Hyperparameter
Qwen3-0.6B-MDLM
SDAR-1.7B-Chat
SDAR-4B-Chat
Base model (HuggingFace checkpoint)
dllm-hub/Qwen3-0.6B-diffusion-mdlm-v0.1
JetLM/SDAR-1.7B-Chat
JetLM/SDAR-4B-Chat
Trainer framework
TraceRL ( Wang et al., 2025c )
StableDRL ( Zhong et al., 2026 )
Training data
GSM8K train
MATH train
Optimizer
AdamW (lr= 5×10−6 )
AdamW (lr= 1×10−6 )
Responses per prompt ( G )
16
Group rollout batch
1024 (64 × 16)
512 (32 × 16)
Appendix
Table 2: Training hyperparameters used for the three model scales.
Figure 6: Calibration and endpoint ablations on SDAR-4B. Top: linear half-decay schedules with upper endpoints 1.0 , 1.08 , 1.5 , and 2.0 . Bottom: full decay versus half-decay at a fixed upper endpoint of 2.0 . Columns correspond to the four benchmarks. Curves show means and one-standard-deviation bars across three inference seeds.
Figure 7: The selected [1.5→0.75] StepRS schedule versus constant risk coefficients on SDAR-4B. Each panel shows the complete pass@ 1 –pass@ 64 profile for one benchmark, with mean and one-standard-deviation bars over three inference seeds.
Figure 8: SDAR-1.7B constant-risk comparison: β∈{0.25,0.5,1.0} versus StepRS [1.0→0.5] . Curves show means with one-standard-deviation bars across three inference seeds.
Figure 9: SDAR-1.7B endpoint comparison: [0.25→0] , [0.5→0.25] , [1.0→0] , and the default [1.0→0.5] . All schedules are linear; curves show three-seed means and one-standard-deviation bars.
Figure 10: Qwen3-0.6B-MDLM constant-risk comparison. The default StepRS schedule is [0.5→0.25] . Curves show three-seed means with capped one-standard-deviation error bars. All three constant coefficients β∈{0.25,0.5,1.0} are shown.
Figure 11: Qwen3-0.6B-MDLM endpoint comparison: [0.5→0] , [1.0→0] , and the default [0.5→0.25] , all with linear interpolation. Curves show means with one-standard-deviation error bars across three inference seeds.
Figure 12: Unique-answer ratios for Base, GRPO, and StepRS across all three backbones. Points show means and capped error bars show sample standard deviations over inference seeds 42 , 43 , and 44 , with 64 responses per problem per seed.
Diffusion large language models are a compelling alternative to autoregressive models, yet existing RL methods for diffusion treat all denoising steps as equally important and rely on biased, high-variance likelihood estimates. We identify two fundamental weaknesses: the absence of temporal credit assignment across the denoising trajectory, and the systematic bias of mean-field likelihood estimates used for policy optimization. To address these, we propose Denoising-Aware Credit Assignment for GRPO (DACA-GRPO), a lightweight, plug-and-play enhancement for any GRPO-style trainer. DACA-GRPO introduces two complementary mechanisms: Denoising Progress Scores, which extract per-token importance weights from intermediate predictions at no additional forward cost, and Stratified Masking Likelihood, which partitions token positions into strata so that each token is predicted with most of the sequence as context, reducing the mean-field bias. Applied on top of three GRPO base methods, DACA-GRPO achieves consistent improvements across seven benchmarks spanning mathematical reasoning, code generation, constraint satisfaction, and constrained generation, with gains of up to 5.6pp on math reasoning, 7.4pp on code generation, 36.3pp on constraint satisfaction, and 5.9pp on JSON schema adherence.
Reinforcement learning (RL) holds immense promise for enhancing the reasoning capabilities of diffusion large language models (dLLMs). However, progress is fundamentally constrained by a dual misalignment between authentic generation trajectory and the gradient update process: (i) Process-reward misalignment. Sparse, terminal rewards are indiscriminately assigned to all intermediate steps of the generation process, failing to provide discriminative credit assignment. (ii) State-trajectory misalignment. Policy updates are often diverted toward artificial, out-of-trajectory states, squandering gradients on less informative samples. To address these limitations, we introduce Process Aligned Policy Optimization (PAPO), a novel framework that holistically aligns the RL update with the dLLM's generative trajectory via Step-Aware Process Rewards (SPR) that transform sparse terminal rewards into dense, step-wise credit, and Entropy-Guided Historical Re-enactment (EHR) that replays authentic trajectories at high-uncertainty steps. Extensive experiments on four benchmarks demonstrate that PAPO significantly outperforms baselines, achieving gains of up to 4.5% on GSM8K, 4.8% on MATH500, 42.2% on Countdown and 16.1% on Sudoku.
Yawen Shao, Jie Xiao, Kai Zhu +6
University of Science and Technology of China · Tongyi Lab · Northeastern University
Diffusion large language models (dLLMs) offer a promising route to parallel and efficient text generation, but improving their reasoning ability requires effective post-training. Reinforcement learning with verifiable rewards (RLVR) is a natural choice for this purpose, yet its application to dLLMs is hindered by the absence of tractable sequence-level log-ratios, which are central to standard policy optimization. The lack of tractable sequence-level log-ratios forces existing methods to rely on high-variance ELBO-based approximations, where high verifier rewards can amplify inaccurate score estimates and destabilize RL training. To overcome this issue, we propose \textbf{R}elative \textbf{S}core \textbf{P}olicy \textbf{O}ptimization (RSPO), a simple RLVR method that uses verifiable rewards to calibrate noisy likelihood estimates in dLLMs. The core of our algorithm relies on a key observation: a reward advantage can be interpreted not only as an update direction, but also as a target for the relative log-ratio between the current and reference policies. Accordingly, RSPO calibrates this noisy relative log-ratio estimate by comparing its reward advantage with the reward-implied target relative log-ratio, updating the policy according to the gap between the current estimate and the target rather than the raw advantage alone. Experiments on mathematical reasoning and planning benchmarks show that RSPO yields especially strong gains on planning tasks and competitive mathematical-reasoning performance.
Zichao Yu, Shengze Xu, Bingqing Jiang +2
University of Science and Technology of China · The Chinese University of Hong Kong · The University of Hong Kong