cs.CLSep 29, 2025

RFG: Self-Improving Diffusion Large Language Models with Reward-Free Guidance

Authors: Tianlang Chen, Minkai Xu, Jure Leskovec, Stefano Ermon

Organizations: Stanford University

Abstract

Diffusion Large Language Models (dLLMs) have shown strong reasoning capabilities, yet further improving them typically requires costly post-training with additional data and supervision. We ask whether a post-trained dLLM can improve itself at inference time without additional training, data, or reward models. This requires a guidance signal from the checkpoints alone that is well-defined on the partially masked intermediate states of dLLMs, which existing methods fail to provide. Here we propose Reward-Free Guidance (RFG), a training-free framework for inference-time self-improvement of dLLMs that derives such a signal from the checkpoints themselves. We theoretically demonstrate that a reward signal for a partially masked state can be parameterized by the log-likelihood ratio between a post-trained policy dLLM and its reference model. In practice, this signal can be obtained using off-the-shelf checkpoints alone. Extensive experiments show that RFG consistently improves state-of-the-art post-trained dLLMs by up to 16.1%. Furthermore, despite being entirely training-free, RFG delivers gains that rival or even surpass resource-intensive reinforcement learning techniques.

Figures & tables

Explore similar work

CardsList
  1. Enhancing Diffusion Language Models with Autoregressive Post-Training Weights

    Oct 6, 2026Yiming Qin, Ke Wang, Amel Abdelraheem +2Diffusion Language ModelsDiffusion Models

  2. Back on Track: Aligning Rewards and States for Reasoning in Diffusion Large Language Models

    Jun 7, 2026Yawen Shao, Jie Xiao, Kai Zhu +6Large Language Model Reinforcement LearningDiffusion Language Models