Discrete diffusion models offer the ability to re-draft, revisiting and correcting earlier tokens throughout generation. This capability depends on the forward corruption process that defines what the denoiser learns to correct. Masked diffusion models fix tokens once they are unmasked, while uniform diffusion permits revisions but relies on uniformly random token substitutions. We instead learn which substitutions are most useful for training the denoiser to re-draft. We introduce Variational Stackelberg Discrete Diffusion (VSDD), a framework for learning a semantically aware corruption process. VSDD formulates training as a leader-follower game: the leader defines a Markovian corruption process parameterized by the denoiser's token embeddings, while the follower optimizes a variational denoising objective with the corruption process held fixed. The leader rewards corruptions based on how much the denoiser improves after learning from them, rather than on how easily the current denoiser can reconstruct them. We measure this improvement under a fixed reference corruption process, approximate the follower's response with a one-step gradient update, and optimize the leader using a score-function estimator. We evaluate VSDD across molecular, text, and playlist generation. VSDD substantially improves molecular validity over uniform and masked diffusion, reduces text perplexity relative to uniform diffusion while remaining competitive with masked diffusion, and achieves sizable improvements in offline playlist recommendation metrics.
Figures & tables
Figure 1: Re-drafting and learned corruption in molecular generation.
Method
Validity (%) ↑
Uniqueness (%) ↑
Novelty (%) ↑
Diversity ↑
MDLM
58.7
100
100
0.8717
D3PM-Uniform
48.1
100
100
0.8674
D3PM-Reinforce
62.0
100
100
0.8727
FLDD
3.0
100
100
0.9761
VSDD (ours)
84.3
100
100
0.8679
Table 1: Chemical validity, uniqueness, novelty and diversity of generated molecules.
Figure 2: Correlation between the similarity of token embeddings and the transition probabilities.
Table 2: Text and playlist generation results.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Story name
Corruption
# occurrences
Corrupted occurrence
Lily and red ball
Lily → Tom
4
2
Tom and Max
Max → Sam
5
1
Sam and blue car
car → dog
4
3
Lily and cat
cat → dog
5
2
Ben and toy train
train → ball
5
3
Appendix
Table 3: Synthetic story corruptions. For each story, an entity is replaced at a selected occurrence.
Model
Exact Repair ↑
Consistency ↑
VSDD
0.880±0.160
0.974±0.033
D3PM-Uniform
0.360±0.265
0.860±0.065
MDLM
0.960±0.080
0.992±0.016
Appendix
Table 4: Aggregate repair performance. Results report the mean and the standard deviation across corruption types.
Example 1: Lily → Tom (2nd occurrence)
Original Once upon a time there was a little girl named Lily. Lily had a big red ball. She loved to play with the red ball every day. One day Lily took the red ball to the park. Lily was very happy.
Corrupted Once upon a time there was a little girl named Lily. Tom had a big red ball. She loved to play with the red ball every day. One day Lily took the red ball to the park. Lily was very happy.
Model
Infilled text
Repair?
VSDD
…named Lily. Lily had a big red ball. She loved …
Yes
D3PM-Uniform
…named Lily. She had a big red ball. She loved … Generates “She” instead of “Lily”.
No
MDLM
…named Lily. Lily had a big red ball. She loved …
Yes
Appendix
Table 5: The 2nd occurrence of “Lily” is corrupted to “Tom”, each model attempts to recover the original token.
Example 2: Max → Sam (1st occurrence)
Original Tom had a small dog named Max . Tom and Max liked to play in the park. Every morning Tom took Max for a long walk. Max was a good dog and Tom loved Max very much.
Corrupted Tom had a small dog named Sam . Tom and Max liked to play in the park. Every morning Tom took Max for a long walk. Max was a good dog and Tom loved Max very much.
Model
Infilled text
Repair?
VSDD
…dog named Max . Tom and Max liked … Successful in 3/5 trials.
Yes
D3PM-Uniform
…dog named Max . Tom and Max liked … Successful in 1/5 trials.
Rarely
MDLM
…dog named Tom . Tom and Max liked … Repairs in 4/5 trials, but occasionally generates “Tom” instead of “Max”.
Partial
Appendix
Table 6: The 1st occurrence of “Max” is corrupted to “Sam”.
Example 3: car → dog (3rd occurrence)
Original There was a little boy named Sam. Sam had a blue car. Sam liked to play with his blue car. One day Sam lost his blue car . Sam was sad but then he found the blue car under the bed.
Corrupted There was a little boy named Sam. Sam had a blue car. Sam liked to play with his blue car. One day Sam lost his blue dog . Sam was sad but then he found the blue car under the bed.
Model
Infilled text
Repair?
VSDD
…lost his blue car . Sam was sad … Successful in 5/5 trials.
Yes
D3PM-Uniform
…lost his blue car . Sam was sad … Successful in 4/5 trials.
Mostly
MDLM
…lost his blue car . Sam was sad … Successful in 5/5 trials.
Yes
Appendix
Table 7: The third occurrence of “car” is corrupted to “dog”.
Example 4: cat → dog (2nd occurrence)
Original Lily had a pretty cat. The cat was soft and white. Lily liked to pet her cat every day. One day the cat found a little mouse. Lily and the cat played in the garden.
Corrupted Lily had a pretty cat. The dog was soft and white. Lily liked to pet her cat every day. One day the cat found a little mouse. Lily and the cat played in the garden.
Model
Infilled text
Repair?
VSDD
…The cat was soft and white … Successful in 5/5 trials.
Yes
D3PM-Uniform
…The a was soft and white … Occasionally generates the article “a” instead of “cat”.
No
MDLM
…The cat was soft and white … Successful in 5/5 trials.
Yes
Appendix
Table 8: The second occurrence of “cat” is corrupted to “dog”.
Example 5: train → ball (3rd occurrence)
Original Ben had a toy train. Ben loved his toy train. Every day Ben played with the toy train in his room. One day Ben took the toy train to show his friend. His friend liked the toy train too.
Corrupted Ben had a toy train. Ben loved his toy train. Every day Ben played with the toy ball in his room. One day Ben took the toy train to show his friend. His friend liked the toy train too.
Model
Infilled text
Repair?
VSDD
…the toy train in his room … Successful in 5/5 trials.
Yes
D3PM-Uniform
…the toy train in his room … Successful in 2/5 trials.
Partial
MDLM
…the toy train in his room … Successful in 5/5 trials.
Yes
Appendix
Table 9: Qualitative example for the train → ball corruption. The third occurrence of “train” is corrupted to “ball”, and each model attempts to recover the original token.
Story
Removed subsequence
subsequence length
Lily & red ball
“red ball. She loved to play with the”
9
Tom & Max
“in the park. Every morning Tom took”
8
Sam & blue car
“liked to play with his blue car. One”
9
Appendix
Table 10: Removed subsequences used for re-drafting. Each example masks a contiguous subsequence of tokens from the original story.
Model
Token Overlap ↑
VSDD
0.243±0.127
D3PM-Uniform
0.163±0.159
MDLM
0.184±0.082
Appendix
Table 11: Span infilling performance. Token Overlap measures agreement between the generated and original removed spans.
Example: Lily story (9 tokens removed)
Original …Lily. Lily had a big red ball. She loved to play with the red ball every day …
Removed span “red ball. She loved to play with the”
Model
Infilled span
Overlap ↑
VSDD
“ red ball that she loved. She took her ”
42.2%
D3PM-Uniform
“ red toy. The red liked red new toy ”
11.1%
MDLM
“ ball with many hair. She played with the ”
8.9%
Appendix
Table 12: Example: Subsequence Removal.
Smooth
Diversity
Corruption
Repair rate
O
C
R
O
C
R
Duplication
78.9%
0.9800
0.9802
0.9807
2477
2453
2476
Transition disruption
86.1%
0.9800
0.9767
0.9792
2510
2497
2520
Energy misplacement
85.3%
0.9800
0.9797
0.9799
2477
2477
2481
Appendix
Table 13: Repair results for different corruption types.
Token
Row Entropy
Top Target
P(i→j)
.
0.023
to
0.998
the
0.062
t
0.996
and
0.095
it
0.993
,
0.096
that
0.993
to
0.111
.
0.992
it
0.115
was
0.991
Appendix
Table 14: Naive joint optimisation of ϕ and θ learns trivial near-deterministic noise processes.
Masked diffusion language models (MDLMs) re-predict every position at each denoising step, but standard samplers commit tokens once revealed, leaving this revision capability unused. Existing approaches either add heuristic or learned mechanisms to revise committed tokens, or remask them back to [MASK] before re-predicting; a principled sampler that directly revises visible tokens without auxiliary modules remains underexplored. We introduce D3IM, a parameter-free sampler derived as a corrector-style reverse update that permits direct visible-to-visible revision without additional modules or auxiliary passes. D3IM also reveals a model-side obstacle we term preservation bias: the model tends to reproduce its own wrong committed tokens rather than correct them. We address this with SCOPE (Self-Conditioned On Prediction Errors), a lightweight post-training procedure that simulates D3IM's sampling process. On LLaDA-8B at 64 denoising steps, SCOPE+D3IM improves over the original LLaDA-8B with standard unmasking by +13.0 on GSM8K (68.3%), +4.8 on MATH-500 (23.6%), +15.3 on HumanEval (29.3%), and +10.4 on MBPP (30.8%), with gains that increase as more denoising steps are used on math and HumanEval.
Discrete diffusion models are a powerful class of generative models with strong performance across many domains. For efficiency, however, discrete diffusion typically parameterizes the generative (reverse) process with factorized distributions, which makes it difficult for the model to learn the target process in a small number of steps and necessitates a long, computationally expensive sampling procedure. To reduce the gap between the target and model distributions and enable few-step generation, we propose Forward-Learned Discrete Diffusion (FLDD), which introduces discrete diffusion with a learnable forward (noising) process. Rather than fixing a Markovian forward chain, we adopt a non-Markovian formulation with learnable marginal and posterior distributions. This allows the generative process to remain factorized while matching the target defined by the noising process. We train all parameters end-to-end under the standard variational objective. Experiments on various benchmarks show that, for a given number of sampling steps, our approach produces a higher quality samples than conventional discrete diffusion models using the same reverse parameterization.
Discrete diffusion models form a powerful class of generative models across diverse domains, including text and graphs. However, existing approaches face fundamental limitations. Masked diffusion models suffer from irreversible errors due to early unmasking, while uniform diffusion models, despite enabling self-correction, often yield low-quality samples due to their strong reliance on intermediate latent states. We introduce IDDM, an Interpolating Discrete Diffusion Model, that improves diffusion by reducing dependence on intermediate latent states. Central to IDDM is a controllable resampling mechanism that partially resets probability mass to the marginal distribution, mitigating error accumulation and enabling more effective token corrections. IDDM specifies a generative process whose transitions interpolate between staying at the current state, resampling from a prior, and flipping toward the target state, while enforcing marginal consistency and fully decoupling training from inference. We benchmark our model against state-of-the-art discrete diffusion models across molecular graph generation as well as text generation tasks, demonstrating competitive performance.
Marcel Kollovieh, Sirine Ayadi, Stephan Günnemann
School of Computation, Information and Technology, Technical University of Munich