Continuous diffusion language models generate all tokens in parallel, yet high-quality generation can still require hundreds of network evaluations (NFEs). We study how distributional distillation can reduce this cost by exploiting the student's probabilistic token outputs. Our unified formulation connects the student's output parameterization to the resulting gradient estimators and yields two methods with the same student architecture and reverse-KL matching objective: Simplex-DMD uses continuous token relaxations and pathwise gradients, while Reinforce-DMD uses categorical sampling and REINFORCE with a learned density ratio. We develop both methods for multi-step generation and investigate the training and sampling choices associated with each parameterization. On OpenWebText, for sequences of 1,024 tokens, Simplex-DMD achieves a generative perplexity of 45.6 at a unigram entropy of 5.44 nats in just 4 NFEs, a 49% reduction relative to the strongest evaluated diffusion baseline at matched entropy and sampling budget. Reinforce-DMD improves the frontier at larger budgets, reaching a generative perplexity of 14.9 at an entropy of 5.00 nats with 256 NFEs, a 20% reduction under the same comparison protocol.
Figures & tables
Figure 1: Gen PPL on OpenWebText against number of function evaluations (NFE), at the matched unigram entropy HNFE (top axis), defined as in Table 2 . Each method’s log Gen PPL is linearly interpolated at HNFE between the two sampling temperatures whose entropies bracket it. LangFlow and MDLM are undistilled models; FMLM, ReDi and D-MMD are distillation methods. A method that does not reach HNFE is shown at the closest measured entropy, with its value in parentheses. The dashed purple line marks the Gen PPL of the data, with its entropy in parentheses.
Criterion
Local discrepancy D
Representative methods
Continuous diffusion
Reverse KL
logptη(zt)−logpt(zt)
Diff-Instruct ( Luo et al., 2023 ) , DMD ( Yin et al., 2024 )
Score / Velocity / Moment matching
∇ztlogptη(zt)−∇ztlogpt(zt)2
SiD ( Zhou et al., 2024 ) ; FGM ( Huang et al., 2024 ) ; Moment matching ( Salimans et al., 2024 )
IDLM ( Li et al., 2026 ) ; D-MMD ( Hoogeboom et al., 2026 )
Posterior f -divergence
ℓ∑Df(px∣tη,ℓ(⋅∣zt),px∣tℓ(⋅∣zt))
DiMO ( Zhu et al., 2025 )
Table 1: Local discrepancies for distributional distillation. Score, velocity and moment discrepancies agree up to time-dependent factors. Definitions and derivations are given in Appendix C .
Figure 2: Generative frontiers on OpenWebText at NFE=4 and 256 . Temperature varies from 0.8 to 1.1 ; stars mark temperature 1.0 and diamonds the data. Gen PPL uses an inverted log scale. Dashed GPT-2 and OPT use 1,024 evaluations. Simplex-DMD is omitted at NFE=256 due to low diversity. Full sweeps are in Section E.2 .
Network evaluations
Method
2
4
8
16
128
256
HNFE
5.65
5.44
5.38
5.20
5.27
5.00
LangFlow ( Chen et al., 2026b )
1077
317
147
58.6
39.1
21.3
MDLM ( Sahoo et al., 2024 )
1909
678
267
72.8
44.3
23.8
FMLM ( Lee et al., 2026 )
142 (5.32)
128 (5.40)
94.7
55.2
22.6 (4.87)
13.5 (4.50)
D-MMD ( Hoogeboom et al., 2026 )
1752
431
135
41.7
27.3
18.6
Table 2: Gen PPL on OpenWebText at matched unigram entropy HNFE ; lower is better. HNFE is the Simplex-DMD entropy closest to the data for NFE=2 – 16 , and the corresponding Reinforce-DMD entropy for NFE=128,256 . Log Gen PPL is linearly interpolated at HNFE . If a curve does not reach the target entropy, its closest point is reported with entropy in parentheses. Bold indicates the lowest matched value.
Figure 4: Sampler comparison for c∈{0,0.25,0.5,0.75,1,1.5} and forward renoising.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Teacher
Teacher updates
Distillation updates
Student params.
SDTT ( Deschenaux & Gulcehre, 2025 )
MDLM
1M
r×10 k r=1,…,7
≈ 170M
IDLM ( Li et al., 2026 )
MDLM
1M
1M
169.63M
FMLM ( Lee et al., 2026 )
FLM
1.5M
1M
≈ 170M
ReDi ( Yoo et al., 2025 )
DUO
∼ 500k
≤ 1M
≈ 165M
Simplex-DMD (ours)
LangFlow
1M
10k
≈ 170M
Reinforce-DMD (ours)
LangFlow
1M
10k
≈ 170M
Appendix
Table 3: Training budgets and model sizes for the distilled checkpoints. Teacher updates are pretraining steps; distillation updates are additional optimizer updates. For our methods, the 10k total includes 8k auxiliary and 2k student updates. SDTT uses successive round checkpoints, indexed by r . The ReDi student budget is an upper bound. Parameter counts refer to one sampling model, including embeddings, and exclude auxiliary training networks.
Component
Simplex-DMD
Reinforce-DMD
Student initialization
teacher
teacher
Auxiliary initialization
teacher
teacher
Renoising kernel
qt∣x
qt∣x
Self-conditioning
on
off
KL anchor β
–
0.5
Sampler
forward
forward
Appendix
Table 4: Training and sampling configuration of each student. A dash marks a component the objective does not have. The last row is applied at inference.
Figure 6: Design space on OpenWebText, one component changed per variant, read as in Figure 2 . Each student is shown at the two budgets where it is competitive.
Figure 7: Auxiliary loss for Simplex-DMD, read as in Figure 2 : the soft target of ( 17 ) (CE soft), the squared error of ( 83 ) (L2) and the sampled-token target of ( 84 ) (CE hard). The three variants trace one frontier and differ in where their temperature- 1.0 point sits on it, the sampled-token target lowest in entropy and squared error highest.
Figure 8: Forward noising qt∣x against the bridge qt∣T,x as the training renoising kernel for Simplex-DMD, read as in Figure 2 . Each variant is a separate training run.