Diffusion language models (DLMs) offer a promising alternative to autoregressive (AR) language generation. Recent advances in continuous DLMs, which apply latent diffusion to continuous text embeddings, raise a practical question: which embedding makes the best latent space, i.e., the most diffusible? To answer this, we search through different embeddings and find that scaling the embedding model to stronger ones within the same family (T5 to T5Gemma-1 to T5Gemma-2) greatly improves generative performance. But the raw T5Gemma-2 embeddings are still not optimal. They are so discriminative that even the embeddings of plausible alternative words are separated, which makes the generation vulnerable to imperfect sampling. Consequently, continuous diffusion often fails to reach any of them and ends up at an invalid embedding instead. To address this, we distill T5Gemma-2 into a student encoder that learns the teacher's decoded probabilities as soft labels. Learning from such soft labels makes the student pull the alternative embeddings closer while maintaining the encoding-decoding mechanism. The distilled embeddings form a more connected and diffusible latent space, improving over the vanilla T5Gemma-2 embeddings. As a result, our medium-sized DLM achieves Gen. PPL 17.8 (against real-text PPL 15.4) at real-text entropy on OpenWebText, outperforming GPT-2-M on Gen. PPL.
Figures & tables
Figure 1 : Overview of our performance and mechanism. Left : the same DLM trained on our embeddings (distilled from T5Gemma-2) reaches real-text entropy with the best quality. Right : T5Gemma-2 maps plausible candidates apart in the embedding space, so a generated embedding often reaches none of those. Our distilled embedding pulls the candidates closer, forming a continuous region where an embedding at least decodes to one of them.
Figure 2 : Performance of the same ELF-B trained with different encoders on OWT-512. T5Gemma-2 reaches the entropy of real text, and its Gen. PPL is about 40% lower than T5-small’s at the same entropy. Within the T5 family, T5Gemma-2 > T5Gemma-1 > T5-small.
Figure 3 : Generated embeddings deviate from real T5Gemma-2’s.
Figure 4 : Visualizing the embedding error in 2D. Top: a sample generated at NFE 32 ; uncertain tokens (top-1 below 0.6 ) in purple coincide with the grammar mistakes. Bottom: four positions in the plane of their candidate words and the decoded probabilities. The disk marks the 0.91 nn range of a candidate, and an invalid one lands outside every disk and is decoded problematically. Distances are computed in the original 640 -dimensional space; the 2D layout is a distance-preserving map for display. See details in App. B.4 .
Figure 5 : Illustration of the distillation pipeline and the effect on the embeddings. Left: The student encoder is trained to match the teacher’s decoded probabilities (soft labels). Right: The student pulls the embeddings of plausible candidates closer while keeping the irrelevant word far. The right panel is an illustration; measurements on real embeddings are in App. A.4 .
Figure 6 : Distilled embeddings are more diffusible with lower embedding error. Same layout as Fig. 4 . The teacher has more uncertain embeddings (in purple ) which are mostly invalid, while the student’s generated embeddings lie in the region formed by candidates that are close to each other. Over 64 sequences, the teacher has 142 uncertain positions per sequence, 109 of which are also invalid; the student 42 and 12 ; see App. A.1 .
Figure 7 : Round-trip denoising from the same sentence. Left: distance from each denoised embedding to the real embedding of the word it decodes to. Right: one sequence denoised from t0=0.3 ; invalid tokens in purple .
Method
Gen.#Params a
Gen.PPL( ↓ )
Entropy( ↑ )
MAUVE b ( ↑ )
Real Text
—
15.4 ± 0.3
5.43 ± 0.01
0.95 ± 0.00
Base scale
GPT-2-S (AR) ( Radford et al., 2019 )
85+39 M
34.1 ± 0.8
5.45 ± 0.02
0.88 ± 0.02
Duo ( Sahoo et al., 2025 )
92+77 M
45.6 ± 0.5
5.44 ± 0.00
0.95 ± 0.02
AURORA-LM † ( Liang et al., 2026a )
130M
98.4 ± 1.4
5.45 ± 0.01
0.85 ± 0.03
LDLM † ( Meshchaninov et al., 2026b )
132+66 M
34.3 ± 0.3
5.46 ± 0.01
0.31 ± 0.05
Table 1 : Generation quality on OWT- 1024 . Best Gen.PPL at or above real-text entropy over the sampling sweep in App. B.3 ; bold = best, underline = second best. † : cited. a : generation-time parameters, trunk +head . b : MAUVE is protocol-sensitive and serves as a secondary metric. Gray rows fail to reach real-text entropy.
Figure 8 : Sampling ablation of ELF-B (left) and ELF-M (right) models. Each curve stands for different NFEs (64–512) under the same self-conditioning. The dotted curve on the right varies the sampler’s noise scale. Our embedding consistently outperforms raw T5Gemma-2.
Method
Gen.PPL( ↓ )
Entropy( ↑ )
MAUVE( ↑ )
Real Text (LM1B)
52.7 ± 0.6
4.29 ± 0.01
0.95 ± 0.01
AR-B (re-trained)
37.7 ± 0.5
4.29 ± 0.00
0.09 ± 0.01
CoBit-S
59.4 ± 0.6
4.31 ± 0.01
—
LDLM †
63.0 ± 0.5
4.37 ± 0.00
0.91 ± 0.02
T5Gemma-2 + ELF-B
60.6 ± 0.3
4.32 ± 0.01
0.36 ± 0.00
Student (ours) + ELF-B
50.5 ± 0.5
4.31 ± 0.00
0.36 ± 0.04
Table 2 : Generation quality on LM1B. Same protocol as Table 1 . Duo has no public LM1B checkpoint and is not included.
Params
FLOPs
Latency
Training
Encoding, T5-small
19+16 M
0.04 T
1.5 ms
Encoding, T5Gemma-2
100+168 M
0.21 T
3.4 ms
Encoding, Student
50+168 M
0.10 T
2.1 ms
ELF-B denoiser step (fwd + bwd)
89 M
0.55 T
11.2 ms
ELF-M denoiser step (fwd + bwd)
327 M
2.0 T
30.6 ms
Sampling
ELF-B denoiser step
89 M
0.18 T
11.8 ms
Table 3 : Cost of scaling the embeddings per 1024 -token sequence (one H100, bf16, batch size 1 ).
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 9 : Student’s and teacher’s generated embeddings. Positions with top-1 below 0.6 from 64 sequences per model (NFE 32 ). Curve: the decoder’s median top-1 per distance bin.
Figure 10 : Ablation on the distillation objective. Gen. PPL–entropy curves on OWT-1024 with ELF-B. The teacher, the KL student and the MSE student are sampled under the same settings ( 256 samples, one seed), which vary the self-conditioning scale ( 1 – 6 ), the sampler’s noise scale ( 1.0 – 3.0 ) and the NFE ( 16 – 256 ) but do not form a full grid; the CE student is sampled on the grid of Table 1 . Each curve is the best PPL at the same entropy. These settings do not contain the teacher’s setting in Table 1 (sc =1 , NFE =64 ), so in this plot the teacher stops at entropy 5.38 .
Figure 11 : Effect of student depth. Gen. PPL–entropy curves for students with 3 to 18 layers and the T5Gemma-2 teacher on the 512 -token OWT setting. All students use the same distillation and evaluation protocol. The vertical dotted line marks the entropy of real text, and the black star shows the real-text reference. Deeper students generally provide a better quality–diversity trade-off, although the improvement is not monotonic beyond 12 layers.
Figure 12 : Distance between a word and its candidate on real text, when a share of the context tokens is replaced. Line: median over 317 positions; band: quartiles.
Generative
OOD reconstruction
Discriminative
Encoder ( + ELF-M)
Gen. PPL@Entropy ( ↓ )
WikiText-103
arXiv
SST-2
STS-B
MRPC
T5Gemma-2
19.3@5.44
99.3
99.3
89.4
71.7
71.6
Student (ours)
17.8@5.45
99.3
99.4
78.1
65.3
70.8
T5-small
32.8@5.36
100.0
100.0
84.4
74.3
74.8
Appendix
Table 4 : Distillation trades discriminative power for diffusibility. Reconstruction and probing results are reported on a 0 – 100 scale. Gray values indicate generation results below the real-text entropy.
MAUVE ( ↑ )
256 -token features
1024 -token features
Real vs. real
.944
.924
GPT-2-S
.878±.020
.183±.026
T5-small + ELF-B
.415±.037
.012±.004
T5-small + ELF-M
.858±.017
.016±.001
T5Gemma-2 + ELF-B
.758±.052
.103±.006
Appendix
Table 5 : Sensitivity of MAUVE to feature length. We vary only the maximum feature length while keeping the generated samples, reference samples, and GPT-2-Large featurizer unchanged. Increasing the feature length substantially lowers the scores of all models. We report three decimal places to distinguish the small scores and standard deviations in this analysis.
Latent diffusion models offer an attractive alternative to discrete diffusion for non-autoregressive text generation by operating on continuous text representations and denoising entire sequences in parallel. The major challenge in latent diffusion modeling is constructing a suitable latent space. In this work, we present the Latent Diffusion Language Model (LDLM), in which the latent encoder, diffusion model, and decoder are trained jointly. LDLM builds its latent space by reshaping the representations of a pre-trained language model with a trainable encoder, yielding latents that are easy to both denoise and decode into tokens. We show that naive joint training produces a low-quality diffusion model, and propose a simple training recipe consisting of an MSE decoder loss, diffusion-to-encoder warmup, adaptive timestep sampling, and decoder-input noise. Ablations show that each component substantially impacts generation performance. On OpenWebText and LM1B, LDLM achieves better generation performance than existing discrete and continuous diffusion language models while being 2-13× faster, indicating that jointly learning the latent space is a key step toward making latent diffusion competitive for text generation.
Viacheslav Meshchaninov, Alexander Shabalin, Egor Chimbulatov +4
Constructor University, Bremen, Germany · HSE University, Moscow, Russia · Applied AI Institute, Moscow, Russia +2
Continuous diffusion language models generate all tokens in parallel, yet high-quality generation can still require hundreds of network evaluations (NFEs). We study how distributional distillation can reduce this cost by exploiting the student's probabilistic token outputs. Our unified formulation connects the student's output parameterization to the resulting gradient estimators and yields two methods with the same student architecture and reverse-KL matching objective: Simplex-DMD uses continuous token relaxations and pathwise gradients, while Reinforce-DMD uses categorical sampling and REINFORCE with a learned density ratio. We develop both methods for multi-step generation and investigate the training and sampling choices associated with each parameterization. On OpenWebText, for sequences of 1,024 tokens, Simplex-DMD achieves a generative perplexity of 45.6 at a unigram entropy of 5.44 nats in just 4 NFEs, a 49% reduction relative to the strongest evaluated diffusion baseline at matched entropy and sampling budget. Reinforce-DMD improves the frontier at larger budgets, reaching a generative perplexity of 14.9 at an entropy of 5.00 nats with 256 NFEs, a 20% reduction under the same comparison protocol.
Paul Le Van Kiem, Dario Shariatian, Umut Simsekli +1
Inria, PSL Research University · Cohere · CMAP, Ecole Polytechnique
Diffusion Transformers (DiT) trained with flow matching in a VAE latent space have unified visual generation across images and videos. A natural next step toward a single architecture for both generation (visual synthesis) and understanding (text generation) is to apply this framework to language modeling. We propose TextLDM, which transfers the visual latent diffusion recipe to text generation with minimal architectural modification. A Transformer-based VAE maps discrete tokens to continuous latents, enhanced by Representation Alignment (REPA) with a frozen pretrained language model to produce representations effective for conditional denoising. A standard DiT then performs flow matching in this latent space, identical in architecture to its visual counterpart. The central challenge we address is obtaining high-quality continuous text representations: we find that reconstruction fidelity alone is insufficient, and that aligning latent features with a pretrained language model via REPA is critical for downstream generation quality. Trained from scratch on OpenWebText2, TextLDM substantially outperforms prior diffusion language models and matches GPT-2 under the same settings. Our results establish that the visual DiT recipe transfers effectively to language, taking a concrete step toward unified diffusion architectures for multimodal generation and understanding.