Diffusion language models (DLMs) offer a promising alternative to autoregressive (AR) language generation. Recent advances in continuous DLMs, which apply latent diffusion to continuous text embeddings, raise a practical question: which embedding makes the best latent space, i.e., the most diffusible? To answer this, we search through different embeddings and find that scaling the embedding model to stronger ones within the same family (T5 to T5Gemma-1 to T5Gemma-2) greatly improves generative performance. But the raw T5Gemma-2 embeddings are still not optimal. They are so discriminative that even the embeddings of plausible alternative words are separated, which makes the generation vulnerable to imperfect sampling. Consequently, continuous diffusion often fails to reach any of them and ends up at an invalid embedding instead. To address this, we distill T5Gemma-2 into a student encoder that learns the teacher's decoded probabilities as soft labels. Learning from such soft labels makes the student pull the alternative embeddings closer while maintaining the encoding-decoding mechanism. The distilled embeddings form a more connected and diffusible latent space, improving over the vanilla T5Gemma-2 embeddings. As a result, our medium-sized DLM achieves Gen. PPL 17.8 (against real-text PPL 15.4) at real-text entropy on OpenWebText, outperforming GPT-2-M on Gen. PPL.
Figures & tables
Figure 1 : Overview of our performance and mechanism. Left : the same DLM trained on our embeddings (distilled from T5Gemma-2) reaches real-text entropy with the best quality. Right : T5Gemma-2 maps plausible candidates apart in the embedding space, so a generated embedding often reaches none of those. Our distilled embedding pulls the candidates closer, forming a continuous region where an embedding at least decodes to one of them.
Figure 2 : Performance of the same ELF-B trained with different encoders on OWT-512. T5Gemma-2 reaches the entropy of real text, and its Gen. PPL is about 40% lower than T5-small’s at the same entropy. Within the T5 family, T5Gemma-2 > T5Gemma-1 > T5-small.
Figure 3 : Generated embeddings deviate from real T5Gemma-2’s.
Figure 4 : Visualizing the embedding error in 2D. Top: a sample generated at NFE 32 ; uncertain tokens (top-1 below 0.6 ) in purple coincide with the grammar mistakes. Bottom: four positions in the plane of their candidate words and the decoded probabilities. The disk marks the 0.91 nn range of a candidate, and an invalid one lands outside every disk and is decoded problematically. Distances are computed in the original 640 -dimensional space; the 2D layout is a distance-preserving map for display. See details in App. B.4 .
Figure 5 : Illustration of the distillation pipeline and the effect on the embeddings. Left: The student encoder is trained to match the teacher’s decoded probabilities (soft labels). Right: The student pulls the embeddings of plausible candidates closer while keeping the irrelevant word far. The right panel is an illustration; measurements on real embeddings are in App. A.4 .
Figure 6 : Distilled embeddings are more diffusible with lower embedding error. Same layout as Fig. 4 . The teacher has more uncertain embeddings (in purple ) which are mostly invalid, while the student’s generated embeddings lie in the region formed by candidates that are close to each other. Over 64 sequences, the teacher has 142 uncertain positions per sequence, 109 of which are also invalid; the student 42 and 12 ; see App. A.1 .
Figure 7 : Round-trip denoising from the same sentence. Left: distance from each denoised embedding to the real embedding of the word it decodes to. Right: one sequence denoised from t0=0.3 ; invalid tokens in purple .
Method
Gen.#Params a
Gen.PPL( ↓ )
Entropy( ↑ )
MAUVE b ( ↑ )
Real Text
—
15.4 ± 0.3
5.43 ± 0.01
0.95 ± 0.00
Base scale
GPT-2-S (AR) ( Radford et al., 2019 )
85+39 M
34.1 ± 0.8
5.45 ± 0.02
0.88 ± 0.02
Duo ( Sahoo et al., 2025 )
92+77 M
45.6 ± 0.5
5.44 ± 0.00
0.95 ± 0.02
AURORA-LM † ( Liang et al., 2026a )
130M
98.4 ± 1.4
5.45 ± 0.01
0.85 ± 0.03
LDLM † ( Meshchaninov et al., 2026b )
132+66 M
34.3 ± 0.3
5.46 ± 0.01
0.31 ± 0.05
Table 1 : Generation quality on OWT- 1024 . Best Gen.PPL at or above real-text entropy over the sampling sweep in App. B.3 ; bold = best, underline = second best. † : cited. a : generation-time parameters, trunk +head . b : MAUVE is protocol-sensitive and serves as a secondary metric. Gray rows fail to reach real-text entropy.
Figure 8 : Sampling ablation of ELF-B (left) and ELF-M (right) models. Each curve stands for different NFEs (64–512) under the same self-conditioning. The dotted curve on the right varies the sampler’s noise scale. Our embedding consistently outperforms raw T5Gemma-2.
Method
Gen.PPL( ↓ )
Entropy( ↑ )
MAUVE( ↑ )
Real Text (LM1B)
52.7 ± 0.6
4.29 ± 0.01
0.95 ± 0.01
AR-B (re-trained)
37.7 ± 0.5
4.29 ± 0.00
0.09 ± 0.01
CoBit-S
59.4 ± 0.6
4.31 ± 0.01
—
LDLM †
63.0 ± 0.5
4.37 ± 0.00
0.91 ± 0.02
T5Gemma-2 + ELF-B
60.6 ± 0.3
4.32 ± 0.01
0.36 ± 0.00
Student (ours) + ELF-B
50.5 ± 0.5
4.31 ± 0.00
0.36 ± 0.04
Table 2 : Generation quality on LM1B. Same protocol as Table 1 . Duo has no public LM1B checkpoint and is not included.
Params
FLOPs
Latency
Training
Encoding, T5-small
19+16 M
0.04 T
1.5 ms
Encoding, T5Gemma-2
100+168 M
0.21 T
3.4 ms
Encoding, Student
50+168 M
0.10 T
2.1 ms
ELF-B denoiser step (fwd + bwd)
89 M
0.55 T
11.2 ms
ELF-M denoiser step (fwd + bwd)
327 M
2.0 T
30.6 ms
Sampling
ELF-B denoiser step
89 M
0.18 T
11.8 ms
Table 3 : Cost of scaling the embeddings per 1024 -token sequence (one H100, bf16, batch size 1 ).
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 9 : Student’s and teacher’s generated embeddings. Positions with top-1 below 0.6 from 64 sequences per model (NFE 32 ). Curve: the decoder’s median top-1 per distance bin.
Figure 10 : Ablation on the distillation objective. Gen. PPL–entropy curves on OWT-1024 with ELF-B. The teacher, the KL student and the MSE student are sampled under the same settings ( 256 samples, one seed), which vary the self-conditioning scale ( 1 – 6 ), the sampler’s noise scale ( 1.0 – 3.0 ) and the NFE ( 16 – 256 ) but do not form a full grid; the CE student is sampled on the grid of Table 1 . Each curve is the best PPL at the same entropy. These settings do not contain the teacher’s setting in Table 1 (sc =1 , NFE =64 ), so in this plot the teacher stops at entropy 5.38 .
Figure 11 : Effect of student depth. Gen. PPL–entropy curves for students with 3 to 18 layers and the T5Gemma-2 teacher on the 512 -token OWT setting. All students use the same distillation and evaluation protocol. The vertical dotted line marks the entropy of real text, and the black star shows the real-text reference. Deeper students generally provide a better quality–diversity trade-off, although the improvement is not monotonic beyond 12 layers.
Figure 12 : Distance between a word and its candidate on real text, when a share of the context tokens is replaced. Line: median over 317 positions; band: quartiles.
Generative
OOD reconstruction
Discriminative
Encoder ( + ELF-M)
Gen. PPL@Entropy ( ↓ )
WikiText-103
arXiv
SST-2
STS-B
MRPC
T5Gemma-2
19.3@5.44
99.3
99.3
89.4
71.7
71.6
Student (ours)
17.8@5.45
99.3
99.4
78.1
65.3
70.8
T5-small
32.8@5.36
100.0
100.0
84.4
74.3
74.8
Appendix
Table 4 : Distillation trades discriminative power for diffusibility. Reconstruction and probing results are reported on a 0 – 100 scale. Gray values indicate generation results below the real-text entropy.
MAUVE ( ↑ )
256 -token features
1024 -token features
Real vs. real
.944
.924
GPT-2-S
.878±.020
.183±.026
T5-small + ELF-B
.415±.037
.012±.004
T5-small + ELF-M
.858±.017
.016±.001
T5Gemma-2 + ELF-B
.758±.052
.103±.006
Appendix
Table 5 : Sensitivity of MAUVE to feature length. We vary only the maximum feature length while keeping the generated samples, reference samples, and GPT-2-Large featurizer unchanged. Increasing the feature length substantially lowers the scores of all models. We report three decimal places to distinguish the small scores and standard deviations in this analysis.