Most continuous diffusion language models process one latent position per token at each sampling step, making generation expensive. Two-stage methods lower the cost by reducing the latent length, but they fix the compressed embedding space before training the diffusion model. Embeddings from the fixed space can be difficult to model with diffusion and decode reliably into tokens, which limits generation quality after compression. To address this problem, we introduce JPEG-DLM (Joint-embedding Prediction for Efficient Generation with Diffusion Language Model), which jointly trains a compressor, a flow matching model and a decoding module. With joint-embedding prediction, JPEG-DLM learns compressed embeddings that are more structured, easier to model with diffusion and reliably decodable into tokens. JPEG-DLM achieves the lowest mean Gen-PPL and highest throughput among recent diffusion and flow models on LM1B and OWT. At a compression rate of 0.5 on OWT, it reaches a Gen-PPL of 34.52 and approximately 2.3 times ELF's throughput. These results suggest that jointly learning compressed embeddings offers a promising path toward efficient diffusion language modeling. Code will be released soon.
Figures & tables
Figure 1: Schematic comparison of fixed and jointly learned compressed spaces, showing more structured embeddings and broader decodable regions with joint training.
Figure 2: Overview of JPEG-DLM . The frozen T5 encoder maps tokens to features. The context compressor maps a corrupted view to zc , while the target compressor provides the stop-gradient clean target embedding ztgt and is updated by an exponential moving average. The shared-weight network predicts the clean target embedding from zc for JEP and from the noisy embedding zt for flow matching. The decoding module maps the predicted clean embedding back to features and tokens using Lfeat and Lce .
Figure 3
LM1B ( L=128 )
OWT ( L=1024 )
Method
Param. (M)
Gen-PPL ↓
Ent.
MAUVE ↑
Speed ↑
Gen-PPL ↓
Ent.
MAUVE ↑
Speed ↑
AR ( Radford et al., 2019 )
130
151.36
4.397
0.947
1.00× ( 1.62× w/ KV cache)
41.06
5.583
0.930
1.00× ( 7.14× w/ KV cache)
MDLM ( Sahoo et al., 2024 )
130
192.80
4.410
0.923
0.52×
41.46
5.287
0.755
0.32×
SEDD-Absorb ( Lou et al., 2024 )
170
182.57
4.405
0.941
0.40×
41.15
5.260
0.728
0.25×
Duo ( Sahoo et al., 2025 )
130
161.80
4.382
0.941
0.41×
75.88
5.536
0.903
0.15×
FLM ( Lee et al., 2026 )
179
176.74
4.402
0.863
0.66×
61.37
5.335
0.373
0.21×
Table 1: Unconditional generation on LM1B and OWT. Speed is relative to AR without KV cache. Among diffusion and flow models, bold and underlined mark the best and second-best Gen-PPL, MAUVE, and speed. JPEG-DLM has 105M DiT and token projection parameters, or 154M including the compressor and decompressor.
Figure 3: Generation quality at different latent lengths (a) and latent widths (b).
Figure 4: Joint and two-stage training. (a) Gen-PPL and MAUVE across compression rates; arrows show relative changes from COSMOS to JPEG-DLM , and diamonds mark ELF-B. COSMOS uses final-layer T5-small features at every rate. (b) Correlation between latent positions at r=0.5 .
Figure 5: Gen-PPL over interpolations between 1,024 pairs of clean target embeddings at a compression rate of r=0.5 , without added noise and at three noise levels. Here, t denotes the COSMOS forward time, with t=0 indicating no added noise. For the three noisy settings, JPEG-DLM uses noisy embeddings at the same SNRs along its linear path.
Variant
Ljep
Lfeat
σ
Lce timesteps
Gen-PPL ↓
Ent.
MAUVE ↑
Base
✓
✓
1.0
t=1
48.00
4.064
0.912
sep. nets
✓
✓
1.0
t=1
92.45
4.117
0.786
w/o Ljep
×
✓
1.0
t=1
51.65
4.016
0.849
w/o Ljep + sep. nets
×
✓
1.0
t=1
131.94
4.188
0.632
w/o Lfeat
✓
×
1.0
t=1
51.57
4.069
0.825
σ=0.5
✓
✓
0.5
t=1
61.43
3.993
0.817
Table 2: Ablation study on OWT-128. Base uses a shared-weight network.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
SIGReg
VICReg
Gen-PPL ↓
MAUVE ↑
Entropy
×
×
63.11
0.856
4.090
✓
×
48.60
0.865
4.062
×
✓
40.77
0.868
3.965
✓
✓
48.00
0.912
4.064
Appendix
Table 3: Regularizer ablation on OWT-128 with r=0.5 , d=512 , 64 sampling steps, and self-conditioning at scale 3.0 . The variance and covariance terms are enabled or disabled as a pair. The shaded row reproduces Base from Table 2 .
Figure 6: Gen-PPL on LM1B across sampling steps. Error bars show standard deviations over five seeds. Standard deviations are unavailable for the 8-step baselines.
Steps
Gen-PPL ↓
Ent.
MAUVE ↑
8
225.24
4.162
0.432
16
137.86
4.270
0.896
32
96.76
4.263
0.951
64
78.64
4.248
0.940
128
69.45
4.238
0.941
256
64.07
4.230
0.933
Appendix
Table 4: Generation quality of JPEG-DLM on LM1B with r=0.25 at different sampling budgets. Results are means over five seeds. The shaded row is the 32-step configuration used in the main comparison.
Figure 7: GPT-2 token lengths of ELF generations and two OWT reference sets. The reference constructed from fixed 1,024-token blocks concentrates near the block length, whereas the reference tokenized and decoded with T5 and the ELF samples vary in length. After discretization, the intersection between the ELF and fixed-length reference histograms is only 0.9%.
Parameter
LM1B
OWT-1024
Data and tokenization
Dataset
LM1B ( Chelba et al., 2014 )
OWT
Token sequence length
128
1024
Sequence construction
Blocks packed across sentence boundaries
Blocks packed across sentence boundaries (no padding)
Model architecture
Frozen text encoder
T5-small
T5-small
Appendix
Table 5: Model and training configurations of JPEG-DLM on LM1B and OWT-1024. Sampling compute counts one multiply–accumulate as one operation. For K solver steps, JPEG-DLM uses K+1 DiT evaluations, including the final clean embedding prediction, followed by one decoding pass. The estimate at full latent length uses the same K+1 evaluation budget.
LM1B
OWT-1024
Method
SC
Samples
Tokenizer
Steps
Tokenizer
Steps
AR
×
1,024
BERT
–
GPT-2
–
MDLM
×
1,024
BERT
128
GPT-2 + mask
1024
SEDD-Absorb
×
1,024
GPT-2
128
GPT-2
1024
Duo
×
1,024
BERT
128
GPT-2
1024
FLM
×
1,024
BERT
128
GPT-2
1024
Appendix
Table 6: Tokenization and sampling configurations for the main unconditional generation comparison. BERT ( Devlin et al., 2019 ) denotes bert-base-uncased , and GPT-2 ( Radford et al., 2019 ) denotes GPT-2 BPE. T5 ( Raffel et al., 2020 ) uses SentencePiece ( Kudo and Richardson, 2018 ) . A checkmark in the SC column indicates self-conditioning.
Parameter
OWT-128
Dataset
linluqiu/openwebtext-len128-packing_padding
Sequence length
128
Text encoder / tokenizer
Frozen T5-small / T5
Encoder feature
Final hidden layer
Latent length / compression rate r
128 / 1.0, 64 / 0.5, 32 / 0.25, 16 / 0.125
Latent width
512
Appendix
Table 7: Configuration of COSMOS on OWT-128.
Layer index
Gen-PPL ↓
Ent.
MAUVE ↑
−1
240.811
3.9033
0.1167
−2
154.551
3.9736
0.3009
−3
163.196
3.9717
0.2677
Appendix
Table 8: COSMOS performance at r=1.0 with T5-small layer indices −1 , −2 , and −3 . Results are means over five seeds using 1,024 samples per seed.