Evaluating text generation requires measuring how well the generated distribution matches the data distribution. For autoregressive models, this is done by the perplexity. Diffusion and flow-based language models can only provide a likelihood bound, whose tightness differs between model families. Sample-based substitutes such as generative perplexity with entropy do not consider the distribution fit. We propose SOL, a distance between text distributions. Each sequence is represented by the empirical measure of its hidden states under a fixed transformer and the distributions of these measures are compared by the double sliced Wasserstein distance. We prove that SOL is a metric if the transformer is injective. Experiments show that SOL detects distributional failures, recovers expected model trends, and provides stable sample-based estimates. We put forward SOL to fill the gap in the current evaluation protocol used for non auto-regressive models. As a first step we use SOL to re-evaluate a variety of models trained on OpenWebText.
Figures & tables
Figure 1: Illustration of one SOL slice for fixed directions ξ and g using a causal transformer. A sequence is represented by the point cloud Tθ(x)∈P2(Rd) of its contextual hidden states, and a corpus becomes a collection of such clouds. The first slice projects each cloud along ξ ; its one-dimensional empirical distribution is then represented by a quantile function and sliced again along g . This maps the corpus to P2(R) , where the final Wasserstein distance is computed. SOL integrates this quantity over ξ and g .
GPT-2
Gen. PPL ( ↓ )
Small
12.5
Medium
11.4
Large
7.5
XL
8.5
Table 1: Gen. PPL across GPT-2 sizes for unconditional generation prefers GPT-2 Large.
Figure 2: For a fixed diffusion language model and dataset, the reported MAUVE score depends on the number of evaluation samples n and clusters K . In particular, for the standard choice K=n/10 (red boxes), even after 5k samples it is not stable. Generation details are given in Appendix E.2.2 .
Figure 3: Replacing random tokens in the sequence beginning is not detected by MAUVE, but by SOL.
Figure 4: Mode collapse is invisible to Gen.PPL, while SOL rises sharply, see Subsection 5.1 .
Figure 5: Left: increasing the sequence length diminishes the generation quality as expected. Middle: behavior under mode-collapse , i.e. going from Q2 to Q1 (blue), and if we introduce additional modes , i.e. going from Q2 to Q3 (orange). Right: SOL under word shuffling for different embeddings.
MAUVE
SOL
Criterion ↑
GPT-2 Large
OLMo2 1B
Dream 7B
GPT-2 Large
OLMo2 1B
Dream 7B
Human-like
0.857
0.881
0.881
0.881
0.929
0.952
Interesting
0.714
0.762
0.762
0.762
0.786
0.762
Sensible
0.762
0.786
0.786
0.786
0.881
0.905
Table 2: Human judgement comparison at sequence length 1024 for different embedding models. Worst-case Spearman rank correlation ρmin between metric-induced model rankings and the Bradley–Terry (BT) human ranking. Results are shown for MAUVE and SOL for a full setup see Section C.4 .
Figure 8
Model
Setting
H
Gen. ppl.
SOL ↓
Real OWT G
—
5.449
14.9
2.066
Real ELF OWT E
—
5.402
17.0
2.649
Autoregressive
AR G
Temperature 1.0
5.582
34.9
22.207
Discrete diffusion
DUO-base G
Ancestral, T=1024
5.537
76.9
32.642
Table 3: OpenWebText baselines using 5,000 samples per distribution and a 1,024-token evaluation horizon. SOL uses Dream-7B as the reference encoder. The complete benchmark is reported in Table 24 .
Appendix figures & tables26 assets
Supplementary material from the paper’s appendix.
Appendix
Parameter
Dream-7B
All other encoders
Description
Lξ
8192
1024
First-level projection directions
Lg
8192
1024
Second-level GP directions
M
64
Resolution of second-level functions
λ
0.1
RBF-GP length scale
α
d
Dimension correction factor
Projection seeds
3
Independent ξ draws
Appendix
Table 4: Default SOL estimator parameters by encoder. Here, d denotes the encoder’s hidden dimension. We evaluate all 3×3 combinations of projection and GP seeds. Experiment-specific deviations are stated in the corresponding sections.
K=L
8
16
32
64
128
256
512
1024
2048
4096
8192
16384
Time [s]
0.38
0.39
0.38
0.39
0.41
0.47
0.63
1.02
2.16
5.87
18.95
61.16
Appendix
Table 5: Median post-hidden DSW runtime across three repeats on a single NVIDIA GeForce RTX 5090, with 5,000 texts per arm and M=64 . Hidden states remain on GPU across direction blocks, the initial CPU-to-GPU transfer is included.
Figure 12
Figure 10: Increasing M for the quantile resolution has a small impact beyond M=32 .
Perturbation
mean-SW
SOL
Δ
Selected-word shuffle
10.66
15.75
0.390
Adjacent word swaps
7.09
9.63
0.305
Short-window shuffle
5.63
7.61
0.301
Reverse short spans
7.77
10.81
0.329
Clause permutation
2.61
3.36
0.252
Sentence permutation
2.12
2.60
0.203
Appendix
Table 6: Clean-normalized ordinary-distance responses of mean-SW and SOL under structural perturbations. Response columns are arithmetic means of paired ratios. Δ is the mean paired log-response contrast.
Encoder
Length
t=0.7
t=0.9
t=1.0
t=1.1
t=1.3
GPT-2-large
512
2.901
1.066
0.994
1.038
1.355
1024
2.872
1.091
1.014
1.070
1.434
Qwen3-0.6B
512
6.755
5.085
5.204
5.440
6.332
1024
6.779
5.179
5.257
5.565
6.538
OLMo-2-1B
512
7.719
4.051
4.864
5.947
8.211
1024
8.562
4.103
4.758
6.051
8.561
Appendix
Table 7: SOL evaluating generated text by GPT-2 Large at different decoding temperatures using differnet embedding models.
Temp.
Metric
vMF
Soft-Masked + ReMDM
MDLM
0.7
SOL
48.6981±0.292
57.6332±0.420
54.3128±0.404
Gen. PPL ( H )
7.92(3.950)
3.24(3.490)
7.26(3.884)
0.8
SOL
37.9781±0.160
41.3254±0.303
42.0520±0.244
Gen. PPL ( H )
15.28(4.656)
5.33(4.360)
14.18(4.657)
0.9
SOL
33.1134±0.076
24.7670±0.161
37.4379±0.149
Gen. PPL ( H )
26.36(5.054)
9.18(4.982)
27.01(5.093)
Appendix
Table 8: Temperature sweep on OpenWebText for vMF, Soft-Masked with ReMDM-loop, and MDLM. Panels (a) and (b) report SOL using Dream-7B and GPT-2 Large, respectively. Gen. PPL is evaluated with GPT-2 Large and H denotes mean per-sequence token-frequency entropy in nats; these quantities are independent of the SOL encoder. SOL uses 5,000 samples per side and a 1,024-token horizon, with mean ± population SD over nine seed pairs. Lower SOL is better and bold marks its column minimum.
Figure 11: SOL under increasing [CLS] insertion (left) and word-position shuffling (right) for four encoders. For all four language-model encoders, SOL increases monotonically with the corruption rate, and the two corruptions produce responses of similar magnitude. In contrast to the text encoders in Figure 12 , no encoder is markedly less sensitive to shuffling than to insertion.
Corruption
[CLS] pollution
Word shuffling
0%
0.9702
0.9702
1%
0.9367
0.9701
2%
0.9009
0.9667
5%
0.7888
0.9527
10%
0.6559
0.8423
20%
0.4558
0.3555
Appendix
Table 9: MAUVE under [CLS] pollution and word shuffling using last hidden representation.
Encoder
Clean
Prefix
Suffix
OLMo-2-1B
0.870
2.048
0.913
Dream-7B
2.493
4.496
4.621
Qwen3-0.6B
0.914
1.937
0.946
GPT-2-large
0.224
0.340
0.226
Appendix
Table 10: Shuffle 20% of the tokens the first or last 5% of the text.
Metric
Q2→Q1
Q2→Q3
SOL
1.000
1.000
mean-SW
1.000
0.9995
Appendix
Table 11: Spearman correlations with the target share along the Q2→Q1 and Q2→Q3 interpolation paths.
Encoder
Length
t=0.7
t=0.9
t=1.0
t=1.1
t=1.3
GPT-2-large
512
2.141
1.299
1.138
1.035
1.032
1024
–
–
–
–
–
Qwen3-0.6B
512
6.083
4.310
4.288
4.347
4.847
1024
6.695
4.513
4.470
4.645
5.311
OLMo-2-1B
512
7.822
5.452
5.016
4.808
5.203
1024
9.413
6.440
5.421
4.916
4.884
Appendix
Table 12: SOL scores using the [x,x] second-pass representation across sampling temperatures.
Region
[x]
[x,x]
[x,<sep>,x]
Clean
0.910
0.732
0.853
Prefix
1.710
1.167
1.926
Middle
1.335
1.347
1.375
Suffix
0.942
1.599
0.958
Appendix
Table 13: OLMo regional corruption, cap 1,024. Mean SOL for 10% replacement within a 5% region, taking the square root per seed before averaging, clean scores use uncorrupted candidate text.
Temp.
[x]
[x,x]
[x,<sep>,x]
0.7
8.562
9.413
8.152
0.9
4.103
6.440
4.060
1.0
4.758
5.421
4.703
1.1
6.051
4.916
5.931
1.3
8.561
4.884
8.205
Appendix
Table 14: SOL: GPT-2-large generations scored with OLMo at cap 1,024. Square root of the mean squared score across seeds. Bold marks the minimum within each method.
Figure 12: SOL under increasing shuffling of tokens (left) and inserting a CLS token (right). Compared to GPT2-Large, the text encoders are less sensitive to shuffling than GPT-2 Large.
Insertion rate
SOL (↓)
20%
1
50%
0.942
Appendix
Table 15: NeoBERT SOL under increasing [CLS] insertion, normalized to the 20% value.
MAUVE
SOL
Criterion ↑
GPT-2 Large
OLMo2 1B
Dream 7B
GPT-2 Large
OLMo2 1B
Dream 7B
Human-like
0.881
0.929
0.905
0.952
0.952
0.976
Interesting
0.643
0.786
0.690
0.762
0.762
0.857
Sensible
0.833
0.881
0.881
0.929
0.929
0.976
Appendix
Table 16: Human judgement comparison at sequence length 256. Worst-case Spearman rank correlation between metric-induced model rankings and the Bradley–Terry (BT) human ranking. Results are shown for MAUVE and SOL using GPT-2 Large, OLMo2 1B, and Dream 7B embeddings. Reported are the minimum correlations ρmin across the evaluated settings at sequence length 256.
Corruption
[CLS] pollution
Word shuffling
0%
0.9725
0.9725
1%
0.9590
0.9710
2%
0.9183
0.9697
5%
0.4709
0.9597
10%
0.0624
0.8236
20%
0.0181
0.1208
Appendix
Table 17: MAUVE under [CLS] pollution and word shuffling using mean pooling.
Figure 13: For two disjoint LM1B corpora, the reported MAUVE score depends on the number of evaluation samples n and clusters K . In comparison to generated vs true data the ratio n/10=K is stable, but increasing n converges to 1.
Model
Middle
Final
Real OWT
6.578
6.572
DUO ancestral
6.455
6.441
DUO greedy
6.317
5.011
Appendix
Table 18: Across-document token entropy Hpos in nats ( N=5000 ), measured at position 512 and the last valid token.
MAUVE (↑)
SOL (↓)
Model
Base
Reweighted
Base
Reweighted
DUO ancestral
0.836
0.843
1.670
1.7418
DUO greedy
0.492
0.591
1.748
1.8538
Appendix
Table 19: Final-token reweighting on matched corrected inputs ( N=5000 ). GPT-2-large SOL values average nine projection-seed combinations.
Model
Original
Remove final 8
ELF-M
0.455±0.011
0.823±0.017
ELF-L
0.287±0.006
0.798±0.013
Appendix
Table 20: MAUVE (↑) for shared ELF sampling. Truncation removes eight tokens from both generated and reference texts. Mean ± SD over three clustering seeds.
Rep.
ΔP
ΔR
ΔC
1
+2.14
+6.70
+3.10
2
+3.22
+4.88
+1.80
3
+2.10
+8.32
+0.84
Mean
+2.49
+6.63
+1.91
SD
0.64
1.72
1.13
Appendix
Table 21: ReMDM-loop minus MDLM at k=4 , in percentage points. SD is across three independent repetitions.
Sampler
MAUVE ( ↑ )
SOL ( ↓ )
Greedy
0.0697±0.0012
3.7755±0.1190
Ancestral
0.9407±0.0028
0.5083±0.0113
Nucleus ( p=0.95 )
0.9607±0.0017
0.2712±0.0065
Appendix
Table 22: GPT-2-XL conditional completion on 5,000 WebText documents, using shared 35 -token prefixes and a maximum total sequence length of 256 tokens. SOL uses GPT-2-large.
Sampler
MAUVE ( ↑ )
SOL ( ↓ )
Ancestral
0.8832±0.0071
0.8358±0.0245
Nucleus ( p=0.95 )
0.8740±0.0073
0.9072±0.0279
Appendix
Table 23: GPT-2-XL unconditional generation evaluated against 5,000 WebText documents. Each of the 5,000 samples starts from the end-of-text token and stops at EOS or 256 generated tokens, the initial seed token is excluded from evaluation. SOL uses GPT-2-large.
SOL (↓)
Model
Setting
H@1024
Gen. ppl.
GPT-2-large
OLMo-2-1B
Dream-7B
Real OWT G
—
5.449
14.9
0.2002±0.0116
0.7800±0.0552
2.066±0.109
Real ELF OWT E
—
5.402
17.0
0.2457±0.0239
0.8857±0.0636
2.649±0.219
Autoregressive model Sahoo et al. (2024)
AR G
Temperature 1.0
5.582
34.9
1.064±0.008
7.450±0.132
22.207±0.099
Discrete diffusion
Appendix
Table 24: OpenWebText baselines using 5,000 samples per distribution and a 1,024-token horizon. The standard SOL hyperparameters are used; see Section C.3 . Superscripts G, E and B identify the evaluation reference family. SOL is reported as mean ± population std over nine crossed direction/GP seed pairs.
Diffusion and continuous flow-based language models have emerged as the leading non-autoregressive alternatives to language modeling. Progress in both paradigms is overwhelmingly tracked by generative perplexity (gen-PPL): the per-token negative log-likelihood of samples under a frozen autoregressive (AR) scorer such as gpt2-large, typically paired with an empirical-entropy guardrail to rule out low-entropy collapse. We argue that this metric is unsound. By construction, gen-PPL measures only predictability under the scoring AR, not grammaticality or semantic coherence -- and the set of predictable but still low-quality sequences is combinatorially large. To make this concrete, we construct a suite of zero-parameter, deliberately naive samplers that achieve state-of-the-art gen-PPL on LM1B and OpenWebText at non-degenerate entropy, surpassing recently published diffusion and continuous-flow models while producing text that is incoherent by construction. We recommend evaluation suites that directly quantify the distributional divergence between generated and reference text, and use such a suite to re-benchmark recent non-autoregressive models, recovering a more faithful picture of the current state of the art.
Continuous diffusion language models generate all tokens in parallel, yet high-quality generation can still require hundreds of network evaluations (NFEs). We study how distributional distillation can reduce this cost by exploiting the student's probabilistic token outputs. Our unified formulation connects the student's output parameterization to the resulting gradient estimators and yields two methods with the same student architecture and reverse-KL matching objective: Simplex-DMD uses continuous token relaxations and pathwise gradients, while Reinforce-DMD uses categorical sampling and REINFORCE with a learned density ratio. We develop both methods for multi-step generation and investigate the training and sampling choices associated with each parameterization. On OpenWebText, for sequences of 1,024 tokens, Simplex-DMD achieves a generative perplexity of 45.6 at a unigram entropy of 5.44 nats in just 4 NFEs, a 49% reduction relative to the strongest evaluated diffusion baseline at matched entropy and sampling budget. Reinforce-DMD improves the frontier at larger budgets, reaching a generative perplexity of 14.9 at an entropy of 5.00 nats with 256 NFEs, a 20% reduction under the same comparison protocol.
Paul Le Van Kiem, Dario Shariatian, Umut Simsekli +1
Inria, PSL Research University · Cohere · CMAP, Ecole Polytechnique
Autoregressive language models are widely used for text evaluation, however, their left-to-right factorization introduces positional bias, i.e., early tokens are scored with only leftward context, conflating architectural asymmetry with true text quality. We propose masked reconstruction as an alternative paradigm, where every token is scored using full bidirectional context. We introduce DiffScore, an evaluation framework built on Masked Large Diffusion Language Models. By measuring text recoverability across continuous masking rates, DiffScore eliminates positional bias and naturally establishes an evaluation hierarchy from local fluency to global coherence. We further provide diagnostic tools unavailable to autoregressive frameworks: multi-timestep quality profiles that decompose scores across masking rates, and bidirectional PMI decomposition that disentangles fluency from faithfulness. Experiments across ten benchmarks show that DiffScore consistently outperforms autoregressive baselines in both zero-shot and fine-tuned settings. The code is released at: https://github.com/wenlai-lavine/DiffScore.
Wen Lai, Yingli Shen, Dingnan Jin +4
1Ant Group · 2Tsinghua University · 3Technical University of Munich