Evaluating text generation requires measuring how well the generated distribution matches the data distribution. For autoregressive models, this is done by the perplexity. Diffusion and flow-based language models can only provide a likelihood bound, whose tightness differs between model families. Sample-based substitutes such as generative perplexity with entropy do not consider the distribution fit. We propose SOL, a distance between text distributions. Each sequence is represented by the empirical measure of its hidden states under a fixed transformer and the distributions of these measures are compared by the double sliced Wasserstein distance. We prove that SOL is a metric if the transformer is injective. Experiments show that SOL detects distributional failures, recovers expected model trends, and provides stable sample-based estimates. We put forward SOL to fill the gap in the current evaluation protocol used for non auto-regressive models. As a first step we use SOL to re-evaluate a variety of models trained on OpenWebText.
Figures & tables
Figure 1: Illustration of one SOL slice for fixed directions ξ and g using a causal transformer. A sequence is represented by the point cloud Tθ(x)∈P2(Rd) of its contextual hidden states, and a corpus becomes a collection of such clouds. The first slice projects each cloud along ξ ; its one-dimensional empirical distribution is then represented by a quantile function and sliced again along g . This maps the corpus to P2(R) , where the final Wasserstein distance is computed. SOL integrates this quantity over ξ and g .
GPT-2
Gen. PPL ( ↓ )
Small
12.5
Medium
11.4
Large
7.5
XL
8.5
Table 1: Gen. PPL across GPT-2 sizes for unconditional generation prefers GPT-2 Large.
Figure 2: For a fixed diffusion language model and dataset, the reported MAUVE score depends on the number of evaluation samples n and clusters K . In particular, for the standard choice K=n/10 (red boxes), even after 5k samples it is not stable. Generation details are given in Appendix E.2.2 .
Figure 3: Replacing random tokens in the sequence beginning is not detected by MAUVE, but by SOL.
Figure 4: Mode collapse is invisible to Gen.PPL, while SOL rises sharply, see Subsection 5.1 .
Figure 5: Left: increasing the sequence length diminishes the generation quality as expected. Middle: behavior under mode-collapse , i.e. going from Q2 to Q1 (blue), and if we introduce additional modes , i.e. going from Q2 to Q3 (orange). Right: SOL under word shuffling for different embeddings.
MAUVE
SOL
Criterion ↑
GPT-2 Large
OLMo2 1B
Dream 7B
GPT-2 Large
OLMo2 1B
Dream 7B
Human-like
0.857
0.881
0.881
0.881
0.929
0.952
Interesting
0.714
0.762
0.762
0.762
0.786
0.762
Sensible
0.762
0.786
0.786
0.786
0.881
0.905
Table 2: Human judgement comparison at sequence length 1024 for different embedding models. Worst-case Spearman rank correlation ρmin between metric-induced model rankings and the Bradley–Terry (BT) human ranking. Results are shown for MAUVE and SOL for a full setup see Section C.4 .
Figure 8
Model
Setting
H
Gen. ppl.
SOL ↓
Real OWT G
—
5.449
14.9
2.066
Real ELF OWT E
—
5.402
17.0
2.649
Autoregressive
AR G
Temperature 1.0
5.582
34.9
22.207
Discrete diffusion
DUO-base G
Ancestral, T=1024
5.537
76.9
32.642
Table 3: OpenWebText baselines using 5,000 samples per distribution and a 1,024-token evaluation horizon. SOL uses Dream-7B as the reference encoder. The complete benchmark is reported in Table 24 .
Appendix figures & tables26 assets
Supplementary material from the paper’s appendix.
Appendix
Parameter
Dream-7B
All other encoders
Description
Lξ
8192
1024
First-level projection directions
Lg
8192
1024
Second-level GP directions
M
64
Resolution of second-level functions
λ
0.1
RBF-GP length scale
α
d
Dimension correction factor
Projection seeds
3
Independent ξ draws
Appendix
Table 4: Default SOL estimator parameters by encoder. Here, d denotes the encoder’s hidden dimension. We evaluate all 3×3 combinations of projection and GP seeds. Experiment-specific deviations are stated in the corresponding sections.
K=L
8
16
32
64
128
256
512
1024
2048
4096
8192
16384
Time [s]
0.38
0.39
0.38
0.39
0.41
0.47
0.63
1.02
2.16
5.87
18.95
61.16
Appendix
Table 5: Median post-hidden DSW runtime across three repeats on a single NVIDIA GeForce RTX 5090, with 5,000 texts per arm and M=64 . Hidden states remain on GPU across direction blocks, the initial CPU-to-GPU transfer is included.
Figure 12
Figure 10: Increasing M for the quantile resolution has a small impact beyond M=32 .
Perturbation
mean-SW
SOL
Δ
Selected-word shuffle
10.66
15.75
0.390
Adjacent word swaps
7.09
9.63
0.305
Short-window shuffle
5.63
7.61
0.301
Reverse short spans
7.77
10.81
0.329
Clause permutation
2.61
3.36
0.252
Sentence permutation
2.12
2.60
0.203
Appendix
Table 6: Clean-normalized ordinary-distance responses of mean-SW and SOL under structural perturbations. Response columns are arithmetic means of paired ratios. Δ is the mean paired log-response contrast.
Encoder
Length
t=0.7
t=0.9
t=1.0
t=1.1
t=1.3
GPT-2-large
512
2.901
1.066
0.994
1.038
1.355
1024
2.872
1.091
1.014
1.070
1.434
Qwen3-0.6B
512
6.755
5.085
5.204
5.440
6.332
1024
6.779
5.179
5.257
5.565
6.538
OLMo-2-1B
512
7.719
4.051
4.864
5.947
8.211
1024
8.562
4.103
4.758
6.051
8.561
Appendix
Table 7: SOL evaluating generated text by GPT-2 Large at different decoding temperatures using differnet embedding models.
Temp.
Metric
vMF
Soft-Masked + ReMDM
MDLM
0.7
SOL
48.6981±0.292
57.6332±0.420
54.3128±0.404
Gen. PPL ( H )
7.92(3.950)
3.24(3.490)
7.26(3.884)
0.8
SOL
37.9781±0.160
41.3254±0.303
42.0520±0.244
Gen. PPL ( H )
15.28(4.656)
5.33(4.360)
14.18(4.657)
0.9
SOL
33.1134±0.076
24.7670±0.161
37.4379±0.149
Gen. PPL ( H )
26.36(5.054)
9.18(4.982)
27.01(5.093)
Appendix
Table 8: Temperature sweep on OpenWebText for vMF, Soft-Masked with ReMDM-loop, and MDLM. Panels (a) and (b) report SOL using Dream-7B and GPT-2 Large, respectively. Gen. PPL is evaluated with GPT-2 Large and H denotes mean per-sequence token-frequency entropy in nats; these quantities are independent of the SOL encoder. SOL uses 5,000 samples per side and a 1,024-token horizon, with mean ± population SD over nine seed pairs. Lower SOL is better and bold marks its column minimum.
Figure 11: SOL under increasing [CLS] insertion (left) and word-position shuffling (right) for four encoders. For all four language-model encoders, SOL increases monotonically with the corruption rate, and the two corruptions produce responses of similar magnitude. In contrast to the text encoders in Figure 12 , no encoder is markedly less sensitive to shuffling than to insertion.
Corruption
[CLS] pollution
Word shuffling
0%
0.9702
0.9702
1%
0.9367
0.9701
2%
0.9009
0.9667
5%
0.7888
0.9527
10%
0.6559
0.8423
20%
0.4558
0.3555
Appendix
Table 9: MAUVE under [CLS] pollution and word shuffling using last hidden representation.
Encoder
Clean
Prefix
Suffix
OLMo-2-1B
0.870
2.048
0.913
Dream-7B
2.493
4.496
4.621
Qwen3-0.6B
0.914
1.937
0.946
GPT-2-large
0.224
0.340
0.226
Appendix
Table 10: Shuffle 20% of the tokens the first or last 5% of the text.
Metric
Q2→Q1
Q2→Q3
SOL
1.000
1.000
mean-SW
1.000
0.9995
Appendix
Table 11: Spearman correlations with the target share along the Q2→Q1 and Q2→Q3 interpolation paths.
Encoder
Length
t=0.7
t=0.9
t=1.0
t=1.1
t=1.3
GPT-2-large
512
2.141
1.299
1.138
1.035
1.032
1024
–
–
–
–
–
Qwen3-0.6B
512
6.083
4.310
4.288
4.347
4.847
1024
6.695
4.513
4.470
4.645
5.311
OLMo-2-1B
512
7.822
5.452
5.016
4.808
5.203
1024
9.413
6.440
5.421
4.916
4.884
Appendix
Table 12: SOL scores using the [x,x] second-pass representation across sampling temperatures.
Region
[x]
[x,x]
[x,<sep>,x]
Clean
0.910
0.732
0.853
Prefix
1.710
1.167
1.926
Middle
1.335
1.347
1.375
Suffix
0.942
1.599
0.958
Appendix
Table 13: OLMo regional corruption, cap 1,024. Mean SOL for 10% replacement within a 5% region, taking the square root per seed before averaging, clean scores use uncorrupted candidate text.
Temp.
[x]
[x,x]
[x,<sep>,x]
0.7
8.562
9.413
8.152
0.9
4.103
6.440
4.060
1.0
4.758
5.421
4.703
1.1
6.051
4.916
5.931
1.3
8.561
4.884
8.205
Appendix
Table 14: SOL: GPT-2-large generations scored with OLMo at cap 1,024. Square root of the mean squared score across seeds. Bold marks the minimum within each method.
Figure 12: SOL under increasing shuffling of tokens (left) and inserting a CLS token (right). Compared to GPT2-Large, the text encoders are less sensitive to shuffling than GPT-2 Large.
Insertion rate
SOL (↓)
20%
1
50%
0.942
Appendix
Table 15: NeoBERT SOL under increasing [CLS] insertion, normalized to the 20% value.
MAUVE
SOL
Criterion ↑
GPT-2 Large
OLMo2 1B
Dream 7B
GPT-2 Large
OLMo2 1B
Dream 7B
Human-like
0.881
0.929
0.905
0.952
0.952
0.976
Interesting
0.643
0.786
0.690
0.762
0.762
0.857
Sensible
0.833
0.881
0.881
0.929
0.929
0.976
Appendix
Table 16: Human judgement comparison at sequence length 256. Worst-case Spearman rank correlation between metric-induced model rankings and the Bradley–Terry (BT) human ranking. Results are shown for MAUVE and SOL using GPT-2 Large, OLMo2 1B, and Dream 7B embeddings. Reported are the minimum correlations ρmin across the evaluated settings at sequence length 256.
Corruption
[CLS] pollution
Word shuffling
0%
0.9725
0.9725
1%
0.9590
0.9710
2%
0.9183
0.9697
5%
0.4709
0.9597
10%
0.0624
0.8236
20%
0.0181
0.1208
Appendix
Table 17: MAUVE under [CLS] pollution and word shuffling using mean pooling.
Figure 13: For two disjoint LM1B corpora, the reported MAUVE score depends on the number of evaluation samples n and clusters K . In comparison to generated vs true data the ratio n/10=K is stable, but increasing n converges to 1.
Model
Middle
Final
Real OWT
6.578
6.572
DUO ancestral
6.455
6.441
DUO greedy
6.317
5.011
Appendix
Table 18: Across-document token entropy Hpos in nats ( N=5000 ), measured at position 512 and the last valid token.
MAUVE (↑)
SOL (↓)
Model
Base
Reweighted
Base
Reweighted
DUO ancestral
0.836
0.843
1.670
1.7418
DUO greedy
0.492
0.591
1.748
1.8538
Appendix
Table 19: Final-token reweighting on matched corrected inputs ( N=5000 ). GPT-2-large SOL values average nine projection-seed combinations.
Model
Original
Remove final 8
ELF-M
0.455±0.011
0.823±0.017
ELF-L
0.287±0.006
0.798±0.013
Appendix
Table 20: MAUVE (↑) for shared ELF sampling. Truncation removes eight tokens from both generated and reference texts. Mean ± SD over three clustering seeds.
Rep.
ΔP
ΔR
ΔC
1
+2.14
+6.70
+3.10
2
+3.22
+4.88
+1.80
3
+2.10
+8.32
+0.84
Mean
+2.49
+6.63
+1.91
SD
0.64
1.72
1.13
Appendix
Table 21: ReMDM-loop minus MDLM at k=4 , in percentage points. SD is across three independent repetitions.
Sampler
MAUVE ( ↑ )
SOL ( ↓ )
Greedy
0.0697±0.0012
3.7755±0.1190
Ancestral
0.9407±0.0028
0.5083±0.0113
Nucleus ( p=0.95 )
0.9607±0.0017
0.2712±0.0065
Appendix
Table 22: GPT-2-XL conditional completion on 5,000 WebText documents, using shared 35 -token prefixes and a maximum total sequence length of 256 tokens. SOL uses GPT-2-large.
Sampler
MAUVE ( ↑ )
SOL ( ↓ )
Ancestral
0.8832±0.0071
0.8358±0.0245
Nucleus ( p=0.95 )
0.8740±0.0073
0.9072±0.0279
Appendix
Table 23: GPT-2-XL unconditional generation evaluated against 5,000 WebText documents. Each of the 5,000 samples starts from the end-of-text token and stops at EOS or 256 generated tokens, the initial seed token is excluded from evaluation. SOL uses GPT-2-large.
SOL (↓)
Model
Setting
H@1024
Gen. ppl.
GPT-2-large
OLMo-2-1B
Dream-7B
Real OWT G
—
5.449
14.9
0.2002±0.0116
0.7800±0.0552
2.066±0.109
Real ELF OWT E
—
5.402
17.0
0.2457±0.0239
0.8857±0.0636
2.649±0.219
Autoregressive model Sahoo et al. (2024)
AR G
Temperature 1.0
5.582
34.9
1.064±0.008
7.450±0.132
22.207±0.099
Discrete diffusion
Appendix
Table 24: OpenWebText baselines using 5,000 samples per distribution and a 1,024-token horizon. The standard SOL hyperparameters are used; see Section C.3 . Superscripts G, E and B identify the evaluation reference family. SOL is reported as mean ± population std over nine crossed direction/GP seed pairs.