Existing open-ended generation metrics measure likelihood, lexical diversity, or distributional similarity in generic representation space, yet can miss fundamental dimensions of quality. A prominent blind spot is global coherence: a generated passage may be locally fluent while remaining globally contradictory, causally inconsistent, or topically disconnected. We identify representation as a central bottleneck in detecting these failures and introduce CHORD (Coherence-aware Hidden-state Open-generation Reference Distance), a coherence-sensitive distributional metric. CHORD encodes generated and human-written corpora in the hidden-state space of a frozen LLM using a coherence-eliciting prompt, and compares the resulting distributions using RBF-MMD. To test coherence sensitivity and selectivity, we construct a counterfactual evaluation suite pairing graded coherence-degrading perturbations with meaning-preserving controls. CHORD selectively detects relation, discourse, structural, and mixture failures that perplexity, entropy, MAUVE, FBD, and MMD-based baselines either miss or cannot separate from benign rewriting. Factorial ablations show that representation is the primary source of coherence sensitivity, while RBF-MMD improves sample efficiency. Larger backbones capture finer-grained distinctions, but coherence prompting improves selectivity only when the backbone can follow the prompt. On unconditional generation and prefix continuation, CHORD yields model rankings that strongly align with human judgments of whether outputs make sense and appear human-written. Together, these results establish representation design as central to reliable distributional evaluation. Code: https://github.com/MAPS-research/CHORD. Experiments: https://github.com/MAPS-research/CHORD-Experiment.
Figures & tables
Figure 1: Motivation and intuition for CHORD . Existing corpus-level metrics may overlook text that is locally fluent yet globally incoherent; CHORD selectively distinguishes coherence degradation from meaning-preserving rewriting.
Figure 2: CHORD overview. A frozen LLM encodes prompted human and generated passages into coherence-oriented hidden states; RBF-MMD compares the resulting corpus distributions.
Relation
Discourse
Structural
Mixture
Control
harder to detect easier to detect
Method
Causal reversal
Contra- diction
Broken transition
Topic drift
Sentence permutation
Word shuffle
Repetition
DLM mix
Document mix
Benign paraphrase
Likelihood / diversity statistics
gen-PPL (GPT-2)
0.8
1.1
0.5
1.8
7.4
46.9
62.4
62.4
35.9
1.4
Unigram entropy
− 0.1
0.6
6.0
0.0
− 0.0
2.1
0.9
22.1
17.4
2.6
Distributional metrics
Table 1: Null-standardized responses zM at the highest severity for each perturbation type. Blue shading marks detected conditions: the lower bound of the paired-bootstrap confidence interval for selectivity Δz exceeds zero (Section 3.2 ). All CHORD backbones use the same coherence prompt and feature-extraction rule. Appendix H reports results at every severity level.
Figure 4
Generator
CHORD ↓ (Qwen3.5-27B)
CHORD ↓ (Qwen3.5-2B distilled)
CHORD ↓ (Qwen3.5-0.8B distilled)
gen-PPL ↓
MAUVE ↑
entropy ↑
Held-out human (packed)
0.17 #1 ( ± 0.04)
0.16 #1 ( ± 0.03)
0.15 #1 ( ± 0.05)
18.94 #3 ( ± 0.31)
0.95 #1 ( ± 0.01)
7.51 #2 ( ± 0.02)
GPT-2-large (774M, nucleus)
19.84 #2 ( ± 1.02)
9.98 #2 ( ± 0.74)
10.40 #3 ( ± 0.77)
6.70 #1 ( ± 0.10)
0.77 #5 ( ± 0.04)
7.03 #7 ( ± 0.02)
GPT-2-medium (355M, nucleus)
28.37 #3 ( ± 0.90)
10.78 #3 ( ± 0.59)
10.36 #2 ( ± 0.66)
10.10 #2 ( ± 0.11)
0.81 #4 ( ± 0.03)
7.06 #6 ( ± 0.01)
ELF-L (652M) ( Hu et al., 2026 )
51.19 #4 ( ± 1.09)
68.20 #4 ( ± 1.33)
69.72 #4 ( ± 1.45)
23.15 #5 ( ± 0.53)
0.09 #7 ( ± 0.01)
7.07 #5 ( ± 0.02)
LangFlow (171M) ( Chen et al., 2026 )
52.58 #5 ( ± 1.12)
70.84 #5 ( ± 1.54)
78.79 #7 ( ± 1.73)
18.98 #4 ( ± 0.49)
0.65 #6 ( ± 0.05)
7.60 #1 ( ± 0.03)
SEDD-small (170M) ( Lou et al., 2024 )
57.05 #6 ( ± 1.02)
74.95 #6 ( ± 1.27)
77.64 #6 ( ± 1.14)
69.21 #6 ( ± 0.96)
0.89 #2 ( ± 0.02)
7.21 #4 ( ± 0.02)
Table 2: Unconditional generation on OpenWebText. We report mean ± std over 10 non-overlapping sample sets, each containing 500 samples of 512 tokens. CHORD reports RBF-MMD 2 ( ×10−2 ); scores are comparable only within the same encoder column. Details are provided in Appendix J .
Figure 5: Prefix continuation on OpenWebText. Each model generates a 128-token continuation from a 128-token human prefix. Dotted lines mark the held-out human value.
CHORD (ours)
Existing metrics
Human dimension
Qwen3.5 27B
Qwen3.5-2B distilled
MAUVE (GPT-2)
MAUVE (ELECTRA)
FBD (BERT)
MMD (MiniLM)
gen-PPL (GPT-2)
Interesting
0.86
0.76
0.00
0.79
0.76
0.90
0.64
Makes sense
0.98
0.93
− 0.19
0.93
0.93
0.90
0.88
Human-like
0.98
0.95
− 0.21
0.90
0.95
0.93
0.88
Table 3: Agreement with human system rankings. Spearman correlation ( ρ ) between each metric and Bradley–Terry rankings derived from the human study of Pillutla et al. (2021) . The study contains 3,240 pairwise judgments over eight GPT-2 generation settings.
Appendix figures & tables28 assets
Supplementary material from the paper’s appendix.
Appendix
Perturbation
Before
After
Contradiction
“…became the worst kind of media free-for-all.”
“…became the most respectful kind of media coverage.”
Causal reversal
“You may opt out at any time.”
“You may opt in at any time.”
Broken transition
“His suicide suddenly made more sense.”
“His suicide suddenly made more sense, but the moon is made of green cheese.”
Topic drift
“…spoke strongly on behalf of the virtues of physical education.”
“…spoke strongly on behalf of the virtues of chess strategy.”
Sentence permutation
clean text sample
selected sentences are reordered within the sample
Word shuffle
clean text sample
words are locally shuffled within selected sentences
Appendix
Table 4: Representative examples from the counterfactual evaluation set. Only the targeted sentence or operation is modified; the surrounding text is preserved. Examples are shortened for readability.
Relation
Discourse
Editor model
Contra- diction
Causal reversal
Broken transition
Topic drift
Benign paraphrase
Detected (of 12)
Qwen3-30B-A3B (default)
84.3
25.6
331
79.7
6.0
10
Mistral-Small-24B
151
90.5
559
406
9.5
11
Appendix
Table 5: Robustness to the editor model. Null-standardized CHORD scores zM at the highest severity of each LLM-edited perturbation type. Shaded entries are detected relative to the benign control under the paired bootstrap criterion (Section 3.2 ). The final column reports detections across the 12 LLM-edited conditions. Scores are calibrated separately for each set and are not comparable across rows.
Method
Causal reversal
Contra- diction
Broken transition
Topic drift
Sentence permutation
Word shuffle
Repetition
DLM mix
Document mix
Benign rewriting
Instruction-tuned embeddings + RBF-MMD
Qwen3-Embedding-8B (none)
12.1
11.5
19.3
17.3
14.5
12.0
15.1
129
142
7.6
Qwen3-Embedding-8B (similarity)
9.3
13.0
22.5
14.8
10.7
8.4
19.0
76.2
73.5
5.0
Qwen3-Embedding-8B (coherence)
10.6
11.5
51.4
19.7
13.2
12.0
19.9
157
182
7.4
e5-mistral-7b-instruct (none)
8.3
7.7
30.3
12.8
13.2
22.0
102
87.6
92.1
5.6
e5-mistral-7b-instruct (similarity)
8.3
10.9
44.4
13.5
12.8
20.7
69.6
92.1
92.6
5.9
Appendix
Table 6: Additional baselines on the counterfactual evaluation set. Entries are null-standardized responses at the highest perturbation severity; blue shading indicates detection relative to benign rewriting, as in Table 1 . Embedding methods use RBF-MMD, with the input instruction shown in parentheses. Scalar evaluators use the absolute candidate–reference difference in mean score.
Figure 6: Hidden representations versus token counts. Under sentence shuffling, unigram selectivity stays near zero, while hidden-state selectivity increases with the shuffle rate. Both methods use the same RBF-MMD, human-reference null, and benign-paraphrase control.
Figure 7: Next-token predictions at the CHORD representation-eliciting position. Bars show changes in the fraction of passages producing each decoded string, relative to benign paraphrases. Coherence errors increase negative descriptions, but the outputs do not consistently express a clear coherence judgment.
You are a strict, careful writing-quality rater. Read the document between the <document> markers and rate its logical consistency, discourse flow, and overall quality on an integer scale from 0 (incoherent, broken, or self-contradictory) to 10 (flawless, fully coherent writing). Judge only the writing itself; do not reward or penalize the topic, opinions, or genre, and ignore truncation at the very end of the document.
<document>
{text}
</document>
Return only JSON, the score first: {"score": <integer 0-10>, "reason": "<one short sentence>"}
Appendix
Table 7: Direct coherence judge prompt. The same prompt is used for all passages, without a system prompt or few-shot examples.
Relation
Discourse
Structural
Mixture
Control
Method
Contra- diction
Causal reversal
Broken transition
Topic drift
Sentence permutation
Word shuffle
Repetition
DLM mix
Document mix
Benign rewriting
LLM judge (same Qwen3.5-27B)
16.2
9.5
26.0
16.9
28.1
40.0
42.6
49.0
53.0
0.3
G-Eval (same Qwen3.5-27B)
9.9
5.2
33.9
18.8
41.4
49.2
62.2
75.4
77.6
0.6
CHORD (Qwen3.5-27B)
84.3
25.6
331
79.7
358
740
857
1150
1082
6.0
Appendix
Table 8: Same-backbone comparison with LLM judges. All methods use frozen Qwen3.5-27B. Entries are null-standardized responses at the highest perturbation severity; shaded entries indicate detection relative to benign rewriting. For each judge, the statistic is the absolute difference between candidate and reference mean ratings.
Generator
Params
Judge mean (0–10) ↑
Judge zM
CHORD ↓ (MMD 2 ×10−2 )
Human (packed, held-out)
–
3.96 #1
− 0.1 ( ± 0.6) (n.s.)
0.17 #1 ( ± 0.04)
GPT2-large (AR)
774M
2.02 #2
19.9 ( ± 1.5)
19.84 #2 ( ± 1.02)
GPT2-medium (AR)
355M
1.67 #3
23.7 ( ± 1.3)
28.37 #3 ( ± 0.90)
ELF-L
652M
1.08 #4
30.1 ( ± 1.3)
51.19 #4 ( ± 1.09)
LangFlow
171M
0.52 #7
36.3 ( ± 1.3)
52.58 #5 ( ± 1.12)
SEDD-small
170M
0.85 #5
32.7 ( ± 1.3)
57.05 #6 ( ± 1.02)
Appendix
Table 9: Direct coherence judging on unconditional generation. Both methods use Qwen3.5-27B on the ten folds of Table 2 ; the CHORD column is the Qwen3.5-27B column of that table. The judge mean is the mean rating over all 5,000 texts of a generator. Judge zM standardizes the absolute difference between the generator and reference fold means against a null of two disjoint 500-text reference subsets, reported as mean ± std over folds. Rows are ordered by CHORD rank. Rank markers (# k ) order the corpora by higher mean rating or lower CHORD score; judge zM is not ranked. (n.s.) denotes a deviation that does not exceed the 95th percentile of the null.
Backbone
Extraction method
Mean Δz
Detected (of 40)
Reversed (of 40)
GPT-2 (0.124B)
Coherence prompt
33.0
11
18
Generic prompt
7.8
10
20
Raw last token
27.2
14
11
Mean pooling
76.2
24
0
GPT-2-medium (0.355B)
Coherence prompt
26.2
9
15
Generic prompt
31.6
10
14
Appendix
Table 10: How the extraction method changes coherence selectivity. Each row summarizes 40 harmful conditions. “Detected” counts conditions whose selectivity is reliably positive; “reversed” counts conditions whose response is reliably stronger for benign rewriting than for coherence damage. Bold marks the strongest extraction method for each backbone.
Figure 8: Target attribute and prompt format. (a) Selectivity for compact prompts naming different target attributes, normalized within each perturbation category. (b) Selectivity for three prompt formats; error bars span three prompt wordings.
Attribute
Template
Coherence
This passage: “ x ”, considering its coherence and ordering of its ideas, means in one word:
Grammar
This passage: “ x ”, considering its grammatical correctness and sentence structure, means in one word:
Quality
This passage: “ x ”, considering its overall writing quality, means in one word:
Topic
This passage: “ x ”, considering its main topic and subject matter, means in one word:
Sentiment
This passage: “ x ”, considering its overall sentiment and emotional tone, means in one word:
Neutral
This sentence: “ x ” means in one word: (PromptEOL template; Jiang et al., 2024 )
Appendix
Table 11: Target-attribute templates. The compact-completion format is fixed; only the target attribute changes.
Format
Template
Compact, wording 1
This passage: “ x ”, in terms of its logical coherence and the order of its ideas, means in one word:
Compact, wording 2
This passage: “ x ”, considering whether its ideas form a logically connected and well-ordered whole, means in one word:
Compact, wording 3
This passage: “ x ”, considering the consistency, organization, and logical flow of its ideas, means in one word:
Scalar rating (wording 1)
This passage: “ x ”. Considering its logical coherence and the order of its ideas, rate it from 1 to 5:
Direct answer (wording 1)
This passage: “ x ”. Considering its logical coherence and the order of its ideas, provide a yes or no judgment:
Appendix
Table 12: Coherence-wording and prompt-format templates. The first three rows list the compact wordings used in both the counterfactual and human-agreement robustness experiments. The rating and direct-answer rows show wording 1; wordings 2 and 3 substitute the corresponding coherence clause from the compact rows.
Figure 9: Score separation versus corpus size. On anchored contradiction, kernel distances separate harmful from null at smaller N than Fréchet distance or k -means KL. The benign curve shows that paraphrase-induced shifts also become detectable at large N .
Figure 10: CHORD responses across perturbation severities. Scores are null-standardized shifts zM on a log scale; open markers indicate conditions not detected relative to the benign control.
Relation
Discourse
Structural
Mixture
Control
#Detected
Method
Contr.
Causal
Broken
Topic
Perm.
Shuf.
Rep.
DLM mix
Doc. mix
Benign
Sem. (12)
Form (28)
Likelihood / diversity statistics
gen-PPL (GPT-2)
0.3 1.1
− 0.1 0.8
0.4 0.5
1.3 1.8
1.9 7.4 ∗
6.2 ∗ 46.9 ∗
9.7 ∗ 62.4 ∗
18.5 ∗ 62.4 ∗
10.7 ∗ 35.9 ∗
0.2 1.4
0/12
26/28
Unigram entropy
0.1 0.6
− 0.2 − 0.1
3.7 6.0
− 0.3 0.0
0.3 − 0.0
0.5 2.1
0.0 0.9
6.9 ∗ 22.1 ∗
7.5 ∗ 17.4 ∗
0.7 2.6
0/12
11/28
Distributional metrics
MAUVE (GPT-2)
32.2 27.6
17.4 23.1
29.1 21.0
24.1 18.9
1.1 24.5
− 0.3 1.1
13.0 56.5 ∗
14.7 35.8 ∗
9.4 44.9 ∗
27.8 27.3
0/12
8/28
Appendix
Table 13: Results across perturbation severities , underlying Table 1 . Each cell reports the null-standardized score zM at the lowest severity (top) and highest severity (bottom) for each perturbation type. All rows use the same calibration as Table 1 . A ∗ marks a condition detected relative to the benign control. The final columns report detected conditions across all severity levels: semantic (relation and discourse; 12 total) and surface-form (structural and mixture; 28 total).
gen.
MAUVE ↑ at truncation length
Corpus
tokens
128
256
384
512
1024
Held-out human (packed)
493
0.95 #1
0.94 #1
0.94 #1
0.95 #1
0.95 #1
GPT-2-large (nucleus)
483
0.70 #7
0.67 #7
0.58 #7
0.83 #5
0.78 #5
GPT-2-medium (nucleus)
486
0.74 #6
0.70 #6
0.62 #6
0.84 #4
0.81 #4
SEDD
493
0.91 #3
0.89 #3
0.88 #3
0.89 #2
0.89 #2
MDLM
493
0.92 #2
0.91 #2
0.90 #2
0.88 #3
0.85 #3
Appendix
Table 14: MAUVE sensitivity to evaluation length. Documents are truncated at sentence boundaries to the stated token budget. Each cell is the mean over the ten folds of Table 2 . Superscripts give ranks among the seven corpora; higher MAUVE is better.
CHORD zM↓ at truncation length
Corpus
128
256
384
512
1024
Held-out human (packed)
0.5 #1
0.4 #1
0.6 #1
0.3 #1
0.0 #1
GPT-2-large (nucleus)
219 #2
361 #2
439 #2
365 #2
291 #2
GPT-2-medium (nucleus)
331 #3
507 #3
611 #3
522 #3
417 #3
SEDD
962 #7
1108 #6
1208 #6
1046 #5
843 #5
MDLM
954 #6
1117 #7
1209 #7
1060 #7
856 #6
Appendix
Table 15: CHORD sensitivity to evaluation length. Documents use the same truncations and folds as Table 14 ; each cell is the mean over the ten folds. Superscripts give ranks among the seven corpora; lower zM is better. The human < autoregressive < diffusion/flow ordering holds at every length.
Continuation
CHORD zM↓
MAUVE ↑
gen-PPL ↓
entropy ↑
Human
0.1 #1
0.92 #2
27.7 #3
7.07 #4
GPT-2-medium (nucleus)
58.3 #2
0.80 #5
16.6 #2
6.81 #5
GPT-2-medium (hot, T=1.6 )
145.3 #3
0.94 #1
68.4 #4
7.28 #1
GPT-2-medium (greedy)
216.9 #4
0.11 #6
2.9 #1
6.12 #6
SEDD (128 steps)
244.1 #5
0.88 #3
113.0 #5
7.09 #3
MDLM (128 steps)
291.1 #6
0.87 #4
155.3 #6
7.20 #2
Appendix
Table 16: Prefix-continuation results plotted in Figure 5 . CHORD scores are null-standardized with n=80 per subset. Superscript rank markers (# k ) order all six rows, including the human continuations, by each metric’s conventional direction.
Annotator 2
Annotator 1
Human better
Tie
Model better
Total
Human better
40
2
3
45
Tie
3
2
1
6
Model better
2
0
7
9
Total
45
4
11
60
Appendix
Table 17: Inter-annotator agreement. Joint distribution of the two annotators’ judgments over the 60 pairs. Rows give annotator 1 and columns annotator 2; the 49 diagonal pairs are agreements.
Generator
Pairs
Human wins
Tie
Model wins
Disagree
GPT-2-large (nucleus)
10
45.0%
25.0%
30.0%
40%
GPT-2-medium (nucleus)
10
50.0%
15.0%
35.0%
20%
SEDD
10
80.0%
5.0%
15.0%
20%
MDLM
10
100.0%
0.0%
0.0%
0%
ELF-L
10
80.0%
0.0%
20.0%
20%
LangFlow
10
95.0%
5.0%
0.0%
10%
Appendix
Table 18: Blind pairwise human evaluation on unconditional generation. Each row compares held-out human text with the named generator. Pairs is the number of independent text pairs; both annotators judge every pair, so the win and tie percentages are computed over twice as many judgments. Disagree reports the fraction of pairs on which the annotators gave different judgments. The final two rows pool both autoregressive systems and all four diffusion and flow systems; brackets give 95% bootstrap confidence intervals over pairs.
Figure 11: Effect of coherence-failure position on detection. One topic-drift sentence is injected at the prefix, middle, or suffix of otherwise clean text samples, on a fixed frozen Qwen3.5-9B backbone. (a) Δz normalized by the prefix value; flatter curves indicate less positional dependence. (b) Suffix-to-prefix Δz ratio. The dotted line at 1 × marks equal sensitivity at the prefix and suffix.
Figure 12: Effect of extraction layer on coherence selectivity. Each point reads the hidden state from a different transformer layer, with the prompt, distance, null calibration, and evaluation conditions fixed. (a) Mean Δz , normalized by each backbone’s peak. (b) Number of the nine perturbation types detected selectively. Both backbones show weak signals in early layers and a broad high-selectivity plateau in late layers. The adopted −3 layer lies on this plateau, and the final layer does not improve on it.
Figure 13: Cost against quality. Peak GPU memory (left) and feature extraction throughput (right) from Table 19 , plotted against the number of selectively detected perturbation types in Table 1 . Orange circles denote CHORD encoders; blue squares denote baselines. Lower memory, higher throughput, and more detected types are better.
Featurize (GPU)
CPU
Metric
Extractor
Params
dim
GFLOPs/doc
batch
wall (s)
docs/s
b=1 (ms)
peak (GB)
weights (GB)
statistic (s)
CHORD (Qwen3.5-27B)
Qwen3.5-27B
26.1B
5120
26207
8
46.5 ± 1.6
10.8
170.8
54.2
48.6
0.5
CHORD (Qwen3.5-9B)
Qwen3.5-9B
8.4B
4096
8377
16
14.0 ± 0.1
35.6
76.8
20.5
15.6
0.4
CHORD (Qwen3.5-2B distilled)
Qwen3.5-2B + LoRA + PS
2.2B
256
2224
32
5.3 ± 0.1
94.1
67.7
7.1
4.1
0.1
CHORD (Qwen3.5-0.8B distilled)
Qwen3.5-0.8B + LoRA + PS
853M
256
857
32
4.4 ± 0.0
112.5
57.9
4.5
1.6
0.1
MAUVE (GPT-2)
GPT-2-large
774M
1280
789
32
2.1 ± 0.4
241.4
17.4
5.1
1.4
0.6 †
Appendix
Table 19: Runtime and memory on 500 packed OpenWebText documents, single H200. Params includes merged LoRA ( Hu et al., 2022 ) weights and the student’s projection head PS (Appendix I ); dim is the embedding dimension. Wall is total feature extraction time; b=1 is per-document latency at batch size 1. Peak is maximum allocated GPU memory; weights is parameter storage at the stated precision. Statistic reports CPU time for ten MMD folds; † marks one MAUVE or FBD comparison of 250 against 250 documents. We include gen-PPL at batch sizes 8 and 32 to show the effect of batching.
Wrong fact
Faithful
Representation
number
entity
negation
paraphrase
CHORD (Qwen3.5-27B)
258.7
112.0
380.5
− 0.4
CHORD (Qwen3.5-2B distilled)
216.4
99.0
504.2
− 0.4
Qwen3.5-9B (layer −3 )
164.6
55.6
237.2
− 0.4
GPT-2-large (last token)
4.9
1.0
5.6
0.7
NeoBERT ( Breton et al., 2025 ) (masked mean)
10.9
5.8
4.4
0.2
Appendix
Table 20: Source-conditioned QA faithfulness on SQuAD. Each entry is a null-standardized score zM ; higher values indicate a larger shift from the correct-answer reference. The faithful paraphrase is an independent correct rewrite and should remain near the baseline.
Interesting
Makes sense
Human-like
CHORD -27B
0.86 (0.64)
0.98 ( 0.93 )
0.98 ( 0.93 )
CHORD -2B (distilled)
0.76 (0.57)
0.93 (0.86)
0.95 (0.86)
MAUVE (GPT-2)
0.00 (0.14)
− 0.19 (0.00)
− 0.21 (0.00)
MAUVE (ELECTRA)
0.79 (0.64)
0.93 (0.79)
0.90 (0.79)
FBD (BERT)
0.76 (0.57)
0.93 (0.86)
0.95 (0.86)
MMD (MiniLM)
0.90 ( 0.79 )
0.90 (0.79)
0.93 (0.79)
Appendix
Table 21: Correlation with human rankings on the human study of Pillutla et al. (2021) . Each cell reports Spearman’s ρ (Kendall’s τ in parentheses) between the metric’s ranking of the eight GPT-2 generation settings and the human Bradley–Terry ranking, computed on the exact 256-token texts shown to annotators. The best value in each column is in bold. Distance-valued metrics ( CHORD , FBD, MMD-MiniLM) are negated before correlation; for gen-PPL and unigram entropy we use the distance to the human-corpus value, −∣m−mhuman∣ , the most favorable orientation for these two-sided statistics.
Role
Size
Content
Reference
5,500
Safe Alpaca-7B responses
Clean candidates
1,000
Safe Alpaca-7B responses before replacement
Matched donor pairs
500
Safe/unsafe siblings from one-safe rows
Generator-shift donors
500
Safe Alpaca3-8B responses
Appendix
Table 22: Construction of the safety-prevalence experiment. “One-safe” rows contain one safe and one unsafe response to the same prompt.
Metric / prompt
1%
2%
5%
10%
25%
50%
Detected
CHORD , safety prompt (w1)
0.09
0.39
2.27
10.23
63.27
238.38
5%
CHORD , safety prompt (w2)
0.13
0.59
2.23
9.62
54.27
204.53
5%
CHORD , safety prompt (w3)
0.10
0.60
2.72
10.56
58.12
237.59
5%
CHORD , coherence prompt
0.08
0.55
2.55
9.43
36.47
135.45
5%
CHORD , neutral prompt
0.00
0.33
1.63
5.64
32.62
122.47
5%
Granite Guardian 3.1 2B
0.31
0.87
2.26
5.15
13.94
28.28
10%
Appendix
Table 23: Selectivity to unsafe-content prevalence. Entries are Δz=zunsafe−zsafe at matched replacement rates. Bold CHORD entries have a paired 95% confidence interval above zero. “Detected” gives the lowest reliably detected rate.
Evaluating open-ended text generation involves understanding how different properties of a continuation relate to its perceived quality. We present a reference-based framework for examining coherence and diversity through three perspectives: aligning their evolution with human trajectories, comparing their summaries with a human continuation of the same prompt, and estimating their likelihood under a human reference distribution. Experiments with human quality ratings suggest that diversity-based alignment and mean-based comparisons capture quality-related variation, although the comparisons do not establish a predictive advantage for temporal alignment over simpler baselines. Reference likelihood also shows positive associations with ratings, with results varying across reference configurations and scoring horizons. Together, these analyses provide a structured way to examine how measured coherence and diversity relate to human judgments, while distinguishing similarity to human references from quality itself. Code and analysis resources are available at https://github.com/EstebanGarces/likely_human.
Esteban Garcés Arias
Department of Statistics, LMU Munich Munich Center for Machine Learning (MCML)
Diffusion and continuous flow-based language models have emerged as the leading non-autoregressive alternatives to language modeling. Progress in both paradigms is overwhelmingly tracked by generative perplexity (gen-PPL): the per-token negative log-likelihood of samples under a frozen autoregressive (AR) scorer such as gpt2-large, typically paired with an empirical-entropy guardrail to rule out low-entropy collapse. We argue that this metric is unsound. By construction, gen-PPL measures only predictability under the scoring AR, not grammaticality or semantic coherence -- and the set of predictable but still low-quality sequences is combinatorially large. To make this concrete, we construct a suite of zero-parameter, deliberately naive samplers that achieve state-of-the-art gen-PPL on LM1B and OpenWebText at non-degenerate entropy, surpassing recently published diffusion and continuous-flow models while producing text that is incoherent by construction. We recommend evaluation suites that directly quantify the distributional divergence between generated and reference text, and use such a suite to re-benchmark recent non-autoregressive models, recovering a more faithful picture of the current state of the art.
Evaluating open-ended outputs from large language models (LLMs) remains challenging due to the absence of ground truth. Existing metrics rely on final-answer accuracy or surface-level statistics, leaving the reasoning process itself unexamined. We introduce TRACE (Toulmin-based Reasoning Assessment through Constructive Elements), a metric that analyzes Chain-of-Thought (CoT) reasoning processes. Rather than judging outcomes, TRACE inspects how arguments are constructed by integrating Toulmin's argumentation theory with Flavell's metacognitive framework to assess reasoning structure. Experiments on 26.3K QA samples across 7 reasoning models show strong correlation with benchmark accuracy (r=0.74). Furthermore, TRACE is effective as a reinforcement learning reward signal, outperforming accuracy-only baselines. Together, these results indicate that logically sound reasoning leads to higher-quality answers. TRACE thus serves as a complementary metric for evaluating open-ended outputs. Code is available at https://github.com/hyyangkisti/trace.
Yundong Kim, Heyoung Yang
Applied Agent Research Center, Korea Institute of Science and Technology Information (KISTI), Republic of Korea · Department of Computer Science and Engineering, University of Seoul, Republic of Korea.