Coherence-Aware Distributional Evaluation of Open-Ended Text Generation
Organizations: New York University · Center for Data Science, NYU Shanghai · Georgia Institute of Technology
Abstract
Existing open-ended generation metrics measure likelihood, lexical diversity, or distributional similarity in generic representation space, yet can miss fundamental dimensions of quality. A prominent blind spot is global coherence: a generated passage may be locally fluent while remaining globally contradictory, causally inconsistent, or topically disconnected. We identify representation as a central bottleneck in detecting these failures and introduce CHORD (Coherence-aware Hidden-state Open-generation Reference Distance), a coherence-sensitive distributional metric. CHORD encodes generated and human-written corpora in the hidden-state space of a frozen LLM using a coherence-eliciting prompt, and compares the resulting distributions using RBF-MMD. To test coherence sensitivity and selectivity, we construct a counterfactual evaluation suite pairing graded coherence-degrading perturbations with meaning-preserving controls. CHORD selectively detects relation, discourse, structural, and mixture failures that perplexity, entropy, MAUVE, FBD, and MMD-based baselines either miss or cannot separate from benign rewriting. Factorial ablations show that representation is the primary source of coherence sensitivity, while RBF-MMD improves sample efficiency. Larger backbones capture finer-grained distinctions, but coherence prompting improves selectivity only when the backbone can follow the prompt. On unconditional generation and prefix continuation, CHORD yields model rankings that strongly align with human judgments of whether outputs make sense and appear human-written. Together, these results establish representation design as central to reliable distributional evaluation. Code: https://github.com/MAPS-research/CHORD. Experiments: https://github.com/MAPS-research/CHORD-Experiment.
Figures & tables
| Relation | Discourse | Structural | Mixture | Control | ||||||
| harder to detect easier to detect | ||||||||||
| Method | Causal reversal | Contra- diction | Broken transition | Topic drift | Sentence permutation | Word shuffle | Repetition | DLM mix | Document mix | Benign paraphrase |
| Likelihood / diversity statistics | ||||||||||
| gen-PPL (GPT-2) | 0.8 | 1.1 | 0.5 | 1.8 | 7.4 | 46.9 | 62.4 | 62.4 | 35.9 | 1.4 |
| Unigram entropy | 0.1 | 0.6 | 6.0 | 0.0 | 0.0 | 2.1 | 0.9 | 22.1 | 17.4 | 2.6 |
| Distributional metrics | ||||||||||
| Generator | CHORD (Qwen3.5-27B) | CHORD (Qwen3.5-2B distilled) | CHORD (Qwen3.5-0.8B distilled) | gen-PPL | MAUVE | entropy |
| Held-out human (packed) | 0.17 #1 ( 0.04) | 0.16 #1 ( 0.03) | 0.15 #1 ( 0.05) | 18.94 #3 ( 0.31) | 0.95 #1 ( 0.01) | 7.51 #2 ( 0.02) |
| GPT-2-large (774M, nucleus) | 19.84 #2 ( 1.02) | 9.98 #2 ( 0.74) | 10.40 #3 ( 0.77) | 6.70 #1 ( 0.10) | 0.77 #5 ( 0.04) | 7.03 #7 ( 0.02) |
| GPT-2-medium (355M, nucleus) | 28.37 #3 ( 0.90) | 10.78 #3 ( 0.59) | 10.36 #2 ( 0.66) | 10.10 #2 ( 0.11) | 0.81 #4 ( 0.03) | 7.06 #6 ( 0.01) |
| ELF-L (652M) ( Hu et al., 2026 ) | 51.19 #4 ( 1.09) | 68.20 #4 ( 1.33) | 69.72 #4 ( 1.45) | 23.15 #5 ( 0.53) | 0.09 #7 ( 0.01) | 7.07 #5 ( 0.02) |
| LangFlow (171M) ( Chen et al., 2026 ) | 52.58 #5 ( 1.12) | 70.84 #5 ( 1.54) | 78.79 #7 ( 1.73) | 18.98 #4 ( 0.49) | 0.65 #6 ( 0.05) | 7.60 #1 ( 0.03) |
| SEDD-small (170M) ( Lou et al., 2024 ) | 57.05 #6 ( 1.02) | 74.95 #6 ( 1.27) | 77.64 #6 ( 1.14) | 69.21 #6 ( 0.96) | 0.89 #2 ( 0.02) | 7.21 #4 ( 0.02) |
| CHORD (ours) | Existing metrics | ||||||
| Human dimension | Qwen3.5 27B | Qwen3.5-2B distilled | MAUVE (GPT-2) | MAUVE (ELECTRA) | FBD (BERT) | MMD (MiniLM) | gen-PPL (GPT-2) |
| Interesting | 0.86 | 0.76 | 0.00 | 0.79 | 0.76 | 0.90 | 0.64 |
| Makes sense | 0.98 | 0.93 | 0.19 | 0.93 | 0.93 | 0.90 | 0.88 |
| Human-like | 0.98 | 0.95 | 0.21 | 0.90 | 0.95 | 0.93 | 0.88 |
Appendix figures & tables28 assets
Supplementary material from the paper’s appendix.
Appendix
| Perturbation | Before | After |
| Contradiction | “…became the worst kind of media free-for-all.” | “…became the most respectful kind of media coverage.” |
| Causal reversal | “You may opt out at any time.” | “You may opt in at any time.” |
| Broken transition | “His suicide suddenly made more sense.” | “His suicide suddenly made more sense, but the moon is made of green cheese.” |
| Topic drift | “…spoke strongly on behalf of the virtues of physical education.” | “…spoke strongly on behalf of the virtues of chess strategy.” |
| Sentence permutation | clean text sample | selected sentences are reordered within the sample |
| Word shuffle | clean text sample | words are locally shuffled within selected sentences |
| Relation | Discourse | |||||
| Editor model | Contra- diction | Causal reversal | Broken transition | Topic drift | Benign paraphrase | Detected (of 12) |
| Qwen3-30B-A3B (default) | 84.3 | 25.6 | 331 | 79.7 | 6.0 | 10 |
| Mistral-Small-24B | 151 | 90.5 | 559 | 406 | 9.5 | 11 |
| Method | Causal reversal | Contra- diction | Broken transition | Topic drift | Sentence permutation | Word shuffle | Repetition | DLM mix | Document mix | Benign rewriting |
| Instruction-tuned embeddings + RBF-MMD | ||||||||||
| Qwen3-Embedding-8B (none) | 12.1 | 11.5 | 19.3 | 17.3 | 14.5 | 12.0 | 15.1 | 129 | 142 | 7.6 |
| Qwen3-Embedding-8B (similarity) | 9.3 | 13.0 | 22.5 | 14.8 | 10.7 | 8.4 | 19.0 | 76.2 | 73.5 | 5.0 |
| Qwen3-Embedding-8B (coherence) | 10.6 | 11.5 | 51.4 | 19.7 | 13.2 | 12.0 | 19.9 | 157 | 182 | 7.4 |
| e5-mistral-7b-instruct (none) | 8.3 | 7.7 | 30.3 | 12.8 | 13.2 | 22.0 | 102 | 87.6 | 92.1 | 5.6 |
| e5-mistral-7b-instruct (similarity) | 8.3 | 10.9 | 44.4 | 13.5 | 12.8 | 20.7 | 69.6 | 92.1 | 92.6 | 5.9 |
| You are a strict, careful writing-quality rater. Read the document between the <document> markers and rate its logical consistency, discourse flow, and overall quality on an integer scale from 0 (incoherent, broken, or self-contradictory) to 10 (flawless, fully coherent writing). Judge only the writing itself; do not reward or penalize the topic, opinions, or genre, and ignore truncation at the very end of the document. |
| <document> |
| {text} |
| </document> |
| Return only JSON, the score first: {"score": <integer 0-10>, "reason": "<one short sentence>"} |
| Relation | Discourse | Structural | Mixture | Control | ||||||
| Method | Contra- diction | Causal reversal | Broken transition | Topic drift | Sentence permutation | Word shuffle | Repetition | DLM mix | Document mix | Benign rewriting |
| LLM judge (same Qwen3.5-27B) | 16.2 | 9.5 | 26.0 | 16.9 | 28.1 | 40.0 | 42.6 | 49.0 | 53.0 | 0.3 |
| G-Eval (same Qwen3.5-27B) | 9.9 | 5.2 | 33.9 | 18.8 | 41.4 | 49.2 | 62.2 | 75.4 | 77.6 | 0.6 |
| CHORD (Qwen3.5-27B) | 84.3 | 25.6 | 331 | 79.7 | 358 | 740 | 857 | 1150 | 1082 | 6.0 |
| Generator | Params | Judge mean (0–10) | Judge | CHORD (MMD 2 ) |
| Human (packed, held-out) | – | 3.96 #1 | 0.1 ( 0.6) (n.s.) | 0.17 #1 ( 0.04) |
| GPT2-large (AR) | 774M | 2.02 #2 | 19.9 ( 1.5) | 19.84 #2 ( 1.02) |
| GPT2-medium (AR) | 355M | 1.67 #3 | 23.7 ( 1.3) | 28.37 #3 ( 0.90) |
| ELF-L | 652M | 1.08 #4 | 30.1 ( 1.3) | 51.19 #4 ( 1.09) |
| LangFlow | 171M | 0.52 #7 | 36.3 ( 1.3) | 52.58 #5 ( 1.12) |
| SEDD-small | 170M | 0.85 #5 | 32.7 ( 1.3) | 57.05 #6 ( 1.02) |
| Backbone | Extraction method | Mean | Detected (of 40) | Reversed (of 40) |
| GPT-2 (0.124B) | Coherence prompt | 33.0 | 11 | 18 |
| Generic prompt | 7.8 | 10 | 20 | |
| Raw last token | 27.2 | 14 | 11 | |
| Mean pooling | 76.2 | 24 | 0 | |
| GPT-2-medium (0.355B) | Coherence prompt | 26.2 | 9 | 15 |
| Generic prompt | 31.6 | 10 | 14 |
| Attribute | Template |
| Coherence | This passage: “ ”, considering its coherence and ordering of its ideas, means in one word: |
| Grammar | This passage: “ ”, considering its grammatical correctness and sentence structure, means in one word: |
| Quality | This passage: “ ”, considering its overall writing quality, means in one word: |
| Topic | This passage: “ ”, considering its main topic and subject matter, means in one word: |
| Sentiment | This passage: “ ”, considering its overall sentiment and emotional tone, means in one word: |
| Neutral | This sentence: “ ” means in one word: (PromptEOL template; Jiang et al., 2024 ) |
| Format | Template |
| Compact, wording 1 | This passage: “ ”, in terms of its logical coherence and the order of its ideas, means in one word: |
| Compact, wording 2 | This passage: “ ”, considering whether its ideas form a logically connected and well-ordered whole, means in one word: |
| Compact, wording 3 | This passage: “ ”, considering the consistency, organization, and logical flow of its ideas, means in one word: |
| Scalar rating (wording 1) | This passage: “ ”. Considering its logical coherence and the order of its ideas, rate it from 1 to 5: |
| Direct answer (wording 1) | This passage: “ ”. Considering its logical coherence and the order of its ideas, provide a yes or no judgment: |
| Relation | Discourse | Structural | Mixture | Control | #Detected | |||||||
| Method | Contr. | Causal | Broken | Topic | Perm. | Shuf. | Rep. | DLM mix | Doc. mix | Benign | Sem. (12) | Form (28) |
| Likelihood / diversity statistics | ||||||||||||
| gen-PPL (GPT-2) | 0.3 1.1 | 0.1 0.8 | 0.4 0.5 | 1.3 1.8 | 1.9 7.4 ∗ | 6.2 ∗ 46.9 ∗ | 9.7 ∗ 62.4 ∗ | 18.5 ∗ 62.4 ∗ | 10.7 ∗ 35.9 ∗ | 0.2 1.4 | 0/12 | 26/28 |
| Unigram entropy | 0.1 0.6 | 0.2 0.1 | 3.7 6.0 | 0.3 0.0 | 0.3 0.0 | 0.5 2.1 | 0.0 0.9 | 6.9 ∗ 22.1 ∗ | 7.5 ∗ 17.4 ∗ | 0.7 2.6 | 0/12 | 11/28 |
| Distributional metrics | ||||||||||||
| MAUVE (GPT-2) | 32.2 27.6 | 17.4 23.1 | 29.1 21.0 | 24.1 18.9 | 1.1 24.5 | 0.3 1.1 | 13.0 56.5 ∗ | 14.7 35.8 ∗ | 9.4 44.9 ∗ | 27.8 27.3 | 0/12 | 8/28 |
| gen. | MAUVE at truncation length | |||||
| Corpus | tokens | 128 | 256 | 384 | 512 | 1024 |
| Held-out human (packed) | 493 | 0.95 #1 | 0.94 #1 | 0.94 #1 | 0.95 #1 | 0.95 #1 |
| GPT-2-large (nucleus) | 483 | 0.70 #7 | 0.67 #7 | 0.58 #7 | 0.83 #5 | 0.78 #5 |
| GPT-2-medium (nucleus) | 486 | 0.74 #6 | 0.70 #6 | 0.62 #6 | 0.84 #4 | 0.81 #4 |
| SEDD | 493 | 0.91 #3 | 0.89 #3 | 0.88 #3 | 0.89 #2 | 0.89 #2 |
| MDLM | 493 | 0.92 #2 | 0.91 #2 | 0.90 #2 | 0.88 #3 | 0.85 #3 |
| CHORD at truncation length | |||||
| Corpus | 128 | 256 | 384 | 512 | 1024 |
| Held-out human (packed) | 0.5 #1 | 0.4 #1 | 0.6 #1 | 0.3 #1 | 0.0 #1 |
| GPT-2-large (nucleus) | 219 #2 | 361 #2 | 439 #2 | 365 #2 | 291 #2 |
| GPT-2-medium (nucleus) | 331 #3 | 507 #3 | 611 #3 | 522 #3 | 417 #3 |
| SEDD | 962 #7 | 1108 #6 | 1208 #6 | 1046 #5 | 843 #5 |
| MDLM | 954 #6 | 1117 #7 | 1209 #7 | 1060 #7 | 856 #6 |
| Continuation | CHORD | MAUVE | gen-PPL | entropy |
| Human | 0.1 #1 | 0.92 #2 | 27.7 #3 | 7.07 #4 |
| GPT-2-medium (nucleus) | 58.3 #2 | 0.80 #5 | 16.6 #2 | 6.81 #5 |
| GPT-2-medium (hot, ) | 145.3 #3 | 0.94 #1 | 68.4 #4 | 7.28 #1 |
| GPT-2-medium (greedy) | 216.9 #4 | 0.11 #6 | 2.9 #1 | 6.12 #6 |
| SEDD (128 steps) | 244.1 #5 | 0.88 #3 | 113.0 #5 | 7.09 #3 |
| MDLM (128 steps) | 291.1 #6 | 0.87 #4 | 155.3 #6 | 7.20 #2 |
| Annotator 2 | ||||
| Annotator 1 | Human better | Tie | Model better | Total |
| Human better | 40 | 2 | 3 | 45 |
| Tie | 3 | 2 | 1 | 6 |
| Model better | 2 | 0 | 7 | 9 |
| Total | 45 | 4 | 11 | 60 |
| Generator | Pairs | Human wins | Tie | Model wins | Disagree |
| GPT-2-large (nucleus) | 10 | 45.0% | 25.0% | 30.0% | 40% |
| GPT-2-medium (nucleus) | 10 | 50.0% | 15.0% | 35.0% | 20% |
| SEDD | 10 | 80.0% | 5.0% | 15.0% | 20% |
| MDLM | 10 | 100.0% | 0.0% | 0.0% | 0% |
| ELF-L | 10 | 80.0% | 0.0% | 20.0% | 20% |
| LangFlow | 10 | 95.0% | 5.0% | 0.0% | 10% |
| Featurize (GPU) | CPU | ||||||||||
| Metric | Extractor | Params | dim | GFLOPs/doc | batch | wall (s) | docs/s | b=1 (ms) | peak (GB) | weights (GB) | statistic (s) |
| CHORD (Qwen3.5-27B) | Qwen3.5-27B | 26.1B | 5120 | 26207 | 8 | 46.5 1.6 | 10.8 | 170.8 | 54.2 | 48.6 | 0.5 |
| CHORD (Qwen3.5-9B) | Qwen3.5-9B | 8.4B | 4096 | 8377 | 16 | 14.0 0.1 | 35.6 | 76.8 | 20.5 | 15.6 | 0.4 |
| CHORD (Qwen3.5-2B distilled) | Qwen3.5-2B + LoRA + | 2.2B | 256 | 2224 | 32 | 5.3 0.1 | 94.1 | 67.7 | 7.1 | 4.1 | 0.1 |
| CHORD (Qwen3.5-0.8B distilled) | Qwen3.5-0.8B + LoRA + | 853M | 256 | 857 | 32 | 4.4 0.0 | 112.5 | 57.9 | 4.5 | 1.6 | 0.1 |
| MAUVE (GPT-2) | GPT-2-large | 774M | 1280 | 789 | 32 | 2.1 0.4 | 241.4 | 17.4 | 5.1 | 1.4 | 0.6 † |
| Wrong fact | Faithful | |||
| Representation | number | entity | negation | paraphrase |
| CHORD (Qwen3.5-27B) | 258.7 | 112.0 | 380.5 | 0.4 |
| CHORD (Qwen3.5-2B distilled) | 216.4 | 99.0 | 504.2 | 0.4 |
| Qwen3.5-9B (layer ) | 164.6 | 55.6 | 237.2 | 0.4 |
| GPT-2-large (last token) | 4.9 | 1.0 | 5.6 | 0.7 |
| NeoBERT ( Breton et al., 2025 ) (masked mean) | 10.9 | 5.8 | 4.4 | 0.2 |
| Interesting | Makes sense | Human-like | |
| CHORD -27B | 0.86 (0.64) | 0.98 ( 0.93 ) | 0.98 ( 0.93 ) |
| CHORD -2B (distilled) | 0.76 (0.57) | 0.93 (0.86) | 0.95 (0.86) |
| MAUVE (GPT-2) | 0.00 (0.14) | 0.19 (0.00) | 0.21 (0.00) |
| MAUVE (ELECTRA) | 0.79 (0.64) | 0.93 (0.79) | 0.90 (0.79) |
| FBD (BERT) | 0.76 (0.57) | 0.93 (0.86) | 0.95 (0.86) |
| MMD (MiniLM) | 0.90 ( 0.79 ) | 0.90 (0.79) | 0.93 (0.79) |
| Role | Size | Content |
| Reference | 5,500 | Safe Alpaca-7B responses |
| Clean candidates | 1,000 | Safe Alpaca-7B responses before replacement |
| Matched donor pairs | 500 | Safe/unsafe siblings from one-safe rows |
| Generator-shift donors | 500 | Safe Alpaca3-8B responses |
| Metric / prompt | 1% | 2% | 5% | 10% | 25% | 50% | Detected |
| CHORD , safety prompt (w1) | 0.09 | 0.39 | 2.27 | 10.23 | 63.27 | 238.38 | 5% |
| CHORD , safety prompt (w2) | 0.13 | 0.59 | 2.23 | 9.62 | 54.27 | 204.53 | 5% |
| CHORD , safety prompt (w3) | 0.10 | 0.60 | 2.72 | 10.56 | 58.12 | 237.59 | 5% |
| CHORD , coherence prompt | 0.08 | 0.55 | 2.55 | 9.43 | 36.47 | 135.45 | 5% |
| CHORD , neutral prompt | 0.00 | 0.33 | 1.63 | 5.64 | 32.62 | 122.47 | 5% |
| Granite Guardian 3.1 2B | 0.31 | 0.87 | 2.26 | 5.15 | 13.94 | 28.28 | 10% |