cs.CLMay 26, 2026

Lost in Sampling: Assessing Lexical Reachability in LLMs via the Word Coverage Score (WCS)

Authors: Samer AwadJavier CondeCarlos ArriagaTairan FuJavier Coronado-BlázquezPedro Reviriego

Abstract

Modern Large Language Models (LLMs) are often criticized for producing repetitive and homogeneous text, despite possessing vast latent vocabularies. While previous research has focused on model knowledge and training data, we investigate the role of decoding mechanics in suppressing linguistic diversity. We introduce the Word Coverage Score (WCS), a metric that quantifies the extent to which contextually appropriate human vocabulary is mathematically pruned by standard sampling filters (e.g., Top-pp, Top-kk, and Min-pp). Rather than assessing static knowledge, the WCS measures the lexical survival rate of low-frequency, high-information human words as a function of sampling parameters. By auditing open-weight models on human-authored corpus fragments, we identify which logical lexical choices are rendered unreachable by the decoder, even when they reside within the probability space. Our results provide quantitative evidence that industry-standard sampling defaults act as unintended censorship mechanisms, smoothing the unique textures of human expression into a homogenized discourse. The WCS offers a rigorous framework for optimizing the trade-off between text coherence and lexical richness, providing a diagnostic tool for preserving the diversity of human language in generative models.

Explore similar work

Jun 21, 2026cs.CL

Breaking the Likelihood Trap: Variance-Calibrated Modulation for Large Language Model Decoding

In open-ended generation, LLMs frequently fall into the "likelihood trap", characterized by repetitive degeneration and vocabulary dullness, resulting in a discrepancy between machine-generated and human-written text. While post-hoc tail truncation (e.g., Top-p, Min-p) avoids sampling from the unreliable tail, it can misalign generation with human lexical preferences by over-sampling from the uncalibrated head; fixed scalar repetition penalties, in turn, ignore how the scale of the logit distribution varies across inference steps, which can disrupt semantic coherence. To address both shortcomings, we propose Variance-Calibrated Modulation (VCM), a training-free pre-decoding intervention. VCM directly reshapes the probability distribution prior to truncation via two dynamic mechanisms: (1) Contextual Searchlight via PMI, which naturally suppresses global stopwords and elevates context-evoked tokens, and (2) Adaptive Self-Debiasing, which utilizes real-time logit standard deviation to provide scale-invariant penalization. In experiments across open-ended generation, factual QA, and mathematical reasoning, we show that VCM consistently mitigates the likelihood trap. With negligible computational overhead, VCM integrates with existing decoding strategies, improving diversity and coherence and, particularly at higher decoding temperatures, reasoning accuracy. Our code is publicly available on GitHub: https://github.com/AetherDing/VCM
Yuanhao Ding, Meimingwei Li, Esteban Garces Arias +3
May 11, 2026cs.CL

Sampling More, Getting Less: Calibration is the Diversity Bottleneck in LLMs

Diversity is essential for language-model applications ranging from creative generation to scientific discovery, yet modern LLMs often collapse into a narrow subset of plausible outputs. While prior work has developed benchmarks for measuring this lack of diversity, less is known about how the step-by-step probability distributions at inference time cause the problem. We introduce a validity--diversity framework that attributes diversity collapse to how an LLM allocates probability mass across valid and invalid continuations during decoding. This framework decomposes the bottleneck into two complementary forms of miscalibration. First, order calibration: valid tokens are not reliably ranked above invalid tokens, so rank-based cutoff rules must trade off between recovering valid continuations and admitting invalid ones. Second, shape calibration: probability mass is overly concentrated only on few valid continuations while having a heavy-tail of mixed valid and invalid tokens, so maintaining high validity limits diversity. We formalize both mechanisms and show that local failures compound across decoding steps, producing strong sequence-level losses in diversity. Empirically, we develop controlled diagnostics for probing these bottlenecks, including tasks with exactly known valid sets and oracle cutoff baselines. Across 14 language models spanning multiple families and scales, we find that diversity collapse is not merely a limitation of particular sampling heuristics, but a consequence of order and shape miscalibration in the LLM distribution.
Amin Banayeeanzade, Qingchuan Yang, Dhruv Tarsadiya +6
Mar 19, 2026cs.CL

The Truncation Blind Spot: How Decoding Strategies Systematically Exclude Human-Like Token Choices

Standard decoding strategies for text generation, including top-kk, nucleus sampling, and contrastive search, select tokens based on likelihood, restricting outputs to high-probability regions. In contrast, human language production prioritizes communicative appropriateness, allowing the use of contextually suitable but statistically rare tokens. This mismatch induces a \emph{truncation blind spot}, whereby such tokens remain accessible to humans but are systematically excluded by likelihood-based decoding. We investigate this phenomenon using over 1.8 million machine-generated texts from eight language models, including large proprietary systems (GPT-3.5-turbo, Claude-3-Haiku), across five decoding strategies and 53 hyperparameter settings, alongside 5,261 human-written references. We find that 8--18% of human-selected tokens fall outside typical truncation boundaries. This exclusion is not random: content-bearing tokens are omitted at rates 2.9×2.9\times higher than grammatical function tokens. As a consequence, simple classifiers based on predictability and lexical diversity separate machine-generated from human-written text with mean AUC-ROC above 0.97. Detectability persists across model scales, architectures, and alignment procedures, and instead tracks the intensity of truncation. A classifier trained only on the oldest model in our study (GPT2-XL, 1.5B) detects outputs from substantially more recent and capable systems at near in-distribution accuracy, indicating that the detection signal is shared across generators rather than model-specific. These results indicate that detectability is a structural consequence of likelihood-based token selection rather than a limitation of model capability. We release code, datasets, and analysis at https://github.com/EstebanGarces/human_vs_machine
Esteban Garces Arias, Nurzhan Sapargali, Christian Heumann +1