cs.CLMar 19, 2026

The Truncation Blind Spot: How Decoding Strategies Systematically Exclude Human-Like Token Choices

Authors: Esteban Garces AriasNurzhan SapargaliChristian HeumannMatthias Aßenmacher

Organizations: Department of Statistics Ludwig Maximilian University Munich Munich Center for Machine Learning (MCML) Munich, Germany · Department of Statistics Ludwig Maximilian University Munich Munich, Germany

Abstract

Standard decoding strategies for text generation, including top-kk, nucleus sampling, and contrastive search, select tokens based on likelihood, restricting outputs to high-probability regions. In contrast, human language production prioritizes communicative appropriateness, allowing the use of contextually suitable but statistically rare tokens. This mismatch induces a \emph{truncation blind spot}, whereby such tokens remain accessible to humans but are systematically excluded by likelihood-based decoding. We investigate this phenomenon using over 1.8 million machine-generated texts from eight language models, including large proprietary systems (GPT-3.5-turbo, Claude-3-Haiku), across five decoding strategies and 53 hyperparameter settings, alongside 5,261 human-written references. We find that 8--18% of human-selected tokens fall outside typical truncation boundaries. This exclusion is not random: content-bearing tokens are omitted at rates 2.9×2.9\times higher than grammatical function tokens. As a consequence, simple classifiers based on predictability and lexical diversity separate machine-generated from human-written text with mean AUC-ROC above 0.97. Detectability persists across model scales, architectures, and alignment procedures, and instead tracks the intensity of truncation. A classifier trained only on the oldest model in our study (GPT2-XL, 1.5B) detects outputs from substantially more recent and capable systems at near in-distribution accuracy, indicating that the detection signal is shared across generators rather than model-specific. These results indicate that detectability is a structural consequence of likelihood-based token selection rather than a limitation of model capability. We release code, datasets, and analysis at https://github.com/EstebanGarces/human_vs_machine

Explore similar work

Apr 17, 2026cs.CL

Spotlights and Blindspots: Evaluating Machine-Generated Text Detection

With the rise of generative language models, machine-generated text detection has become a critical challenge. A wide variety of models is available, but inconsistent datasets, evaluation metrics, and assessment strategies obscure comparisons of model effectiveness. To address this, we evaluate 15 different detection models from six distinct systems, as well as seven trained models, across seven English-language textual test sets and three creative human-written datasets. We provide an empirical analysis of model performance, the influence of training and evaluation data, and the impact of key metrics. We find that no single system excels in all areas and nearly all are effective for certain tasks, and the representation of model performance is critically linked to dataset and metric choices. We find high variance in model ranks based on datasets and metrics, and overall poor performance on novel human-written texts in high-risk domains. Across datasets and metrics, we find that methodological choices that are often assumed or overlooked are essential for clearly and accurately reflecting model performance.
Kevin Stowe, Kailash Patil
May 7, 2026cs.CL

Log-Likelihood, Simpson's Paradox, and the Detection of Machine-Generated Text

The ability to reliably distinguish human-written text from that generated by large language models is of profound societal importance. The dominant approach to this problem exploits the likelihood hypothesis: that machine-generated text should appear more probable to a detector language model than human-written text. However, we demonstrate that the token-level signal distinguishing human and machine text is non-uniform across the hidden space of the detector model, and naively averaging likelihood-based token scores across regions with fundamentally different statistical structure, as most detectors do, causes a form of Simpson's paradox: a strong local signal is destroyed by inappropriate aggregation. To correct for this, we introduce a learned local calibration step grounded in Bayesian decision theory. Rather than aggregating raw token scores, we first learn lightweight predictors of the score distributions conditioned on position in hidden space, and aggregate calibrated log-likelihood ratios instead. This single intervention dramatically and consistently improves detection performance across all baseline detectors and all datasets we consider. For example, our calibrated variant of Fast-DetectGPT improves AUROC from 0.630.63 to 0.850.85 on GPT-5.4 text, and a locally-calibrated DMAP detector we introduce achieves state-of-the-art performance across the board. That said, our central contribution is not a new detector, but a precise diagnosis of a significant cause of under-performance of existing detectors and a principled, modular remedy compatible with any token-averaging pipeline. This will serve as a foundation for the community to build upon, with natural avenues including richer distributional models, improved calibration strategies, and principled ensembling with hidden-space geometry signals via the full Bayes-optimal decision rule.
Tom Kempton, Viktor Drobnyi, Maeve Madigan +1
Jul 5, 2026cs.CL

Telescope: Improving Zero Shot Detection of LLM Generated Content By Measuring Token Repetition Probability

Distinguishing Large Language Model (LLM) generated text from human writing is a critical and difficult challenge. While LLMs are trained to write like humans, we hypothesize that this training leaves an indelible mark. LLMs develop a particularly strong aversion to token repetition very early in training. This bias persists as a ''Vestigial Heuristic'' (a developmental artifact) that is activated in LLM-generated text, separating LLM from human writing. To probe this phenomenon, we introduce Telescope Perplexity, a metric that evaluates the token repetition of the model, P(sis1:i)P(s_i | s_{1:i}) . Our empirical investigation reveals that the Telescope Perplexity signature emerges early in pre-training, and Telescope Perplexity empirically enables highly effective zero-shot LLM detection. We show state-of-the-art or competitive performance across diverse datasets (including modern evaluation sets we introduce), reference models, and perturbation schemes with greater efficiency than other methods.
Christopher Nassif, Josh F. Cooper