stat.MLSep 24, 2026

Robust Detection of LLM-Generated Text under Contamination

Authors: Jiaxun Li, Saptarshi Chakraborty, Ambuj Tewari

Organizations: Department of Statistics, University of Michigan

Abstract

We study the detection of LLM-generated text under editing and contamination. Modeling human and machine text as finite-order Markov processes with Huber contamination, we characterize an exact boundary for reliable detection under our assumptions. Detection is impossible when contamination is sufficiently large relative to clean-source separation. Below this boundary, a collection of clipped likelihood-ratio tests achieves vanishing worst-case errors. This construction motivates clipping as a simple modification of existing statistical detectors. For a broad class of additive scores, we identify conditions under which the clipped test is consistent while the raw test's worst-case power tends to zero. We evaluate seven detectors across three datasets and three generation models, and on the RAID benchmark. Clipping improves robustness in both studies, with gains varying across detectors and contamination settings. For example, at a target false-positive rate of 5%, clipping improves the log-likelihood--log-rank ratio (LRR) detector's true-positive rate by a median of 8.3 percentage points in the controlled study and 2.1 and 4.3 points in rate- and attack-specific RAID evaluations, respectively.

Figures & tables

Appendix figures & tables56 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. EVIL-Detect for NLPCC 2026 Shared Task 6: LLM-Generated Text Detection

    Aug 11, 2026Hongrui Bao, Hangyu Rong, Zhuoshang Wang +2Machine-Generated Text DetectionLarge Language Model Evaluation

  2. LLM Output Detectability and Task Performance Can be Jointly Optimized

    May 2, 2026Koshiro Saito, Ryuto Koike, Masahiro Kaneko +1Large Language Model WatermarksMachine-Generated Text Detection