cs.CLOct 2, 2026

Ontological Instability and Statistical Amplification: The Paradox of "Humanizing" LLM-Generated Text

Authors: Claudiu Creanga, Liviu Dinu

Organizations: Interdisciplinary School of Doctoral Studies University of Bucharest, Romania · Faculty of Mathematics and Informatics University of Bucharest, Romania

Abstract

Supervised AI-text detectors report high benchmark accuracy, but it is not clear what their decisions are based on. We analyze a RoBERTa-based detector under semantic, structural, and tokenizer-level perturbations, using the M4 dataset (N = 10,000) and controlled generations (N = 300). When Mistral-7B-Instruct was asked to make machine text sound more human, Verb Diversity rose from 0.77 to 0.92 and the outputs became easier to detect. Detection scores appear to track statistical complexity, which also leads to a 76.3% false-positive rate on formal human writing. As a control, we evaluate event-based Latent Space detection. Paraphrasing changed 87% of its event sequences (Jaccard = 0.067), and homoglyphs altered 70% of the extracted verbs even though extraction still ran (Jaccard = 0.30). Its best domain AUC was 0.577. RoBERTa's robustness seems specific to the features it uses, and structural abstraction did not make detection more robust.

Explore similar work

CardsList
  1. READER: Reasoning-Enhanced AI-Generated Text Detection

    May 24, 2026Pingfan Su, Kai Ye, Shijin Gong +4AI-Generated Text Detection

  2. ARB: A Matched Authorship-Rewriting Benchmark Dataset for AI-Text Detector Evaluation

    Jul 31, 2026Gaetano Perrone, Simon Pietro RomanoAI-Generated Text DetectionRobustness of Machine-Generated Text Detection

  3. Base Models Look Human To AI Detectors

    May 19, 2026Yixuan Even Xu, Ziqian Zhong, Aditi Raghunathan +2AI-Generated Text DetectionLanguage Model Generation Evaluation