cs.CLOct 8, 2026

Large Language Model Turnover Undermines Screening for Artificial Intelligence-Assisted Scientific Writing

Authors: Kazuki Nakajima, Takayuki Mizuno

Organizations: Tokyo Metropolitan University · National Institute of Informatics

Abstract

Journals and conferences have begun to screen submitted manuscripts for text written using large language models (LLMs). The reliability of this screening rests on benchmark evaluations against a fixed set of LLM versions, while the versions in actual use keep changing. Here we quantify how this LLM turnover affects the screening of scientific manuscripts. We paired 4,000 pre-ChatGPT abstracts from the Proceedings of the National Academy of Sciences with their rewrites by 23 LLM versions from three vendors, released between June 2023 and August 2026. We then trained detectors under maintenance scenarios ranging from a detector retrained on every new version to one trained once and never updated. Detectors trained only on a vendor's past versions can collapse at the boundaries between model generations: calibrated to falsely flag 1% of human-written abstracts, they catch above 99% of rewrites just before the sharpest boundary and 3.8% just after it. Detectors trained on later versions can also miss rewrites of earlier ones. Vocabulary differences between versions largely track where detection transfers and where it fails. In the two screening scenarios we simulated, screens covering all 23 versions either flagged one in eight human-written abstracts or missed one in three rewrites of the newest version. Indeed, a commercial detector missed most rewrites of the version just after the sharpest boundary while flagging almost no human-written abstracts. Research-integrity policy should therefore treat the benchmark accuracy of a detector as provisional, to be re-verified with every LLM release, including earlier versions.

Figures & tables

Explore similar work

CardsList
  1. Policies Permitting LLM Use for Polishing Peer Reviews Are Currently Not Enforceable

    Mar 20, 2026Rounak Saha, Gurusha Juneja, Dayita Chaudhuri +3LLM EvaluationAI-Generated Text Detection

  2. LLM-as-a-Reviewer: Benchmarking Their Ability, Divergence, and Prompt Injection Resistance as Paper Reviewers

    May 25, 2026Lingyao Li, Junjie Xiong, Changjia Zhu +5LLM EvaluationLLM-as-a-Judge

  3. ARB: A Matched Authorship-Rewriting Benchmark Dataset for AI-Text Detector Evaluation

    Jul 31, 2026Gaetano Perrone, Simon Pietro RomanoAI-Generated Text DetectionRobustness of Machine-Generated Text Detection