Large Language Model Turnover Undermines Screening for Artificial Intelligence-Assisted Scientific Writing
Organizations: Tokyo Metropolitan University · National Institute of Informatics
Abstract
Journals and conferences have begun to screen submitted manuscripts for text written using large language models (LLMs). The reliability of this screening rests on benchmark evaluations against a fixed set of LLM versions, while the versions in actual use keep changing. Here we quantify how this LLM turnover affects the screening of scientific manuscripts. We paired 4,000 pre-ChatGPT abstracts from the Proceedings of the National Academy of Sciences with their rewrites by 23 LLM versions from three vendors, released between June 2023 and August 2026. We then trained detectors under maintenance scenarios ranging from a detector retrained on every new version to one trained once and never updated. Detectors trained only on a vendor's past versions can collapse at the boundaries between model generations: calibrated to falsely flag 1% of human-written abstracts, they catch above 99% of rewrites just before the sharpest boundary and 3.8% just after it. Detectors trained on later versions can also miss rewrites of earlier ones. Vocabulary differences between versions largely track where detection transfers and where it fails. In the two screening scenarios we simulated, screens covering all 23 versions either flagged one in eight human-written abstracts or missed one in three rewrites of the newest version. Indeed, a commercial detector missed most rewrites of the version just after the sharpest boundary while flagging almost no human-written abstracts. Research-integrity policy should therefore treat the benchmark accuracy of a detector as provisional, to be re-verified with every LLM release, including earlier versions.
Figures & tables
| Scenario | Definition | Real-world situation being simulated |
| Current version | Trained on rewrites from the target version itself. | An idealized detector built with full access to the exact version being screened. |
| Past and current versions | Trained on rewrites from the target version and all versions released before it. | A detector retrained immediately whenever a new version appears. |
| Past versions only | Trained on rewrites from the versions released before the target version, but not the target itself. | The common situation: a new version’s text arrives for screening before any training data for that version exist. |
| Frozen at a specific version | Trained on rewrites from one designated version only and never updated (in our experiments, the earliest version we study from the vendor). | A legacy detector built once and left in service without maintenance. |
| Union | One detector per version, each trained under the current version scenario; an abstract is flagged when at least one detector fires. | An institution that stacks every available specialized detector so as not to miss any version. |
| Pooled | A single classifier trained on rewrites pooled from all 23 versions. | A generalist detector trained on a broad corpus of AI text, a design common in academic and commercial tools. |
| Version | Identifier | Date used for ordering | Generation settings |
|---|---|---|---|
| GPT-4 | gpt-4-0613 | 2023-06-13 | Standard API |
| GPT-4 Turbo | gpt-4-turbo-2024-04-09 | 2024-04-09 | Standard API |
| GPT-4o (May 2024) | gpt-4o-2024-05-13 | 2024-05-13 | Standard API |
| GPT-4o (Aug 2024) | gpt-4o-2024-08-06 | 2024-08-06 | Standard API |
| GPT-4o (Nov 2024) | gpt-4o-2024-11-20 | 2024-11-20 | Standard API |
| GPT-4.1 | gpt-4.1-2025-04-14 | 2025-04-14 | Standard API |
| Version | Refusals | Truncated | Empty | Retained | Openings or closings removed |
|---|---|---|---|---|---|
| GPT-5 | 0 | 13 | 0 | 3,987 | 0 |
| Llama 3.1 | 0 | 2 | 0 | 3,998 | 222 |
| Llama 4 Maverick | 0 | 0 | 0 | 4,000 | 88 |
| Muse-Glimmer | 0 | 22 | 9 | 3,969 | 0 |
| Qwen3 | 0 | 0 | 0 | 4,000 | 35 |
| Qwen3.8 | 0 | 1 | 0 | 3,999 | 4 |
| Pattern | GPT-5 | Llama 4 Maverick |
|---|---|---|
| The rewrite contains more words than the original. | 21 | 7 |
| The rewrite uses the word ‘we’ less often than the original. | 7 | 28 |
| The rewrite uses the word ‘our’ more often than the original. | 8 | 21 |
| The rewrite uses the word family ‘show’, ‘shown’, and ‘showed’ less often than the original. | 14 | 15 |
| The rewrite uses the word ‘findings’ more often than the original. | 13 | 14 |
| The rewrite uses the word family ‘analyse’, ‘analyze’, and ‘analytical’ more often than the original. | 2 | 1 |
| Pattern | Definition | Examples |
|---|---|---|
| Undefined abbreviation | The rewrite uses an abbreviation without the full name that the original introduces it with. | The original introduces “mitochondrial superoxide dismutase (MnSOD)” with its full name, whereas the rewrite uses “MnSOD” throughout without the full name. |
| Changed notation | Terms, units, symbols, letter case, or hyphenation are changed while the meaning is kept. | • The original writes “ ” in symbols, whereas the rewrite spells it out as “pi to pi*”. • The original writes “200 mA/cm2”, whereas the rewrite writes “200 mA cm -2 ”. |
| Dropped details | Specific information of the original, such as numbers, method names, species, or examples, is missing. | • The everyday illustration that the original gives in parentheses does not reappear in the rewrite. • The method name “Amino acid-coded mass tagging (AACT)” of the original does not reappear in the rewrite. |
| Added or altered content | Procedures, parameters, mechanisms, or numbers that the original does not state are added as if factual, or the object of a statement is changed. | • The original attributes a finding to Eurasians as a whole, whereas the rewrite narrows it to “Europeans and East Asians”. • The rewrite describes a titration procedure that the original does not mention. |
| Added background | Textbook-style background or general statements are added. | The rewrite expands “the hippocampus” of the original into “the hippocampus, a brain region critical for learning and memory”. |
| Added significance | Significance, outlook, or implications absent from the original are added, typically at the end. | The rewrite adds the closing clause “thereby shedding new light on the complex history of human migration and settlement”. |
| LLM version | OpenAI detector | RADAR | Fast-DetectGPT | Binoculars |
|---|---|---|---|---|
| GPT-4 | 0.1 | 3.1 | 0.1 | 0.1 |
| GPT-4 Turbo | 0.1 | 4.0 | 1.3 | 2.6 |
| GPT-4o (May 2024) | 0.4 | 6.5 | 4.4 | 7.2 |
| GPT-4o (Aug 2024) | 0.4 | 6.8 | 2.9 | 5.0 |
| GPT-4o (Nov 2024) | 0.2 | 2.9 | 8.5 | 9.4 |
| GPT-4.1 | 0.1 | 0.6 | 7.7 | 4.5 |
| Two-stage prompt | First | Second | ||
| Main text | 500 abstracts | alternative | alternative | |
| Papers retained for all 23 versions | 3,955 | 495 | 498 | 494 |
| TPR of current-version detectors, median (range) over 23 versions | 99.9 (97.4–100.0) | 97.3 (74.3–100.0) | 88.1 (54.5–99.9) | 75.9 (36.6–99.3) |
| GPT-5: current version past versions only | 100.0 26.6 | 99.0 10.2 | 93.7 5.8 | 95.1 7.9 |
| Muse-Glimmer: current version past versions only | 97.4 3.8 | 84.6 2.6 | 77.1 15.3 | 75.3 7.1 |
| Spearman between JSD and averaged TPR | ||||
| LLM version | Union, per detector | Union, whole screen | Pooled |
|---|---|---|---|
| GPT-4 | 100.0 0.0 | 99.9 0.1 | 94.4 1.2 |
| GPT-4 Turbo | 100.0 0.0 | 100.0 0.0 | 100.0 0.1 |
| GPT-4o (May 2024) | 100.0 0.0 | 99.7 0.2 | 98.4 0.6 |
| Llama 3.1 | 100.0 0.1 | 99.6 0.2 | 98.4 0.4 |
| GPT-4o (Aug 2024) | 100.0 0.0 | 99.9 0.1 | 99.7 0.2 |
| Qwen2.5 | 99.9 0.1 | 97.7 0.5 | 91.2 1.2 |