Using LLMs to Detect LLM-Generated Texts: A Cross-Generation Analysis
Organizations: Institute of Cyber Security for Society (iCSS) & School of Computing, University of Kent, UK · School of Cyber Science and Engineering, Shanghai Jiao Tong University, China · School of Computer Science, University of Birmingham, UK
Abstract
Automated detection of LLM-generated texts (LGTs) is critical, yet dedicated detectors often struggle to generalize across domains and models. While general-purpose LLMs offer flexible zero-shot authorship classification with explanatory rationale, their detection behavior, especially regarding self-detection versus cross-detection across model generations, remains poorly understood. We systematically evaluate 15 LLMs spanning three model generations as both generators and detectors. Using a benchmark of 1,000 human-written texts and 15,000 LGTs (1,000 per model), we collected over 233,000 binary classifications alongside natural-language explanations. Our results reveal that detection efficacy is primarily driven by detector capability rather than generator provenance, although outputs from newer generators remain notably harder to detect. Crucially, statistical comparisons show no systematic advantage or disadvantage for self-detection across models. Error analysis further exposes generational bias shifts: first-generation detectors under-detect LGTs (high false-negative rates), second-generation detectors over-flag human texts (high false-positive rates), and the latest models achieve balanced trade-offs. Finally, we highlight significant inconsistencies in how different LLMs apply textual cues to justify their decisions. Code: https://github.com/hyyuan/detect-llm-generated-texts.
Figures & tables
| Generation | Selected models |
|---|---|
| G1 | GPT-3.5 Turbo Instruct ( OpenAI, 2023 ) ; Mixtral 8 22B Instruct ( Mistral AI, 2024 ) . |
| G2 | GPT-4o ( OpenAI, 2024 ) ; Qwen2.5-72B Instruct ( Qwen Team, 2024 ) ; DeepSeek-V3 ( DeepSeek-AI, 2024 ) ; Llama 3.3 70B Instruct ( Meta AI, 2024 ) ; GLM-4 32B 0414 ( Z.AI, 2025 ) . |
| G3 | DeepSeek-V4-Pro ( DeepSeek-AI, 2026 ) ; Qwen3.6-Max-Preview ( Qwen Team, 2026 ) ; GLM-5.1 ( Z.AI, 2026 ) ; MiMo-V2.5-Pro ( Xiaomi MiMo Team, 2026 ) ; GPT-5.5 ( OpenAI, 2026 ) ; Claude Opus 4.7 ( Anthropic, 2026b ) ; Gemini 3.1 Pro Preview ( Google, 2026 ) ; Kimi K2.6 ( Moonshot AI, 2026 ) . |
| Generation | Generator | Metric | GPT-3.5 | Mixtral | DS-V3 | GLM-4 | GPT-4o | Llama-3.3 | Qwen2 | Claude | DS-V4 | Gemini-3.1 | GLM-5.1 | GPT-5.5 | Kimi-K2.6 | MiMo-V2.5 | Qwen3.6 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| G1 | GPT-3.5 | F1 | 3.0 | 42.6 | 62.7 | 59.1 | 80.6 | 77.1 | 1.7 | 97.0 | 93.1 | 99.9 | 84.7 | 99.0 | 99.3 | 91.9 | 99.0 |
| Accuracy | 47.4 | 56.5 | 60.6 | 61.2 | 76.6 | 71.1 | 48.3 | 96.9 | 93.0 | 99.9 | 86.5 | 99.0 | 99.3 | 91.9 | 99.0 | ||
| Mixtral | F1 | 2.3 | 38.6 | 62.0 | 51.3 | 75.4 | 70.0 | 0.6 | 96.8 | 82.9 | 99.7 | 80.2 | 98.9 | 99.2 | 86.1 | 99.1 | |
| Accuracy | 47.3 | 54.7 | 60.1 | 56.2 | 71.6 | 64.2 | 47.9 | 96.7 | 84.1 | 99.7 | 83.2 | 98.9 | 99.2 | 86.8 | 99.1 | ||
| G2 | DS-V3 | F1 | 2.4 | 39.5 | 51.9 | 53.8 | 76.5 | 73.0 | 1 | 97.0 | 88.1 | 99.8 | 84.5 | 99.0 | 99.2 | 88.3 | 99.2 |
| Accuracy | 47.4 | 55.0 | 52.9 | 57.8 | 72.6 | 67.0 | 48.0 | 96.9 | 88.4 | 99.8 | 86.4 | 98.9 | 99.2 | 88.7 | 99.2 |
| Generation | Detector | F1 (pp) | Accuracy (pp) |
|---|---|---|---|
| G1 | GPT-3.5 Turbo Instruct | ||
| Mixtral 8x22B | |||
| G2 | DeepSeek-V3 | ||
| GLM-4-32B | |||
| GPT-4o | |||
| Llama 3.3 70B |
| FNR by Generator Generation | FPR on HGTs | |||
| Detector Generation | G1 | G2 | G3 | Shared Corpus |
| G1 | 84.1 | 83.3 | 81.9 | 13.0 |
| G2 | 39.9 | 39.6 | 48.1 | 36.5 |
| G3 | 7.2 | 6.1 | 13.4 | 3.6 |
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
| Group | Model | Default temperature | Default top- | Matches |
| G1 | GPT-3.5 Turbo Instruct | 1.0 | 1.0 | |
| Mixtral 8 22B Instruct | 0.3 | 1.0 | ||
| G2 | GPT-4o | 1.0 | 1.0 | |
| Qwen2.5-72B Instruct | 1.0 | 1.0 | ||
| DeepSeek-V3 | 1.0 | 1.0 | ||
| Llama 3.3 70B Instruct | 1.0 | 1.0 |
| Panel A | Panel B | ||||||||
| F1-score | Accuracy | F1-score | Accuracy | ||||||
| Detector | Mean | SD | Mean | SD | Generator | Mean | SD | Mean | SD |
| GPT-3.5 | 4.8 | 3.9 | 48.0 | 1.1 | GPT-3.5 | 72.7 | 33.3 | 79.2 | 19.9 |
| Mixtral | 41.9 | 9.0 | 56.4 | 4.6 | Mixtral | 69.5 | 33.3 | 76.6 | 20.2 |
| DeepSeek-V3 | 56.1 | 10.8 | 56.2 | 7.1 | DeepSeek-V3 | 70.2 | 33.7 | 77.2 | 20.7 |
| GLM-4 | 53.5 | 6.3 | 57.8 | 4.1 | GLM-4 | 69.7 | 33.3 | 76.8 | 20.3 |
| Accuracy | F1-score | ||||||
| Generation | Detector | (pp) | 95% CI (pp) | (pp) | 95% CI (pp) | ||
| G1 | GPT-3.5 Turbo Instruct | 0.0632 | 0.0686 | ||||
| G1 | Mixtral 8x22B | 0.0632 | 0.1338 | ||||
| G2 | DeepSeek-V3 | 0.0015 | 0.0018 | ||||
| G2 | GLM-4-32B | 1.0000 | 1.0000 | ||||
| G2 | GPT-4o | 0.0015 | 0.0015 | ||||
| Accuracy | F1-score | ||||||
| Generation | Detector | (pp) | 95% CI (pp) | (pp) | 95% CI (pp) | ||
| G1 | GPT-3.5 Turbo Instruct | 1.0000 | 1.0000 | ||||
| G1 | Mixtral 8x22B | 0.0999 | 0.1863 | ||||
| G2 | DeepSeek-V3 | 0.0015 | 0.0015 | ||||
| G2 | GLM-4-32B | 1.0000 | 1.0000 | ||||
| G2 | GPT-4o | 0.7551 | 0.3856 | ||||
| Panel A: Detector agreement by text type | ||||
|---|---|---|---|---|
| Text type | Texts | Disagreement | Pairwise comparisons | Label agreement % |
| HGT | 20 | 8 | 300 | 83.7 |
| LGT | 120 | 120 | 1,800 | 48.4 |
| Panel B: Classification outcomes by detector | ||||
| Detector | TP | TN | FP | FN |
| GPT-3.5 Turbo Instruct (G1) | 0 | 20 | 0 | 120 |
| Generator | Prompt ID | Shared text excerpt | Detector 1: judgement and explanation | Detector 2: judgement and explanation | Observed difference |
|---|---|---|---|---|---|
| GPT-3.5 | 40 | “Unreliable narration is a literary technique where the narrator’s credibility is called into question.” | GPT-3.5: H (FN). Detailed literary analysis is described as beyond an LLM’s capability. | GPT-5.5: M (TP). Generic organisation, repeated claims, standard transitions, and the lack of concrete examples support the machine label. | The detectors draw opposite conclusions about the depth of the same text. |
| Human | 43 | Factual history of Brentford Football Club. | GPT-3.5: H (TN). Specific records, statistics, and recent updates are treated as human evidence. | GPT-5.5: H (TN). Article cross-references and an incomplete phrase are treated as signs of a copied human-edited source. | The detectors agree on the label but rely on different reasons. |
| GPT-3.5 | 113 | “Medical malpractice occurs when a healthcare provider fails to follow the accepted standard of care.” | DeepSeek-V3: H (FN). Specific examples, coherence, and a detailed approach are treated as human evidence. | DeepSeek-V4-Pro: M (TP). Predictable transitions, generic coverage, and uniform tone are treated as LGT cues. | The same organisation and detail support opposite labels and reasons. |
| Mixtral | 146 | “I’ve always found that the true essence of a city lies beyond its well-trodden tourist paths.” | Mixtral: H (FN). Sensory details and personal recollection are interpreted as lived experience. | GPT-5.5: M (TP). Familiar travel phrases and a decorative anecdote are interpreted as generated personalisation. | The same personal voice is treated as genuine or simulated. |
| Human | 425 | “Indian Mary Park is part of the Josephine County Parks system.” | DeepSeek-V4-Pro: M (FP). Missing distances are interpreted as model generation errors. | GPT-5.5: H (TN). The same omissions are interpreted as extraction or formatting artefacts. | The same textual defect supports opposite provenance inferences. |
| Detector | Prompt ID and topic | Generators and output excerpts | Output 1: judgement and explanation | Output 2: judgement and explanation | Observed difference |
|---|---|---|---|---|---|
| GPT-4o | 43: academic pressure | GPT-3.5: “Hey there! How have you been dealing with all the academic pressure lately?” GPT-4o: “Alice: Hey Ben, you’ve been looking a bit stressed lately.” | GPT-3.5 output: H (FN). Empathy, emotional support, and conversational markers are treated as human evidence. | GPT-4o output: M (TP). Sequential advice and a deliberate narrative arc are treated as machine evidence. | Different dialogues lead the same detector to opposite labels and reasons. |
| DeepSeek-V3 | 40: unreliable narration | GPT-3.5: “Unreliable narration is a literary technique where the narrator’s credibility is called into question.” Mixtral: “Unreliable narration is a narrative technique where the credibility of a narrator is compromised.” | GPT-3.5 output: M (TP). Uniform structure, systematic coverage, and repetition support the machine label. | Mixtral output: H (FN). Nuanced analysis and integrated literary examples support the human label. | The judgement changes with the depth and form of the generated text. |
| DeepSeek-V4-Pro | 34: climate-responsive architecture | GPT-3.5: “Climate change is a global issue that affects every aspect of our lives, including the design of our buildings.” GPT-5.5: “Future architectural designs are likely to become more climate-responsive, flexible, and resource-efficient.” | GPT-3.5 output: M (TP). Formulaic organisation, generic language, and predictable progression support the machine label. | GPT-5.5 output: H (FN). Progressive examples and the absence of awkward phrasing support the human label. | Outputs from different generator generations lead to opposite judgements. |
| Cue category | Interpretation in detector explanations | Supporting literature |
|---|---|---|
| Structure and formatting | Organisation, paragraph structure, formatting, or ordering of ideas | Previous studies found that human evaluators use formatting, textual structure, and sentence organisation to make a judgment ( Clark et al., 2021 ; Russell et al., 2025 ) . |
| Fluency and polish | Grammar, spelling, fluency, formality, smoothness, or tonal consistency | Grammar, fluency, formality, and clarity are commonly used as detection cues, although they may support either a human or machine judgement ( Clark et al., 2021 ; Mitrović et al., 2023 ; Russell et al., 2025 ) . |
| Generic or abstract language | Broad, generic, shallow, or insufficiently original content | LGTs are often perceived as focusing on general concepts and lacking detail or originality ( Mitrović et al., 2023 ; Russell et al., 2025 ) . |
| Hedging and uncertainty | Qualified claims and expressions of uncertainty or epistemic caution | Human evaluators sometimes use expressions of uncertainty to make a judgment, although uncertainty is not a consistently reliable detection cue ( Clark et al., 2021 ) . |
| Personal voice | Personal pronouns, feelings, anecdotes, lived experience, or individual perspective | Personal expression and emotional language are often associated with human writing and may cause LGTs to be misclassified as human-written ( Clark et al., 2021 ; Mitrović et al., 2023 ; Russell et al., 2025 ) . |
| Coherence and transitions | Logical flow, consistency, repetition, cohesion, or formulaic connections | Previous studies found that evaluators frequently consider coherence, repetition, and sentence structure when making a judgment ( Clark et al., 2021 ; Russell et al., 2025 ) . |
| Cue category | Predefined keywords and phrases |
|---|---|
| Structure and formatting | structured ; format ; paragraph ; introduction ; conclusion ; essay ; organization |
| Fluency and polish | polished ; fluent ; smooth ; grammatically ; well-formed ; well formed ; even tone |
| Generic or abstract language | generic ; abstract ; high level ; encyclopedic ; surface-level ; surface level ; broad |
| Hedging and uncertainty | hedge ; hedged ; hedging ; in most cases ; tends to ; may ; might ; depending on |
| Personal voice | personal voice ; idiosyncr* ; anecdote ; lived experience ; individual perspective |
| Coherence and transitions | transition ; coheren* ; flow ; parallel ; balanced ; systematic* |
| Detector | Label ( ) | SF | GAL | CT | SE |
| GPT-3.5 Turbo Instruct | FN (120) | 10.0 | 2.5 | 20.0 | 41.7 |
| Mixtral 8x22B | TP (30) | 93.3 | 73.3 | 83.3 | 40.0 |
| FN (90) | 61.1 | 46.7 | 85.6 | 65.6 | |
| GPT-4o | TP (93) | 91.4 | 20.4 | 77.4 | 34.4 |
| FN (27) | 33.3 | 11.1 | 66.7 | 48.1 | |
| DeepSeek-V3 | TP (70) | 90.0 | 57.1 | 78.6 | 32.9 |
| Generation | Model | Public evidence for text watermarking on the evaluated endpoint | Acc (pp). |
| G1 | GPT-3.5 Turbo Instruct | Text watermark researched by provider; no deployment confirmed | |
| Mixtral 8 22B Instruct | No public model- and endpoint-specific claim identified | ||
| G2 | DeepSeek-V3 | No public model- and endpoint-specific claim identified | |
| GLM-4 32B 0414 | No public model- and endpoint-specific claim identified | ||
| GPT-4o | Text watermark researched by provider; no deployment confirmed | ||
| Llama 3.3 70B Instruct | No public model- and endpoint-specific claim identified |