Organizations: Guilin University of Electronic Technology, China · Jinan Unaiversity, China · Jilin University, China · Nanjing University of Science and Technology, China
As LLM-generated text becomes increasingly human-like, accurately localizing LLM-authored spans in human-LLM co-authored documents is important for attribution and accountability in cases involving copyright infringement, fraud, and other harmful uses of AI-generated content. Sentence-level detectors provide local authorship evidence, but content variation can cause score fluctuations even among sentences from the same source, creating spurious boundaries. Recovering a coherent document partition therefore remains challenging when both the number and locations of authorship transitions are unknown. We propose Local-Evidence-Aware Change-Point Detection (LA-CPD), a structured method that transforms noisy sentence-level score sequences into coherent authorship segments. Given scores from a frozen local detector, LA-CPD combines a length-weighted within-segment residual with a windowed two-mean contrast to capture segment consistency and sustained changes around candidate cut points. Dynamic programming optimizes cut locations for each candidate count, while an AIC-style criterion selects the final partition, yielding sentence labels, authorship boundaries, and maximal LLM-authored spans. On a held-out human-LLM co-authored test set, LA-CPD outperforms WCP+AIC, increasing sentence-level accuracy from 0.747 to 0.796 while improving boundary localization and LLM-span delineation.
Figures & tables
Figure 1: LA-CPD training. Human–LLM sentence pairs are used to adapt a LoRA-augmented scoring model with a frozen reference model through a sampling-discrepancy separation loss. The adapted scorer is then frozen to produce sentence-level scores qi=Φ(ui) for structured segmentation.
Figure 2: LA-CPD inference. Top left: a frozen local detector produces ordered sentence-level scores, while sentence lengths determine the fitting weights. Top right: V(s,e) and E(c;h) define the candidate-partition objective. Bottom right: dynamic programming optimizes cut locations for each fixed cut count k . Bottom middle: an AIC-style criterion selects the final statistical cut set. Bottom left: two-cluster grouping of segment means delineates Human/LLM labels, authorship boundaries, and contiguous LLM-authored spans.
Methods
Overall
By condition
Acc ↑
WD ↓
AI-F1 ↑
H0 Acc ↑
B1 Acc / AI-F1 ↑
B2 Acc / AI-F1 ↑
B3-2 Acc / AI-F1 ↑
B3-3 Acc / AI-F1 ↑
SenPred
0.598
0.874
0.086
0.404
0.690 / 0.086
0.696 / 0.091
0.527 / 0.072
0.673 / 0.135
Voting
0.711
0.581
0.317
0.466
0.836 / 0.419
0.829 / 0.402
0.642 / 0.208
0.784 / 0.434
TextTiling
0.724
0.413
0.506
0.534
0.839 / 0.726
0.805 / 0.667
0.681 / 0.334
0.760 / 0.567
PaLD-scores
0.713
0.880
0.037
0.753
0.707 / 0.041
0.700 / 0.036
0.730 / 0.034
0.672 / 0.048
VCP (MAD)
0.729
0.257
0.427
0.891
0.699 / 0.642
0.725 / 0.673
0.711 / 0.106
0.617 / 0.335
Table 1: Performance comparison on the GPT-5.5 human–LLM co-authored test set. H0 reports accuracy, whereas mixed conditions report accuracy and AI-F1. Best and second-best results among methods with unknown boundary counts are shown in bold and underlined , respectively. Oracle- K , which uses the ground-truth boundary count, is reported only for diagnostic purposes and is excluded from ranking.
Figure 3: Cross-model generalization performance of LA-CPD in terms of sentence-level accuracy across different generation models.
Methods
Acc ↑
WD ↓
AI-F1 ↑
B3-2 F1 ↑
LA-CPD (Full)
0.796
0.231
0.600
0.338
LA-CPD w/o E(c;h)
0.715
0.279
0.519
0.241
LA-CPD w/o Length Weighting
0.781
0.223
0.516
0.185
LA-CPD w/o AIC
0.789
0.250
0.605
0.325
Table 2: Ablation study of LA-CPD. Best and second-best results are shown in bold and underlined , respectively.
Metric
h
α
β
r
1
3
6†
8
0.25
0.75†
1.5
0.75
1.25†
2.0
1.0
2.0
2.5†
3.0
Acc ↑
0.702
0.766
0.796
0.793
0.794
0.796
0.786
0.790
0.796
0.793
0.769
0.795
0.796
0.786
AI-F1 ↑
0.494
0.561
0.600
0.567
0.592
0.600
0.575
0.577
0.600
0.588
0.614
0.607
0.600
0.544
WD ↓
0.301
0.277
0.231
0.247
0.255
0.231
0.259
0.257
0.231
0.255
0.321
0.267
0.231
0.244
Table 3: Hyperparameter sensitivity of LA-CPD on the GPT-5.5 human–LLM co-authored dataset. † denotes the default configuration. Best and second-best results within each parameter group are shown in bold and underlined , respectively.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Component
Setting
Scoring and reference backbone
Gemma-3-1B
LoRA target modules
q_proj , k_proj , v_proj , o_proj
LoRA rank
4
LoRA scaling parameter
16
LoRA dropout
0.05
Continued-adaptation training pairs
4,000
Appendix
Table 4: Configuration of the continued local-detector adaptation.
Operation
Time
Main storage
Prefix statistics
O(n)
O(n)
Interval residual cache
O(n2)
O(n2)
Windowed evidence
O(nh)
O(n)
Shared dynamic programming
O((Keff+1)n2)
O((Keff+1)n)
Candidate backtracking and residual accumulation
O((Keff+1)2)
Does not change the overall bound
Authorship delineation
O(nlogn+m2)
O(n)
Appendix
Table 5: Time and storage costs of structured segmentation, excluding detector forward passes.
C←{i∈{1,…,n−1}:yi=yi+1};
Appendix
Algorithm 1 Full LA-CPD Inference
Stratum
Relative position
Target proportion
Early
25%–40%
Approximately one third
Middle
40%–60%
Approximately one third
Late
60%–75%
Approximately one third
Appendix
Table 6: Target boundary-position strata during corpus construction.
H0
B1
B2
B3-2
B3-3
Total
12,000
20,000
19,999
10,998
6,000
68,997
Appendix
Table 7: Condition counts in the cleaned candidate pool, before removing the two B1 records with zero-length AI spans.
Split
Total
H0
B1
B2
B3-2
B3-3
Training
48,297
8,434
14,073
14,018
7,639
4,133
Validation
6,900
1,198
1,924
1,996
1,169
613
Test
13,798
2,368
4,001
3,985
2,190
1,254
Total
68,995
12,000
19,998
19,999
10,998
6,000
Appendix
Table 8: Statistics of the complete GPT-5.5 corpus after removing the two invalid records and grouping examples by source family.
Variant
Acc ↑
WD ↓
AI-F1 ↑
B3-2 F1 ↑
LA-CPD (Full)
0.796
0.231
0.600
0.338
LA-CPD w/o E(c;h)
0.715
0.279
0.519
0.241
LA-CPD w/o Length Weighting
0.781
0.223
0.516
0.185
LA-CPD w/o AIC
0.789
0.250
0.605
0.325
WCP+AIC (reference)
0.747
0.270
0.552
0.280
Appendix
Table 9: Complete component ablation results on the held-out GPT-5.5 test subset of 1,000 documents. Best and second-best results among the reported methods are shown in bold and underlined , respectively.
Figure 4: Complete cross-distribution generalization results on Hybrid Essays, RoFT, Claude–SQuAD, Claude–WritingPrompts, and Claude–XSum. The top and bottom panels report sentence-level accuracy and AI-fragment F1, respectively. Bars are grouped by method, and colors identify the evaluation datasets. Missing results are left blank. Higher Acc and AI-F1 indicate better performance. LA-CPD oracle- K uses the ground-truth boundary count and is included only as a diagnostic setting.
Metric
ρ=1.00
ρ=0.75
ρ=0.50
ρ=0.35
Acc ↑
0.756
0.747
0.737
0.731
H0 Acc ↑
0.852
0.840
0.829
0.810
AI-F1 ↑
0.519
0.512
0.510
0.513
B3-2 AI-F1 ↑
0.237
0.249
0.264
0.276
Appendix
Table 10: Sensitivity to the regime-return discount ρ under the MAD-based complexity penalty. Other parameters are fixed to h=3 , α=0.5 , β=1 , and c=2 .
The rise of large language models (LLMs) has created an urgent need to distinguish between human-written and LLM-generated text to ensure authenticity and societal trust. Existing detectors typically provide a binary classification for an entire passage; however, this is insufficient for human--LLM co-authored text, where the objective is to localize specific segments authored by humans or LLMs. To bridge this gap, we propose algorithms to segment text into human- and LLM-authored pieces. Our key observation is that such a segmentation task is conceptually similar to classical change point detection in time-series analysis. Leveraging this analogy, we adapt change point detection to LLM-generated text detection, develop a weighted algorithm and a generalized algorithm to accommodate heterogeneous detection score variability, and establish the minimax optimality of our procedure. Empirically, we demonstrate the strong performance of our approach against a wide range of existing baselines. The python implementation of our proposal is available at https://github.com/Mamba413/DetectLLMSegmentation.
Mengchu Li, Jin Zhu, Jinglai Li +1
School of Mathematics University of Birmingham Birmingham, UK · Department of Statistics London School of Economics and Political Science London, UK
The rise of human-AI collaborative writing has created a growing need for fine-grained detection methods that support localizing likely LLM-generated content in mixed-authorship documents. Existing methods for detecting LLM-generated text mainly focus on document-level classification and cannot identify which parts of the text are generated by LLMs. This paper introduces a new method to address this urgent need. Our method operates at the token level, the natural unit of modern language models, and builds on existing token-level detection scores. The key idea is to smooth adjacent token scores to reduce their variability, while using an adaptive Lepski-type rule to select the bandwidth according to the local authorship structure. Our method is simple to implement and does not require token-level labeled data for training. Theoretically, we characterize this trade-off and show that the proposed method achieves favorable mean square error performance in estimating the underlying signal. Empirically, we demonstrate strong performance of our method against a wide range of baselines in both synthetic datasets and a realistic dataset. We deploy a publicly accessible website that implements the methods as well.
Yangjun Lu, Hongyi Zhou, Fabian Spill +3
School of Mathematics, University of Birmingham, Birmingham, UK · School of Statistics and Data Science, Shanghai University of Finance and Economics, Shanghai, China · Department of Statistics, The London School of Economics and Political Science, London, UK
Mixed human-LLM documents require locating authorship transitions from detector scores whose reliability varies across text units. Existing weighted mean contrasts are vulnerable to extreme scores, while directly replacing means with robust centers obscures how a misplaced boundary changes the population objective. We propose Robust Weighted Profile-Loss Change Point Detection (RWCP), which combines capped reliability weights, Huber profile gains, and narrowest-over-threshold search in reliability coordinates. Our key analysis expresses the population gap between a true and a displaced split as a merge cost, avoiding a closed-form solution for the nonlinear center of a mixed segment. Under explicit curvature, spacing, and dependence conditions, core RWCP recovers the number of changes and localizes their boundaries; its quadratic-loss limit recovers squared weighted CUSUM. We also study RWCP-R, a separately evaluated decoder that shares source centers across nonadjacent passages. Across five retrospective cached-score benchmark families, core RWCP reduces family-macro WindowDiff by 17.6% relative to weighted change-point detection, and RWCP-R lowers it further. Boundary recovery improves most clearly for isolated changes, while both fixed configurations miss changes in collaborative and densely alternating text.
Wan Tian, Zhongyi Li, Yawen Li +3
Peking University · Beihang University · Capital Normal University +2