LA-CPD: Local-Evidence-Aware Change-Point Detection for Human-LLM Authorship Segmentation
Organizations: Guilin University of Electronic Technology, China · Jinan Unaiversity, China · Jilin University, China · Nanjing University of Science and Technology, China
Abstract
As LLM-generated text becomes increasingly human-like, accurately localizing LLM-authored spans in human-LLM co-authored documents is important for attribution and accountability in cases involving copyright infringement, fraud, and other harmful uses of AI-generated content. Sentence-level detectors provide local authorship evidence, but content variation can cause score fluctuations even among sentences from the same source, creating spurious boundaries. Recovering a coherent document partition therefore remains challenging when both the number and locations of authorship transitions are unknown. We propose Local-Evidence-Aware Change-Point Detection (LA-CPD), a structured method that transforms noisy sentence-level score sequences into coherent authorship segments. Given scores from a frozen local detector, LA-CPD combines a length-weighted within-segment residual with a windowed two-mean contrast to capture segment consistency and sustained changes around candidate cut points. Dynamic programming optimizes cut locations for each candidate count, while an AIC-style criterion selects the final partition, yielding sentence labels, authorship boundaries, and maximal LLM-authored spans. On a held-out human-LLM co-authored test set, LA-CPD outperforms WCP+AIC, increasing sentence-level accuracy from 0.747 to 0.796 while improving boundary localization and LLM-span delineation.
Figures & tables
| Methods | Overall | By condition | ||||||
| Acc | WD | AI-F1 | H0 Acc | B1 Acc / AI-F1 | B2 Acc / AI-F1 | B3-2 Acc / AI-F1 | B3-3 Acc / AI-F1 | |
| SenPred | 0.598 | 0.874 | 0.086 | 0.404 | 0.690 / 0.086 | 0.696 / 0.091 | 0.527 / 0.072 | 0.673 / 0.135 |
| Voting | 0.711 | 0.581 | 0.317 | 0.466 | 0.836 / 0.419 | 0.829 / 0.402 | 0.642 / 0.208 | 0.784 / 0.434 |
| TextTiling | 0.724 | 0.413 | 0.506 | 0.534 | 0.839 / 0.726 | 0.805 / 0.667 | 0.681 / 0.334 | 0.760 / 0.567 |
| PaLD-scores | 0.713 | 0.880 | 0.037 | 0.753 | 0.707 / 0.041 | 0.700 / 0.036 | 0.730 / 0.034 | 0.672 / 0.048 |
| VCP (MAD) | 0.729 | 0.257 | 0.427 | 0.891 | 0.699 / 0.642 | 0.725 / 0.673 | 0.711 / 0.106 | 0.617 / 0.335 |
| Methods | Acc | WD | AI-F1 | B3-2 F1 |
|---|---|---|---|---|
| LA-CPD (Full) | 0.796 | 0.231 | 0.600 | 0.338 |
| LA-CPD w/o | 0.715 | 0.279 | 0.519 | 0.241 |
| LA-CPD w/o Length Weighting | 0.781 | 0.223 | 0.516 | 0.185 |
| LA-CPD w/o AIC | 0.789 | 0.250 | 0.605 | 0.325 |
| Metric | ||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | 3 | 8 | 0.25 | 1.5 | 0.75 | 2.0 | 1.0 | 2.0 | 3.0 | |||||
| Acc | 0.702 | 0.766 | 0.796 | 0.793 | 0.794 | 0.796 | 0.786 | 0.790 | 0.796 | 0.793 | 0.769 | 0.795 | 0.796 | 0.786 |
| AI-F1 | 0.494 | 0.561 | 0.600 | 0.567 | 0.592 | 0.600 | 0.575 | 0.577 | 0.600 | 0.588 | 0.614 | 0.607 | 0.600 | 0.544 |
| WD | 0.301 | 0.277 | 0.231 | 0.247 | 0.255 | 0.231 | 0.259 | 0.257 | 0.231 | 0.255 | 0.321 | 0.267 | 0.231 | 0.244 |
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
| Component | Setting |
|---|---|
| Scoring and reference backbone | Gemma-3-1B |
| LoRA target modules | q_proj , k_proj , v_proj , o_proj |
| LoRA rank | 4 |
| LoRA scaling parameter | 16 |
| LoRA dropout | 0.05 |
| Continued-adaptation training pairs | 4,000 |
| Operation | Time | Main storage |
|---|---|---|
| Prefix statistics | ||
| Interval residual cache | ||
| Windowed evidence | ||
| Shared dynamic programming | ||
| Candidate backtracking and residual accumulation | Does not change the overall bound | |
| Authorship delineation |
| Stratum | Relative position | Target proportion |
|---|---|---|
| Early | 25%–40% | Approximately one third |
| Middle | 40%–60% | Approximately one third |
| Late | 60%–75% | Approximately one third |
| H0 | B1 | B2 | B3-2 | B3-3 | Total |
| 12,000 | 20,000 | 19,999 | 10,998 | 6,000 | 68,997 |
| Split | Total | H0 | B1 | B2 | B3-2 | B3-3 |
|---|---|---|---|---|---|---|
| Training | 48,297 | 8,434 | 14,073 | 14,018 | 7,639 | 4,133 |
| Validation | 6,900 | 1,198 | 1,924 | 1,996 | 1,169 | 613 |
| Test | 13,798 | 2,368 | 4,001 | 3,985 | 2,190 | 1,254 |
| Total | 68,995 | 12,000 | 19,998 | 19,999 | 10,998 | 6,000 |
| Variant | Acc | WD | AI-F1 | B3-2 F1 |
|---|---|---|---|---|
| LA-CPD (Full) | 0.796 | 0.231 | 0.600 | 0.338 |
| LA-CPD w/o | 0.715 | 0.279 | 0.519 | 0.241 |
| LA-CPD w/o Length Weighting | 0.781 | 0.223 | 0.516 | 0.185 |
| LA-CPD w/o AIC | 0.789 | 0.250 | 0.605 | 0.325 |
| WCP+AIC (reference) | 0.747 | 0.270 | 0.552 | 0.280 |
| Metric | ||||
|---|---|---|---|---|
| Acc | 0.756 | 0.747 | 0.737 | 0.731 |
| H0 Acc | 0.852 | 0.840 | 0.829 | 0.810 |
| AI-F1 | 0.519 | 0.512 | 0.510 | 0.513 |
| B3-2 AI-F1 | 0.237 | 0.249 | 0.264 | 0.276 |