Repeated training on model-generated data can degrade later models. One possible response is to use provenance when deciding which generated examples to reuse. We test both how reliably that provenance can be recovered and whether it helps identify better training data. Using financial-risk text, we first identify the source of generated passages and then repeat the test after rewriting them. Generator attribution is 98.7% accurate on the original passages but falls to 53.1% after paraphrasing and 29.0% after style rewriting. Generated-versus-human detection remains close to perfect against the tested human comparison set. We then compare two ways of selecting generated examples over three rounds of generation and retraining. One uses source information. The other uses a score from a separate reference model. The two rules select different examples, but the planned comparison does not detect a stable difference in the degradation of the resulting models. The results show that identifying where data came from and identifying which data are useful for training are separate problems. The experiment therefore separates source identity, criterion-facing selection, and recursive training outcome: neither the provenance score nor the tested criterion-facing proxy is established as sufficient for future recursive behaviour.
Figures & tables
Fig. 1: Generator-identification accuracy before and after four forms of rewriting. Values are means over seven source models. The word-TFIDF classifier remains strong after roundtrip translation but falls after paraphrasing, style transfer, and generic evasion. A separately trained transformer classifier performs the same generator- identification task and falls to 0.094 after style transfer. Generated-versus-human detection is a separate task and is not plotted.
Fig. 2: Held-out corpus perplexity across recursive generations. Lower is better. Thin lines show independent lineages and thick lines show rule means with sample-standard-deviation error bars. The arm labelled “Criterion” in the frozen experiment is the reference-model-perplexity rule described in the text. Provenance and perplexity use five lineages each. Random and fixed-draw use three each. All rules share G0. Mean slopes are 7.543 for perplexity and 7.367 for provenance, a difference of +0.176 perplexity points per generation. The exact two-sided permutation test gives p=0.746 and the test-inversion 95% interval is [−0.794,+1.267] . The interval does not establish equivalence.
Evaluation
Measure
Result
Seven-generator attribution before rewriting
Accuracy
0.987
Style transfer, word / transformer
Accuracy
0.290 / 0.094
Generated vs. human baseline
Minimum AUROC
>0.9999
Unseen source, one roster model held out
AUROC range
0.770–0.992
Five sources outside detector development
AUROC range
0.971–0.993
Style transfer with conformal rejection
Wrong singleton rate
0.408
TABLE I: Selected source-identification and robustness results
Condition
Word TF-IDF
Transformer
Original text
0.987
0.960
Roundtrip translation
0.967
0.787
Paraphrase
0.531
0.547
Style transfer
0.290
0.094
Generic evasion
0.469
0.489
TABLE II: Generator-attribution accuracy before and after rewriting
Rewrite
Original only
Included
Held out
Roundtrip
0.963
0.998
1.000
Paraphrase
0.510
0.998
0.998
Style transfer
0.490
0.993
0.834
Generic evasion
0.447
0.996
0.997
Mean
N/A
N/A
0.957
TABLE III: Character-TFIDF accuracy when rewrite types are represented or held out
Condition
Coverage
Singleton
Wrong singleton
Original text
0.942
0.947
0.004
Roundtrip
0.723
0.727
0.004
Paraphrase
0.055
0.058
0.003
Style transfer
0.138
0.547
0.408
Generic evasion
0.042
0.045
0.003
TABLE IV: Conformal source rejection at α=0.05 , mean over seven source models
Tier
Generator
AUROC
Roster holdout
range 0.770–0.992, mean 0.926
Test-only
glm-5.2-fp8 (zai)
0.971
Test-only
kimi-k2.7-code (moonshot)
0.982
Test-only
qwen3-32b (qwen)
0.993
Test-only
deepseek-v4-flash (deepseek)
0.987
Test-only
Gemma2-27B-it (gemma)
0.978
TABLE V: Unseen-source AUROC by evaluation and generator
Lineage
Slope (perplexity points / generation)
Perplexity s2
7.673
Perplexity s3
7.372
Perplexity s5
7.218
Perplexity s7
7.258
Perplexity s11
8.194
Provenance s2
5.991
TABLE VI: Per-lineage held-out perplexity slopes and primary comparison
Diagnostic
Value
Random reference
Overlap G1 (of 80)
36.0
40.0
Overlap G2 (of 80)
37.1
40.0
Overlap G3 (of 80)
36.2
40.0
Maximum absolute score correlation
0.255
N/A
TABLE VII: Retained-set overlap and score correlation
True difference (perplexity points / generation)
Estimated power
0.00
0.048
0.50
0.157
1.00
0.476
1.48
0.799
2.00
0.964
TABLE VIII: Estimated power by true slope difference