Web text makes up the majority of pretraining data and is increasingly AI-generated. After applying FineWeb quality filtering, we find that 27.5% of tokens from June 2026 web data are labeled as AI-generated by Pangram, rising to 31.1% by August. Unlike synthetic data or model-collapse setups, this wild AI text comes from many models, is written for human readers, and arrives unlabeled in pretraining corpora. How does AI text in the wild affect language model pretraining? To answer this question, we pretrain 800 language models, varying the ratio of added AI tokens to human tokens, and fit scaling laws to held-out losses on both human and AI-generated text. For data-starved models, adding AI tokens to pretraining data initially lowers loss on human text, but the benefit saturates as more are added and quickly reverses into harm. For models trained on high budgets of human text, AI tokens raise loss almost immediately, while the same number of fresh human tokens keeps lowering it. Scaling laws such as Hoffman et al. (2022) fail to predict this behavior. We propose a new scaling law with separate benefit and harm terms that allows the value of an AI token to change sign while also reducing to Chinchilla in the absence of AI text. When fit on smaller models, our scaling law predicts the effect of AI text on held-out human-text loss for models up to 3.6x larger with 41% lower error than the best existing law over all AI ratios. We recommend filtering AI text when the target is human text, repeating human text before expanding the training dataset with AI-generated web text, and reporting validation loss on human and AI text separately AI text remains valuable when the target is AI text. We release WildAI, an 83B-token corpus with AI, topic, and format labels, all 800 models and code at https://github.com/pangramlabs/WildAI.
Figures & tables
Figure 1: AI-generated web text helps only data-starved models and costs compute otherwise. Left: share of web tokens in documents Pangram 3.3.2 labels AI or Mixed, by month. The AI share alone is 10% in June 2024, 16% in June 2025, 28% in June 2026, and 31% in August 2026. Middle: our law’s value of one added AI token at 268M, in human tokens, from 5 to 100 TPPh : positive at 5 TPPh , below zero from 20 TPPh once r exceeds about 0.08. Right: compute an unfiltered crawl needs to match training on its human subset at 20 TPPh (268M) for its AI share: 1.6 × at August 2026’s 31%, and 2.1 × and 3.0 × at the 42% and 51% forecasts for the ends of 2027 and 2028.
Figure 2: AI-labeled web text leans toward a few formats and survives quality filters more often than human text. Left: each format’s share of AI-labeled tokens over its share of human-labeled tokens, January to June 2026, for the five most over- and under-represented formats. Right: share of AI- and human-labeled documents in a 2026 Common Crawl sample left after each FineWeb stage.
Figure 3: Added AI text lowers loss only at small human budgets. Change in C4 loss against each model’s human-only control when AI tokens or the same number of fresh human tokens are added to a fixed human corpus, at 5, 20 and 40 TPPh .
Figure 4: Our law fits the trained models and predicts larger ones. Left: observed change in C4 loss with our law. Middle: the held-out 477M and 973M runs at 20 TPPh against our law and Chinchilla. Chinchilla predicts that AI text keeps lowering loss; the runs and our law turn upward. Right: predicted against observed change for every AI run.
Law
k
C4
FW22
FW26-H
Paloma
Chinchilla ( Hoffmann et al., 2022 )
5
4.32
4.65
5.40
9.15
Muennighoff et al. (2023)
7
3.96
4.21
4.94
9.15
CD ( Qin et al., 2026 )
8
3.68
3.63
4.17
9.15
Lovelace et al. (2026)
9
3.55
3.74
4.14
8.38
ATLAS ( Longpre et al., 2026 )
6
4.36
4.72
5.51
9.15
He et al. (2025)
6
6.54
6.55
6.62
8.96
Table 1: Paired RMSE ×103 on the held-out sizes, lower is better.
Figure 5: Filtering AI text pays off at larger human budgets, and repeating human text beats adding AI text. Left: change in loss from removing the AI documents of a 22.3% AI mix without replacing them, predicted by our law (color; black line: no change) and measured (circles; values for 973M). Right: change in loss when the human corpus is repeated or the same number of new AI tokens is added, for models at 20 TPPh .
Figure 6: Whether AI text helps depends on the evaluation set. Left: the AI share of training tokens that minimizes our law’s predicted loss with the human corpus fixed, per evaluation set: near zero for human text (C4, FW22), above 90% for AI text (Cosmo), and falling from 98% to 34% for mixed FW26; the drop near 10 TPPh is where the optimum moves between two local minima of the predicted loss. Right: for the 243 AI additions that raise loss on human-labeled FW26 text, the median change in loss that a validation set reports as the AI-labeled share of its text grows. The sign flips at 5.1% AI; at our 2026 crawl’s 22.3% the median run looks 1.9% better.
Appendix figures & tables30 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: Knowledge articles, tutorials and listicles are the most AI-labeled formats in nearly every topic. Share of documents Pangram 3.3.2 labels AI in each WebOrganizer topic and format of our training pool documents.
Figure 8: The web shifts toward the formats and topics AI writes most. We plot January 2021 to August 2026, Pangram 3.3.2 and WebOrganizer labels by calendar quarter. Each format’s and topic’s share of web tokens and the share of each format’s and topic’s tokens labeled AI.
Figure 9: DCLM’s quality classifier keeps AI-labeled documents far more often than human ones. Survival of a 10,000-document 2026 Common Crawl sample through each stage of the FineWeb (top) and DCLM (bottom) pipelines, split by Pangram 3.3.2 label.
Figure 10: The AI share of web tokens, measured each crawl month and forecast to December 2030 by a random walk with drift fitted to December 2022 through August 2026, with 80 and 95% prediction intervals.
December
Forecast
80% interval
95% interval
2026
33.9%
30.0 to 37.8%
27.8 to 40.0%
2027
42.3%
33.5 to 51.1%
28.7 to 55.9%
2028
50.7%
38.0 to 63.4%
31.0 to 70.4%
2029
59.1%
42.7 to 75.5%
33.7 to 84.5%
2030
67.5%
47.5 to 87.4%
36.5 to 98.4%
Appendix
Table 2: Forecast AI share of web tokens each December. The intervals are Student- t prediction intervals that include the uncertainty of the drift.
Model
Layers
Width
Heads
N (with input embedding)
Without
19.9M
4
256
2
19.9M
11.5M
35.8M
6
384
3
35.8M
23.2M
86.2M
9
640
5
86.2M
65.2M
135M
12
768
6
135.3M
110.1M
268M
16
1024
8
268.4M
234.9M
477M †
20
1280
10
477.1M
435.2M
Appendix
Table 3: The seven model sizes. All use the nanochat architecture. N in the scaling laws counts the input embedding; the last column subtracts it. Reserved sizes ( † ) enter no fit.
Additions
Model
TPPh
Controls
AI
Human
Total
In Table 5
Other seeds
19.9M
2.9–87.5
20
113
20
153
61
25
35.8M
3.2–87.5
16
96
16
128
60
3
86.2M
3.8–87.5
17
88
8
113
54
7
135M
4.1–87.5
18
110
45
173
73
30
268M
4.4–87.5
17
95
47
159
69
32
Appendix
Table 4: Every model in the paper. The 800 models of the scaling-law cohort by size and arm: human-only controls, AI additions and fresh-human additions, each addition trained on exactly its control’s human documents. Sizes marked † are held out from every fit. TPPh : range over the controls. In Table 5 : runs in the matched-budget grid, controls included; the others belong to control groups at per-size budgets from earlier run generations and enter every fit and score in the same way. Other seeds : runs with a seed other than 1337. Below the rule, the 44 filtering and repetition models, which enter no fit.
Model \TPPh
5
20
40
55
75
90
Other
19.9M
12 + 3
11 + 4
5
9 + 1
5
5
14: 66 + 12
35.8M
12 + 3
11 + 4
5
8 + 1
5
5
10: 50 + 8
86.2M
11 + 3
9 + 3
5
7
5
5
11: 46 + 2
135M
14 + 7
12 + 7
6 + 3
5 + 1
5 + 1
5 + 1
12: 63 + 25
268M
11 + 7
11 + 6
6 + 4
5 + 1
5 + 1
5 + 1
11: 52 + 27
477M †
7 + 3
7 + 3
5
3
3
–
3: 17 + 6
Appendix
Table 5: The training grid. Each cell is one human-only control plus the number of AI-addition runs and, after the plus sign, fresh-human-addition runs on the same human corpus. Sizes marked † , below the rule, are reserved: every law is scored on them and none is fitted on them.
Label in figures ( TPPh )
5
20
40
55
75
90
Actual DH/N
4.7
18.9
37.8
54.7
73.0
87.5
Appendix
Table 6: TPPh of the aligned budgets. 477M has no 90 budget; 973M has no 55, 75 or 90 budget.
Human text
Mixed
AI text
Coefficient
C4
FW22
FW26-H
Paloma
FW26
FW26-AI
Cosmo
E
0.7115
0.6683
0.6001
0.5003
0.5565
0.1827
0.2886
A
0.2548
0.2759
0.2728
0.9020
0.2655
0.1725
0.2152
α
0.3027
0.2900
0.2919
0.1198
0.2993
0.4269
0.3714
B
0.08480
0.09414
0.09822
0.02689
0.1063
0.4261
0.2761
β
0.4722
0.4485
0.4350
0.9614
0.4149
0.1209
0.1815
Appendix
Table 7: Coefficients of Equation 2 fitted separately on each evaluation target, on the same 726 models (19.9M to 268M) as Table 1 .
Figure 11: The credit and harm terms of our law. The two terms of Equation 2 and their net effect on C4 at 268M, for human budgets from 5 to 100 TPPh .
Form
k
477M and 973M
The credit AI tokens earn (harm as in our law)
No credit: AI tokens add no data
8
3.29
Linear, ηr (no saturation)
9
1.87
Exponential window, η=1 (CD’s credit)
10
0.97
Rational window, free η
11
0.98
Tanh window, free η
11
0.83
Appendix
Table 8: Our law on C4 with the harm or benefit changed one at a time: paired RMSE ×103 , lower is better, scored on the held-out sizes.
Figure 12: The window in which AI text helps closes as the human budget grows. Left: the fitted credit window R⋆=Ktρ , the rate at which the benefit of AI text saturates, against the human budget for each evaluation set; it is orders of magnitude wider for the mixed and AI-text sets (FW26, Cosmo) than for human text. Right: the largest reduction in C4 loss that added AI text achieves in each fitted group and under our law: up to 3.0% at 5 TPPh , under 0.4% at 20 TPPh and none beyond.
Figure 13: The harm of AI text decelerates. Change in C4 loss against r from 1 to 64: each doubling of AI text adds the same increase in loss. Dashed: an accelerating power-law penalty in the style of Lovelace et al. (2026) (exponent 1.8, for illustration)
Figure 14: Effective human tokens place the AI runs on the Chinchilla data term. Counting human tokens only, AI runs scatter off the Chinchilla data term traced by the human-only runs; counted as effective tokens, they collapse onto it.
Figure 15: Our law across every fitted size and budget. Change in C4 loss against r for every fitted control group with our law’s curves for added AI text and added fresh human text.
Figure 16: Our law’s value of an AI token against measured values. Value of an added AI token in human tokens under our law and measured from paired AI and fresh-human additions to the same control, at 5 and 20 TPPh .
Figure 17: Only our law meets all five criteria at low error. Left: whether each benchmarked law meets the criteria of § 4.2 : always (check), partly or only for some coefficient values (half circle), or for no coefficient values (cross). For criterion C4 (a finite first-token value), a half circle marks a value that is finite but fixed, as in Chinchilla, where every AI token is worth exactly one human token, or finite only when a fitted exponent reaches one. Right: each law’s paired RMSE on the held-out runs for every evaluation set, the numbers of Table 1 and Table 9 . The law of Sedova et al. (2026) also meets every criterion, at 11.6 times our error on C4.
477M and 973M
Law
k
FW26
FW26-AI
Cosmo
General
Chinchilla ( Hoffmann et al., 2022 )
5
19.74
103.86
78.76
Repetition and paraphrase
Muennighoff et al. (2023)
7
17.67
99.83
75.30
CD ( Qin et al., 2026 )
8
16.06
98.08
72.83
Appendix
Table 9: Table 1 for the mixed and AI-text targets: paired RMSE ×103 on the held-out sizes, lower is better. On FW26-AI and Cosmo, the exponent of the human-share multiplier of He et al. (2025) sits at its bound of zero, where the law is exactly Chinchilla: the multiplier can only raise the loss as AI text is added, and added AI text lowers the loss on these targets.
Human-text targets
477M and 973M
Law
k
C4
FW22
FW26-H
Paloma
General
Chinchilla ( Hoffmann et al., 2022 )
5
26.93
28.41
29.74
30.91
Repetition and paraphrase
Muennighoff et al. (2023)
7
21.70
22.24
22.76
31.66
Appendix
Table 10: Absolute BPB RMSE ×103 on the held-out sizes, lower is better: the error in the predicted loss level itself rather than in the paired change against the human-only control, so the ranking can differ.
477M and 973M
Law
k
C4
FW22
FW26-H
Paloma
General
Chinchilla ( Hoffmann et al., 2022 )
5
2.93
3.18
3.82
7.75
Repetition and paraphrase
Muennighoff et al. (2023)
7
2.47
2.63
3.21
7.75
CD ( Qin et al., 2026 )
8
2.22
2.12
2.60
7.75
Appendix
Table 11: The paired errors of Table 1 restricted to the held-out runs with r<1 , the range a web crawl can reach (our 2026 crawl is r=0.29 ; an even split of AI and human text is r=1 ).
Draws where ours is (%)
Target
AI runs
Ours
Joint law
Joint − ours
below joint
best of 12
C4
all
0.83 [0.64, 0.99]
1.41
+0.58 [+0.32, +0.84]
99.99
99.99
r<1
0.64 [0.53, 0.77]
0.98
+0.34 [-0.03, +0.65]
96.5
96.5
FW22
all
0.85 [0.74, 0.96]
1.50
+0.65 [+0.32, +0.94]
100
100
r<1
0.71 [0.56, 0.87]
1.01
+0.30 [-0.10, +0.65]
92.6
92.6
FW26-H
all
1.28 [0.97, 1.64]
2.25
+0.97 [+0.58, +1.35]
100
100
Appendix
Table 12: Our law’s lead holds over all AI ratios on human text, but not at r<1 on C4 and FW22. Paired RMSE ×103 on the held-out runs for our law and the joint law of Shukor et al. (2025) , with 95% intervals for ours and for the gap from 10,000 bootstrap draws that resample the 51 held-out AI runs (43 with r<1 ) by their 10 control groups, every fit held fixed. The last two columns give the share of draws in which our law has lower error than the joint law, and than all eleven comparators.
Figure 18: The joint law of Shukor et al. (2025) extrapolates inconsistently. Our law and the joint law on C4 at 20 TPPh , fitted through 268M or 135M. Left: 268M, a trained size (points: trained models). Middle: extrapolated to 8B, where the joint law’s predicted effect of AI text changes sign with the sizes it is fitted on, while every fit of ours predicts harm. Right: the loss floor each fit of the joint law implies for a training mix, relative to human-only training; our law has one floor for every mix.
Figure 19: Less AI text is always more compute-efficient for C4, and more AI text for Cosmopedia. Compute-equivalent gain over our 2026 crawl mix (22.3% AI) of 268M models at 20 TPPh as the AI share of training tokens varies, on C4 and Cosmopedia.
Figure 20: Repeating human text beats adding AI text on human-text evaluation sets. Change in loss on each evaluation set when the human corpus is repeated or the same number of new AI tokens is added, against tokens added per human token. On the human-text sets (C4, FW22, FW26-H) repetition lowers loss where AI text raises it; on the AI-text sets (FW26-AI, Cosmo) AI text lowers loss more.
Figure 21: On AI-generated text, added AI tokens lower loss at every budget. Counterpart of Figure 3 for Cosmopedia: change in Cosmo loss when AI tokens or the same number of fresh human tokens are added, at 5, 20 and 40 TPPh for every fitted size. AI text lowers the loss at every budget and size, by up to 25%; fresh human text lowers it far less.
Figure 22: Mixed validation data hides harm to human text. Left: every model’s loss on the AI-labeled against the human-labeled partition of FW26. Human-only models have 12 to 27% lower loss on AI text, and models trained with AI text fall further below the equal-loss line. Right: of the 243 AI additions that raise loss on human-labeled text, the share a validation set reports as improvements as the AI-labeled share of its text grows, for the fitted sizes and the held-out sizes: 49.4% at 5% AI and 95.5% at our 2026 crawl’s 22.3%.
Figure 23: Added AI tokens raise CORE about as much as fresh human tokens. Change in CORE against each model’s human-only control, as AI tokens or the same number of fresh human tokens are added. MMLU ( Hendrycks et al., 2021 ) is not drawn because it stays at chance accuracy for every model, as expected at these sizes.
Figure 24: Models trained with AI text use more AI-typical phrases. Rate of 26 AI-typical phrases per 1,000 words in generations from models trained with AI text added at ratio r .
AI phrases per 1,000 words
Stories labeled AI (%)
r
AI share (%)
5 TPPh
20 TPPh
20 TPPh
0
0.0
0.16
0.16
18.6 [15.4, 22.0]
0.025
2.4
0.24
0.18
22.4 [18.6, 26.1]
0.05
4.8
0.18
0.21
21.1 [17.7, 24.7]
0.25
20.0
0.26
0.24
26.8 [23.0, 30.6]
0.5
33.3
0.23
0.24
22.7 [19.1, 26.7]
Appendix
Table 13: The 268M models trained with more AI text write more like AI. Story continuations of 500 WritingPrompts prompts from the 268M models at 5 and 20 TPPh , one model per r . AI-typical phrases per 1,000 words over eight continuations per prompt, and the share of one continuation per prompt that Pangram 3.3.2 labels AI, with 95% intervals from a bootstrap over prompts.
Upweighting high-quality data in LLM pretraining often improves performance, but in datalimited regimes, especially under overtraining, stronger upweighting increases repetition and can degrade performance. However, standard scaling laws do not reliably extrapolate across mixture recipes or under repetitions, making the selection for optimal data recipes at scaling underdetermined. To solve this, we introduce InfoLaw (Information Scaling Laws), a data-aware scaling framework that predicts loss from consumed tokens, model size, data mixture weights, and repetition. The key idea is to model pretraining as information accumulation, where quality controls information density and repetition induces scaledependent diminishing returns. We first collect the model performance after training on datasets that vary in scale, quality distribution, and repetition level. Then we build up the modeling for information so that information accurately predicts those model performance. InfoLaw predicts performance on unseen data recipes and larger scale runs (up to 7B, 425B tokens) with 0.15% mean and 0.96% max absolute error in loss, and it extrapolates reliably across overtraining levels, enabling efficient data-recipe selection under varying compute budgets.
LLM pretraining is shifting from a compute-bound to a data-bound regime, where available human (organic) text falls far short of scaling demands. However, reaching the data-bound regime does not mean the model has fully utilized its organic corpus. In this paper, we introduce SynPro, a synthetic data generation framework that helps LLMs more thoroughly learn from limited organic data. SynPro applies two operations, rephrasing and reformat, that present the same organic source in diverse forms to facilitate deeper learning without introducing external information. Both generators are optimized via reinforcement learning with quality, faithfulness, and data influence rewards, and are continuously updated as pretraining plateaus to target content the model has yet to absorb. We pretrain 400M and 1.1B models with 10% of their Chinchilla-optimal tokens (0.8B and 2.2B) from DCLM-Baseline, reflecting a realistic data-bound regime in frontier pretraining. Our results reveal that organic data is significantly underutilized by standard repetition: SynPro unlocks 3.7-5.2x the effective tokens of repetition, even surpassing the non-data-bound oracle that trains on equivalent unique data at the 1.1B scale. Analyses confirm that faithful, model-aware synthesis sustains data-bound scaling without causing distribution collapse. We open-source our code at https://github.com/cxcscmu/SynPro.
Zichun Yu, Chenyan Xiong
Language Technologies Institute, Carnegie Mellon University · Xlue
Although it is generally agreed that AI-generated text poses a broad societal risk, there is no common understanding in the AI-generated text detection literature on what constitutes harmful use. Rather, existing datasets and approaches often define their own criteria and make their own assumptions, sometimes implicitly, and often only loosely related to real-world needs and applications. To address this gap, we here systematically define various notions of AI-generated text and their characteristics. To study these, we collect AITDNA - a new benchmark of human-machine co-constructed texts that is annotated with detailed genesis information, such as the entire edit and AI-interaction history. We benchmark various machine-generated text detectors and find that they often only perform well for specific notions but not as broad detectors. We release code and data publicly.
Nils Dycke, Marina Sakharova, Nico Daheim +1
Ubiquitous Knowledge Processing Lab (UKP Lab), Department of Computer Science, Technical University of Darmstadt · National Research Center for Applied Cybersecurity ATHENE, Germany · Zuse School ELIZA