cs.CLSep 30, 2026

How Much Is an AI Token Worth? Scaling Laws for Wild AI-Generated Web Text

Authors: Jenna Russell, Ben Glickenhaus, Katherine Thai, John Wieting, Mohit Iyyer, Max Spero, Bradley Emi

Organizations: University of Maryland · Pangram Labs

Abstract

Web text makes up the majority of pretraining data and is increasingly AI-generated. After applying FineWeb quality filtering, we find that 27.5% of tokens from June 2026 web data are labeled as AI-generated by Pangram, rising to 31.1% by August. Unlike synthetic data or model-collapse setups, this wild AI text comes from many models, is written for human readers, and arrives unlabeled in pretraining corpora. How does AI text in the wild affect language model pretraining? To answer this question, we pretrain 800 language models, varying the ratio of added AI tokens to human tokens, and fit scaling laws to held-out losses on both human and AI-generated text. For data-starved models, adding AI tokens to pretraining data initially lowers loss on human text, but the benefit saturates as more are added and quickly reverses into harm. For models trained on high budgets of human text, AI tokens raise loss almost immediately, while the same number of fresh human tokens keeps lowering it. Scaling laws such as Hoffman et al. (2022) fail to predict this behavior. We propose a new scaling law with separate benefit and harm terms that allows the value of an AI token to change sign while also reducing to Chinchilla in the absence of AI text. When fit on smaller models, our scaling law predicts the effect of AI text on held-out human-text loss for models up to 3.6x larger with 41% lower error than the best existing law over all AI ratios. We recommend filtering AI text when the target is human text, repeating human text before expanding the training dataset with AI-generated web text, and reporting validation loss on human and AI text separately AI text remains valuable when the target is AI text. We release WildAI, an 83B-token corpus with AI, topic, and format labels, all 800 models and code at https://github.com/pangramlabs/WildAI.

Figures & tables

Appendix figures & tables30 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Generating Pretraining Tokens from Organic Data for Data-Bound Scaling

    May 18, 2026Zichun Yu, Chenyan XiongSynthetic DataSingle-Cell Foundation Models

  2. 'Your AI Text is not Mine': Redefining and Evaluating AI-generated Text Detection under Realistic Assumptions

    Jun 3, 2026Nils Dycke, Marina Sakharova, Nico Daheim +1Machine-Generated Text DetectionSri Lanka