Research
How Much Is an AI Token Worth? Scaling Laws for Wild AI-Generated Web Text
Overview Research area: Natural language processing / large language model pretraining data curation and scaling laws. Technical level: Intermediate (the paper uses scaling-law notation and reports fi

- arXiv
- 2609.40295
- Published
- 2026-09-30
- Authors
- Jenna Russell, Ben Glickenhaus, Katherine Thai, John Wieting, Mohit Iyyer, Max Spero, Bradley Emi
AI summary
Overview
Research area: Natural language processing / large language model pretraining data curation and scaling laws. Technical level: Intermediate (the paper uses scaling-law notation and reports fit coefficients, but its core claims are explained in plain terms). Scope: A large-scale empirical study of how AI-generated web text ("wild AI text") affects language model pretraining, with a new scaling law and practical data-curation recommendations.
What This Paper Is About
Web text, the main ingredient of LLM pretraining data, is increasingly written by AI rather than people: the authors find that 27.5% of tokens in June 2026 web data that survive FineWeb quality filtering are labeled AI-generated, rising to 31.1% by August 2026. Unlike synthetic data (written on purpose to help training) or model-collapse setups (a model trained recursively on its own output), this "wild" AI text comes from many models, is written for human readers, and arrives unlabeled and mixed into pretraining corpora. The paper asks how much one such AI token is worth compared to a human token, and proposes a scaling law that predicts the answer.
Key Contributions
- A measurement of AI text on the web. The authors label a 310,000-document monthly sample (5,000 documents per month, January 2021 through August 2026, 62 crawl months) and report the rising AI share, plus the finding that FineWeb keeps AI-labeled documents 2.3× as often as human ones and DCLM 9.8× as often.
- The Wild AI dataset. An 83.31B-token corpus (96.04M documents) drawn from FineWeb v1.4.0 extended to June 2026, split into a human-written corpus of 58.91M documents (42.19B tokens) and an AI-generated corpus of 32.54M documents (35.32B tokens), with AI, topic, and format labels, plus all 800 trained models and code.
- A new scaling law. A law combining a saturating benefit term with a logarithmic harm penalty that reduces to Chinchilla when no AI text is present, fit on 726 models of 19.9M to 268M parameters and tested on 74 held-out models of 477M and 973M parameters (1.8× and 3.6× the largest fitted size).
- Actionable recommendations. Filter AI text when the target is human text, repeat human text before adding AI text, and report validation loss on human and AI text separately.
Main Findings
- The AI share of filtered web data is rising fast. Less than 0.1% of FineWeb tokens were AI-generated in June 2021, 10.1% in June 2024, 16.1% in June 2025, and 27.5% in June 2026, reaching 31.1% in August 2026 (3.6 points above June). A random-walk forecast projects 42.3% by the end of 2027 and 50.7% by the end of 2028.
- Quality filters prefer AI text. FineWeb's pipeline keeps 29.3% of AI-labeled documents versus 12.8% of human-labeled ones (2.3×, 95% interval 2.1 to 2.6×) and DCLM's full pipeline keeps 14.5% versus 1.5% (9.8×, interval 7.5 to 12.7×). The DCLM gap opens at its learned classifier, which keeps 40% of AI documents that reach it and 9.5% of human ones.
- AI tokens help only data-starved models. Adding AI tokens lowers loss on human text below about 10 human tokens per parameter, but the benefit saturates and reverses; from the Chinchilla-optimal 20 tokens per parameter, AI tokens raise loss, while the same number of fresh human tokens keeps lowering it.
- Chinchilla cannot model this. On reserved model sizes, Chinchilla predicts fresh-human additions to a paired error of 1.32 × 10⁻³ on C4 but AI additions only to 4.32 × 10⁻³.
- The new law predicts larger models with lower error. It reaches a paired error of 0.83 × 10⁻³ on C4, versus 1.41 × 10⁻³ for the best existing law (the joint law of Shukor et al., 2025) and 4.32 × 10⁻³ for Chinchilla, and it predicts models up to 3.6× larger with 41% lower error than the best existing law over all AI ratios. It also has the lowest error on FW22, FW26-H, and Paloma among the eleven existing laws benchmarked.
- Not filtering costs compute. Training on unfiltered web text at August 2026's AI share (31.1%) requires 1.6× as much compute as training on its human subset at 20 tokens per parameter, rising to 2.1× at the 42% forecast for the end of 2027 and 3.0× at the 51% forecast for the end of 2028.
- Filtering AI text helps at realistic budgets. Across 19 paired comparisons at a 22.3% AI share, models above roughly 10 to 15 tokens per parameter (less at larger sizes) had lower C4 loss with the AI text removed; the law gets the direction right in 18 of the 19 pairs (paired error 3.28 × 10⁻³), while Chinchilla predicts removing AI text raises loss in all 13 pairs where it actually lowers it.
- Repeating human text beats adding AI text. At 20 tokens per parameter, C4 loss after one added epoch was 2.3% lower when repeating human tokens than when adding fresh AI tokens, a gap that expands to 6.3–6.4% after eight added epochs (nine in total).
- The optimal AI share depends on the target. The loss-minimizing AI share on C4 is 37% at 5 human tokens per parameter but 1.0% by 20; on AI-generated text (Cosmopedia) it stays above 90% at every budget, and the law values the first AI token at more than 20 human tokens.
- Mixed validation sets hide the harm. Because AI text is easier to predict, all 553 models with added AI text lower loss on AI-labeled FW26 text while 44% raise loss on human-labeled text. For the 243 AI additions that raise loss on human-labeled FW26 text, the reported change in loss flips sign at a 5.1% AI share of the validation set; at the 2026 crawl's 22.3% the median harmful run looks 1.9% better, hiding the harm in 95.5% of harmful runs.
- Training on AI text changes how models write. At 268M parameters, AI-typical phrases per 1,000 words rise from 0.16 without AI text to 0.36 at r = 1 and 0.54 at r = 4, and at r = 2 Pangram labels twice as many generated stories AI (37.2% vs. 18.6%).
Methodology in Plain English
The authors first build a labeled web dataset: they extend FineWeb v1.4.0 past its June 2025 cutoff by rerunning the FineWeb pipeline (DataTrove) on Common Crawl data from July 2025 to June 2026, then label documents with EditLens Llama-3.2-3B to narrow the pool and Pangram 3.3.2 for the final human/AI labels, keeping only documents where all sampled windows agree and discarding "Mixed" ones. They then pretrain 800 nanochat models from 19.9M to 973M parameters, varying human tokens per parameter (2.9 to 87.5) and the ratio of added AI tokens to human tokens (0 to 64), and evaluate held-out loss on human sets (C4, pre-2022 FineWeb, Paloma), mixed 2026 FineWeb data (22.3% AI), its AI and human splits, and Cosmopedia. They fit scaling laws on 726 smaller models and test them on 74 larger held-out models, scoring laws by the RMSE of the predicted change in log loss against each run's own human-only control ("paired error"). The proposed law keeps the Chinchilla backbone, borrows a saturating credit window for the benefit of AI tokens, and adds a logarithmic harm penalty whose marginal cost peaks when AI and human tokens are equal.
Why This Matters
Impact on research. The paper reframes AI-generated web text as a distinct, measurable property of pretraining data rather than a generic "synthetic data" problem, and shows that widely used scaling laws (Chinchilla) and quality filters (FineWeb, DCLM) both misjudge it. It also shows that evaluation design matters: a mixed validation set can conceal harm in 95.5% of harmful runs at a 22.3% AI share.
Real-world applications:
- Deciding whether to filter, keep, or repeat data when assembling pretraining corpora from fresh web crawls.
- Budgeting compute for training runs, using the headroom implied by AI contamination (1.6× at August 2026's share, up to 3.0× by 2028).
- Designing evaluation suites that report human-text and AI-text loss separately.
- Building models that read AI-generated input, as in agentic workflows where models prompt and consume other models' outputs.
Industry relevance. Every organization that trains on web-scale crawls faces the same choice, and the paper argues most current models train far past 20 tokens per parameter (for example, Qwen3-32B at roughly 1,100 tokens per parameter), the regime where AI text hurts. The released Wild AI corpus, labels, code, and 800 checkpoints let practitioners regenerate the fits and figures without retraining.
Future Directions
- Filter wild AI text and supplement it with targeted synthetic data, as the authors note Microsoft AI does.
- Identify the specific domains or tasks in which wild AI text still helps.
- Measure how the AI text already present in existing pretraining corpora has affected the quality of current models.
- Train models that can read AI-written input without writing like it, for example by tagging AI text or masking its loss.
- Extend the laws beyond the tested range: the law is fit on models up to 268M parameters and tested up to 973M, and the authors note that larger models memorize more and may be hurt more by model-generated data.
Target Audience
Researchers and engineers working on pretraining data curation, scaling laws, and dataset filtering; teams planning large pretraining runs and compute budgets; evaluation designers concerned with contamination from AI-generated text; and dataset builders interested in the released Wild AI corpus of 83.31B tokens with AI, topic, and format labels.
Authors’ abstract
Web text makes up the majority of pretraining data and is increasingly AI-generated. After applying FineWeb quality filtering, we find that 27.5% of tokens from June 2026 web data are labeled as AI-generated by Pangram, rising to 31.1% by August. Unlike synthetic data or model-collapse setups, this *wild* AI text comes from many models, is written for human readers, and arrives unlabeled in pretraining corpora. How does AI text in the wild affect language model pretraining? To answer this question, we pretrain 800 language models, varying the ratio of added AI tokens to human tokens, and fit scaling laws to held-out losses on both human and AI-generated text. For data-starved models, adding AI tokens to pretraining data initially lowers loss on human text, but the benefit saturates as more are added and quickly *reverses* into harm. For models trained on high budgets of human text, AI tokens raise loss almost immediately, while the same number of fresh human tokens keeps lowering it. Scaling laws such as Hoffman et al. (2022) fail to predict this behavior. We propose a new scaling law with separate benefit and harm terms that allows the value of an AI token to change sign while also reducing to Chinchilla in the absence of AI text. When fit on smaller models, our scaling law predicts the effect of AI text on held-out human-text loss for models up to 3.6x larger with 41% lower error than the best existing law over all AI ratios. We recommend filtering AI text when the target is human text, repeating human text before expanding the training dataset with AI-generated web text, and reporting validation loss on human and AI text separately AI text remains valuable when the target is AI text. We release WildAI, an 83B-token corpus with AI, topic, and format labels, all 800 models and code at https://github.com/pangramlabs/WildAI.