How Much Is an AI Token Worth? Scaling Laws for Wild AI-Generated Web Text
- Type
- paper
- Venue
- arXiv 2026-09-30 (cs.CL, cs.LG)
- Year
- 2026
- Source
- arxiv
- Access
- public
- Language
- en
- Added
- 2026-10-01
- Verified
- 2026-10-01
Summary
Measures how 'wild' AI-generated web text affects LM pretraining: after FineWeb quality filtering, 27.5% of June 2026 web tokens are labeled AI-generated by Pangram, rising to 31.1% by August. Pretrains 800 language models varying the ratio of added AI tokens to human tokens and fits scaling laws to held-out losses on both human and AI-generated text. For data-starved models, adding AI tokens initially lowers loss on human text, but the benefit saturates and quickly reverses into harm; for models trained on high budgets of human text, AI tokens raise loss almost immediately while the same number of fresh human tokens keeps lowering it. Standard scaling laws (e.g. Hoffman/Chinchilla 2022) fail to predict this behavior, so the authors propose a new scaling law with separate benefit and harm terms that lets the value of an AI token change sign and reduces to Chinchilla in the absence of AI text; fit on smaller models it predicts the effect on models up to 3.6x larger with 41% lower error than the best existing law over all AI ratios. Recommendations: filter AI text when the target is human text, repeat human text before expanding the training dataset with AI-generated web text, and report validation loss on human and AI text separately; AI text remains valuable when the target is AI text. Releases WildAI, an 83B-token corpus with AI, topic, and format labels, plus all 800 models and code.
Keywords
AI-generated text · wild AI text · pretraining · scaling laws · Chinchilla · synthetic data · data filtering · WildAI
Topics
AI-generated text, wild AI text, pretraining, scaling laws, Chinchilla, synthetic data, data filtering, WildAI, Pangram
Research notes
- Discovery: shared directly in chat (2026-10-01)
- License: CC BY-NC-SA 4.0
- Associated dataset WildAI (83B tokens with AI/topic/format labels) logged in the Datasets sheet
- Code and all 800 models released per the abstract (linked from the paper page)
- Directly relevant to the user's own pretraining work: quantifies AI-text contamination and how to handle it.