The Curse of Recursion: Training on Generated Data Makes Models Forget
- Type
- paper
- Venue
- arXiv
- Year
- 2026
- Source
- arxiv
- Access
- free
- Language
- en
- Added
- 2026-08-14T18:28:00Z
- Verified
- 2026-08-14T18:28:00Z
Summary
Asks what happens to GPT-n once web-scale training mixes in LLM-generated text. Across Gaussians, GMMs, and language models, recursive training on generated data erases the tails of the original distribution and concentrates on a narrowing mode they call model collapse. Early generations already lose diversity; later ones can forget original-task performance. Mixing some real data delays the effect in their setups but does not remove it. Discord posted the PDF of arXiv 2305.17493.
Keywords
model-collapse · synthetic-data · recursive-training · generated-data · llms
Topics
model collapse, synthetic data, LLM training
Research notes
- Primary: arxiv abs (cs.LG; also cs.AI, cs.CL, cs.CR, cs.CV). Discord posted PDF. License not stated on the Atom API; left blank per catalog convention. Later Nature 2024 journal version (AI models collapse when trained on recursively generated data) is not the posted URL. No official code on the abs query. Not a new hosted corpus, so no datasets_local row.