← Back to explorer

The Curse of Recursion: Training on Generated Data Makes Models Forget

Type
paper
Venue
arXiv
Year
2026
Source
arxiv
Access
free
Language
en
Added
2026-08-14T18:28:00Z
Verified
2026-08-14T18:28:00Z

Summary

Asks what happens to GPT-n once web-scale training mixes in LLM-generated text. Across Gaussians, GMMs, and language models, recursive training on generated data erases the tails of the original distribution and concentrates on a narrowing mode they call model collapse. Early generations already lose diversity; later ones can forget original-task performance. Mixing some real data delays the effect in their setups but does not remove it. Discord posted the PDF of arXiv 2305.17493.

Keywords

model-collapse · synthetic-data · recursive-training · generated-data · llms

Topics

model collapse, synthetic data, LLM training

Research notes

  • Primary: arxiv abs (cs.LG; also cs.AI, cs.CL, cs.CR, cs.CV). Discord posted PDF. License not stated on the Atom API; left blank per catalog convention. Later Nature 2024 journal version (AI models collapse when trained on recursively generated data) is not the posted URL. No official code on the abs query. Not a new hosted corpus, so no datasets_local row.