At this point in human history, it is clear that artificial intelligence (AI), like large language models (LLM), are here to stay.
Models like GPT-3, GPT-4 and Stable Diffusion have made it easy to create text, images and music with just a few prompts, transforming entire sectors such as entertainment, education and marketing.
However, this proliferation of AI-generated content has brought with it a problem that many researchers are already warning about: training new AIs with data generated by other AIs could generate cumulative errors.
This phenomenon is known as “model collapse” and, according to experts, it can deteriorate the quality of the models and seriously compromise the future of artificial intelligence.
The danger of “synthetic” data
Large volumes of data are needed to effectively train artificial intelligence (AI), especially in the case of advanced models such as deep neural networks or large-scale language models (LLMs).
The main reason is that these models need to learn complex patterns and generalize knowledge from data, and to do so they require a wide range of examples.
Some of this data can be obtained from the Internet, and from large public and private databases. While other data can be generated artificially.
These synthetic data sets are generated through algorithms or simulations rather than being collected directly from the real world.
This type of “synthetic” data has become increasingly popular in the research and development of machine learning and artificial intelligence (AI) models for several reasons.
These include the scarcity or inaccessibility of real data, privacy problems, and the need to have controlled data sets.
What is model collapse?
The term “model collapse” describes the progressive deterioration of artificial intelligences that are trained with data generated by other AIs.
As more AI-generated content begins to flood the internet, new models, which rely heavily on data available online, may end up training on data that was not produced by humans, but by their generative predecessors.
According to research by Ilia Shumailov and his colleagues, published in the journal “Nature” in July this year, this process can lead to irreversible degeneration in the models.
Why do models collapse?
The problem is that when an AI is trained with content generated by another AI, it begins to lose precision and variability.
This especially affects the “extremes” of the data distribution, that is, those less common examples or those that represent rare situations in reality.
Over time, these extremes disappear, resulting in a loss of diversity in the data and a convergence toward a very limited and predictable distribution.
Shumailov and his team demonstrated this collapsing process in several types of models, including large language models (LLMs), Gaussian mixture models (GMMs) and variational autoencoders (VAEs).
In each case, they watched successive generations of AI become less accurate, to the point where they produced incoherent or absurd results.
The risk of training AI with contaminated data
One of the main dangers of training AI with data generated by other AIs is that this data is contaminated with errors that accumulate with each generation. As an example, we can imagine an AI that has been trained primarily on texts generated by GPT-4.
If a new version of GPT (say GPT-5) is trained largely using texts that come from GPT-4, those texts willcontain subtle errors that its successor will inherit and amplify.
Over time, these errors will not only propagate, but will end up corrupting AI’s ability to generate accurate and useful responses.
This is what Shumailov and his colleagues called “late model collapse,” a point at which the model deviates so much from reality that it loses all connection to the original distribution of the human data.
A study carried out by Sarkar and other researchers in Edinburgh and Madrid showed a similar phenomenon in image diffusion models.
By repeatedly training a model with data generated by other versions of the same model, the resulting images became progressively blurrier and unrecognizable.
By the third generation, images of flowers or birds, which at first were clear and defined, had transformed into meaningless abstract blurs.
The analogy with the “nuclear age” and steel
The phenomenon of model collapse has been compared to a historical problem of the 20th century: the contamination of steel by nuclear radiation.
After World War II, nuclear testing released radioactive particles into the atmosphere, contaminating the steel made thereafter.
For radiation-sensitive applications such as Geiger counters, modern steel was no longer useful, and scientists had to turn to steel from sunken ships before the nuclear age to obtain a “clean” material.
Pre-AI data will be valuable
Similarly, some researchers believe that in the future it will be necessary to use “pre-AI” data to train models that are not contaminated by content generated by other artificial intelligences.
However, unlike steel, data cannot simply be “recycled” from the past. As researcher Ilia Shumailov points out, historical data, although useful in some cases, does not reflect the constantly changing world.
Social norms, language, and technology evolve over time, making old data unsuitable for training models that need to understand current dynamics.
Therefore, the solution is not simply to look to the past, but to find ways to filter AI-generated content and ensure that the training data is genuinely human.
Impacts on diversity and bias
Language models already struggle to accurately represent marginalized or minority groups, and the use of AI-generated data could make this situation worse.
As models lose the ability to capture the extremes of the data distribution, they also lose the ability to adequately represent people and experiences that fall outside of what the algorithms consider “normal” or frequent.
This is an especially serious problem in sensitive contexts such as medicine, criminal justice or labor recruitment, where AI-based decisions can have direct consequences on people’s lives.
If AI models are trained on contaminated data that does not adequately represent society as a whole, they could make more biased decisions, amplifying existing inequalities.
Collapse is not inevitable, but it will be difficult
To prevent model collapse, researchers are exploring several solutions. One of them is the creation of curated and standardized data sets, where humans verify that the content has not been generated by AI.
These “clean” data sets could be shared openly between AI developers, ensuring that models are trained with genuine and diverse data.
Another option is to develop tools that can identify whether content has been generated by AI or not. Although this technology is still in its early stages, it would be a crucial advance in filtering data and preventing contamination in the training of future models.
Will AIs eat themselves?
Using AI-generated data to train new models may seem like an efficient and inexpensive solution, but it poses serious risks for the future of artificial intelligence.
Model collapse is a real phenomenon that threatens to deteriorate the accuracy, diversity and usefulness of generative models.
If we want artificial intelligence to continue advancing in beneficial ways, it is essential that we address these issues, ensuring that training data is as representative and “human” as possible.
Only in this way can we prevent AIs from “eating themselves” and maintain confidence in this revolutionary technology.
This post is also available in: