In recent years, artificial intelligence (AI) has impressed with its ability to autonomously generate text, images and even code. However, this advance faces a fundamental challenge: the increasing scarcity of real data to train these models.

We’ve elaborated on the role of this issue in model collapse, but in the early days of 2025, figures like Elon Musk, owner of xAI, and other experts have recognized the depletion of real-world data needed for AI learning and improvement.

What does this mean for the future of AI? “Are we facing the end of an era or the beginning of a new era?”

The real data crisis

During a conversation broadcast on

This statement aligns with the findings of a study by the MIT Data Provenance Initiative, which warns of the rapid decline in data available on the web for this purpose.

The study reveals that, in recent years, access to 5% of the data of the most used data sets in AI has been restricted, and some websites have blocked access to their content using the robots.txt protocol.

Additionally, many publishers have begun monetizing their content or blocking web trackers used by AI companies.

Data is increasingly valuable

The Data Provenance Initiative highlights an “emerging crisis in consent” as data owners resist its use without compensation.

Platforms like Reddit and StackOverflow have started charging for access to their data, and even legal action has been taken, such as The New York Times suing OpenAI and Microsoft for unauthorized use of their content.

This crisis deeply impacts the industry. For giants like OpenAI, Google and Meta, a shortage of high-quality data could slow the development of their models.

The limitation of web data deprives AIs of the information necessary for their continuous learning. The problem is even more acute for small AI companies and academic researchers, who rely on free, public data sets.

Without these resources, many projects would come to a standstill, concentrating access to technology in large corporations with the capacity to acquire data through exclusive agreements or payments.

Synthetic data as a solution

Faced with this shortage, Elon Musk and other experts propose a solution: synthetic data. This data is generated by AI models, offering an alternative to collecting real-world data.

Musk suggests that “the only way to supplement [real data] is with synthetic data, where AI creates [the training data],” allowing for “self-training” of models.

The use of synthetic data, although not new, has become very relevant. Companies like Microsoft, Meta, OpenAI and Anthropic already use them.

Gartner estimates that currently 60% of the data used in AI and analytics projects is synthetically generated. Examples of this are Microsoft’s Phi-4 model, Google’s Gemma models, and Meta’s Llama models.

All that glitters is not gold

Synthetic data offers advantages such as reduced collection and storage costs.

AI startup Writer, which developed its Palmyra

However, synthetic data also presents challenges. Studies suggest that overuse can lead to “model collapse,” where systems become less creative and more biased, replicating the limitations of the original data.

Future risks and challenges

The future of AI will depend on the balance between real and synthetic data. Overreliance on the latter could stagnate models, limiting creativity and increasing susceptibility to bias.

Furthermore, it could exacerbate the gap between large companies and smaller players, concentrating power in those with the resources to generate large volumes of high-quality synthetic data.

Another crucial challenge is the lack of a clear regulatory framework on the use of data for AI training. The lack of consensus on the legitimate use of web data has led to legal disputes and discontent among content creators.

We are still waiting for a system that allows owners to control the use of their content, differentiating between academic/non-profit purposes and commercial purposes.

This post is also available in: Español Français Русский Italiano