For years, generative artificial intelligence has been a spectacle of innovation. Models like ChatGPT, capable of writing texts or solving complex problems, marked a before and after in the interaction with technologies.
This advancement seemed unstoppable, fueled by a simple formula: train increasingly larger models with more data and more GPUs. However, something has changed.
Today, influential voices within the industry, such as Ilya Sutskever, co-founder of OpenAI, warn that this path is reaching its limit. The new models barely improve compared to the previous ones, despite the colossal investments.
So what is going wrong? More importantly, what comes next? Discover why training models as before is no longer enoughand how the industry is looking for a new direction to continue innovating.
The rise and limits of the traditional approach
Training AI models gave spectacular results in the last decade. Tools like GPT-3 and GPT-4 amazed the world by typing like humans, answering complex questions, and performing creative tasks with great fluidity.
However, this success came at a price. Training a model with these characteristics requires enormous time and resources.
The laboratories behind these technologies invest tens of millions of dollars and months of work to complete a training cycle, with no guarantee that the results will be significant.
Additionally, as the models grew, profits began to decline. Improvements between model generations have become less impactful, even though costs continue to rise.
This is known as “diminishing returns.” It is no longer enough to add more data or computing power. Current models are hitting a performance ceiling, and what was once amazing now seems incremental.
Signs of stagnation in the industry
Although generative artificial intelligence has revolutionized entire sectors, signs of stagnation are increasingly evident:
Disappointing releases
Models like OpenAI’s Orion or Google’s Gemini, which in the past would have generated great excitement, are now perceived as minor advances.
These tools reportedly offer incremental improvements over their predecessors, but not much. Companies are postponing launches to polish their models, but the truth is that the impact is no longer as striking.
High costs and diminishing returns
Training these models remains extremely expensive and time-consuming. Each cycle can cost tens of millions of dollars and take months to complete, with no guarantee of success.
This elevated risk has caused companies to reconsider whether the current approach is sustainable.
Changes in approach
Industry leaders such as Ilya Sutskever and Yann LeCunagree that the method of scaling resources has been exhausted.
This reinforces the idea that the industry must find new strategies to overcome current limits and keep innovation moving.
The new direction: models that reason
The AI industry is exploring a new path: models that “reason.” This aims to improve the models’ ability to analyze their own responses in real time before delivering them, which would increase their accuracy.
A key technique in this strategy is test-time compute. Unlike traditional training, this technique allows the model to evaluate multiple options during inference and choose the most appropriate answer.
OpenAI, for example, is implementing this approach in its O1 model, which reviews and filters its possible answers before providing a definitive one. This transition is not only being explored by OpenAI.
Companies like Anthropic, Google and Microsoft are also working on similar models. The goal is to provide AI with something resembling critical thinking, which represents a fundamental change in how we understand how it works.
From GPUs to inference hardware
The stagnation in mass training of generative AI models has driven a shift in focus towards the inference stage, where models generate responses or execute tasks.
This change is not only technical, but strategic: instead of investing in GPUs to train giant models, companies are developing specialized hardware to optimize performance in inference.
A key figure in this change is Jensen Huang, CEO of Nvidia, who noted that we have entered a new era of scaling, focused on improving efficiency when using AI.
Nvidia has introduced Blackwell, a new generation of chips designed specifically for inference, which promises massive demand in data centers.
Why does this new approach promise to be better?
This approach allows models to perform complex tasks in less time, which is essential for reasoning models. Additionally, specialized hardware could reduce long-term operating costs by optimizing power consumption and performance.
Other companies are also working on similar solutions, taking advantage of this change to compete in a market that demands faster and more accurate AI.
Thus, the transition from mass training to inference could redefine how artificial intelligence systems are built and used.
The only constant is change
After a decade marked by impressive advances in AI, the traditional approach of scaling resources to infinity is showing its limits. However, this apparent stagnation is not the end, but rather an invitation to change.
The move toward reasoning models and specialized hardware for inference indicates that the industry is learning to look beyond quantity, focusing on quality and efficiency.
This shift not only promises to overcome current challenges, but also opens new opportunities for the development of more intelligent and useful systems.
The future of AI will not be a linear path and that is precisely its strength. Innovation always finds new routes, and constant change is what ensures that this technology will continue to change the world in ways we can barely imagine.
This post is also available in: