Reaching new levels of capability and performance appears to be more challenging and costly than ever for Artificial Intelligence models.
The emergence of new models such as o3 by OpenAI highlights how improvements in AI are increasingly difficult to achieve and, above all, more expensive.
As models advance, the “scaling laws” that guided progress are showing signs of exhaustion. But even worse, new techniques to overcome these limitations bring with them new challenges.
This raises crucial questions about the future of AI, its practical uses and the sustainability of advanced models. Let’s explore this interesting topic that will surely set a trend in this newly released 2025.
The second era of the Escalation Laws
Historically, scaling laws have guided the development of AI models: more data, more parameters, and more computational power translated into better results.
However, this approach is showing diminishing returns. Researchers have begun to refer to this stage as “the second era of scaling laws,” in which traditional methods are no longer sufficient to generate significant advances.
Instead, a new strategy known as “inference-time scaling” or “test-time scaling” is emerging.
Inference Time Scaling: What is it and how does it work?
Inference time scaling consists of increasing the use of computational resources in the inference phase, that is, when the model responds to a request.
In the case of OpenAI’s o3, this means using more chips, more powerful chips or working on them for longer periods to generate responses.
The initial results of this approach are impressive. For example, the o3 model scored 88% in the ARC-AGI benchmark, a test intended to evaluate advances towards artificial general intelligence (AGI).
In comparison, OpenAI’s previous model, o1, only achieved 32%. Additionally, on a complex math test, the o3 achieved 25%, well above the 2% of other models.
However, these achievements come at a considerable cost. To achieve the highest score on ARC-AGI, the o3 used more than $1,000 in computing resources per task, compared to $5 per task for the o1 model.
The high-efficiency version of the o3, which sacrificed only 12% performance, used 170 times less computing, but still represents a significant cost compared to older models.
What are the causes of these excessive costs?
First of all, we have the dependency on massive computing. The increase in the use of chips and the need for more powerful processors directly increases operating costs.
For example, o3’s high-performance tests used more than $10,000 worth of resources for a single task.
Marginal improvements in AI models now require disproportionate investments in infrastructure and time. The difficulty in improving performance is because the easy advances have already been achieved and the next improvements demand increasingly sophisticated approaches.
Finally, we must not rule out the possibility of inefficiencies in implementation. Although scaling inference time produces results, it also increases calculation times. In some cases, it takes 10 to 15 minutes for the model to produce an optimal response.
Sometimes, o3 is like killing flies with cannons
Given the magnitude of the costs, models like the o3 do not seem to be designed for day-to-day tasks. Using it for most common tasks, such as composing messages or summarizing texts, is like killing flies with a cannon.
While tools like GPT-4o or Google Gemini 1.5 can answer simple questions quickly and efficiently, o3 focuses oncomplex queries that require a high level of reasoning.
For example, the model could be useful for institutions with deep pockets, such as research teams or corporations seeking to solve specific, highly relevant problems.
However, even in these cases, o3 still has limitations. According to François Chollet, creator of the ARC-AGI benchmark, o3 is not AGI and still fails on simple tasks that a human could easily solve.
Furthermore, like most language models, o3 still suffers from problems with hallucinations, that is, providing incorrect but convincing-sounding answers.
Solutions and the future of scaling
To address cost and efficiency challenges, researchers are exploring several strategies:
- Development of more efficient chips: Companies such as Groq, Cerebras and MatX are working on creating processors that optimize performance during inference. These chips could dramatically reduce the costs associated with inference-time scaling.
- Algorithm improvements: Optimizing AI algorithms could allow models to generate similar results using fewer resources. This would include techniques to dynamically adjust the level of computation needed depending on the complexity of the task.
- Integrating hybrid approaches: Jack Clark, co-founder of Anthropic, suggests that combining traditional scaling and inference-time scaling could maximize returns. This would allow the strengths of both approaches to be leveraged for specific tasks.
- Cost reduction through economies of scale: As these technologies become popular, economies of scale could emerge that reduce production and operating costs.
Is it worth it? Open questions about the development of AI
The future of scaling in AI raises important questions. Is it sustainable to continue increasing costs to achieve marginal improvements? What types of tasks would justify using models like o3?
Furthermore, how much further can AI advance without a fundamental change in the way we understand and develop these systems? The race for AI is far from over, but each step forward seems to require a higher price.
This post is also available in: