On December 20, 2024, the OpenAI platform presented its most recent model, known as “o3”, in a video announcement that generated great expectation.

This model has captured the attention of the technology sector due to a notable achievement: it outperformed humans for the first time in the ARC-AGI test, a test designed to measure the ability to reason and adapt to new challenges.

However, this advance raises questions about its architecture, cost and possible implications, especially given that there are rumors that a high-computation task with this model could cost up to $2,000.

What is ARC-AGI?

The ARC-AGI (“Abstraction and Reasoning Corpus for Artificial General Intelligence”) is a benchmark developed by François Chollet, a Google scientist, to measure adaptability and abstract reasoning in intelligent systems.

Instead of focusing on tasks based on memorized knowledge, like many other benchmarks, the ARC-AGI evaluates the models’ ability to acquire new skills.

The test consists of solving problems based on visual patterns. Each challenge presents a series of examples in the form of grids of colored pixels, followed by a similar question that the system must solve by completing the correct grid.

Humans, for reference, score an average of just over 75% on the ARC-AGI, while o3 achieved an astonishing 76%.

This represents a “turning point” in the capabilities of artificial intelligence, according to Chollet himself, who described this result as a “qualitative change.”

What makes o3 different?

One of the most intriguing features of the o3 is that its architecture appears completely different from previous models in the GPT series. Although OpenAI has not shared precise details, Chollet speculates that the system uses a “test-time intensive search” approach.

This could involve extensive Chains of Thought analysis, an emerging approach that allows models to break down their processes into intermediate steps.

o3’s performance also suggests extensive use of computing, possibly similar to the Monte Carlo search tree strategy used by AlphaZero, the famous DeepMind program that dominated chess.

This means that the model not only processes information, but also explores different possibilities before arriving at a solution, something that has not been seen in previous GPT models.

Is this a step towards AGI?

The concept of artificial general intelligence (AGI) refers to a level of machine intelligence capable of equaling or surpassing human capabilities in any cognitive task.

Although o3’s success in ARC-AGI is significant, Chollet emphasizes that it should not be confused with having achieved AGI. “O3 still fails at very simple tasks that humans solve easily,” he noted, highlighting the model’s current limitations.

The ARC-AGI is a research tool, not a definitive test of general intelligence. Still, the o3 results represent an important step forward toward systems that are more intelligent and humane in their behavior.

Cost and computational resources

One of the most debated aspects about o3 is the cost associated with its operation. OpenAI has not revealed how many computational resources were needed to train and run the model on the ARC-AGI.

However, Chollet hinted that o3’s approach could be considered a form of brute force, in that it uses massive amounts of computing to solve relatively simple problems.

This processing-intensive method comes with a high price tag, with compute-intensive tasks potentially costing up to $2,000 each.

This raises questions about the sustainability and efficiency of these types of models. While the results are impressive, the cost in terms of energy consumption and calculation time could limit its practical large-scale application.

Furthermore, reliance on large amounts of data and resources could deepen inequalities in access to technology.

Only organizations with significant budgets could afford to develop and maintain models like o3, which could further centralize technological power.

When can we try o3?

OpenAI plans to release a “mini” version of o3 in late January 2025, with a full version scheduled for later.

Although this initial version may not include all of the capabilities demonstrated in the ARC-AGI, it will provide an opportunity for researchers and developers to explore the model’s potential.

The development of o3 marks an important advance in artificial intelligence, showing adaptation and reasoning capabilities that previously seemed unattainable.

However, it also raises fundamental questions about the costs, both economic and ethical, of pursuing these types of technologies.

Is this the path to AGI or simply a demonstration of how difficult and expensive it is to overcome current limits? The debate is served.

This post is also available in: Español Français Русский Italiano