Online chatbots, such as OpenAI’s ChatGPT and Google’s Gemini, often face difficulties solving simple mathematical problems and generate buggy or incomplete computer code.
To overcome these challenges, the OpenAI intelligent platform has launched a new model that promises significant advances in reasoning and the ability to tackle complex tasks.
This new model, called “o1”, is still recent and has not been widely evaluated. However, it has shown signs that it could represent a turning point in the evolution of artificial intelligence.
What is OpenAI o1?
o1 is a new artificial intelligence model introduced on September 12, 2024 by OpenAI, designed to improve reasoning and problem-solving capabilities in complex tasks.
Unlike its predecessors, such as GPT-4o, this model does not simply answer questions immediately, but rather “reflects” on problems, breaking them down into smaller steps and seeking solutions more methodically.
This approach represents a major shift in the traditional dynamics of AI models, which until now have struggled in tasks that require deep logical analysis, such as solving mathematical problems or accurately generating code.
##How does OpenAI o1 work?
The core of OpenAI o1 lies in a training method known as reinforcement learning.
This process allows the model to learn through trial and error, honing its skills as it repeats and evaluates thousands of different problems and situations.
In the context of mathematics, for example, the model works on several complex problems and tests different solutions, identifying patterns in the methods that lead to the correct answer.
In this way, it improves its ability to solve mathematical and scientific problems with greater precision than previous versions of AI.
The reinforcement learning process takes place over weeks or even months. During this time, the model is exposed to an enormous amount of data and problems, and learns from its mistakes and successes.
However, it is important to note that although OpenAI o1 can offer more accurate answers than its predecessors, it is still susceptible to making errors and generating “hallucinations”, that is, incorrect or fabricated answers.
Reason through the chain of thought
One of the most notable advances of OpenAI o1 is its ability to apply what is known as chain of thought, a technique that allows the model to break down a problem into smaller, more readable steps.
This approach not only improves the accuracy of the model, but also makes it more transparent. By observing how the model reaches a conclusion, users can have a better understanding of its reasoning process.
This ability is particularly useful in coding and mathematics tasks, where decomposing a problem into smaller subproblems is crucial to arriving at the correct solution.
OpenAI o1 has demonstrated this ability in several tests, including one in which the model was able to diagnose a disease based on a detailed report of a patient’s symptoms, highlighting its potential in medical applications.
##How does o1 compare to other AI models?
One of the main questions about OpenAI o1 is how it compares to other existing AI models, both from OpenAI and other technology companies.
In recent tests shared by OpenAI, o1 outperformed GPT-4 in the ability to solve mathematical problems and generate accurate code, two areas where GPT-4 still showed limitations.
For example, in the qualification exam for the International Mathematics Olympiad (IMO), GPT-4 could only get 13% correct, while OpenAI o1 achieved 83%.
This considerable jump demonstrates o1’s ability to approach problems in a more logical and methodical way.
Code generation with OpenAI o1
Regarding coding, OpenAI o1 has also shown a significant improvement.
Solving the problems of the 2024 International Olympiad in Informatics (IOI), the model scored 213 points and ranked in the 49th percentile, putting it in the middle of human competitors.
This is notable, given that IOI participants are some of the best programmers in the world.
Additionally, when competition restrictions were relaxed, the model was able to improve its performance even further, scoring above the threshold for a gold medal.
Models like Google’s Gemini and Meta’s open-source developments (like Llama 3.1) have also made progress in improving reasoning in AI, as we saw recently with the Reflection 70B model.
However, the advantage of OpenAI o1 lies in its deliberate approach towards problem solving and chain of thought integration, giving it an advantage in complex tasks that require multiple steps of reasoning.
Is it dangerous to have an AI that can reason?
An AI with reasoning capabilities can address complex problems in fields such as medicine, engineering and science, offering innovative and efficient solutions.
However, it is also true that advanced AI could be used maliciously if it falls into the wrong hands. It is therefore essential to implement robust security measures to prevent abuse.
According to OpenAI, a crucial aspect in the development of o1 was security. OpenAI has implemented behavioral policies in the model’s chain of thought to ensure that it follows human values and principles in potentially dangerous contexts.
This approach has improved the model’s ability to safely reject inappropriate requests, such as encouraging violent or illegal behavior.
In the most challenging tests, still according to OpenAI, o1 achieved a substantial improvement in identifying and handling sensitive situations compared to GPT-4.
One of the most intriguing challenges is how to integrate this chain of thought in a way that is useful for monitoring without compromising the user experience.
OpenAI has decided not to show the raw chain of thought to users, but the model can generate a summary explaining how it arrived at an answer. This allows a balance between transparency and functionality.
How to test OpenAI o1?
OpenAI o1 is now available for testing, but access is limited. Users who have subscriptions to the ChatGPT Plus and ChatGPT Teams services can start using this new technology.
Additionally, OpenAI is also offering o1 to companies and software developers who are interested in integrating it into their AI applications.
However, general access to the o1 model is not yet widely available to all users, as OpenAI is initially focusing on offering it to subscribers and commercial partners.
This post is also available in: