OpenAI has launched GPT-4o mini, its most efficient model to date, designed to make artificial intelligence more accessible and affordable. This model is predicted to be a fundamental component within the OpenAI collaborative ecosystem.
According to the company, this model costs $0.15 per million input tokens and $0.60 per million output tokens, being more than 60% more efficient than GPT-3.5 Turbo.
Starting July 19, 2024, all ChatGPT users, whether on the Free, Plus, or Team plans, use this model instead of GPT-3.5. Enterprise users will have it available starting next week.
What can we expect from the new model? What does it mean that it is more economical than its predecessors? How do you measure the efficiency of a language model? We will discuss all this and more below.
GPT-4o mini: the “small” of the family
OpenAI has introduced GPT-4o mini as a “small” language model that maintains high performance, while significantly reducing computational cost and energy consumption.
This model promises to be more accessible and sustainable, which could facilitate its adoption and use in a variety of applications.
According to benchmark results released by OpenAI, GPT-4o mini is positioned as a more efficient alternative to large models like GPT-4o, maintaining competitive performance in natural language processing tasks.
GPT-4o mini is presented in a market where other small models such as Microsoft Phi-3 mini and Google Gemini Nano are already present. Despite seeming like a late arrival, the efficiency and power of the GPT-4o mini can surpass full versions of previous generations of models from other suppliers.
What does it mean that GPT-4o mini is a “small” model?
A Large Language Model (LLM) like GPT is described as “small” or “large” primarily based on the number of parameters it has. These are parts of the model that are adjusted during training to improve performance on specific tasks.
In general, a greater number of parameters means a better ability of the model to understand and generate natural language effectively. But it also means that the model needs more computational power (and therefore more energy) to run.
GPT-4o mini has significantly fewer parameters compared to models like GPT-4 or GPT-3.5. This makes it less computationally intensive and faster to run compared to larger models.
How good is GPT-4o mini relative to the competition?
According to benchmarks conducted by OpenAI, GPT-4o mini scores 82% in the MMLU benchmark, outperforming GPT-4 in chat preferences and other small models in reasoning and multimodal tasks.
Supports a 128K token context window and up to 16K output tokens per request. While its low latency makes it ideal for applications that require multiple model calls, large volumes of context, or fast real-time interactions.
It currently supports text and vision, with future capabilities for text, image, video and audio. Regarding its multimodality, GPT-4o mini is superior in textual and multimodal reasoning, with outstanding performance in mathematics and coding.
How are LLMs and multimodal models measured and compared?
There are tools and methods used to test and measure the performance of generative AI models, such as text and image generation models.
These are essential to understanding the effectiveness, accuracy and creativity of these systems. LLM benchmarks are standard datasets and tasks widely adopted by the research community to evaluate and compare the performance of various models.
These are used to evaluate how well an AI model can generate coherent, relevant and creative content, including text, images, music and others.
They also allow different AI models to be compared to identify which ones are more efficient or creative in specific tasks. By testing models on various benchmarks, it is possible to identify biases and limitations, such as the ability to generate text in multiple languages.
Additionally, benchmarks provide essential feedback during AI model development, helping to refine and improve the algorithm. Some of the most well-known and popular benchmarks for LLMs are HellaSwag, MMLU and TruthfulQA.
Frequently Asked Questions
Here are some frequently asked questions related to measuring the efficiency of large language models (LLM) and about the new GPT-4o mini:
What does it mean for a model to be more efficient or more “economical”?
It means that the model can generate results comparable to larger, more expensive models, but using fewer computational and energy resources. This also has an effect on the average time it takes for the model to generate the response, or latency.
How does GPT-4o mini compare to GPT-4?
GPT-4o mini is a more efficient version of GPT-4o, designed to offer similar performance with a lower computational cost. All this achieving a performance just inferior to the largest and most complex model.
What is the advantage of “small” language models, like GPT-4o mini?
Small models like GPT-4o mini allow for further democratization of AI technology. Having fewer parameters means the model can generate answers more quickly, which is crucial for real-time applications.
Additionally, training and deployment costs are lower, making the technology more viable for a broader range of commercial and academic applications.
Finally, lower energy consumption contributes to a reduction in environmental impact, which is an important consideration in the implementation of large-scale technologies.
GPT-4o mini and multiple API calls
This new AI from OpenAI supports multiple API calls, which translates into low latency and the possibility of performing several tasks in parallel.
Currently this lighter and cheaper version supports text and vision through the OpenAI API and in the near future it will incorporate input and output of text, images, video and audio.
Is the new GPT-4o mini model free?
Well of course yes, in fact it has come to replace GPT-3.5 in the free versions. In fact compared to GPT-3.5 turbo it beats it in all areas, so it is great news for free users.
Until what date is GPT-4o mini trained?
This free version of ChatGPT has been trained with data dating back to October 2023, according to Sam Altman, the CEO responsible for OpenAI, the company will develop this new language model.
GPT-4o mini, a new safer model
Security is also worked on in the creation of each model and this new AI model with multimodal reasoning filters more hate speech, adult content or content related to spam sites, at the same level as GPT-4o.
The GPT-4o mini API is the first to apply the OpenAI instruction hierarchy that helps improve the model’s ability to resist information leaks, message injections, and system message extractions.
This post is also available in: