The new compact models of Mistral, presented to the public at the beginning of the week, stand out for their ability to work effectively on everyday devices, such as laptops and smartphones.
This trend towards smaller, more optimized models is revolutionizing the AI landscape, allowing developers and scientists to access powerful tools without the need for high-end infrastructure.
In this article, we will explore in detail Mistral’s new models, NVIDIA’s previous innovation with the same technology, and the growing inclination towards running AI locally.
Ministerial: Power within everyone’s reach
Recently, Mistral AI has announced the launch of its new AI models, Ministral 8B and Ministral 3B, designed to run on consumer devices such as laptops and phones.
The French company has designed these models based on the Mistral 7B, presented last year. The purpose is to optimize memory usage and response speed, especially in applications that require real-time interactions.
The essence of these new models lies in their ability to bring AI to a wide variety of devices.
This allows small businesses, researchers and users in general to take advantage of the advantages of artificial intelligence without having to invest in expensive servers or cloud infrastructure.
By running these models locally, users can benefit from greater control over data and lower latency in responses, which is crucial in situations where processing speed is critical.
NVIDIA and the push towards miniaturized models
The launch of the small Mistral models comes shortly after NVIDIA presented its Mistral-NeMo-Minitron 8B in August 2024. This model represents an effective miniaturization of the Mistral NeMo 12B, which had been previously revealed in collaboration with Mistral AI.
Through techniques such as pruning and distillation, NVIDIA managed to reduce the number of parameters from 12 billion to 8 billion, while maintaining an accuracy comparable to that of the original model.
Bryan Catanzaro, vice president of applied research at NVIDIA, highlighted at the time that this optimization not only makes it easier to run AI on GPU-powered workstations, but also opens the door to broader access for organizations with limited resources.
Small models allow companies and developers to implement generative AI capabilities into their infrastructures, optimizing costs and operational efficiency.
Additionally, small models like the Mistral-NeMo-Minitron 8B are ideal for running locally, preventing sensitive data from being sent to external servers and providing an additional level of security and privacy.
This trend towards local AI is becoming increasingly attractive, especially in fields such as healthcare, where handling personal data is critical.
The trend towards small models and local execution of AI
In recent years, multiple technology companies and research labs have developed “open weights” versions of language models that can be downloaded and run on consumer hardware.
This has allowed researchers and developers to use advanced AI tools without relying on cloud services, which also gives them greater control over their projects.
For example, the Ollama platform allows users to download open source models, such as Llama 3.1, Phi-3, Mistral and Gemma 2, and access them through a command line interface.
This ability to run models locally is especially valuable for those seeking to preserve data privacy or working in environments with limited connectivity.
The ability to run AI on local devices ensures that users do not have to compromise the confidentiality of information, a critical aspect in fields such as medicine and biotechnology.
Additionally, the trend toward smaller, more efficient models is helping to democratize access to AI.
Researchers and developers around the world are using models like Alibaba’s Qwen and Meta’s Llama to create custom applications that respond to specific needs in their respective fields.
Use cases and advantages of local AI
Deploying small AI models and running locally offers numerous advantages. In the scientific field, researchers can use models to process data and perform analysis without depending on third-party services.
This translates into faster access to results and greater work efficiency.
A prominent example is the use of local models to extract diagnoses from medical reports, allowing scientists and doctors to analyze data without compromising patient privacy.
On the other hand, local execution also allows customization of models for specific tasks. For example, a researcher in New Hampshire tuned a Qwen model to help summarize scientific articles and write manuscripts.
This level of adaptation would not be possible with cloud-based models, where updates and changes to the model can affect the performance and consistency of results.
In short, the trend toward small language models and the ability to run them locally is transforming the way researchers and developers use artificial intelligence.
This post is also available in: