From personalizing experiences on mobile devices to creating advanced language models, artificial intelligence (AI) is transforming every aspect of modern technology.
These models power the most sophisticated chatbots, taking interaction with technology to a new level. However, for AI to be closer to us, it must work on all our devices with reasonable power consumption.
A fundamental pillar of this advance is the development of Neural Processing Units (NPUs), specialized chips that optimize AI workloads.
These NPUs are redefining the architecture of processors in both PCs and mobile devices, improving efficiency and performance. Let’s see what those devices of tomorrow will be like, with integrated NPUs.
What is an NPU and how does it work?
An NPU (Neural Processing Unit) is a specialized processor designed to execute artificial intelligence tasks efficiently.
Unlike CPUs and GPUs, which traditionally handle multiple types of computational operations, NPUs focus exclusively on processing AI workloads, especially those related to neural networks.
NPUs excel at matrix multiplication, a critical operation for most AI applications, such as training machine learning models or inference in neural networks.
This focus on parallelization and optimization of power usage allows NPUs to perform large amounts of calculations with significantly lower power consumption than CPUs or GPUs.
A GPU can perform similar operations, but with greater energy consumption. This makes them less suitable for devices where energy efficiency is key, such as laptops and smartphones.
The evolution of NPUs
The concept of an NPU is not entirely new. In fact, its origin can be traced back to early attempts to develop neuromorphic computing, which sought to emulate the architecture of the human brain and nervous system to improve processing efficiency.
Although many of these early attempts failed, some led to more viable technologies, such as DSPs (Digital Signal Processors), which over time have been adapted for AI applications.
Companies like Qualcomm and Xilinx were pioneers in this transition. Qualcomm, for example, has been refining its Hexagon DSP, transforming it over time into what is known today as Hexagon NPU, capable of offering performance of up to 45 TOPS (trillion operations per second) in AI.
This computing power allows devices equipped with NPUs to execute complex AI tasks, such as voice and image recognition, locally and without depending on cloud computing.
This is a crucial advance in the race to achieve greater independence on mobile devices and PCs.
Advantages of NPUs over CPUs and GPUs
One of the main problems facing the AI industry today is the dependence on cloud processing.
Although cloud computing offers almost unlimited processing power, it also raises issues with latency, bandwidth consumption, and potential security risks.
NPUs allow many of these AI calculations to be performed directly on the device, which not only reduces latency but also improves security by preventing data from leaving the device.
Additionally, NPUs are much more energy efficient. Although GPUs are popular for cloud AI, due to their parallelization capabilities, they are extremely power demanding.
NPUs, on the other hand, achieve similar performance for certain AI tasks, but with significantly lower energy consumption. This is key for portable devices like laptops and smartphones, where battery life is a priority.
NPUs in the new processors: Intel, Qualcomm and AMD in the lead
As demand for artificial intelligence continues to grow, leading processor manufacturers are integrating NPUs into their new architectures. Intel, Qualcomm and AMD are three of the main players in this race.
Intel Meteor Lake and Lunar Lake
Intel has been one of the pioneers in integrating NPUs into its processors. With the launch of its Meteor Lake architecture, Intel has introduced a “hybrid 3D performance architecture” that includes not only a CPU and GPU, but also a dedicated NPU.
The function of this NPU is to handle AI tasks without burdening the CPU or GPU, improving overall system performance and efficiency.
Meteor Lake offers 11 TOPS performance in AI, making it an attractive solution for PCs looking to take advantage of artificial intelligence without resorting to the cloud.
In addition, Intel is already working on the next generation of processors, known as Lunar Lake, which promise 100 TOPS of AI performance, further consolidating its position in the race to dominate the AI market in PCs.
Qualcomm: Snapdragon X Elite
Qualcomm, which has been working for years on the integration of AI in its SoCs (systems on chip) for mobile devices, has gone one step further with its Snapdragon X Elite, a chip based on ARM architecture that has an NPU capable of offering 45 TOPS.
The Snapdragon
AMD: Ryzen with AI
AMD has also joined this race with its Ryzen processors. With the addition of its XDNA NPU to the Ryzen 8040 series, AMD delivers 16 TOPS of AI performance.
However, the company has already announced that its future processors, such as Strix Point, will be equipped with NPUs capable of reaching 40 TOPS, making them direct competitors to Qualcomm and Intel.
Copilot+ PCs: the future of AI in Windows
One of the most exciting moves in this evolution is the development of Copilot+ PCs by Microsoft.
These devices, equipped with high-performance NPUs, are part of a new class of hardware designed specifically to execute intensive AI tasks locally.
Windows 11 is already optimized to take advantage of NPUs, allowing users to access advanced AI features without relying on the cloud.
Featured features of Copilot+ PCs include:
- Windows Studio Effects: A series of AI-accelerated effects, such as background blur, voice focus, and eye contact, that enhance the video conferencing experience.
- Super Resolution: Technology that uses AI to improve the visual quality in video games, making them faster and improving graphics.
- Co-creation with Paint: A new feature that allows users to transform images into AI-generated art directly in the Microsoft Paint program.
Microsoft has also included a Copilot button on some devices, which facilitates quick access to AI functions, allowing users to perform tasks such as real-time translations or generate images without having to resort to external services.
The expansion of NPUs to other operating systems
Although Microsoft has been a pioneer in integrating NPUs into its Windows ecosystem, this trend is likely to spread to other operating systems in the near future.
macOS, for example, already has a Neural Engine in its M1, M2 and M3 chips, although its AI performance is currently lower than that of Qualcomm and AMD NPUs.
However, as demand for local AI computing continues to grow, we will almost certainly see more devices from Apple and other companies adopting more powerful NPUs.
In short, NPUs represent the next big leap in processor evolution, and their integration into PCs, mobile devices and other systems will usher in a new era for local AI computing.
With the arrival of processors like Intel’s Lunar Lake and the moves of the competition, we are about to witness a revolution in how we use AI in our daily lives.
This post is also available in: