Artificial intelligence (AI) is everywhere, from medicine to automotive, and from creating original texts to videos that seem like a dream (or even a nightmare).
However, for these AI applications to work efficiently, they require specialized hardware capable of handling the intensive parallel processing and memory demands.
For this reason, specialized AI chips have been developed, optimized to run machine learning algorithms and neural networks with unprecedented speed and efficiency.
Do you want to know how they work, or find out what the most modern AI chips are, what they can do and how much they cost? Below, we will examine all this and much more in detail.
What are AI chips?
Artificial intelligence (AI) chips are integrated circuits specifically designed to handle the complex computational tasks required by AI algorithms.
These chips are optimized to process large volumes of data at high speed and perform complex calculations efficiently, which is essential for applications such as speech recognition and natural language models (LLMs).
Although there are many types of AI chips, they all share one fundamental characteristic: parallel processing. This means that they can run multiple calculations simultaneously, and with the lowest possible energy consumption.
How do AI chips work?
When AI was an emerging technology it had to adapt to existing hardware. For that reason, the first AI applications took advantage of graphics cards, and specifically their processors; the GPUs.
GPUs (Graphics Processing Units)
Originally, GPUs were developed for graphics processing in video games and visual applications. Something they continue to do on our computers.
But as AI began to demand more processing power, GPUs proved to be ideal for training deep neural networks due to their parallel architecture.
In the early days, frameworks like Nvidia’s CUDA played a key role in allowing AI developers to take full advantage of the power of GPUs.
CUDA enabled parallel programming on Nvidia GPUs, which greatly accelerated the training of AI models. It also cemented Nvidia’s current hegemony in this field.
TPUs (Tensor Processing Units)
TPUs were introduced by Google in 2016 in response to the growing demand for processing for AI tasks, particularly in the training and inference of deep neural networks.
A TPU is designed to perform matrix processing, which involves decomposing input data into multiple tasks called vectors.
This allows TPUs to solve large mathematical tasks much faster and with less power consumption than traditional processors. TPUs can be 30 to 80 times more efficient per watt than current CPUs and GPUs.
Nvidia includes similar technology with its TensorCores, which are in both its RTX graphics cards and GPUs designed for data centers.
NPUs (Neural Processing Units)
Unlike GPUs and TPUs, NPUs are specialized in inference and executing neural networks with low latency and high performance on power-constrained devices.
For this reason, NPUs are found in mobile devices, such as smartphones and cameras, as well as in some advanced processors for autonomous vehicles.
For example, both Apple’s mobile processors (since the A11 Bionic of the iPhone 8) and computer processors (Apple Silicon M1/M2) include NPUs, which are known there as Neural Engine.
The same can be said of Qualcomm’s “top of the range” since the Snapdragon 855. For example, the Snapdragon 8 Gen 3 has a Hexagon NPU, which improves its performance in AI tasks.
How is performance measured in AI chips?
To measure performance and compare different AI chips, several metrics and techniques have been used that evaluate the ability of chips to handle specific AI tasks.
Initially, FLOPS were used, which are floating point operations per second. This measure gave an idea of the raw computing power of one or more CPUs, but does not focus on specific AI tasks.
That’s why we started using TOPS (tensor operations per second). As we have already described, these operations are fundamental in the training and inference of AI models, and that is why they are the most used performance metric.
Another important metric is latency, which measures the time it takes for a chip to complete a specific task. This is crucial in real-time applications, such as speech recognition.
There are specific benchmarks for AI, such as MLPerf, DAWNBench or SPEC AI (among others). These are a set of standard tasks that allow you to measure hardware and software performance on AI tasks.
What are the most powerful AI chips?
Below are some of the most advanced AI chips available today. Some of them can be purchased (if you have enough), but others can only be accessed in the cloud.
Google TPU v4 Pods
Designed specifically for deep learning tasks, these TPU clusters can include up to 4,096 TPU v4 chips per pod. With a capacity of up to 275 TFLOPS (teraFLOPS) per chip, TPU v4 Pods are extremely powerful.
This massive performance is ideal for large-scale AI model training and advanced research.
NVIDIA H100
This is NVIDIA’s latest and most powerful GPU, designed for high-performance computing and AI tasks. With a cost of up to $35,000 euros per unit, it has 14592 CUDA Cores and 456 Tensor Cores.
This gives the Nvidia H100 up to 700 TFLOPS of AI performance, significantly outperforming its predecessors (such as the A100).
The H100 is ideal for data centers, supercomputing and advanced research, offering exceptional performance in AI model training and inference tasks.
Intel Havana Gaudi2
Habana Gaudi2 is Intel’s bet to steal a piece of the pie from Nvidia. This is an AI accelerator designed to optimize performance in AI model training tasks.
This chip has 24 tensor processing cores (TPCs). In terms of performance, the Gaudi2 offers performance of up to 32 TFLOPS per chip. This is up to 2.5 times higher than its predecessor
The Gaudi2 is very powerful and is mainly used in data centers. The price of this chip may vary depending on the supplier.
AMD Instinct MI325X
The MI325X designed to be used in data centers and supercomputing applications, where high performance and processing capacity is required.
With just over 2 months on the market, the data available on the MI325X places it in the same segment as the Nvidia H100. Specifically, it offers 1.3 times the performance of the H100.
These chips represent the cutting edge of specialized AI hardware. What is clear is that these are highly specialized products, and accessing them directly is not an easy task.
This post is also available in: