0:01
2026 will be a massive year for AI hardware, no matter if GPUs, ASICs or beyond.
2026 will be a massive year for AI hardware, no matter if GPUs, ASICs or beyond.
And I’m not sure we can still talk about an emerging industry.
Googles first TPU was released in 2015, over a decade ago and Nvidia will celebrate its 10-year AI anniversary next year.
Because Volta, the first Tensor Core GPU, launched in 2017. So, it’s settled!
Not a new industry anymore, but plenty of new chips.
And that’s what we will take a look at in this video.
Here are, in no particular order, the most interesting AI chips for 2026.
Let us know which chip you think will come out on top and leave a comment if we missed a chip, you think deserves to be in the next video.
Let’s start with a company and a chip that recently announced its first major deployment.
What was deployed, you ask?
Three-year-old hardware based on an almost six-year-old architecture.
I’m talking about Qualcomm finally deploying their AI100 chips at scale.
Even if the scale was small.
Installing a cluster of 1,024 AI100 chips doesn’t matter in 2026.
But what might matter is the new AI200 chip Qualcomm announced in October last year.
With about 70 billion transistors produced on TSMCs N3E and 768 gigabytes of LowPowerDDR5x memory, the AI200 ASIC is clearly designed for inference.
Betting on LPDDR5x, instead of the supply constrained HBM, might have been a great idea one year ago, but memory prices are skyrocketing across the board, and that includes LPDDR5x.
It’s unlikely that AI200 will make major waves this year.
But its successor AI250 is supposed to come with a new “compute near memory” architecture.
According to Qualcomm, that results in a 10x increase of effective memory bandwidth.
Paired with fast next-gen LPDDR6 memory, there might be something worthwhile on the horizon.
So don’t put on your party hats just yet, but do keep Qualcomm in mind.
If Nvidia is Goliath, this company would be David.
At least I’m sure that’s what AMD would like to hear.
AMD is the second major player in the GPU space and 2026 could be a turning point.
If MI455X and the Helios rack are on time.
MI455X is based on the new CDNA 5 architecture and packs 320 billion transistors into a mix of twelve 2nm and 3nm logic chiplets, connected via advanced 3. 5D packaging.
But the biggest selling point could be the massive memory setup.
MI455X comes with 432 gigabytes of next-gen HBM4 and with a bandwidth of almost 20 terabytes per second.
There’s much more to MI455X and the Helios rack than we have time for in this video.
If you want to know how AMD is challenging Nvidia, check out the SemiAnalysis article on AMD’s AI strategy.
It not only offers a deep-dive into MI455X, but also explains the Helios rack architecture.
It’s definitely worth a read; I’ll drop a link in the video description below.
Google is truly the grandfather of AI.
Not only was “Attention is all you need” basically a Google paper, but TPUs have been kicking it since 2015.
While everyone was hyped about blockchains, Google already prepared for AI.
Ironwood, aka TPUv7, is built on TSMCs N3E with very likely over 100 billion transistors across two large compute chiplets and comes with 192 gigabytes of HBM3e.
It’s specifically designed to run Gemini inference.
But Goole’s superpowers are the so-called Optical Circuit Switches.
These are tiny physical mirrors that allow for a fast and super-efficient optical interconnect.
Google uses them to create “Superpods” that connect up to 9,216 TPUs.
My further reading recommendation is, as usual, the SemiAnalysis TPUv7 deep-dive.
TPUs offer a whole different approach to AI workloads and have a very different TCO metric.
Link in the description below.
TPUs always have been at the forefront, but so far, they were only used by Google.
This changes with Ironwood.
The question is, if external customers can use them as well as Google can.
If so, Ironwood will be a heavyweight for efficient and low-cost inference.
Another chip that isn’t really brand new, but still hot for 2026, is Cerebras Wafer Scale Engine 3 or WSE-3.
It was launched back in 2024, but new clusters have just recently been announced.
WSE-3, as the name implies, uses an entire silicon wafer to create a massive chip.
It contains an unimaginable amount of SRAM.
No, not megabytes, I’m talking about 44 gigabytes.
And a theoretical memory bandwidth of 21 petabytes per second.
But remember, we are talking about an entire wafer.
Based on the slightly older N4P node, it contains a whopping 4 trillion transistors.
Everything is extraordinary with WSE-3.
But in 2026, 44 gigabytes aren’t really cutting it anymore. Even if it’s SRAM.
Still, a very unique concept and competitive when it comes to super-fast service.
While WSE-3 could be the odd one out in 2026, we might see an announcement of WSE-4. Fingers crossed.
Groq is an interesting one.
Not to be confused with Elon Musks Grok AI model.
The one I’m talking about has a q, not a k.
And that Groq with q got acquired by Nvidia.
We don’t know for sure if there will be a 2026 product or if it’s more of a foundation for the future.
But I still wanted to mention it, because it’s a very cool concept.
The Groq Language Processing Unit, or LPU for short, has relatively unspectacular specs.
55 billion transistors on a 14nm GlobalFoundry node.
And no external memory at all.
Nothing to write home about it seems.
But the LPU doesn’t need any of that.
It has 230 megabytes of ultra-fast SRAM placed super close to the compute cores.
The chip is designed in a way that every step, every calculation, runs like clockwork.
This is called deterministic execution and it basically eliminates the latency problems that GPUs have to deal with.
But because you have so little memory, you need to connect a lot of LPUs to run even smaller models.
Everything is a trade-off.
No matter what, Nvidia thought it was worth about $20 billion dollars.
And I’m excited to see what the 2nd gen LPU, based on Samsungs 4nm process, can do.
And what Nvidia does with it.
Ok, I can’t avoid the elephant in the room any longer.
Nvidia is number one for a reason.
And with Vera-Rubin, Nvidia is taking the next step, moving to TSMCs N3P and HBM4.
A single VR200 can crank out an insane 35 petaflops of FP4 per package.
Combined with 288 gigabytes of ultra-fast HBM4 at 22 terabytes per second of bandwidth, at least that’s the speed Nvidia is targeting, a single Superchip is already a force to be reckoned with.
But combining 72 of these in the NVL72 rack, using Nvidia’s exquisite NVLink scale-up network, Vera-Rubin will very likely once again top the charts. Or will it?
MI455X does have a memory advantage.
At least until Rubin Ultra, which will come with a clean terabyte of HBM.
No matter what, there’s no doubt that VR200 is the most anticipated release for 2026.
And it’s also the topic of the latest SemiAnalysis article, just in case you really want to know how Vera-Rubin is built and how it works.
Mark Zuckerberg is on a massive spending spree.
Meta is basically buying everything they can and that includes GPUs from Nvidia and AMD and even TPUs from Google.
Which makes it easy to forget that Meta has its own silicon.
The Meta Training and Inference Accelerator, MTIA for short, is now in its third iteration.
We don’t know all the specs yet, but it will be produced in TSMCs N3P with very likely over 100 billion transistors and HBM memory.
Previous MTIA version still used LPDDR5x.
Meta knows exactly what their internal workload needs.
And that’s not only AI chatbots, but AI recommendation models that run the algorithms for Facebook, Instagram and Threads.
MTIAv3 won’t be Metas chip for AGI, but it offers good margins for their business model.
And that means, they can use all the external hardware to train new models, while internally they switch to their own silicon for inference.
Looking at the roadmap, we will see a lot more MTIA in the future.
There is a lot of AI silicon, but only a few chips see truly large-scale deployment.
One of them is Amazon’s Trainium.
There are hundreds of thousands of Trainium2 chips deployed inside the AWS AI datacenters in Canton and New Carlisle.
Which makes Trainium3 a top contender for future large-scale deployment.
Trainium3, despite its name, combines training and inference capabilities in a single chip.
Built on TSMCs N3P, it consists of about 125 billion transistors and comes with 144 gigabytes of fast HBM3e.
You can find much more, including TCO numbers, in the SemiAnalysis Trainium3 deep-dive.
And yes, it’s an actual technical deep-dive.
Anthropic’s amazing Claude Code was trained and runs on Trainium, it can only get better with Trainium3.
And Anthropic isn’t the only fan.
OpenAI will use 2 gigawatts of Trainium compute starting this year.
If you are looking for an ASIC with truly large-scale deployment, look no further than Trainium3.
Microsoft’s first ASIC, Maia 100, was a little bit of a mystery.
Like every new in-house design, it has to prove itself first.
Does it warrant a second generation?
With Maia 200 we have the answer.
At around 825 square millimeters, the almost reticle busting chip contains 140 billion transistors, is manufactured in TSMCs N3P and comes with a pretty large 216 gigabytes of HBM3e.
Maia 200 will be used for inference and is optimized for FP8 and FP4, providing over 5 and 10 petaflops respectively.
Microsoft will use it for its in-house models but Maia 200 will also run future ChatGPT models.
More and more inference is moving away from Nvidia to custom silicon.
Last, but certainly not least, we have Intel, who is having another go at an AI GPU.
But it doesn’t seem like Jaguar Shores is targeting a 2026 release.
But with so many ASICs, I had to talk about one more GPU, even if it might be a 2027 product.
Jaguar Shores comes with a pretty strong spec sheet.
18A process node, 175 billion transistors and 288 gigabytes of HBM4.
On paper, it's competitive. But paper is patient.
After many attempts, Intel not only has to prove it can build and produce an AI GPU, but also provide the proper software support.
But no matter what, a third player in the GPU space would be very welcome.
And Jaguar Shore will show if Intel can compete again.
The chip wars are heating up and what we just talked about is only a small selection of the entire market.
If you compare this list with what the SemiAnalysis Accelerator & HBM model tracks, you will know what I mean.
Right now, all lines are only going up. But for how long?
If you are working in or with the industry and are interested in a clear picture view of the entire AI silicon market and where it’s headed, the SemiAnalysis Accelerator & HBM model is your best choice.
Not only does it offer a comprehensive overview of all current and future chips, it also provides shipment numbers and ASPs.
Plus, deep insights into the HBM market.
You can find the link in the description below.
Right next to all the deep-dive articles we talked about in the video.
And if you are less interested in spec, but more in real-world performance, InferenceX offers comprehensive benchmarks across a wide range of accelerators.
Because what really matters in the end is how a chip performs.
Check out the SemiAnalysis InferenceX dashboard, it’s free.
You know where you can find the link.
I hope you enjoyed this video and see you in the next one!