AI Chip & Silicon Round-up 2026

0:01

2026 will be a massive year for AI hardware,  no matter if GPUs, ASICs or beyond.

0:01

And I’m not sure we can still talk about an emerging  industry.

0:09

Googles first TPU was released in 2015, over a decade ago and Nvidia will celebrate its  10-year AI anniversary next year.

0:16

Because Volta, the first Tensor Core GPU, launched in 2017. So, it’s settled!

0:24

Not a new industry anymore, but plenty of new chips.

0:32

And that’s what  we will take a look at in this video.

0:37

Here are, in no particular order, the  most interesting AI chips for 2026.

0:43

Let us know which chip you think will come out  on top and leave a comment if we missed a chip, you think deserves to be in the next video.

0:49

Let’s start with a company and a chip that recently announced its first major  deployment.

0:54

What was deployed, you ask?

0:59

Three-year-old hardware based on an  almost six-year-old architecture.

0:59

I’m talking about Qualcomm finally deploying their AI100  chips at scale.

1:05

Even if the scale was small.

1:12

Installing a cluster of 1,024 AI100 chips  doesn’t matter in 2026.

1:12

But what might matter is the new AI200 chip Qualcomm  announced in October last year.

1:18

With about 70 billion transistors produced on TSMCs  N3E and 768 gigabytes of LowPowerDDR5x memory, the AI200 ASIC is clearly designed  for inference.

1:33

Betting on LPDDR5x, instead of the supply constrained HBM,  might have been a great idea one year ago, but memory prices are skyrocketing across  the board, and that includes LPDDR5x.

1:51

It’s unlikely that AI200 will make major  waves this year.

1:51

But its successor AI250 is supposed to come with a new “compute near memory”  architecture.

1:57

According to Qualcomm, that results in a 10x increase of effective memory bandwidth.

2:02

Paired with fast next-gen LPDDR6 memory, there might be something worthwhile on the horizon.

2:10

So don’t put on your party hats just yet, but do keep Qualcomm in mind.

2:15

If Nvidia is Goliath, this company would be David.

2:20

At least I’m sure that’s what AMD  would like to hear.

2:20

AMD is the second major player in the GPU space and 2026 could be a turning  point.

2:26

If MI455X and the Helios rack are on time.

2:35

MI455X is based on the new CDNA 5 architecture  and packs 320 billion transistors into a mix of twelve 2nm and 3nm logic chiplets, connected via  advanced 3. 5D packaging.

2:43

But the biggest selling point could be the massive memory setup.

2:51

MI455X  comes with 432 gigabytes of next-gen HBM4 and with a bandwidth of almost 20 terabytes per second.

3:01

There’s much more to MI455X and the Helios rack than we have time for in this video.

3:08

If you  want to know how AMD is challenging Nvidia, check out the SemiAnalysis article on AMD’s AI  strategy.

3:13

It not only offers a deep-dive into MI455X, but also explains the Helios rack  architecture.

3:20

It’s definitely worth a read; I’ll drop a link in the video description below.

3:26

Google is truly the grandfather of AI.

3:26

Not only was “Attention is all you need” basically  a Google paper, but TPUs have been kicking it since 2015.

3:38

While everyone was hyped about  blockchains, Google already prepared for AI.

3:44

Ironwood, aka TPUv7, is built on TSMCs N3E with  very likely over 100 billion transistors across two large compute chiplets and comes with 192  gigabytes of HBM3e.

3:53

It’s specifically designed to run Gemini inference.

4:00

But Goole’s superpowers  are the so-called Optical Circuit Switches.

4:00

These are tiny physical mirrors that allow for a fast  and super-efficient optical interconnect.

4:07

Google uses them to create “Superpods” that connect up to  9,216 TPUs.

4:13

My further reading recommendation is, as usual, the SemiAnalysis TPUv7 deep-dive.

4:20

TPUs offer a whole different approach to AI workloads and have a very different TCO  metric.

4:26

Link in the description below.

4:31

TPUs always have been at the forefront, but so  far, they were only used by Google.

4:31

This changes with Ironwood.

4:37

The question is, if external  customers can use them as well as Google can.

4:43

If so, Ironwood will be a heavyweight  for efficient and low-cost inference.

4:49

Another chip that isn’t really brand new, but  still hot for 2026, is Cerebras Wafer Scale Engine 3 or WSE-3.

4:55

It was launched back in 2024, but  new clusters have just recently been announced.

5:01

WSE-3, as the name implies, uses an entire silicon  wafer to create a massive chip.

5:01

It contains an unimaginable amount of SRAM.

5:08

No, not megabytes,  I’m talking about 44 gigabytes.

5:08

And a theoretical memory bandwidth of 21 petabytes per second.

5:16

But  remember, we are talking about an entire wafer.

5:22

Based on the slightly older N4P node, it  contains a whopping 4 trillion transistors.

5:28

Everything is extraordinary with WSE-3.

5:28

But in 2026, 44 gigabytes aren’t really cutting it anymore. Even if it’s SRAM.

5:35

Still, a very unique concept and competitive when it comes to super-fast service.

5:41

While WSE-3 could be the odd one out in 2026, we might see an announcement  of WSE-4. Fingers crossed.

5:51

Groq is an interesting one.

5:51

Not to be confused  with Elon Musks Grok AI model.

5:51

The one I’m talking about has a q, not a k.

5:57

And that Groq with q  got acquired by Nvidia.

5:57

We don’t know for sure if there will be a 2026 product or if it’s more of  a foundation for the future.

6:04

But I still wanted to mention it, because it’s a very cool concept.

6:10

The Groq Language Processing Unit, or LPU for short, has relatively unspectacular specs.

6:15

55  billion transistors on a 14nm GlobalFoundry node.

6:23

And no external memory at all.

6:23

Nothing to write home about it seems.

6:28

But the LPU doesn’t need any of that.

6:28

It has 230 megabytes of ultra-fast SRAM placed super close to the compute cores.

6:35

The  chip is designed in a way that every step, every calculation, runs like clockwork.

6:40

This is  called deterministic execution and it basically eliminates the latency problems that GPUs have to  deal with.

6:47

But because you have so little memory, you need to connect a lot of LPUs to run even  smaller models.

6:53

Everything is a trade-off.

6:59

No matter what, Nvidia thought it was worth  about $20 billion dollars.

6:59

And I’m excited to see what the 2nd gen LPU, based on Samsungs 4nm  process, can do.

7:04

And what Nvidia does with it.

7:12

Ok, I can’t avoid the elephant in the room  any longer.

7:12

Nvidia is number one for a reason.

7:17

And with Vera-Rubin, Nvidia is taking the  next step, moving to TSMCs N3P and HBM4.

7:24

A single VR200 can crank out an insane 35  petaflops of FP4 per package.

7:24

Combined with 288 gigabytes of ultra-fast HBM4 at  22 terabytes per second of bandwidth, at least that’s the speed Nvidia is targeting, a  single Superchip is already a force to be reckoned with.

7:44

But combining 72 of these in the NVL72 rack,  using Nvidia’s exquisite NVLink scale-up network, Vera-Rubin will very likely once again top  the charts. Or will it?

7:52

MI455X does have a memory advantage.

7:58

At least until Rubin Ultra,  which will come with a clean terabyte of HBM.

8:04

No matter what, there’s no doubt that VR200 is  the most anticipated release for 2026.

8:04

And it’s also the topic of the latest SemiAnalysis  article, just in case you really want to know how Vera-Rubin is built and how it works.

8:15

Mark Zuckerberg is on a massive spending spree.

8:21

Meta is basically buying everything  they can and that includes GPUs from Nvidia and AMD and even TPUs from Google.

8:26

Which makes it easy to forget that Meta has its own silicon.

8:32

The Meta Training and Inference  Accelerator, MTIA for short, is now in its third iteration.

8:38

We don’t know all the specs yet, but  it will be produced in TSMCs N3P with very likely over 100 billion transistors and HBM memory.

8:45

Previous MTIA version still used LPDDR5x.

8:53

Meta knows exactly what their internal workload  needs.

8:53

And that’s not only AI chatbots, but AI recommendation models that run the algorithms  for Facebook, Instagram and Threads.

8:59

MTIAv3 won’t be Metas chip for AGI, but it offers good  margins for their business model.

9:05

And that means, they can use all the external hardware to  train new models, while internally they switch to their own silicon for inference.

9:17

Looking at the roadmap, we will see a lot more MTIA in the future.

9:21

There is a lot of AI silicon, but only a few chips see truly large-scale  deployment.

9:25

One of them is Amazon’s Trainium.

9:32

There are hundreds of thousands of Trainium2 chips  deployed inside the AWS AI datacenters in Canton and New Carlisle.

9:39

Which makes Trainium3 a top  contender for future large-scale deployment.

9:46

Trainium3, despite its name, combines training and  inference capabilities in a single chip.

9:46

Built on TSMCs N3P, it consists of about 125 billion  transistors and comes with 144 gigabytes of fast HBM3e.

9:59

You can find much more, including TCO  numbers, in the SemiAnalysis Trainium3 deep-dive.

10:06

And yes, it’s an actual technical deep-dive.

10:06

Anthropic’s amazing Claude Code was trained and runs on Trainium, it can only get better with  Trainium3.

10:12

And Anthropic isn’t the only fan.

10:17

OpenAI will use 2 gigawatts of Trainium  compute starting this year.

10:17

If you are looking for an ASIC with truly large-scale  deployment, look no further than Trainium3.

10:29

Microsoft’s first ASIC, Maia 100, was a little  bit of a mystery.

10:29

Like every new in-house design, it has to prove itself first.

10:36

Does  it warrant a second generation?

10:41

With Maia 200 we have the answer.

10:41

At around 825  square millimeters, the almost reticle busting chip contains 140 billion transistors, is  manufactured in TSMCs N3P and comes with a pretty large 216 gigabytes of HBM3e.

10:53

Maia 200 will be used for inference and is optimized for FP8 and FP4, providing over  5 and 10 petaflops respectively.

10:59

Microsoft will use it for its in-house models but Maia  200 will also run future ChatGPT models.

11:12

More and more inference is moving  away from Nvidia to custom silicon.

11:17

Last, but certainly not least, we have Intel,  who is having another go at an AI GPU.

11:17

But it doesn’t seem like Jaguar Shores is targeting  a 2026 release.

11:23

But with so many ASICs, I had to talk about one more GPU, even if  it might be a 2027 product.

11:29

Jaguar Shores comes with a pretty strong spec sheet.

11:34

18A  process node, 175 billion transistors and 288 gigabytes of HBM4.

11:41

On paper, it's competitive. But paper is patient.

11:41

After many attempts, Intel not only has to prove it  can build and produce an AI GPU, but also provide the proper software support.

11:52

But no matter what, a third player in the GPU space would be very welcome.

11:58

And Jaguar  Shore will show if Intel can compete again.

12:04

The chip wars are heating up and what we just  talked about is only a small selection of the entire market.

12:09

If you compare this list with what  the SemiAnalysis Accelerator & HBM model tracks, you will know what I mean.

12:16

Right now, all  lines are only going up. But for how long?

12:22

If you are working in or with the industry and are  interested in a clear picture view of the entire AI silicon market and where it’s headed, the  SemiAnalysis Accelerator & HBM model is your best choice.

12:33

Not only does it offer a comprehensive  overview of all current and future chips, it also provides shipment numbers and ASPs.

12:39

Plus, deep insights into the HBM market.

12:39

You can find the link in the description  below.

12:46

Right next to all the deep-dive articles we talked about in the video.

12:49

And if you are less interested in spec, but more in real-world performance, InferenceX  offers comprehensive benchmarks across a wide range of accelerators.

12:59

Because what really  matters in the end is how a chip performs.

13:05

Check out the SemiAnalysis InferenceX dashboard,  it’s free.

13:05

You know where you can find the link.

13:11

I hope you enjoyed this video  and see you in the next one!