Ep. 027 - OpenAI Jalapeño: Better Than Nvidia Blackwell (Accelerators)

0:05

Hello everyone.

0:05

Welcome back to Semi analysis weekly.

0:07

Uh today I'm joined by Brian and Myron.

0:09

We're going to talk about the new article we put out on OpenAI Jalapeno veteran Blackwell their self-designed ASIC which we compared with Ruben. We analyzed the TCO.

0:18

uh we assessed the throughput per megawatt claims a little bit more of the details in terms of the micro architecture system architecture how they used AI to design it and write kernels and why the speed uh from tape out to actually like first working system with first real benchmarks run on it has been so impressive.

0:40

So guys excited to dig in. This is a fun one.

0:46

Yeah, [laughter] >> you could say something in response there, Brian.

0:55

>> Sorry, I >> Is this a Is this a >> Is this a virtual background that you got like pink cotton candy background going on or what? >> Of course. Of course.

1:05

I think it's the most similar color to what uh I think Open AI's belonged.

1:11

I don't remember off the top of my head, but they're good.

1:14

They got this like nice designs for their blocks. Yeah.

1:16

But yeah, but anyway, Jalapeno has been great.

1:18

Uh quite a surprise on the last day of hot chips for those attending.

1:25

Uh one of the most surprising talks in my opinion.

1:28

Uh and yeah, I believe it caught the performance caught everyone by surprise.

1:33

We all knew that OpenAI chip was in the works and they have announced uh deals with Broadcom but I first look at performance results and we yeah and it's really looking quite good for OpenAI a brief summary for those that didn't read the article which is kind of unlikely is that on per megawatt basis and perf per TCO basis.

1:57

So in other words, uh how much it cost to actually run the chip.

2:03

Uh actually Jalapeno actually beats Vera Rubin's July results and we are comparing against July results as we believe it's kind of fair because the software is in like a similar state uh between these two points. Yeah.

2:21

So as you shown as on the picture, this is Jalapeno compared against GB300.

2:27

So JP300 it beats out of the water completely.

2:31

Uh but as we mentioned during our in our article it's a bit unfair to compare it against Blackwell because Blackwell uses HBM3 while uh Jalapeno uses HBM4.

2:42

So actually my can talk about this a bit later on the differences between the HBM but it's a bit unfair to compare HBM 4 against HBM3.

2:55

So we're comparing it against Rubin instead.

2:59

And yeah, as in the results show, he does beat Rubin on output token throughput per oil utility megawatt.

3:05

But then again, these are July results.

3:08

And where Rubin right now is likely uh much better than July.

3:14

But Jalapeno of course will improve as time goes on.

3:17

As we shown in another diagram below Jordan, I think >> uh we show how jalapeno has improved from uh 25 days. >> 25 days.

3:32

>> Yeah, I'll go grab that one.

3:32

But um maybe like the >> the key point you were making if you can explain a little bit more about >> the benchmarks that they're running which is Deepse CR1 8K1K.

3:46

um not really the like absolute biggest model and not the most demanding inference workload because it's uh just random data.

3:53

It's not agenic workflows.

3:56

Um but really quickly they got this up.

4:01

Clearly the performance is strong as you're yes implying with this new chart.

4:07

um performance is like improving by the week at this point even by the day and um maybe you can you can also explain the denominator there like why they're choosing to divide by the power consumption.

4:22

Yeah, actually uh Jensen brought up this point in Computex 2026 during his keynote that uh data centers nowadays are getting power limited and that power is starting becoming the constraint.

4:36

If you have money, you can always get more servers, more chips, but the the constraint is starting to become the data centers power.

4:43

And like companies have tried getting through this using a BTM or behind the meter power but end of the day the power tends to be your your constraint for building a data center. Yeah, absolutely.

4:59

So like if you have 100 megawatts of power you can only fit so many chips in there.

5:04

It doesn't really matter if these chips are more expensive uh on a per megawatt basis.

5:11

if they're producing more tokens per megawatt, you think you can make more money off of the tokens they produce and you justify the extra expense on the actual chips.

5:22

And while it's a valid metric for some customers, it may not be valid for others.

5:29

And so it's it's not the most typical way.

5:31

Most typically we see just token throughput per GPU compared like per package.

5:38

Um, but in this case, OpenAI doesn't have customers for Jalapeno.

5:44

They just run it for themselves and so they don't really care tokens per package.

5:49

They care about tokens per like megawatt that's going into the system, right? Yeah.

5:54

And actually going off of this uh topic uh someone did mention I forgot one of us mentioned that uh like performance per chip is at the end of the day is just a imaginary thing.

6:08

You can just glue two chips together and call it you double your performance per chip.

6:13

So looking at per chip is not really a wrong metric.

6:18

just it can be easily like misled cuz you can just which is what Ruben Ultra and Blackware Ultra does if I'm not wrong.

6:26

You just glue two chips together and bing bang boom, you got double your true per chip. Yeah. >> Yeah.

6:33

I mean, you can do this if you're cerebrous, too, right?

6:35

You can just put three of the wafers in a rack and then you can put them all on a chart where three is better than one, right? >> Yeah. Or or Yeah.

6:43

Just say your chip is the whole wafer, right?

6:45

the whole wafer, right? Um yeah um yeah but I mean uh it it doesn't mean that um they don't care about performance per per cost right because I'd say that um you know performance per what and performance cost they're pretty closely closely linked right I mean I think um the more power the chip

7:08

consumes it probably means the chip is more expensive as well generally and I think we see that um yeah we we did our um you know TCO analysis of what we think open would pay for a jalapeno system and you know they still pretty much win or or very close on uh performance for TCO as well, right? >> Yeah. Here's that chart on screen. I >> Yeah.

7:32

Here's that chart on screen.

7:32

I mean, this one's just more interesting because obviously the TCO calculation in terms of like how many tokens you're going to get per dollar um depends on that input for per dollar.

7:47

And so we've got some variables on screen there about, you know, what's the total cost per hour to own, for example, a GV300 or to own Jalapeno.

7:56

And we're assuming that OpenAI is going to pay 279 an hour for a GB300.

8:05

They're buying them themselves.

8:05

They're running the data centers themselves, for example, in that case.

8:08

Um they're not playing paying the current going NeoCloud prices, which are like six bucks an hour for GB300 right now.

8:14

And put Vera Rubin at um 361 and then Jalapeno at $1. 56.

8:21

And so let's just compare it per package.

8:25

Like clearly they're going to like the thesis of designing a chip in house is that you want to pay the Broadcom margins only, not the Nvidia margins or the Broadcom plus Google TPU margins or some equivalent there, right?

8:40

Um and that's that's being borne out on this chart that you can see. >> All right.

8:53

And I guess the the next part of this is like um where could like how could these curbs move?

9:02

Um I think with jalapeno they're actually somewhat sandbagging these results significantly. Right.

9:14

Uh yeah, I mean this is so the caveats I guess on the performance is that I I said it earlier that um the tests they ran were deepseek 1 uh that's based on the V3 architecture that came out in January of 2025.

9:29

So it's not the most current DeepC4 but it is a relatively big model 600 billion total parameters. They ran Kimmy K 2. 5 which is 1. 5 trillion.

9:40

So that's a actual big model or just 1 trillion, sorry, not 1. 5.

9:43

And then they ran GPTOSS 12B, their own open source model, relatively small.

9:49

Uh they had solid performance on all three.

9:53

Um but I mean solid performance.

9:58

They're beating Farah Rubin [laughter] on uh on all three today.

10:01

And they're beating them with single token prediction and no pre-filled decode disagregation.

10:08

So Brian, maybe you can explain single versus multi-token prediction and specifically in these performance claims when compared directly to Bar Rubin, which is using MTP.

10:23

Um, this is like OpenAI kind of fighting with one hand behind their back. >> Yep.

10:30

>> Like MTP is a really big optimization >> for interactivity. >> Yeah, of course.

10:35

Uh actually going back on uh the deepse uh one it reminded me of an X post I saw like some time ago that someone said the closed lab open AI and entropic probably have an internal version of MLA and they probably discovered uh MLA like long before deep but it's quite interesting that uh we were told that the MLA kernels like openi didn't have any internal MLA kernel uh implementation.

11:07

So I'm not saying like the open source tritons one they don't have any internet optimized MLA kernels which is quite interesting because that well that means none of the open AI models uses MLA which is quite interesting to me and uh yeah and on the AK1K and Deepc R1 like model choice we did release an AX uh benchmark case but the timing was very bad so I think we didn't get uh the OpenAI Jalapeno team to actually run it on Ethernet X.

11:39

But yeah, as Jordan says, uh measures stuff like prefix cache which uh exposes a lot more areas for optimization.

11:49

And in our we describe like a lot of stuff there's loadbearing that is being tested for Agent X and it's starting to show like what break what's breaking what's not and areas for improvement that AK1K didn't properly show.

12:08

And uh and yeah going so going back to the MTP question uh for those unaware MTP stands for multi token prediction and it's a form of speculative decoding where you kind of guess the next future tokens and then verify them in a single forward pass because uh interestingly the LM doesn't output just the next token's probability but the token probability of every single token position before the last token.

12:37

And we were not given an MTP results.

12:39

And it's my personal guess that OpenAI has some internal speculative decoding technique that's not MTP and not DSpark or any open source uh configs.

12:47

or any open source uh configs. So they didn't give us speculative decoding results because there's no way to actually verify it through open source which is also quite interesting in my opinion means this spark isn't the best that we can do and

13:06

this park actually gives quite a a crazy advantage over MTP by doing like uh guessing all the tokens at once uh instead of doing MTP which is just one layer of a model trying to uh guess future tokens one by one instead of doing them all at once. Yeah. And with Yeah. And we were told Yeah. And with Yeah.

13:25

And we were told that the internal speculative decoding method gives like three to five times uh improvement on production models which yeah would just knock the jalapeno versus ferubin comparisons out of the water once again.

13:41

So you can imagine that graph shifting like three to five times. Yeah, it's quite insane.

13:48

So Kudo mode's gone, Brian.

13:52

>> I mean, everyone on X is is saying, "Oh, it's just an ASIC.

13:56

That's what you expected to do."

13:57

And yeah, in some sense, that's what we that's what we expect an ASIC to do.

14:00

But the software has come to such a point that like this line doesn't really matter if your software is good enough that you can get a motor up like very quickly.

14:09

What's the difference between a general purpose or like they say GPGPU but N6 if software can just bridge this gap? >> Yeah. >> Yeah. Yeah.

14:23

I mean this chip has I mean it's a toss up with Ver Rubin because we haven't seen stuff since they published that uh note in uh June or July.

14:30

Um so maybe they've had a month extra to develop on the these chips.

14:37

But I mean, conceptually, like we've never seen anybody else put out a chart where there's a curve showing they're beating Nvidia on every point of the curve and uh like a real test, right?

14:53

So, it's I mean it's shocking that they did this.

15:01

We can talk about the timeline a little bit later, but we we keep talking about the point in the curve and I think we we maybe haven't explained this in great detail or we've done it on previous podcast and people aren't, you know, familiar with this.

15:12

So, the I'm going to put this chart back up on screen and try to explain a Fredo curve here for the purposes of understanding the performance claims made by Jalapeno here.

15:25

And uh to do that we need to explain that the y- axis is how many tokens you can produce per megawatt you're putting into the system and the y- axis is how fast or sorry the x-axis is how fast the tokens appear to each individual user.

15:39

And so whether you're optimizing for the y- axis or the x-axis jalapeno is beating the GB300 right now on the deepseek model which implies something like if you were to fix the interactivity per user.

15:54

So everybody sees 100 tokens per second or 50 tokens per second.

15:58

And I zoom in on like exactly that part of the curve.

16:01

We're basically seeing that Jalapeno has double the amount of tokens that it can produce per megawatt implying two times more revenue, two times more profitability, whatever you want to say from an inference endpoint serving provider.

16:19

And then if you look at the far right side of this curve across the x-axis and you see that at a very low batch size they can go all the way to 700 tokens per second per user.

16:29

And you compare that to where the others peak out at 350.

16:35

This is once again double the performance.

16:39

And I think the conclusion is that they're basically winning on both sides of the curve.

16:49

So, both fast tokens and cheap tokens, which is it's just so interesting because we've seen so many other um companies make claims about how they're going to beat Nvidia and they just pick one of those, right?

17:09

right? Croc or Cerebrus or any of the other startups that are going to focus on SRAMM like call it a Dmatrix that's coming up with stuff or Samonova where we've even seen some results on bench you know on the infertex benchmark or something similar to it they're saying they're going for fast tokens just decode speed low batch size they don't

17:27

worry about throughput and then you've got other guys that are worried about throughput call it AMD just as a simple example but even TP or tranium could be in this bucket of just like accelerators that are going for throughput and they go, "Yeah, but we're not going to be able to compete with the other guys at high interactivity." And OpenAI

17:41

And OpenAI has a chip that can do both for them. Really well.

17:45

I mean, at a minimum, this thing is is doing really well right now and is going to serve real production tokens for them. >> Yeah.

17:55

And yeah, it's actually quite surprising that I mean I was surprised that OpenAI was the first chip that actually uh like non Nvidia non AMD chip that actually appear on our public infrance X.

18:09

We were like expecting some Nova or like Cerebras or even TPU trainium to be one of the first.

18:15

one of the first. But yeah, maybe this really come put into perspective like the difference in uh like how this time might be different from the rest like this the first chip that actually poses a real threat to the Kuda mode because like in open source we welcome uh results from anyone like you believe his chip is good run the results run the curves run the benchmarks show us what your chips does and we gladly put it on

18:44

our dashboard and we gladly like write write an article about it if it's good but yeah and we have extended this offer to like edged very recently on X and of course etched uh didn't get back to us on that but yeah if your chip is good

19:00

just run the benchmark show us the results and we let the results talk uh yeah so >> yeah it's it's >> we love to see more completion from the chips yeah exactly >> so did you guys see um Jensen's response to this I mean I level. I think the

19:14

to this I mean I level. I think the Cudamote has been uh getting drained slowly over the last couple of years like top models in the world like Claude and Gemini are trained without Nvidia GPUs to you know TPU right um Anthropic use lots of tranium it's not like you can only use GPUs but Nvidia is incredibly valuable company they're

19:39

going to keep shipping all these GPUs and I think there's such a thing as like an Nvidia mode which includes everything in the supply chain the whole develop for ecosystem like all of the um availability to purchase and support and how you're going to like do the logistics of deploying these data centers and monitoring them over time. I

19:53

I mean, OpenAI's now got to go figure out how to turn on 100 megawatts of these chips, not just like three test racks, which is a monumental challenge to get over as if taping out a chip of this this uh you know quality is easy.

20:07

Um which is not like the next phase will be pretty hard for them as well.

20:13

Even Cerebrous is going through this themselves right now.

20:16

So I guess the the question for me uh that well I think Jensen kind of answered in the style like I was saying there when he was on Mad Money with Jim Kramer, [laughter] everybody's favorite.

20:34

Um, and he was basically like, I'm not bothered.

20:40

And uh, yeah, like what's like Myron, what's your take on this? You think? Well, okay.

20:47

What's your take on that?

20:49

And just the whole like just the timeline to go from, yeah, we're tired of buying Nvidia GPUs for everything.

20:55

We're going to go build it decision made at OpenAI to actually having a chip that can run in a sex benchmark.

20:59

It's like under two years for concept to like real chip in lab and under nine months to actually get it taped out, right? >> Yeah.

21:12

Uh yeah, I I have several thoughts on this.

21:14

on this. Um I guess starting from uh CUDA mo eroding um I think you know a lot of um how can you success both in terms of the silicon design as well as uh bringing up the software is it's it's been AI assisted right um and you know I guess somewhat the irony is that um this was all done on Nvidia

21:42

GPUs training these models to bring up the capability to a point where um you know AI is able to you know program kernels um and that's been I think probably the big shift um in terms of um making it easier for the labs to and anyone to adopt um you know alternative systems for their their serving stack. Um I think you know one of the big

22:07

Um I think you know one of the big reasons that uh Anthropic you know has decided to uh you know bring in AMD as as one of their hardware providers is because um you know Aentic programming as it allows them to sort of get around the the challenges uh of like using the AMD software stack for instance.

22:31

So um you know somewhat ironically uh it's it's Nvidia's hardware has enabled um moving off of Nvidia's hardware.

22:40

Um, and you know, in terms of I guess you jalapeno, it's it's such a big surprise because I mean we always knew that the team is capable, you know, they have they've had experience building um the other main successful um as program in the form of TPU.

23:04

TPU. Um I think a lot of the yeah the the hardware team behind this is from their former TPU people as as we see in a lot of other um I guess whether it be AI accelerator startups or other A6 A6 teams they tend to come from you know former people with TPU backgrounds um

23:24

and um so for what it's worth sorry sorry sorry to interrupt but there's like basically no chip starters I can point to where it's a bunch of guys who are ex Nvidia But there's a lot of ex Google people out there doing stuff which is interesting. >> Yeah,

23:39

>> Yeah, >> I I do wonder sort of why why that is the case.

23:42

But anyway, um another another topic.

23:45

Um so uh and I think you know I mean designing an AI uh accelerator that is competitive with um Nvidia is is not easy, right?

23:57

U it's such a huge market.

24:00

Um, of course, everyone wants to try it, but um, you know, time and time again, we've seen uh, you know, as as as you guys have mentioned, we've seen a lot of entrance, but they haven't really been able to do it.

24:10

Um, so, you know, we I think the expectation was that open would deliver a decent effort with their first generation.

24:19

Um, and it turns out it was much better than decent.

24:21

they, you know, as we said, they've already come out with something that's pretty much competitive or better than what the best of, uh, what Nvidia has to offer.

24:32

So, that's surprise number one.

24:32

Um, and then I think, you know, that also says something probably about um, other ASIC programs like especially, you know, Meta, Microsoft.

24:43

Is it open air that's really good?

24:46

Is, you know, are the silicon teams at Meta and Microsoft, do they have skill issues?

24:50

It's probably a bit of both, right?

24:52

I think um they're the guys that look the worst from from this announcement.

24:57

Um but you know, going back to okay, where next?

25:00

going back to okay, where next? Um so designed a chip uh obviously you're scaling up the supply chain to um you deliver systems uh of mass like you talking about delivering like millions of these chips um you know thousands of racks um deploying them in data centers you know gigawatts of power um that's that's going to be not easy but also I think you know other people have have done it successfully I think the hardest

25:29

part is really having that system design and then I think um you know OpenAI has partnered with um with people who have experience scaling this up right basically it's Broadcom and and Celestico on the system side uh and they've had experience you know doing

25:46

this with TPU dig on the comment you made about the difference between like a in-house silicon program that's been going for years and years like MTIA at Meta or Maya at Microsoft if we just literally look at the specs of Jalapeno. Like it

26:00

Like it doesn't look super fancy on paper.

26:06

I mean, it is an HBM4 chip, so it's going to have great HBM bandwidth.

26:14

>> Um, it's got lots of >> FP4 flops, but like still >> less than >> Mhm.

26:20

>> half of what Reuben's got on FP4.

26:24

Less HPM capacity and less TDP.

26:24

So I mean like under like almost three times less TDP per chip than Reuben, right?

26:35

So if you just it's this weird thing where um when people are making the bull case for AMD, they go just look at MI450, right?

26:48

It's going to have more FP4 flops.

26:51

It's going to have more HPM capacity.

26:52

It's going to have more HPM bandwidth.

26:54

It's going to be more TDP.

26:56

and therefore it's going to be better.

26:59

And they never want to look at the benchmarks.

27:00

They never want to look at the like results of the chip, right?

27:01

Or the MI355 when it was coming into production for the first time, right?

27:10

But now you got an open AI chip which on papers is like objectively worse specs than everything on Reuben.

27:14

Um comparable to GB300 on everything except for HPM bandwidth.

27:24

uh and yet it's way outperforming GB300 and outperforming Reuben so far.

27:28

So like clearly this is due to the use of it's due to something about the micro architecture which we can get into and the software.

27:43

So, like what's your take on what it takes to design a chip now?

27:45

Because is it purely having access to the latest models and being willing to like yolo trust them on RTL and kernels?

27:53

Like this is the barrier.

28:06

Yeah, that's a good question and I I don't know the answer to that, but I think you know to your point, right?

28:10

think you know to your point, right? um you know everyone can deliver um great you know stacks on paper right you know when when we look at these stacks they're all you know pe theoretical um and I think a lot of emphasis on theoretical because I think that some of the flops numbers like no matter like how how you try to like reach them they're like impossible to reach so um

28:39

it's it's somewhat determined by um you know the the chip company's market teams marketing teams but um yeah basically it's you know there are these flops that can you actually realize them in an actual workload same with HP you know

28:55

bandwidth um I mean one of the big tenets of uh open's design philosophy forio is that um you know everyone can deliver raw hpn bandwidth you just like look you buy the the best HPM and then you put more stacks of it. Um, which I mean it's

29:14

Um, which I mean it's [snorts] it's buying HPM itself is is not easy these days, but um you know it's basically not super hard, right?

29:25

It's not like it doesn't take um tremendous design skill.

29:27

Um, so but really, you know, what's really throttling a lot of these chips is that they can't realize anything close to the raw HBM bandwidth because there's so many other uh things in the micro architecture that stop or stop you from doing that.

29:44

And I think a lot of that is just um whether that be like very complicated memory subsystems or um or or just a lot of data movement required.

29:56

Um whereas um the OKAI team has uh really been able has focused on um with mic microitecture that um reduces data movement so that um they can realize a lot of this you know the the the massive amount of page being damage they have.

30:15

So I think that's really the main I guess skill um in all this and why for the competitors it's been really challenging to um you know they can deliver these all these all these like great specs but it's really difficult to actually realize them in an actual workload. >> Yeah. Yeah.

30:36

And I agree with this like the fact that you can't actually reach this.

30:40

I know there's there's a saying somewhere from someone I'm not sure who uh that the figures are like not numbers that you can reach and are numbers that the manufacturer can guarantee you never exceed like those flops which way to put it. Yeah. >> Yeah.

30:58

Those are numbers that you are guaranteed not to exceed.

30:59

And I think in one of the previous uh articles, I think in the cerebrous articles, they talk about the roofline models of chips compared to Nvidia.

31:12

And although Nvidia's like flops are huge and like crazy especially the FT FP4 ones at the end of the day they are extremely uh right words on the roof lines in the computer bound region and a lot of workloads would rarely even reach those roof lines.

31:32

So it doesn't matter at the end of the day the flops like the those flops in the table doesn't really matter because oops Jordan dipped again because the workloads most of the time does not actually hit those roof lines. Yeah, Jordan is gone.

31:48

So maybe have some something curious to ask you about like the HPM.

31:52

There's been talk about like Samsung HBM being better than the other HBM and like Halipino was uh lucky maybe to or like maybe it was a decision. I'm not so sure. I'm not a highway guy.

32:05

Like what makes Samsung HBM like better or like uh better quality than the rest? Yeah. Yeah.

32:12

So Samsung um for a long time you know from the HBM 3 and 3 generation um Samsung's HBM you know was was really quite inferior to um Hinx which you know who dominated who you know still dominates HPM market share is you the leading supplier for Nvidia for instance um and you know really 3e that generation is really bad from Samsung um part of it was it was built um an inferior process, right?

32:43

Um so the ite um Samsung, you know, they realized this and they really went all out on on their HBM core technology.

32:58

HBM core technology. Um so the DRAM dies are built on a more advanced 1C process whereas you know highix and micron are using 01B process the same which is the same process that HPN 3 is built on um they're also using there's a logic based in HBM cube that has the fi u and Samsung is built built this on a advanced logic process which is you know SF4 um Samsung Foundry 4 nanometer node

33:34

um whereas Hinx is using 12 nanometer TSMC um and Micron is still using their own you know DAM process um for this base D even though you know the the bandwidth um requirements are significantly higher for HKM4 and that's sort of been um that's that's why you know Micron has had some issues with um achieving the this, you know, the highest speeds for HP4 and and similar with with Clinics, they've had some issues. um they've had to say redesign

34:06

um they've had to say redesign the base die and this is why um you know it it was only until like they've had to like delay shipments for um HPF core for for Nvidia's uh you know rub right um so Samsung has ended up being you know because of these you know I think it it seems to be clear that Samsung has actually the best technology for HM4 um this is why you know this um Jalapeno is HPM4. It can deliver 15.

34:38

4 terabytes a second of you know HPN bandwidth um which means the 10 GB per second uh pins speed the HPM4 they've gotten um this is a little bit higher than say the 9.

34:50

6 six that we think basically your uh Nvidia will will ship Ruben with.

34:56

Um and yeah, I think it probably does end up being that it's because of Samsung being um the supplier here.

35:05

Um and you know the reason is also that whether it's it's luck or skill I think um we can debate about it but you know guess I guess traditionally um Broadcom's HBN has mostly come from Samsung right so I think it it's partly sort of that luck um this is this has hurt Broadcom for 3 but um you know for four this is been I guess that's turned out pretty for that. >> Yeah.

35:38

Yeah, that's very interesting.

35:38

And yeah, this this story of like Samsung HBM and like uh because if I'm not wrong, SKH Highix invented HBM, right? >> Yeah.

35:48

So SKH Highix um along with AMD were sort of they realized that um very correctly that um you know memory bandwidth is doesn't isn't scaling um with along with uh you know logic performance.

36:07

So they needed something to really like bring out bandwidth and they came up with this HBM concept and originally it was for for gaming GPUs.

36:17

So the first product with HPM was on a gaming GPU um with on on one of the uh AMD gaming GPUs.

36:26

And but this this turned out to be a bit bit overkill but uh thankfully um you know they had this technology um for for another very memory bandwidth intensive application which is AI. >> Yeah. Yeah.

36:43

It's interesting that like SK Hanx and AMD invented HBM, but now like I mean Samsung is doing better in HBM 4 and Nvidia is uh doing better than AMD.

36:56

But on the topic of the HBM bandwave, so bandwidth is a problem, right? And is not capacity.

37:02

Is that why like companies are going towards like if I'm not wrong six high and four high instead of a high?

37:12

Um so I'd say yeah the the primary I mean it's in the name high bandwidth memory the the appeal is is the bandwidth um capacity is important but um I think the main sort of benefit is really bandwidth because you know there there are cases where you pay for the additional capacity but you don't um you don't you might not need it um whereas I think for memory bands um you can always just reuse that into serving tokens faster, right?

37:45

Um so um [snorts] and bandwidth is really the key and the thing is like whether it's a 8 high, 12 high or four high stack um the bandwidth is the same but um because you pay for capacity and and that makes sense because you know more capacity more layers is what um adds to the cost of the supplier.

38:03

um you know the sort of you pay for the extra capacity but the dollar per bandwidth um gets much worse.

38:12

So um for some some some companies if they want to optimize um and say actually we just want to like get best dollar per bandwidth then going lower stacks is is the right optimization.

38:23

Of course you want that balance between capacity and and bandwidth but uh you know I think there is a a philosophy especially now that HTM is getting much more expensive because um you know we have very limited um supply of HPM wafers.

38:39

Um so I think that the trade-off is starting to look more in favor of going to lower stack heights rather than just increasing them further and further.

38:52

>> Yeah, it's interesting.

38:52

Yeah, it makes sense with me especially like uh increasingly more Rex scale architectures >> but yeah like capacity is not becoming as much of an issue anymore. >> Yeah, exactly. >> Yeah.

39:06

And this is um I think specifically borne out in the per watt argument as well where um if you look at the raw HPM bandwidth comparing the specs of the chips, it's like okay, it's up there with the other ones, but then if you divide by the amount of uh power consumed by the chip, the fact that they're getting 15.

39:25

4 turbines per second on a chip with a TDP of 700 watts is is incredible.

39:30

Like this bandwidth per watt ratio like arbitrary units here of 22 is just so much bigger than anything else.

39:38

It's literally double Reuben.

39:38

Even at the max Q like low power setting option on Reuben to like run it at 1,800 watts, that's still like it's got a little bit more um HPM bandwidth, you know, whatever that is 25% more 20 versus 15 terabytes per second, but it's literally more than double the power.

40:00

So I mean that's that's where your if you can realize that bandwidth with that much power that's where your token output per watt advantage comes from right there.

40:14

Um I guess the the like big question of course my is like B 0 is in the fab right now.

40:22

They should just be able to step up the power and get even more bandwidth from this, right?

40:29

Or are they at some limit?

40:34

I think um I think on the HVM bandwidth they are probably at um limit.

40:37

I think the uh and you know yeah I think that's sort of just the constraint of the memory itself.

40:49

Um but the B 0 um it should deliver more flops um at at you know the same power basic.

40:56

So um that you know I think depending on the workload um if there are compute constraints then then that should benefit um for the B zero stepping uh the A zero stepping is what's uh what the I guess all the current um results are. Yeah. Yeah.

41:17

There are many different we're seeing so many chip startups explore the surface area of possible uh configurations of chips right now.

41:27

But if you are optimizing for HPM bandwidth per watt, a metric that seems pretty relevant in LLM inference, this is the design to go with at this point, right?

41:38

There's nothing else that compares that we've seen we've seen specs on um or that are public specs on, let's say.

41:49

uh not giving away too much there.

41:49

So, um maybe just to talk about realizing it a little bit more.

41:59

Uh I don't know, Brian, do you want to talk a little bit about the software programming model and the micro architecture?

42:06

You want me to talk about that?

42:10

>> I think I think you're more knowledgeable in the aspect, right?

42:12

But but actually be before I let you answer your own question, the another very big interesting point is like the role of AI in all of this right now you can you can make you can make an argument that open AI's biggest advantage is that it's able to access its own Astra models uh the new GP Astra models before anyone else.

42:35

So like and that's that's that's like the biggest difference between I would say between open AI and one of the other new chips companies sova cerebras etc.

42:45

companies sova cerebras etc. Uh yeah the question is how much did at least my question is how much did Astra actually contribute to this like if you look at it from a differences point of view like this is only one of the only differences between uh uh Jalapeno and the other chips like so did Astra really contribute to like most of these performance differences or just a bit yeah but that's just that's a

43:17

tangent of AI work on development and yeah actually Kim K3 when Kim K3 was released there was a part on the blog about its designing of a chip I forgot what what what was it about maybe some yeah I'm not sure what the architecture was about but yeah they did talk about K3 developing a chip and yeah I would guess GPT extra does have similar capabilities and I would say to a better aspect or to higher degree Yeah. Sorry. Back back to you Jordan. Sorry. Back back to you Jordan. Yeah.

43:50

On the architect micro architecture. >> Yeah.

43:53

Well, let I mean, let me comment on the AI assistance on the architecture cuz I think it's there's two ways in which AI clearly assisted the design and then the bring up of the chip, namely design and then bring up.

44:08

So, on the design side, this clearly wasn't Astra because the RCL freeze was in July of last year.

44:16

So, you know, from February to July of last year, when they claimed that AI assistance helped them get an 8% reduction in SIMD area and then a 10% reduction in the matrix engine area during design, I mean, this is preGPT5 that we're talking about.

44:33

So, um, conceptually like the models have to be getting better at RTL in the meantime, but they were already good enough to rapidly like accelerate the um really tedious human driven work that is RTL before a tape out, right?

44:51

Um and so I I think you know maybe that's the biggest claim here.

44:59

Uh which is that I think I think a lot of people understand you can use these models for kernels or and just software engineering because like you put it in a codeex harness you put it in a loop uh you let it test the thing and then you just set goal and like people have had that experience they can kind of understand it but I don't think a lot of people can like have had the experience of designing a chip.

45:23

It's not like traditional software programming and um this was done with an older model.

45:28

So I think that's like point number one.

45:31

Now uh on the actual bring up yeah like kind of like what I just said the work is iterative and it's a verifiable domain.

45:43

So like this is this kind of exactly what RL should be good at.

45:48

You should be able to give a model a task of improving the performance of a kernel or getting the kernel to be functionally correct against some u like verification script um against some test cases, right?

46:00

And then just let the model rip, let it try.

46:07

And Dylan made this podcast on made this point on the Dores Cesh podcast that he was on recently, which is that for years I think now these uh the companies that were getting the most value out of using AI were not actually the companies providing AI.

46:23

like OpenAI and Anthropic were not profitable for a very long time and now they're like just turning a profit but I mean still running incredibly high margin businesses but they're just like realizing the profitability of training these advanced models meanwhile you know Jane Street's going out there and

46:40

printing $15 billion in a quarter clearly using AI for trading or something like that right and there's many other companies that are being started based on the use of AI um this is a very clear example of open AI keeping the benefits of having access to a model before everybody else for themselves. They can take out a chip and

47:02

They can take out a chip and their competitors can't.

47:04

And it's a sign of what's to come.

47:06

I think that's the simplest way to put it.

47:09

Um they're going to be able to go into many domains that are tangentially related to software where it's like the model needs to be able to control a computer, but it's not explicitly like the thing you're training it for.

47:28

You build RL environments, you spend enough tokens, you spend enough like time on reasoning and enough rollouts, you know, enough like attempts at the um problem and you're going to get a good result.

47:41

That seems to be the lesson here. >> Yeah. Yeah. Exactly.

47:44

And I I think Entropic is also realizing the same thing.

47:47

They are starting to hire like silicon people.

47:50

that's on the same part as open and they are doing a lot of stuff like in the laboratory they got like LMS to control microscopes and whatnot recently and they're going seems to be going quite long into this like biological sciences field. Yeah.

48:05

So I really agree on with you Jordan on the point of like uh these frontier model companies are realizing what this oops lots of interference from Jordan but yeah lots of good like downstream impacts of having a good model first can have not just like making money from inference revenue.

48:30

Yeah, lots of interesting developments.

48:40

>> What's your thoughts, dude? >> Yeah, I agree.

48:43

I think um the the progress that Yeah, Labs is only like getting faster, right?

48:50

And that's really because they're using their own models um really effectively to drive product innovation um much faster, right?

48:59

Um, I remember like I think was it earlier this year like Anthropic was releasing a new product like every week or something.

49:09

Um, and I think it was like everything was basically on autopilot.

49:11

They were just using code for everything, right?

49:16

Um, so yeah, I think this I agree with with Yeah.

49:19

with what you guys have said. >> Yeah. Yeah. Um, yeah.

49:24

Yeah, I mean like the cynical view of a program like this for both OpenAI and Enthropic and even Meta and Microsoft is that it's kind of like a head fake that gets them a discount on the Nvidia GPUs and therefore it pays for itself.

49:38

Like you only need to spend a few hundred million on a program to like tape out a chip and you know like the year or two to do it to to potentially like um help with the negotiations and if those negotiations are measured to the tune of hundreds of billions of dollars then it pays for itself pretty quickly.

50:02

But the nonsynical view is that like this is a real thing and it's only going to they're only going to do more of this in the future, right?

50:14

Like um there's no reason that they're going to be less vertically integrated and less interested in developing chips this time next year.

50:24

And there's no reason to say that the RTL time from freeze to or from like initial to freeze to tape out can't go even shorter than 9 months.

50:35

Um and I I think you just need to to think about where to go from there.

50:41

Maybe the other thing that was was kind of interesting here.

50:45

So like if we go through the architecture, I mean there's lots to say about it, right?

50:49

Um the maybe the quick high level is just like it looks like a TPU with much smaller um systolics uh systolic array being like the the way a TPU has its processing elements laid out.

51:07

processing elements laid out. Um, and maybe the the criticism of chips like a TPU, tranium, uh, even some of the ones that are like TPU inspired, let's say like an etched or a Maddox X or something like that, is that these when they go with these

51:26

really big systolic arrays, um, they can have these weird cliffs where small batch dimensions like the the M dimension in your MNK for a matrix multiplication gets all which you know is what happens when you have lots of experts and you have very few requests like low concurrency. Um

51:46

Um you know the these like skinny gems, skinny matrix multiplications um can waste a lot of resources, right?

51:56

or um even just odd numbers where if you go like slightly over 256 or slightly over 128, now you're spending an entire kernel launch on the device side just to run one little skinny gem.

52:08

And uh all of these like uh processing elements on the systolic array are not being used and so therefore it's like inefficient and you don't actually maximize the flops on the um chip itself.

52:23

Uh the trade-off here of course is that to get more efficiency with tiling to you know reduce the issues with like padding overhead or or uh alignment on the matrix dimensions is that you just use smaller systolics and that's what they've done here.

52:40

And so I think that's been really smart clearly for for efficiency across the curve.

52:44

The argument the other way is that they're going to miss out on some I think power efficiency and and like data movement efficiency because well you have to have more small elements instead of one big element. So what do you do then?

52:57

And I guess the way that they've solved this is by um being really smart about how they place weights and KBs um using synchronization between cores like selectively and then saving the collective network the like knock the like network on chip that connects the HPM slices and the computing elements together really really sparingly.

53:20

And the I mean the results is like well the results speak for itself and you see the performance there but the results is that this might be both a chip chip that's simpler to reason about than a GPU and a chip that's a little bit more flexible for some of these weird changing uh dimensions over time than a TPU.

53:42

And so I mean clearly the guys who have experience using GPUs which OpenAI has plenty of experience programming GPUs mixed with the guys who have experience designing a TPU has resulted in a pretty well balanced system here.

53:57

Um maybe the other thing is that it has an L1 cache which is quite funny like all of these accelerators do not have L1 caches now.

54:06

They rely so much on L2 um namely SRAM.

54:09

namely SRAM. we hear SRAMM all the time and um I mean the the reason why I guess is that uh they like the the companies designing other accelerators don't want to use a um scratch pad uh sorry they do want to use like a software manage scratch pad and OpenAI is not using a

54:36

scratchpad cache here um and so that makes the chip like potentially harder to reason about when you think about like barrier latencies and where you're going to move data like you have to be able to amvertise the data movement by doing work on the CPU itself. But that

54:49

But that ties in with, you know, what I was saying at the very beginning, which is like the second phase of using AI, which is um actually programming kernels on like meaning software that runs on the device side.

55:02

And um the the to do this, I mean, we haven't really been able to verify this other than scrolling through a a 30,000line file with some of the OpenAI engineers when we went on site with them.

55:15

It's like um literally it's just sloth.

55:22

Like it's it's not sloth cuz it performs, but it's literally just AI generated assembly basically puked out in gluon this uh low-level like kernel programming languages that they built on top of Triton which uses this you know really interesting programming model.

55:40

Like the the point is like the guys who were scrolling through this this code with us, it was kind of clear that like they know a whole bunch about hardware.

55:48

They know a whole bunch about the concepts in the system and they just like have no idea what this MLA kernel that they're showing us for DeepSeek actually does.

55:55

Like like you can go line by line and it's like nope nope nope.

55:58

Um but it doesn't matter, right?

56:02

And uh the AI understands it.

56:06

the I tests it and you see the results.

56:09

It it produces correct kernels that perform really well.

56:12

And uh I think this is like just a sign of what's to come again, right?

56:15

The um the like that actual code is not necessarily something that a human has to reason about deeply if the AI knows how to manipulate the data movement, the processing elements on the hardware that you've given it. Um, okay.

56:37

We're kind of running out of time here.

56:39

We've been going for a while.

56:41

Uh, we got three things that I had in my notes that we wanted to talk about.

56:43

Uh, they don't do PD disag.

56:49

Brian, maybe you can rant about that because you spend like all of your time debugging PD disag.

56:57

[laughter and gasps] Two, we didn't really talk about the system architecture.

57:00

We can talk a little bit about how they do scale up, scale out domains.

57:04

I mean they don't call it scale out but uh whatever it is scale out the multi-ter scale up stuff which is just like a total mess to try to understand probably can't communicate on a podcast go read the article and then the third thing is it's a generalized inference chip it's not co-designed with their models like they keep saying and the proof of that is it runs Doom 36 frames per I can [laughter] um [clears throat] >> Yeah. [laughter] Yeah.

57:42

We we I remember we were talking to these guys and we were like the like angry disappointed mother who comes in and they show us this like groundbreaking chip that's so fast and runs this stuff and they're like but does it run Agent X? No.

57:55

[laughter] 96% on the test.

58:00

What four questions did you get wrong?

58:11

Um anyway, anything you guys feel is left unsaid on the chip?

58:14

Uh pretty exciting release, eh?

58:21

Yeah, very um I'm yeah really excited to see where the road map goes next and also um you're excited to see I mean Brian mentioned this earlier but Brian uh Anthropic is your is building it or as um they're hiring it they're building a team to do it.

58:39

Um I think this really sets a pretty high benchmark for Anthropic um to to meet or beat.

58:45

Um but I think you know there's every reason to believe that um your anthropic could achieve a similar outcome.

58:55

So very excited to see that as well.

58:59

>> And they're they're hiring the team now, right?

59:01

So it's only like what a month and a half until the RTL freeze and then like what four or five more months for the tape.

59:12

[laughter] So like we should be able to get anthropic custom you know chip tokens um after like one you know 9 month cycle right like one [laughter] one pregnancy term [gasps] you know think yeah anthropics in anthropics hiring people right now so

59:41

it's like the first trime trimester and then like the second one they they get the RTL freeze so they go to the second trimester and then like goes off to the fab and then like tape out it comes back third trimester all done right >> a little bit longer. Yeah,

59:59

Yeah, >> you guys don't like that one.

1:00:02

Okay, we're going to end on one other joke.

1:00:04

So Brian, you you like the one about converting energy usage to calories, right?

1:00:10

So I want to finish this one.

1:00:14

We uh we converted human we converted to human speech and compared uh some of this stuff on a on a calories, right?

1:00:21

Because uh a calorie or like how much energy it takes to burn a what a cubic centimeter of water I think is uh the equivalent of uh whatever one jewel is 0239 food calories.

1:00:40

So, we will uh we'll throw this this one up on screen to lead everybody off with a a little joke.

1:00:44

Uh if you convert all of this efficiency of some of these DeepSeek results that we have on Asian X, uh we've identified that human speech is roughly 20 times 22 times more energy efficient than the uh concurrency one B300 configuration that we were testing it against there. Right. A human speaks at 3.

1:01:07

3 tokens per second, but these uh batch one configurations are going up at 180 tokens per second, much faster than the human brain can work, consume calories, and produce speech.

1:01:18

So anyway, >> I think the cav the caveat is that um not all human spoken tokens are very high quality.

1:01:30

[laughter] I mean, [laughter and gasps] >> I got to caveat some of my interactions with Claude, too.

1:01:38

Then, meant some of this nonsense that's been spitting back at me recently, I've I've I've not been too pleased with either.

1:01:43

[laughter] Can't see all the thinking traces anyway, but the like clawed version of English it's given me has not been that great. Um, [gasps] yeah.

1:01:53

Well, uh, if we're comparing the machines on how many calories they're consuming per token and we're comparing it to our speech, we are really in competition with the machines at this point then, huh?

1:02:05

Hopefully Jeb's paradox continues and and uh, everybody that produces chips wants to consume more tokens and produce more chips and hire more people and everybody gets to come have fun. All right, guys. Uh, good job. Long episode this time.

1:02:23

I hope everybody enjoyed the overview of OpenAI Jalapeno. More to come.

1:02:27

Last joke before we leave, cuz I just saw it.

1:02:31

The best cover image in a while, I'd say. Right.

1:02:37

[laughter] We'll leave that one on screen.

1:02:38

Lisa and Jensen enjoying a nice spicy pot of uh chana or or katsu, whatever the uh code names were for the trays and the rocks and stuff in there.

1:02:50

That's the motivation behind that. Vindaloo. Yeah, >> right. >> Sorry about that.

1:02:55

All right, we'll sign off with that image in everybody's brain who's watching online.

1:02:59

[laughter] Thanks for listening, guys.