0:05
Here ye, here ye, long live the short king.
Here ye, here ye, long live the short king.
We are back with another episode of semi- analysis weekly.
I got Myron laughing already.
We're going to talk about why four high HBM wins in this episode.
Um, we put out an article just yesterday, I think two days ago probably when this episode comes out and uh called a few pretty big changes um in the industry.
Everything from the uh HBM suppliers all the way to the Nvidia GPUs that this is going to impact and potentially other accelerators.
we will discuss and and we'll dig into some of the uh motivations for this change and why uh yeah people are revisiting the balance between capacity and bandwidth. My how you doing man? >> I'm doing very well.
Um so yeah this was a fun article to write and I think it's um the thesis has sort of challenged a lot of conventional assumptions on HBN content per you know per accelerator.
Um so to give listeners some background um you know the the trend over the last few years has been every major chip or AI chip generation over generation you know it packs in more HPM content per chip.
Um so they add more you know stacks of HPM and then one of the big changes is also they make the stack taller or with more layers.
Um so back in the I guess Hopper or Ampia era um 8 high was expanded and then with uh Blackwell Ultra they moved to 12 high so it's having you know 12 um layers of DRAM dies in a HPM stack so you get your basically you increase your capacity by 50% over eight high um and then the the presumption from there was that you you keep going taller.
Uh so um the industry was contemplating 16 high um for you know going to 16 high.
16 DRAM laser DRAM in a in a cube um with HPM4E and Nvidia when they originally previewed Reuben Ultra um to to people at GTC last year um they said that the Reuben Ultra would have one terabyte of of HPN per per package per chip.
Um and to get there they'd have um you know four compute dies uh and you know 16 stacks of HPM 4E uh that's and then each stack of each you know 4E HPM 4 D has 32 Gbits or or 4 GB of of DM capacity um and then they make it 16 high so that's how you get to one terabyte um but now what we understand is you fast forward Um, a year and a half later, Rubin Ultra is going to just have 192 GB of of HPM4.
Um, so how do you get from 1 TB to 192 GB is um, first of all, you know, it's not long it's no longer um, going to be for, you know, compute D.
Um, it's it's just two comput D per package.
Um, it's no longer going to be HPM4E.
There might be a HPM48 version, but it's going to be HPM4 at the start with um and HBM4 is only, you know, three three gig bytes per per die and it's going to be eight high.
Um and then this you and and you'll also notice that um this is even lower than um HPM content than say Blackwell Ultra and and conventional you know what we call vanilla Ruben which were both at um 288 because they use 12 high HP.
Um so why is this happening?
Um the big reason it it's really primarily motivated by supply.
Uh so um as you know I think many readers are aware um in a very big shortage of of memory supply um in in DRAM and DAND um but you know let's talk about DRAM um there's very little you know wafer capacity being added and then as the I guess number of accelerator shipments go up um the number of um sort of DN HPM demand keeps going up and it's you know very hard to just keep up with that demand.
Um so what the suppliers and and also customers like Nvidia uh Google Broadcom etc have realized is that you know we have x amount of we have we've secured this much HPM uh but the logic we have uh secured at co that you know we've secured at TSMC is that you know if we take the HPM supply we have make it you know ship it in 12 12 high cubes it's like it's not enough to like uh it's on enough supply to like ship all that logic.
So let's ration our DM supply or our HDM supply and we ship it in eight high cubes.
Um so you get more from you get more cubes from the I guess the implied wages you get.
Um and then that helps balance the equation more and and that's what Nvidia what Nvidia has realized.
Um but of course you know like going downgrading from uh 12 high hpm to 8 high hpm I mean does that um have performance impacts like you know how does that affect your um affect performance in TCO?
Um and I think you know one of the big things is the I guess in terms of like how important is HPM capacity um for to inference um specifically it's you need there's a threshold that you need uh like you you need to have enough HBM so you can you know hold the model weights um hold enough have enough capacity to support KV cache for you know a lot of users so that you can you batch effectively.
But once you get beyond that threshold, um you what we find is that the the returns to that are diminishing and and because HPM has gone, you know, is getting much more expensive from next year.
Um as the the shortage gets reflected into prices, um the that that sort of cost of additional HBM that you might not need um is also much higher.
Um, so the penalty to having too much HPM capacity becomes yeah much worse.
>> Okay, let's take a step back and just talk about capacity versus bandwidth and the trade-offs there in a little bit more detail.
It seems to me like going into the Reuben generation um for from like a design perspective and when the announcements happened, people were just starting to use Blackwell and there was a lot of talk of 10 trillion plus parameter models.
you know, like going to 20, people talking about NVL72, going to 144, NVL576, all sorts of other like rack configurations, which really seem to be about the models getting bigger.
And >> to me, models get bigger, that means you need more capacity.
But that assumption may not necessarily be true in today's day and age because we can get a lot more performance out of existing models with same total parameter counts by using uh a lot of you know different strategies in in terms of like um reasoning and looped transformers and other >> you know stuff both in training and inference.
So, can you talk through a little bit more details of the trade-offs between capacity and bandwidth and why maybe previously people Nvidia just designed Ruben to be the biggest maxed out of everything, but we think they're going the direction of making the exact right tool for the job that people >> Yeah, exactly. So, um Yeah.
So I guess if we take a step back into um you know what models look like or compared to what systems look like um a few years ago in the let's say in the hopper era right so um if you have a hopper hx node which is 8 GPUs 80 gigabytes of HPN per GPU that's like 650 40 GB uh in that in that server um whereas at the time you I guess this was quite a long time ago, but you know 3.
1 405b that was I guess the best open source model at the time.
Um that would take you 45 GB uh you quantized FB8.
Um so that would take up like around 60% of the the server available HP.
Um assuming you just kept a deployment at that one one scale up mode.
Um but now I think what's happened is the sizes of the you know parameter counts have scaled um but probably not as aggressively as we've thought.
Um so you know the biggest to today is K3 at 2.
8 8 8 trillion um which is you know almost almost six uh a bit more than six times the size of llama 3.
1 but also um you know we can quantize that at MXFP4 which reduces the the capacity um by needs by half and then also you know if you compare it to the I guess biggest system um right now it's MDL72G GB200 which is 72 GPUs um with 288 GB each.
Um so one set of weights for you know for Kimmy K3 stored on one scale up system is like is around about 8% of the available HPM capacity.
Um so basically you know what we're seeing is you know we have um you know par parameter counts have scaled but you know not as much.
Meanwhile, the I guess the sort of deployment sizes or you know when we if you want to serve something on let's say one scale up high band scale up domain that scale up domain has grown a lot um so that the the capacity pressure is much less.
So if you go back to again with the llama and and the hopper HDX system um you know the upgrade to to the H200 which uses used HPM3e um which had more die density um that was really valuable in terms of it gave a lot more valuable breeding room um for to serve a you know one to serve llama on one HX system.
Um so that was why back then it was much more important to just increase um HPN capacity whereas now it's like you you have a lot of breeding um and you know whether you can effectively sort of whether you really need to use all of that um the it's it's not as sort of clearcut and again going back to the the decision that Nvidia made has with uh Ribbon Ultra they Um they're also going to scale the system to MVL576.
So that's another like eight times increase in the scalable load size.
Um so even though they cut go from 12 high HPM to 8 high HPM um they you know increase this scale size by eight times um that's a lot of aggregate HPM capacity within that domain. >> Yeah. Yeah.
And um just again to to harp on that point about capacity versus bandwidth uh seems like today the a lot of people are focused on increased performance from existing models at existing sizes like they want faster versions of KK3 they want faster versions of 5. 6 six soul.
And to be clear, just stacking more capacity on the existing bus does not address that problem of making models go faster on the >> Yeah. Yeah. Pretty. Um Yeah. Yeah. Yeah.
So, speeding up tokens is uh it's all from bandwidth basically.
Um the capacity doesn't really uh add anything.
Um and to your point on you know why have and then you we didn't really talk about why model parameters model parameter sizes haven't scaled as much because yeah there are all these other techniques um that you can get performance um from from models rather than just pure like size scaring.
than just pure like size scaring. Um so yeah more reasoning um >> a lot a lot of a lot more like post- training and reinforcement learning right um and and that's the other point is that um post training is very inference- like um so it also you know tilts to more bandwidth so um we can see that the the amount of compute that's
more inference sensitive which is in the blue and the white bars post training and just your standard inference um is has really quickly dominated um whereas tra classic pre-training which does require more more capacity um is is much is has become a much less lower share of the total pie of compute where it goes in terms of the frontier labs. >> Yeah. Yeah. And our tokconomics model is >> Yeah. Yeah.
And our tokconomics model is is tracking this in great detail.
um sourced from a lot of uh a lot of work from the data center model that the guys have tracking sites.
But it's pretty clear to tell if a given site is um targeted for pre-training or not because you know of the implications on the network and how big the individual building needs to be.
building needs to be. like if you're going to have more than 100,000 GPUs in an individual building or or like an interconnected campus um of multiple buildings like that's a very different data center design compared to you know 2,000 4,000 8,000 GPUs spread across
four data halls and then connect them all around the world and um that's you know it's clear that as the labs bring on more compute in a capacity constrained world right now that they are the marginal data center that they bring into their fleet is not going towards pre-training. It's going towards
It's going towards >> research, post- trainining, inference, some collection of other that is uh >> yeah not not um not pre-training runs for that 20 trillion parameter model that we were thinking about. >> Yeah.
>> Just two three years ago, right? >> So um >> exactly.
So why don't we talk about the maybe the the specifics on like the uh relative price?
I think uh this is something you mentioned right at the beginning.
If maybe you could just explain it once again in a different way.
The um the idea that going from 12 high to 8 high to four high has a different ratio of bandwidth to capacity that causes people to >> you know if they have a constraint on the total amount of bits of memory that they can >> into a given, you know, in let's say an unlimited number of packages.
Uh well then marginally you're not going to want to stack more capacity or want to have more bandwidth.
And so therefore the individual stuff is smaller.
And so just to finish the um thought experiment physically we can't go lower than four high, right? We can't go to one high. Can't go to zero high. [gasps] So what is it? What's the limit here? >> Yeah. Yeah.
So, um the HPM bandwidth is pretty much I guess fixed within each stack or queue.
Um regardless of the stack height or or set another way, you need at least four four layers of DRAM to get the full bandwidth out of a cube.
Um because the bandwidth is you know driven by the interface between the HBM and the the SOC.
Um and you know for HBM 4 and 40 there are you 248 um IO's um data wires between the cube and and the the comput um and then you know depending on whether it's four high eight high 12 high these like you know IO's are kind of divided across um up the stack um so each cube you can support up to 512 12 of these I guess uh signal IO's.
Um so you need four to sort of max that full like utilization of that that whole you know48.
Um but of course you know because uh your um the suppliers they they charge you for for the capacity right um so they um so customers pay on a you know effectively on a dollar per gigabyte basis for HBM that's how they you know price it um so you can have uh but the bandwidth is the same so if you're paying for more capacity with 8 high or 12 high um you're paying you much more
you know almost almost double or triple um what you pay for the same cube before high HBM um but if you're really bandwidth focused um you're getting the same bandwidth from from each of these cubes regardless so the dollar per band pan proposition is much better after four high um so so paying for four high only um when all you really need is that you know bandwidth and four high capacity gives you enough. Um, that's that's almost a
Um, that's that's almost a free launch.
>> So, can you talk a little bit about the different hardware architectures that may be taking advantage of this?
Obviously, the headline for this entire podcast is the fact that, you know, we're making it very clear to everybody with a nice big chart up front that Reuben Ultra was going to be a terabyte and it's now being revised to 192 gigs, which is 192 GB is less than the 288 GB that's shipping today that people are using today.
I've used today in GB300, NVL72 and what's coming in Reuben this year, which is like the first time in the history of me keeping track of Nvidia that the next Frontier GPU, like their flagship GPU, is going to have less capacity than the previous one.
>> So, that's just, you know, a little bit mind-blowing.
>> Are you expecting that you're going to see this?
I mean, you didn't call any of this in the article, right?
But just high level, you're expecting to see other accelerator vendors follow suit if Nvidia is doing this. >> Um, in short, yes.
Um, I think the the capacity challenges that or the supply chain challenges that India faces are are the same for for everyone.
Um, and you know, again, I think a lot of uh road maps have been reset to factor this into mind, right?
Um so um they want to make they want to be able to sh maximize the number of accelerators that they can ship.
Um and you know if they can make the compromise on KPN capacity per accelerator um to make this happen um I think it's yeah it's it's the most logical path going forward.
Um especially given that um you know look that it's it's really bandwidth that's important.
Um and especially given the uh the you know KPM has always been expensive and it's going to get more expensive from next year as well.
Um so yeah in summary a lot of uh customers road maps have been reset um to favor lower height HDM instead of uh instead of that 12th high.
So everybody is looking obviously um for how this is going to impact both the accelerator vendors and the memory vendors.
I know that we've uh on the memory supplier side we've kept some of that behind the pay wall.
So, we won't get into that on in detail, >> but um you made a couple of of things public that are not behind the payw wall >> specifically saying that um this is going to have an impact on DRAM supply in like the typical server supply chain because effectively the amount of HBM cubes that are harvestable compared to 8 high could more than double. That's what the call is.
And I I think doubling makes sense, but why would there be more than double the amount of HBM cubes that are harvestable on four high?
And maybe you can talk a little bit about what the downstream impact of on the rest of the supply chain is on the at a high level like logic wipers and substrate and PCBs and stuff like that. >> Yeah.
So uh HBM4 like um basically the taller the stack um it is harder to manufacture um that you know finished cube.
Um basically you know it as you stack uh as you stack each uh die in that that module um there is going to be each stack has that you know each uh layer additional layer has some some yield loss right um so it's not 100% yield for the processor like building this um this this this module um so um if you compound that um by 8 or 12 times Um it's say if it's like 99% yield for each each layer.
Um this is a simple example that 99 that 1% yield loss compounded eight times or 12 times is much higher than the loss if you compound that um only four times which is what you get from eight high.
So um there's less yield loss.
Um there are also other I guess um in terms of just like being able to send power up that stack um it's much easier doing it um than a four four DM die stack than say eight or 12.
So overall basically you you get more than double because um the yields are much better um than than eight high or 12 high.
Um and then in terms of um you know what this means for >> it's it's it like a high level for high is a product that has a longer manufacturing history and is simpler to manufacture like they're going back to >> it's simpler to manufacture.
There's less uh there's less steps to making the cycle. >> Fair enough. Okay.
Um so moving the bottleneck away.
So like you have a bunch more um HPM cubes but the bottleneck is now no longer wafers and there may be other bottlenecks in logic wafers or substrate or PCB that you might fork on. >> Yeah.
>> Yeah. So um let's say we so everything goes to four high um and then we can get double the cubes than if it was eight high uh or more than double the cubes than if high um then I think you the bottleneck moves away from probably moves away from HPM into uh probably you know logic right leading edge logic in terms of okay can TSMC uh support enough logic wafers to um to you know to co- package alongside all this HBM um and then there are other
chain like >> right double the HP >> yeah yeah yeah can double the number of logic that you know capacity required >> you know that's uh probably not quite um and then there's also other parts of the supply chain right like um you know one of the tightest areas is in substrates um And again, if you can you double the sort of amount of substrates um through the baseline um compared to baseline um which is going to be a big ask. Um and
Um and then there's also of course you know things like power uh which we've spoken about a lot as being a big big constraint.
Um but it doesn't matter if we we can't say maximize the use of of all this HPM.
Um you know it's it's not only HPM that's tight.
Conventional DRM is very tight.
Um and we're seeing a lot of like um especially in for for servers we're seeing a lot of dspecing there just like taking down the per socket DRAM.
Um again because there's not enough DM going around um and then you know could see you know a lot of that is driven by um DM wafers being cannibalized for HBM production.
So if if we relax some of the HPN uh wafers that we need uh for more server DRAM that's that's also um great for the the whole industry um because again um you know when it comes to Gent AI uh it's not just um it's not just uh AI accelerates that we need we need a lot of CPUs just to perform all the tool calls um etc.
Um, and a lot of the time we're also bottlenecked by just like that waiting for a CPU task, right?
Um, and that's why CPU demand is so high, but that's also also the CPUs are running it probably going to be, you know, poorly utilized because they don't have the the right enough DRAM to support them.
Um, so that's that's also uh another benefit of having sort of relaxing HBM constraints.
um you can use them for uh commodity DRAM.
Um but I think the another way to look at this is that um you know at a you know we're in a DAP constraint world um and in terms of maximizing tokens per HPM wafers um this is like going to fall high is how you sort of best optimize for that.
Um, in the same way we talking about like tokens per watt, tokens per dollar, uh, because watts for dollars are valuable and and scarce resources.
Um, HPM wafers are valuable and scarce resources.
Well, DM wafers are valuable and scar scarce resources too.
And if you want to, you know, deliver the most tokens on aggregate, um, go for high is how you do that as well.
So maybe we can end on um some a like burning question that I always get here which is related to how do people actually expand supply of memory it's in the news a lot right like um whenever people say we are constrained by memory not by logic from TSMC not always from data center space
although there are local constraints it's like the it's clear that industrywide uh all accelerator vendors would be able to produce more than they're producing today if they could get their hands on more HPM assuming they have HP in the >> generally speaking HPM or DRAM for the service. Um,
Um, however, you have to jutapose this with the fact that we're also seeing stories about SKHEX printing, you know, massive profits, people going to the sole Lamborghini dealerships the day after the paychecks are cut and buying out all the cars and driving around and a Samsung's doing well, and Micron's doing well.
So when people say, "Okay, do you have do you have clear like a clear um answer for why can these memory vendors that have raised prices into the massive demand shock made a bunch of money?
Why can't they just produce more?
can't they just produce more? Why can't they increase supply and respond to the demand signal faster than you know it taking your >> I mean on that >> I think it it all comes down to like >> uh semiconductor manufacturing is fairly long lead time um in terms of like there are so many things that you need to do
um before you can just you can't just like click your fingers and and add waste capacity Um if if your lines are fully utilized um [clears throat] you need number one you clean rooms uh which has been probably the the number one reason why um why we haven't had you sort of short-term capacity being added. Um you
Um you need to you know build these clean rooms then you need to like fill them up with equipment.
equipment. Um and then that equipment is also very limited because um you know there are there's only you need at the very least you need like EUV like EV tools or tools from ASML and they can only manufacture so many of them a year
because they have their own supply chain that's very long and and to add more tools you know they have to like you bring up like um get their suppliers who do like a very specializ iz thing for them like like make very smooth um mirrors and things um and that takes a while as well. Um so unfortunately it's
Um so unfortunately it's just these physical constraints um that prevent that capacity from coming on um on demand as we need it.
>> So what's the high level forecast for everybody keeps saying when's the HPM or memory you know crunch going to ease off.
What's what's the current timeline you're telling people is the earliest time at which we could start to produce a lot more memory and have everybody >> I mean we don't uh I think we subscribe to the memory model to find more but um
I'd say the TLDDR is not uh not within this decade >> not within this decade 2030 here we go all right you heard it here first folks Um, all right, Myin, anything you think we uh we missed talking through this article? >> Um,
>> Um, oh, I think the the counter argument is um you know what um like at what point does full high is full high not enough?
Uh and I think that's basically again it it comes back to model sizes and you know how many parameters they are.
parameters they are. um if model sizes um explode a lot then the economics start to really favor you know it's less favorable to fall high um that sort of additional capacity that API gets you is much much better is much more valuable
and the reason it is it comes down to like batching economics in that um you know with batching you you know for every user you only need one read of um one read of the weights and the larger the weights are I the more efficiency gain, extra currency gives you. Um, so
Um, so that's when um, you know, additional capacity is more valuable.
Um, so I think we we did a we also did a um, same way for K3.
We did a like we modeled something like 3x K3.
Um, if you have a model that's three times the size of K3, then the foot gain from 8 high and 12 high is much better and and is probably worth the cost.
Um so um so the counter would be is like okay um if if everyone goes to you know four high um do they shoot themselves in the foot and have a like suboptimal system if in a world where these model sizes grow to be huge um again this goes back to uh um I guess model I guess yeah what we're observing is parameter parameter sizes aren't really increasing that much.
Um there are so many techniques that are like trying to keep that footprint small.
Um loop transformers is a new is a relatively recent one and it's basically confirmed that you know GPD6 Astra uses loop transformers.
loop transformers. um which means that um you know instead of like adding parameter count and input you know goes through the the layers sort of more than once right so you add compute depths but you don't add increase the size of the
model um and then the other point is that you know we're seeing that the the loudest cries for high are coming from the labs um and you know the labs are probably the people that are best positioned to to to say, you know, which way models are scaling. And I think that's also a tell
And I think that's also a tell that they're not seeing, you know, parameter sizes scaling as aggressively um in their road mapaps, right, in you know, what they find in their research.
Um >> yeah, look at at a simple level, I've made this point previously, but it would be very embarrassing for OpenAI and Anthropic if Kim K3 was this close to performance with with them.
>> And it's sitting there at 2.
8 trillion parameters and they're at >> 10. Exactly. >> Right. Yeah.
>> It's kind of implied that the models are of similar size to Kimmy. >> Yeah. Yeah. um today.
Now, let me uh let me kind of push on this a little bit further.
So, when you think about the future of these accelerators, um do you think it's possible that we reenter this world where there's multiple SKUs uh for an accelerator where you you know it starts at a certain uh amount of memory capacity and then maybe a revision comes later that's bigger.
Like I'm thinking back to V100 launches at 16 gig and then goes to 32.
A100 launches at 40, goes to 80.
H100 starts at 80 and H200 is basically a rev of it to 144.
There's this trend of of like having these minor revisions to accelerators, but look, they're producing so many of them and the customers are so big.
Can't there just be multiple SKs?
Can't somebody have a four high skew and then somebody else buy an eight high skew of the same, you know, base logic die?
You've got the meta flavor of the Reuben Ultra and then you've got the OpenAI flavor of the Reuben Ultra that have different >> Yeah, absolutely. Um, absolutely.
Um, we're already seeing that, right?
Um I know that for instance Meta has had a lot of uh custom or semi-custom SKUs where they'd have uh you know different memory from the mainstream configuration.
Um and and we're seeing you know for for instance for MI450 there's a custom meta version with uh that has eight eight high hpm instead of 12 high.
Um so absolutely I think you know a sort of more skewing um in a world where where you know supply chain capacity is really tight um makes sense in that you know it becomes much more expensive or costly to like overprovision you know certain resources uh if customers don't need them um so yeah I think more skewing is absolutely on the cards >> makes sense man Well, we're on a trend.
Everybody wants to understand more about the accelerators and the models and how everything all works.
So, going to be a fun couple of years as this stuff gets produced.
I'm actually fascinated that uh it's only happening in the Reuben Ultra generation.
Like, everybody kind of got away with these 288 gig Reubins and nobody is pushing that hard for the current generation to change at all.
It's just the stuff that's coming this time next year or like into 2028 that we're actually going to see this all, which maybe is an important thing.
We didn't make that clear up front.
Reuben Ultra is the, you know, if there was an R200, it would be the R300.
>> Um, >> so it's, you know, it's the it's the next version after the the stuff that's getting put in right now. >> Okay, man. Uh, good job. >> Okay, cool. Thanks.