howdy ho this is dylan with semi-analysis and i'm here to bring you some interesting and new information on nvidia's ada lovelace including some of the leaked specifications some die sizes architecture and cost analysis nvidia was the victim of a cyber attack at the end of february where they were hacked for a vast sum of data this hack was not only a disaster for nvidia but all chick companies and the national security of
0:30
all western countries among the hack data was detailed specifications and simulation data for nvidia's next generation gpus hopper and ada hopper is now sampling and was unveiled by nvidia gtc the specs of hopper match
0:46
exactly to this leak but ada which is named after the famous computer scientist and mathematician ada lovelace is still many months away from launch as such these specifications are very interesting based on these leaked specifications and
1:05
simulations semi-analysis and lacusa have teamed up to analyze the architecture die sizes and do a cost analysis even as well as product positioning we did not download any of the leaked files from the lapsis hack however many people online shared excerpts which is what these are based upon we were able to determine that these were the specs based upon these leaked specs and they are quite interesting here is a table comparing ampere and ada
1:35
lovelace the specifications of both the rest of this video will go through and talk about the block diagrams of each chip the architecture and an estimated die size we're also going to explain how we arrived at this die size given both lacusa and semi-analysis are directly supported by our subscribers we would greatly appreciate it if you could read the you know watch the rest of this video subscribe to my free or paid
2:00
newsletter and subscribe to lacuza on patreon so starting off the top dog in the ada lovelace architecture is the ad102 die and we have estimated it to be 611.3 millimeters squared it's a huge jump over the previous generation ga102
2:23
in terms of performance right we've got 70 percent more cuda cores coming from five additional gpcs the memory bus width remains at 384 bit bus though so we can expect memory speeds to improve slightly to somewhere in the region of
2:40
21 gigabytes per second maybe a little bit higher uh but mostly around there so so this increase in memory bandwidth is not enough to feed the beast that is ada lovelace 102. as such nvidia has seems to have inc added 96 megabytes of l2 cache which is way larger than ampere's ga 102's 6 megabytes of l2 cache now we can use this and and point and look at an amd's navi 22 gpu right and amd's navi rdna 2 gpu architecture has a feature
3:17
called infinity cache which is a large l3 cache we hope nvidia names their large l2 infinity cache just as a troll but yeah that's probably not going to happen the marketing folks probably have a better idea but you know based on this we can see that you know there's there's a difference in cash hierarchy between the two vendors but we expect the general trend of hit rates to be largely the same as such you know amd has hit rates of 78
3:47
1080p 69 at 1440p and 53 in 4k so these these hit rates are basically when you don't need to go to memory to uh in the process of rendering something and so they reduced the memory bandwidth requirements nvidia's l2 work should work largely in the same manner and will help feed ad-102 despite the small increase in memory bandwidth the top-end configuration of ad-102 should probably be 24 gigabytes of gddr6x for the gaming market
4:22
but we expect it to be cut down from this further the next gpu in the lineup would be ad103 ad103 is quite interesting as a configuration and we've estimated it at 379.69 millimeters squared versus ad102 it is a huge downgrade this may be the largest gap between the top die and the second die in a gpu generation with ad102 having more than 70 percent cuda cores versus ad103 the other interesting thing is that the cuda cores are the
4:56
exact same as the current generation ga102 despite this the memory bus comes in at 256 bit so this is the exact same as ga104 in the current generation ampere so with a much smaller memory bus than ad102 at 384 um this this means that gaming gpus based on ad103 would max out probably at 16 gigabytes but cut down variants will likely exist despite memory bandwidth being much lower than ga102 the inclusion of 64 megabytes of l2 cache
5:32
should probably allow this gpu to be fed given nvidia is utilizing a custom tsmc 4n node we ex which is a variant of tsmc's n4 weeks or n4p we expect that they will be able to clock much higher than ga102 the increase in
5:52
clocks combined with some of the architectural enhancements that we expect which we'll get to later in this video will allow 8103 to perform better than the current generation flagship rtx 3090ti if they bring it that is if they bring
6:07
it to the desktop with a high amount of power consumption it's important to note that ga103 never came to desktop and is really only available in the top end of laptop gpus so this could happen again with the ada generation ad104 is estimated at 300 millimeter 3.45 millimeter squared this is the sweet spot for ada it's this is due to its performance and cost effectiveness the gpu is quite small which is nice and it also only has
6:39
192 megabyte bit bus uh 192 bit bus which would lead to it having 12 gigabytes of memory for the gaming gpu market which is you know high enough capacity for you know current use cases and developing games but uh or newer games but it also keeps the bill of materials the bomb down to a reasonable level right if you go to 16 gigabytes that would just be too costly simultaneously nvidia gpus tend to have their 104 designs have similar
7:11
performance to the prior 102. if this trend keeps up the cost to performance would be excellent in fact it may even have more performance than the current generation ga 102 given nvidia is likely to pump clocks quite a bit to hit performance above the 30 90. we expect nvidia to go as high as 350 watt or maybe even 400 watt for ad104 desktop gpu with gddr6x as such we expect this gpu to be the one that most enthusiasts end up purchasing because
7:44
it's still within a reasonable price range the gpu will also be highly efficient right you know with with g6x you know that there is some inefficiencies but when we move down to the laptop world it can be very
7:59
efficient if you you know clock it to the right levels and keep it with regular g6 memory at you know 14 or 16 gigabit a second in the and the total power consumption you know around the 90 to 135 watt arena ad106 is the true mass market gpu of
8:16
this generation it's estimated at a tiny 203.21 millimeter squared it's likely going to be the highest volume gpu in the lineup as the 106 gpus are always the highest end at least they were for pascal turing and ampere due to the tiny 128-bit bus you know the memory will most likely max out at eight gigabytes in the top configuration we expect it to perform very similarly to ga104s 3070. it might even beat the 3070 and get up
8:50
to 30 70 ti levels that assumption might be optimistic given there's only three gpc's in ad106 versus the six in ga 104 but nvidia might be able to solve this you know with higher clocks and with their large 32 megabytes of l2 cache you know with those high 32 megabytes of l2 cache the gpu cache rates are going to be 55 hit rates for their cash is going to be 55 in 1080p 38 and 1440p and 27 in 4k uh so this would this would match you
9:28
know this is based on amd's navi 23 which also has a similarly sized 32 megabytes of l2 cache so before we move on to the you know the baby of the generation ad107 we want to give a bit of background the data posted on twitter from the leak files does not specify a cache size for this gpu the prior gpus in this lineup have a 16 megabyte you know l2 cache per 64 bits of memory controller or frame buffer partition with that that would mean ad-107
10:02
wouldn't make much sense because the gpc count and the bus width are identical and the tpcs per gpu only fell by four so if the l2 cache stays the same you know with that same 16 megabytes per 64 bits of memory controller or fbp
10:20
then the die size would only fall to 203 millimeter squared um from 203 to 184.28 millimeters squared this is a tiny decrease and would not be enough to separate the two gpus in the stack instead you know we can we can look back
10:37
at turing and turing actually had an interesting phenomenon where tu-106 and tu116 had a similar relationship right where you know nvidia nvidia pulled a couple things different things out of this right they made the bus with slightly smaller yes and they pulled ray tracing and tensor cores out but you know one of the things that people often overlook is that they also reduced the amount of cash per fbp um they went with
11:03
0.5 megabytes or 512 kilobytes of l2 cache instead of one megabyte um as the rest of the turing you know the higher end turning lineup did um and so if we apply that same pattern of half the l2 cache per fbp then ad107 ends up being around 145.54 millimeter squared so that's a much more reasonable die size and it would make sense given you know product positioning and and cost and you know having to hit every point in the market
11:32
for this so moving on to ga to ad-107 you know they're also going to cut down the pcie lanes to 8x um and you know cut down the cash and cut down the course count slightly um as such you know you know they they always cut down this 107 die to eight lanes because more are unneeded um there's no need to have 16x and plus this this bottom gpu is generally in the mobile world anyways where 8x makes more sense from a power perspective as well
12:02
so this gpu would still blow the socks off of you know intel's meteor lake integrated graphics or even you know amd's you know our dna two based rembrandt or even their phoenix um you know next generation phoenix apu so this would still be you know cheap you know at such a small die size of 145.54 millimeter squared but it would also be you know faster than any integrated graphics so it would make sense in some of these you know lower
12:29
cost laptops that still want you know where the buyers still want a game um overall ada is quite an interesting lineup at the top end there's quite the interest increase in performance and probably power consumption and ad102 is a similar die size to ga 102 you know so but this is on a much more expensive customized tsmc 4n process node rather than the very cheap you know customized samsung 8n process node the density increase from a tsmc and 4
13:00
derivative is quite large relative to samsung's you know 8 nanometer derivatives which would justify the cost increase per wafer but you know likewise um you know despite being a much newer node tsmc n4 actually has better
13:16
parametric yield than samsung's eight nanometer uh despite you know similar catastrophic catastrophic yields so catastrophic yields are one a feature within the the silicon die does not function you know parametric yield is yes it
13:32
functions but it functions at the level we want it to the issue with samsung's nodes is often not defects but param you know catastrophic yields but parametric yields in that they do not you know they they can get the transistors to work but oh crap this transistor is really bad and it caught or this area of the dye is very bad and it causes the entire chip to have to have much lower clocks so you know this parametric yield
13:57
being higher is is you know very important as you know gpus are very yield harvestable design and very large dies so you know they can keep more parts of the die enabled you know which is something that we've seen on ga 104 where you know a lot of a lot of the die was disabled in the case of something like a 3080 um and a 3060 ti even so the rest of ada comes away a lot tamer than ad102 in terms of die sizes and overall bill of materials cost to
14:29
manufacture performance should generally be above that of amperes at the same power which is with a decently lower cost to fabricate despite the much higher wafer costs of tsmc's 4 process node we played around with the wafer count with a wafer costs and die calculations and we came up some with some estimations and we even you know even included some some you know various costs related to passive such as vrms and mem and you know other active
14:57
components like the memory um and in the end you know the prices are are shouldn't be too bad um i think consumers will actually win um with this generation nvidia sells the die with a markup right everyone knows that and then you know
15:11
this they negotiate a bulk memory pricing for original device manufacturers and add-in board partners to use so these add-in board partners still you know buy the buy the die from nvidia and then they buy the memory which is
15:25
already has a negotiated price from the vendors of memory and they integrate that as alongside all the power components and cooling and etc so you know the the die cost uh has a different effect you know dollar for dollar you know what how much more a die cost versus how much more memory or vrms cost and so nvidia seems to have balanced their l2 caches with the memory size and memory bus width optimally right you don't want to
15:53
go too wide on the memory bus or too narrow because then that's going to influence your memory capacities on these gpus memory sizes seem to stay very reasonable and most gpus will probably have 16 gigabit dies of g6x or g6 um in in general ad104 is going to be replacing ga102 in terms of performance you know actually beating it some while also costing less while ad106 is going to replace you know the current ga 104
16:24
based dies in the performance tier right so 37 e's they think of that that level uh and so in in when you when you think about it from that perspective the memory cost is identical right because ad104 should have similar memory sizes to ga 102 and 8106 will have similar memory sizes to uh ga 104 8106 to geo 104 um you know they will move from 8 gigabit to 16 gigabit um but the cost to fab the die is actually going to be less despite the
16:53
more expensive node because of a much larger you know much much smaller die uh board components such as packaging cooling power components are actually going to end up being cheaper generally due to the better efficiency
17:06
lower power consumption and smaller boards that these gpus will require so you know when we compare dyes in the stack you know ga 104 versus ad104 there is going to be a memory size increase and there is going to be a slight price
17:20
increase in base msrp and but you know the issue with that is ga 104 is only eight gigabytes right or you can go you can go with 16 gigabit dies or double sided and get 16 gigabytes of memory but then you end up with it being too expensive so you know ga 104 is really not in a sweet spot nowadays for memory because of this issue with the bus width um whereas ad104 you know moves up to 12 gigabit 12 gigabytes due to their 192 bit bus
17:51
fears of higher power consumption should definitely be taken into consideration nvidia is likely pumping power for each die similar to what they did the prior generation um in fact we imagine that they're going to push power to what one die higher in the stack did you know so by that me we mean ad104 will likely reach 3080 level power levels of power consumption and ad-106 likely reaches 30-70 levels of power consumption and
18:20
rumors point to the top ad-102 breaking records for power consumption um so so next we'll talk about how we arrived at these die size estimates the first step in the die size analysis was to gather architectural changes
18:34
regarding ada and comparing them to ampere so the sm architecture is 8.9 versus 8.6 on ampere so this is an improvement but it's a generational improvement it's not you know a whole new sm architecture as a result we can assume that the sm
18:51
die size increased you know sm as area 4 and sm increased about 10 percent um and while we aren't sure what these sm architectural changes are they likely could be you know something like a 192 bit kilobit l1 cache or new tensor cores um the highest probability change in our minds is actually the addition of a new third generation ray tracing core you know other architectural changes are on the i o front right so the the leak
19:18
indicates that nv link has been removed entirely from the lineup which indicates nvidia is not going to push the ada lineup for multi-gpu sort of data center and professional visualization applications right there's you know the
19:31
rtx a6000 which can do multiple gpus and be used for these professional visualization and even data center applications um but that's not going to be possible with ada where the 102 die looks like it doesn't have any nv link
19:49
we expect pci 5.5 5.0 and a better memory controller for the slightly higher speed gddr6x and also displayport 2.0 uh you know all these changes to be included there's also the high likelihood that nvenk and nv deck uh you know the encoding engine and decoding engine are upgraded to include you know av1 and and you know some more feature levels around that the biggest change with ada is of course the l2 cache at least on an area perspective
20:25
instead of the small l2 cache you know nvidia's taken a picture out of amd's infinity cache book and used a much larger cache across the board given we have most of the specifications architectural changes and configurations for the lineup we can actually use ampere's ga 102 ip blocks to recreate a hypothetical gpu with the same specifications as ad102 on samsung's eight nanometer so this would be using you know the exact same cache sizes gpc
20:58
sizes memory memory controller sizes etc as ga 102 and you know we would get we would have we arrived at a die size of you know sixteen hundred and twenty nine point six millimeter squared that's if if level it if lovelace ad102 had the same architecture and configuration on uh on samsu on samsung's eight nanometer um this doesn't contemplate you know the changes in sm architecture the larger encoder decloader blocks pcie 5.0
21:32
displayport 2.0 or you know the tweaked memory controller and some of these other changes you know the larger but it does contemplate you know using the same configuration with this die size of 1629.6 millimeters squared you can immediately notice that the l2 cache is titanic right this is huge this is larger than you know amd has their large l3 cat infinity cache on navi 21 but they don't allocate such a large area dedicated to
22:04
the cache and yes while amd is using a denser tsmc n7 node this is only a small part of the puzzle actually most of the difference in density comes from a difference in layout and configuration of that cache so ga 102 uses 48 128
22:23
kilobyte slices of sram with with a one megabyte you know l2 cache per 64 bits of memory controller or frame buffer partition um if we look at other places in nvidia's lineup you know ga 100 that uses 80 512 kilobyte slices
22:43
so these larger slices have vastly improved density um you know in as seen with the comparison to amd's l2 cache the density increase of ga 100 is far more than that just that of the you know process node shrink and yes while amd's not as good as nvidia and many elements of design they are undoubtedly better in areas of cash and packaging we believe a lot of this stems from their cpu teams and the pedigree there so amd is very very good
23:17
at making extremely dense high performance caches and they applied this to their gpus with infinity cache in in fact in our final die size estimates nvidia's 96 megabytes of l2 is still not as dense as amd's 96 megabytes of l3 infinity cache um and you know so regardless of the shrink from samsung's 8 to tsmc4 that alone wouldn't make ga 100 102's building blocks reach a reasonable die size um instead there is a architectural
23:53
rework from the ground up required for the cash design the leaks indicate there is now 16 megabytes of l2 per 64 megabytes of fbp which is you know very different from ga 102's one and so we estimate nvidia is going to actually use 48
24:11
24 and 48 kilobyte slices hood 48 two megabyte slices of sram for their s for their l2 cache rather than 48 128 kilobytes so a huge increase in you know size slice uh size which would greatly increase density you know with this cache
24:32
configuration we can actually even calculate the theoretical cache bandwidth uh with these figures so amd has 1.99 terabytes a second of infinity cache bandwidth on their navi 21 gpu when it's clocked at 1.94 gigahertz if we assume nvidia is running at the same you know 1.94 gigahertz on ad102 they would actually be able to achieve 5.96 terabytes a second of bandwidth with their l2 so clocks will differ in the end products but we expect clocks
25:05
for lovelace to you know ada lovelace to be somewhere around 2.25 gigahertz which would be realistic in a desktop platform right whereas with our dna 32 you know it can clock higher than you know this clock speeds that we indicated before and in fact rdna 3 probably on desktop goes above 2.5 gigahertz how much we're not sure but most likely it goes above that figure nvidia has you know always made the design choice of density over
25:33
um clock speeds whereas amd's designs tend to be a little bit more relaxed in comparison but that doesn't apply to the cache actually you know nvidia could have introduced a higher density you know 8 or 16 megabyte per slice but then they would have to reduce the number of slices vastly which would reduce the bandwidth which you know isn't tenable given the way you know l2 works um on amd's architecture with the delta color
26:01
compression and other other aspects we came up with estimates for what this different cache architecture would do to the area of ad102 you know using the ga 102 building blocks and so once we accounted for the
26:16
architectural changes we also applied shrink factors to tsmc's n7 node and then another one to tsmc's n4 node and you know these these shrink factors are not homogeneous across the die right analog shrinks differently than sram which shrinks
26:30
differently than logic so the sram appears to be about a 60 40 split of sram ii logic which influences the sram shrink figure that we used that's for the l2 cache and then for the rest of the die we applied a 10 gross up
26:47
factor to the sms to account for any architectural changes there and we also had some different shrink factors with the various pieces of digital logic because their mix of sram to logic was different the sram percentage on these other
27:01
pieces of logic was generally around 30 percent and 70 percent logic so this made the shrink factor higher um lastly we kept the analog portions of the die identical you know there is a small shrink you know it um from from you know samsung's eight to seven tsmc7 to tsmc4 but it's not that large um and these would probably most likely be equalized with any potential upgrade to pcie 5.0 5.0 um a slightly better you know gddr6
27:28
memory controller and phi and displayport 2.0 so you know the analog portions we kept pretty much identical obviously nv link was still removed um and that's how we arrived at six eleven point three millimeters squared and actually that that you know nicely lines up with a you know what copied seven kimi the famous nvidia leaker has stated you know he stated the die size was around 600 millimeters squared um you know us getting six 11.3
27:55
seems pretty pretty pretty good and it helps us feel uh confident in our estimate so after gathering you know this small overview um we could use the configurations for the rest of the lineup right gpcs tpc counts l2 sizes you know the command buffers the various differences in fives right whether it's bus width or pcie size you know display controllers you know the crossbar sizes all of these can be dynamically scaled based upon the gpu
28:20
configuration and so you know these shrink factor figures can be used um that we we you know that we determine based on some statements from tsmc plus what we've measured in real products and in the end you know it is a bit of a shot in the dark we're not going to lie to you but we think it's an educated guess onto the die sizes for ad107 we will note that we backed off on the architectural shrink uh due to the smaller amount of cash per fbp
28:50
frame buffer partition so you know the the density gains from you know larger larger l2 cache you know slices do back off slightly ad107 in our estimate overall ada lovelace isn't a massive departure architecturally from the current ampere architecture but the changes it brings such as improved retracing cores and you know improved encoding encoders for av1 larger l2 cache you know these changes bring performance up considerably but they
29:19
also keep cost down despite being on the expensive you know more expensive custom tsmc and four node nvidia's keeping in is keeping with the tradition of keeping memory sizes well balanced across the stack with a mod and moderate increase in memory size per level right they're not going to you know too large of memory sizes which would influence the cost of manufacture and and essentially the cost to the end consumer instead they're you know
29:44
keeping it modest and they're balancing the size of their l2 with that memory bus with um so you know that's that's a good good thing to point out um versus amd you know rumors point to very high performance uh but also you know very
30:00
high cost at the top end of the stack due to you know this hybrid bonding that they want to introduce um allegedly but we're much more interested in navi 33 um as a chip which we uh assume that it will fall somewhere in
30:16
between ad-104 and ad-106 probably closer to 8-104 but the range is large because you know that we don't know the exact you know architectural changes but leaks point to it being a very good competitor in the mass market segment and you know while amd is currently massively behind in gaming performance in you know current or next generation games depending on how you want to see it due to their you know lackluster tray
30:40
tracing performance um and you know they lack many differentiated software features such as dlss and rtx studio um and those hurt competitiveness but you know navi 33 is going to make this the most competitive gpu generation in a
30:54
decade you know nvidia is the reason they're pumping clocks is partially because they want to optimize cost and performance and maximize their margins and performance that the end consumer gets but also because they're you know
31:07
frankly scared of what amd is going to produce um and so you know with gpu prices falling fast as ethereum 2.0 slams on the brakes from writing demand you know the you know fault you know the slight fall and you know the crypto market has caused you know mining to become a little bit less profitable and then you know consumers are shifting their spending away from a mix of you know goods more towards services and all of these things
31:32
combined with you know higher inflation uh across the whole broader market mean that demand has gone down and prices are coming more in line and as such we predict ada lovelace and in rd and a3 gpu prices uh at least for navi 33 to to be quite good in the price to pro you know performance market of 400 to a thousand dollars you know there'll be a couple gpus scattered through there and the top you know the top end of the stack will be well above
32:00
a thousand right both nvidia and amd are going to increase prices on the top of the stack but you know which is great for you know you know the you know hyper enthusiasts but you know even for normal people that can't afford you know such an expensive gpu i think this is going to be the best generation for consumers in a long time right much better than you know the short-lived and pure uh much better than turing much
32:23
better than ampere right you know so so um you know the short-lived amp here in the sense that prices skyrocketed a bit after release but in short you know price to performance is going to go up um and at the top of the stack costs will but so will performance move up you know considerably and consumers will win um so thanks for watching and uh please subscribe to the channel to my free or paid newsletter and to lacusa's patreon thanks bye