Nvidia's Acquisition of Mellanox, Plan for the Future

0:00

hi this is the baseless speculation podcast with Spartacus and Dylan and today we're gonna talk a little bit about Nvidia one of the interesting things that has happened to Nvidia recently is they've announced that they've acquired a networking firm called Mellanox for almost seven billion dollars given what Mellanox does some have been a little bit uncertain as to

0:29

what this new acquisition will do for graphics focused Nvidia however we think we have a pretty good idea on how they'll use it Dylan do you want to talk a little bit about that yeah so if we look back into the past we look at what kinds of acquisitions NVIDIA has done they've acquired three companies of note worthiness one was portal player and so

0:53

portal player used to make application processors for iPods and Nvidia acquired them and then Apple switched to Samsung you know a little bit afterwards so not too successful but it made sense where they were trying to go the next thing they acquired was a gia the physics company and that was in response to Intel acquiring havoc an invidious whole thing there was they wanted to have GPU

1:19

physics engines and while physics was successful GPU physics was not the third company that Nvidia acquired was ikura Akira I'm not sure but they produced modems and that made perfect sense along with portal player in terms of their whole Tegra SOC phone ambitions and they continued the development for a while but eventually they laid off everyone in

1:46

that department because it just wasn't successful you mentioned all those acquisitions all the major acquisitions of Nvidia in a pretty brief amount of time you can really summarize those because a lot of the links to what any video was doing or trying to do at that time we're pretty clear so one of the things that we wanted to talk about was

2:13

how the usefulness of Mellanox is gonna be a little bit more involved there's a little bit more to to unpack and so Dylan could you go through and just start by rolling through what what what is Mellanox do and then maybe move into how in video will actually be using them yes so Mellanox makes networking solutions so they're really popularly known for InfiniBand which is networking

2:41

between nodes of servers the other thing they do is make Ethernet NICs network in your integrated controllers the high-end ones and so their roadmap for those is quite interesting though they currently just make normal NICs but soon they're gonna start making smart NICs if you will and you didn't see me doing the quotation marks but the reason they call them smart NICs is because they're

3:06

going to begin integrating arm CPUs onto them and the whole purpose of that is the intelligent routing of data one of the things they've talked about a lot is the nvme of which is basically fabric storage you have a bunch of nvme drives connected to the network and the smart NICs helps smart Nick helps request that data from that storage that pool of nvme

3:35

that's connected to the storage in a distributed way that's similar to what I think Nvidia wants to do with Mellanox they want to create the GPU of basically a fabric GPU accelerator network so recently the OCP created the OEM so that's two acronyms but the important one is OAM which is open accelerator module and so if you've seen the voltage

4:05

chips out there they have this this sort of you know they have the PCI form factor then they have this other form factor the open accelerator module is that other mezzanine style form factor with the open accelerator module which is supported by Baidu Microsoft and most importantly Facebook the biggest contributor to this module you can do two main things you can scale and you

4:30

can decrease latency and this is done because the open accelerator module connects into the network the fabric better than past designs of accelerator cards in the past accelerators have had to go through the host aka the server CPU to get to the network integrated controller which then attaches you to the fabric what the open accelerator module allows you to do is it takes the

5:01

accelerator and it actually connects the accelerator to the host to the other accelerators and then also the network so currently NVIDIA has the env switch which connects the accelerator to other accelerators what we think Mellanox is going to allow them to do is do that third piece the expansion piece connect the accelerator to the network and this

5:27

like I said lets you do the two main things scalability and latency with scalability Nvidia can take their huge 816 GPU designs their systems and scale them even higher instead of having eight GPUs connected or sixteen GPUs and a couple CPUs or one CPU you can now do hundreds of GPUs connected together in this fabric and Mellanox is technology

5:53

of smart Nix is vital for routing data around between these various GPUs the other thing is latency so scalability dealt with training with latency you're talking about inferencing with inferencing latency is critical you have to have lower latency and that's why most latency is still done on CPUs or FPGAs and video wants two GPUs to take that crown as well like they have the

6:18

crown lip training and the way they want to do that I think is by bypassing the CPU so instead of having the data come in from the network integrated controller you know from the fabric go to the CPU and then go to the PCIe device or the mezzanine accelerator it can instead go from the network and integrated controller for to the accelerator and then the

6:41

accelerator does the processing and sends it back out to the network and then another key thing to mention on all of this is that while there's a lot of neat things that happen once you kind of bypass the host the other thing that's important Nvidia is that hosts generally is the CPU you know made by Intel and you know some other companies but not NVIDIA and Nvidia doesn't really have a

7:09

CPU that can operate in this environment in the same way that you know Intel or some others do so by solving this problem in this manner it kind of overcomes a limitation that NVIDIA has so that's that's another really interesting thing to to mention with how this whole Mellanox arrangement could work now we talked a lot about Mellanox and that's touching on that network

7:44

networking aspect we can't help but talk just a little bit about what Nvidia does on the actual accelerators so the actual processing of data and so the rumors are suggesting that that's going to be pushed forward with a next gen Volta successor called ampere so Volta was announced basically two years ago at around GTC 2017 and so we're approaching

8:19

dgt C 2019 and it's about time to see ampere at some point here we anticipate the the first thing that we're going to see with respect to ampere is gonna be that G a 100 type of part so it's gonna be maybe we might call that the big ampere it might be somewhere around 500 600 millimeters squared pretty pretty big for a 7 nanometer part based on Nvidia's guidance you know we

8:56

we expect probably more than 32 gigs of HBM too and you know quite a bit of processing power with you know next-gen and vlink all that jazz it's gonna be pretty interesting one of the really neat things that was noted by Nvidia was this guy is actually gonna be using HBM 2 based on most recent guidance and if you end up doing the math on HBM 2 you only

9:34

get so many different options if you kind of hold Nvidia to somewhere around you know 4 6 HB M 2 stacks somewhere in there and look at to be currently available you know roughly 2 to 2.4 gigabit per second data rates and it's gonna be interesting it's gonna be somewhere around you know one terabyte per second up to not quite 2 terabytes per second there's there's a little bit

10:05

of variability in there we're not quite sure what exactly that's going to be looking like but it's certainly getting close to the point where we could almost see them moving to more than 4 stacks of HBM 2 depending on how bandwidth hungry this new processor ends up being and speaking of new processors one of the things that we're hearing some rumors

10:33

about are an entirely new GPU just for inferencing and so we're hearing that this might get attached to kind of a ga 101 type GPU rather than a G a 100 so the 101 code name really hasn't been used effectively at all or at least in recent history by Nvidia so this would be a new GPU in a different space so if it looks like other modern inferencing options

11:14

then it would be relatively low power so you know it could be somewhere around 75 watts if they're really pushing for a high-density option like a lot of the other things from you know Microsoft and Google and other inferencing FPGAs and all that good stuff you know right now what Nvidia is doing is they've got its there 104 type part like the tu-104 is in the tesla t 4 you know they take a

11:51

GPU that really lives in that 200 odd watt space and they clock it down so it can live in that 75 watt space and suddenly it gets to be a lot more efficient and that improves density quite a bit I'm sure Nvidia probably doesn't like that that's not quite the area that that they want to be doing it's it's a big GPU to be to be clocking down that much

12:22

so what we're really interested to see what this kind of thing could look like it's probably going to be cutting out a lot of the FP 64 Hardware probably some of the FP 32 hardware as well if it's really going hardcore on just inferencing so it's gonna be a lot of those tensor units and we're sure it's gonna almost certainly be ready to harness a lot of the networking options

12:51

that Dylan was mentioning it's gonna be pretty fascinating that's for sure and then again we we just love GPUs over here and so we can't really stop this podcast without talking about one more thing so you want to talk a little bit about organ yeah so this definitely kind of doesn't fall into the whole Mellanox acquisition but Nvidia is almost sure to

13:18

announced orin orin being the successor to Xavier and Oran should be about five times the performance of Xavier based on some slides and roadmaps that NVIDIA has given out in the past Oran being 5x of Xavier is kind of hard to imagine if they keep it in the same power envelope and it's very likely that they'll actually increase the power envelope Xavier was you know 30 watts and that

13:48

was a huge increase from prior and prior Tegra processors you know they slowly just evolved from being phone to only tablet to now only 30 watts and I think they're gonna continue this trend at Oran will likely be hundred to 150 watts the way I arrived at that is a is pretty crude you know they're taking xavier which is 12 nanometer 16 nanometer or

14:11

whatever you want to call it it's okay you take those shrink okay that's 2x efficiency so now you're at 15 watts for Xavier performance you get you give you given video credit for you know sizeable architecture change to but you're still nowhere near where you need to be to get that 5x performance so the only way is to scale it up even further and since Nvidia thinks that the best

14:37

way to do these is to connect two of them it's probably connected via that NV link 3 that we're gonna see on these other ampere GPUs and if you look at a block diagram of Xavier it's pretty clear that this thing's going to be even more full of application-specific integrated circuits right now half the Dyess is application-specific integrated circuits we'll see will it what ampere

15:03

with a xavier brain i mean uh with orange rings and on the CPU side of things i'm expecting them to retain that 10 wide on the v liw CPU side but I expect the arm decoder to go wider probably three maybe four wide in terms of arm decoding and that's about all I have thanks Dylan as you can see we're pretty excited for GT c and we look forward to seeing what happens thanks guys