0:07
Hey, hi. I'm Karthik and Rachel.
Hey, hi. I'm Karthik and Rachel.
So, this is Wega from SemiAnalysis.
Um and then we have um the the two um partners from AWS today with us.
Um we're doing a uh you know, sort of discussion on on your um partnership with Nvidia.
Um and then, you know, sort of your statement on the 1 million GPU deployment.
Um and you could you maybe you could elaborate more into that.
And then uh lastly on a discussion on training and the progress there and and the partnership with Cerebras. Okay. Okay. Okay. Yeah, I can go first.
Uh well, thanks for it's great to be here at GDC uh the GDC 2026 and thanks for having us here.
Um yeah, so we I mean, our partnership with Nvidia goes back almost 15 years.
We were one of the first cloud providers to sort of start offering GPUs before when, you know, the GenAI was was really a thing, right?
So, a long long uh partnership where we have a very deep deep partnership.
Um I mean, um um we've been offering around like, you know, uh uh if you look at our infrastructure today with the new GPUs, we offer about 2 million GPUs in in in in the cloud, the largest missile provider to offer 2 million GPUs via cloud.
And you know, and we and we mis- recently launched a blog saying that we're going to add another 1 million GPUs just this calendar year. Yeah.
So, basically 50% of what our entire footprint we've built in the last 15 years.
And I think that says sort of like uh that says many two things.
One, the the huge customer demand we're seeing on AWS for these for these um infrastructure.
And also the testament to our ability to scale Yeah.
so fast and in such such a short period of time.
Um so, I think that's what I said we deeply value the Nvidia partnership.
Um and uh looking forward to sort of bringing the next generation of Rubin GPUs this year or the Rubin uh systems.
And our customers are excited uh to to have them.
So, could you go more into like maybe um on like what are the challenges to scale and how is AWS you know, get better at that um or what are your advantages there? Yeah.
And also, when do you expect Rubin to come online? Yeah.
Yeah, so I can start with that.
Uh so, I I am Randa, just Rachel, and I lead product marketing for AI infrastructure at AWS.
Um so, a common problem that we're seeing with customers is that it's easy to build a very cool demo on GenAI in days or even hours, as you can see it from all these trade shows and industry events.
But when customers really hit a wall is really when they deploy these AI solutions in production, um it's not just about deploying application on GPU instances, but it's also everything around it, like the data pipeline, uh the cost control, and the security, and compliance around it.
So, and that's really when AWS can help with our customers with our or like 20 years of uh operating infrastructure at massive scale. So.
Yeah, I mean, we we offer uh the largest uh the broadest and the deepest uh sort of services that uh customers need really to bring these applications to production versus, you know, doing a POC uh at a much smaller scale. Yeah.
But I think back to what Rachel was saying.
So, you know, let's go more into the training inside, right?
Um So, you have uh Trainium announcement at re:Invent last year. That's right.
Um that was a very Trainium 3 uh yeah, Trainium 3 announcement.
Uh with a brand new Scalapack architecture, right? That's right.
Um So, you know, like and then you recently announced a partnership with Cerebras um to sort of do this disaggregated um pre-filled decoding um architecture.
Could you elaborate more on that?
And then, yeah, how does that how how are your customers uh viewing these options? Yeah.
And and um what sort of applications do you see Trainium bring to your customers?
Yeah, so uh I can go first, Rachel.
So, one of the core, you know, strategic strategic principles for AWS when it started 20 years ago was that to, you know, to uh to bring down cost for customers to be able to access infrastructure.
And also give them the the wider selection of uh different hardware platforms available. Yeah.
And and and that's the reason why we've been investing our custom silicon going back to our Graviton generation of CPUs uh some 8 9 years ago.
And the same thing that's sort of why we've invested in Trainium as the the AI custom silicon chips.
We want to bring down cost for our customers and also bring them the uh the opportunity to have a broad selection where, you know, the the capacity is still available to these guys so that it's not an issue for them to scale up.
Uh so, from a cost perspective, I think we're we're uh Trainium line of products offer 30 to 40% better price performance compared to the alternative accelerators available um in the cloud.
And and and and that's an and also from the capacity availability to free up customers to sort of like access to an alternative um infrastructure, uh especially in in this environment where um you know, the supply supply constraints both on the chip side and the power side and everything in between, uh customers appreciate and value that.
And I think that's why we've been investing heavily in Trainium and Trainium 3 is our latest generation one um offering uh two to three x better performance compared to the previous edition Trainium 2.
Um we plan to scale it to a million plus chips this year and and and and more next year.
And and some of the recent uh market customers are, you know, it's opening eyes.
And obviously, Anthropic has been asked heavily for training and inference.
So, we're sort of super excited about partnering with AI labs, but at the same time, I think we're also sort of like seeing a lot of interest from a broader set of customers, especially in the startup arena, uh looking into looking at using Trainium for their uh for their workloads.
Is that for their um chip development workloads?
Yeah, so it's from AI labs.
That's that's Yeah, so for in in terms of the AI labs >> Oh, AI labs, yeah.
Yeah, the OpenAI labs probably I mean, they're the they're they are and will be using it for um uh large-scale training and and and and industry inference workloads.
Um um and uh and and and and recently the announcement of the the Cerebras and uh and the Trainium uh product with uh with supporting disaggregated inference. Yeah.
And the team is also similar there, like bringing down the cost for our customers, especially for inference, bringing down the dollar per token. Yeah.
So, that I mean, that's something that we see customers sort of like tell us, "Hey, you know, for for us to drive deploy more uh GenAI applications, is we want to see uh cost come down and and uh the partnership is sort of uh is is the primary goal is sort of try bring down cost for our customers. Okay. Thank you.
Um Yeah, I think um thank you very much uh for for going through the the efforts we uh you know, helping the ecosystem build and uh the efforts you're putting into the infrastructure um and then, you know, offering the widest and the broadest range of products for our customers. Um yeah. Great. Thanks for having us.
Thanks [music] for having us.