Ep. 034 - The Fight for Fast Tokens, TPU v7, Vera Rubin, and Engrams (AI Supply Chain, InferenceX)

0:04

Hello everyone.

0:04

Welcome back to Semi Analysis Weekly.

0:07

I'm joined this week by the InferenceX team.

0:08

It's been about two months since we've had them on the podcast to give a little bit of an update.

0:11

this episode we're gonna talk about engrams, offloading AgentX, TPUs, and anything else that we think of.

0:19

Guys have been hard at work, so I got Cam, Bryan, and Alec. Guys, welcome to show. So, what's up? What's up? Nice be here. Awesome. Okay.

0:29

So the latest article that came out, we're going start with that one.

0:31

It came out September eighteenth, so almost two weeks ago at this point.

0:35

it's all about n-grams and what happens when people co-design the model architecture for a capability of offloading certain parts of the model or the KVs to DRAM or SSD.

0:50

So how about we just start by talking about what are n-grams and give a little bit of an introduction there. Yeah.

0:58

Yeah, I I so I think, you know, n-grams basically have been n-gram models and the the concept of sort of like transferring more things to the embeddings have been around for a while.

1:14

Bryan can touch on more on like the history, but basically when you have a transformer architecture, each token has an embedding, right?

1:22

And the embedding is some, you know, latent space of like meaning about what that token means and as You know, you have a sequence of tokens and it you know makes its way through the layers of the transformer.

1:35

You can think of each layer as sort of building more information about like grounding the facts of what of what that sequence means, right?

1:44

as the tokens like relate to each other and and so on and so forth.

1:48

So like as it flows through the layers, it more meaning more meaning flows into the the hidden state.

1:55

So the idea with n-gram is basically instead of having these layers you can basically during training time you can have the meaning of multiple so at two grams and three grams and token embeddings kind of learned at training time so then these these can be stored in a hash table and then looked up at inference time and this is basically like for free.

2:27

I mean, not for completely free, but if you go down to the next graph or a couple graphs down, you know, the DeepSeek ran some experiments, right?

2:37

And basically on the right side at 100%, right, it's the the loss for all MoE layers. So no n-gram layers.

2:46

And then on the left side it's all all n-gram layers and no no MoE layers, right?

2:53

So you can see that the loss is actually minimized at like a a middle point, right?

3:00

So when you have some n-gram layers and then some some MoE layers and it it turns out that the the performance is actually quite good.

3:08

So that's like a high level overview, but I think Bryan, maybe you want to talk a little bit more about the motivations for n-gram and the history.

3:16

Yeah, so the paper that came out introducing DeepSeek earlier this year actually talked about the intuition for adding it, which was that models had to build up this representation of Mm-hmm.

3:32

the meaning of words basically in later layers.

3:36

So they tested this by comparing the next token prediction ability of the hidden state.

3:46

of the tokens at different layers.

3:47

And there was an earlier paper on this, I forgot what they called.

3:50

And basically with n-grams, if you calculate the KL divergence between the hidden states of the earlier layers and the later layers, you notice with n-grams, the earlier layers have lower KL divergence with the later layers, which kind of show that the model is moving towards a next token prediction earlier.

4:14

in the model instead of later, which is good because it means that you're not wasting parameters on memorizing stuff, memorizing or trying to build up the meaning of words.

4:26

And the co-design of this model, which is one thing that DeepSeek really specializes in this neck for considering inferencing while designing the architecture of the model, is that the ingredient models can be asynchronously overlapped.

4:44

with other parts of the model, because the inputs for the n-gram modules are just the token IDs.

4:50

So in their paper, the original paper, DeepSeek also talked about the trade-offs of n-grams being put earlier versus later, because if you have n-gram layers put earlier in the model, they can help the most.

5:07

And the ablation shows this, believe.

5:11

the engram layers help the most as it can bring in the memory of the meanings earlier on.

5:14

But the problem with having the engram layers earlier on is that there's really nothing to overlap it right because you have to straight away start waiting for the engram retriever to end before continuing.

5:28

So this actually shows up in Qwen 3. 8 Flash Next I think.

5:35

They put their engrams in the second layer because the first layer gives it time for retriever from DRAM or SSD to actually come in first so that there's no there's very little overhead.

5:48

But interesting part for V4.

5:48

1 Flash is that they put it at layer one so there was still space for overlap but not as much.

5:59

Yeah and this brings me to the results I guess comparing we have graphs in our article comparing the DRAM and SSD offload versions of Engram.

6:12

Jordan, could you pull up the chart? Yeah, trying to find it. Yeah, no worries.

6:18

Yeah, but while Jordan is finding the charts, maybe I'll talk about how the offload works.

6:26

So the offload uses UVA, which stands for United Virtual Architecture, think.

6:30

Yeah, Unified Virtual Addressing, And basically, it just moves the n-gram to CPU pin host memory. And yes, thanks, Jordan.

6:40

And yeah, we can see that DRAM move.

6:45

We actually tested HBM, within HBM, engraved in the OOS.

6:48

I don't think we talk the word in the article because the results aren't very interesting. Well take yeah But HBM.

6:56

take a step back here, Bryan.

6:56

Like the implication of this is just that there's a significant HBM constraint and people are innovating around that constraint by introducing engrams to make it easier to offload to DRAM or to SSD, right? And Yes.

7:11

this is just maybe a natural evolution of that pro like progression over time. Sapphetis. Yeah. Yep.

7:22

I guess one part of the thing. I was muted.

7:22

Bryan, Bryan, maybe maybe you have context on this.

7:29

Like I f I I don't know if necessarily like n-gram was a specific co-design for a restriction of HBM. Was it?

7:37

Like was that the original intention?

7:37

I mean, definitely partially.

7:40

A lot of deep-seq staff are addressing like limited HBM.

7:45

Like last year, even their whole MLA thing was about reducing KV cache for HBM usage.

7:51

And I think like this is along the same lines as well.

7:56

Because if you think about it from like first principles, you want something, you want a technique that can like be overlapped.

8:04

So either retriever or, so yeah, so just retriever.

8:09

And the stuff you get at the front of the model is basically the token IDs.

8:12

So basically, what can you do with token IDs, with only the token IDs that can be added to the hidden states and overlap?

8:20

And I think n-gram is a very natural consequence that you get to.

8:24

Actually, the long cat team from Meituan claim to be concurrently working with the same technique.

8:34

So it's really quite cool that like, I'm not sure if they collaborated or what.

8:39

TwoLabs arrived at this conclusion independently that n-grams was a good technique. Yeah.

8:49

Anyway, going back to Jordan's point about HBM, I think we talked about this in our previous Fall High podcast.

8:57

yeah, HBM is really a big constraint right now and we're trying to offload everything off of HBM.

9:04

We're heavy offloading firstly.

9:07

think Kem and Alec are quite experts in Yeah.

9:11

this field, given that we did this for AgentX.

9:11

Well Maybe they can explain about it.

9:11

yeah, I mean a big part of that twelve high, eight high, four high Rubin D spec conversation around HBM was the fact that the American labs are optimizing for HBM bandwidth over capacity or a bandwidth per bit.

9:31

in some ways the Chinese labs have already had that constraint forced upon them by the chips that they have access to.

9:42

And so Yeah, maybe you can talk a little bit about the like Chinese chips, like the domestic accelerated ecosystem, as well as the sort of already de spec' American chips and how if you just don't have the bandwidth already available, well, you need to do something about it in order to improve serving at high concurrency or like high interactivity, I guess.

10:14

Yeah, so the Chinese chips are very interesting.

10:17

In our earlier DeepSeek V4 Pro article, the DayZero article, we did talk about the Ascent 950 chip design and the HBM they use, they don't really call it HBM.

10:30

It's high-Q, I'm not sure what it stands for, but it's their own version of HBM that they manufacture in China.

10:38

And of course the bandwidth is lower than...

10:43

the stuff we get but they're very creative workarounds for their hardware restrictions and we're actually writing an article on this which will come out very very soon maybe right after this podcast comes out I'm not sure when this podcast will come out but yeah maybe I'm breaking the fourth wall Yeah, b basically the the trend the trend, I mean, at a high level, the trend forever has been.

11:05

I mean, China has been restricted in terms of HBM even more than you know the United States on a chip to chip level.

11:15

So they've just been trying to make you know the KVs smaller and more efficient, compression KV.

11:21

and then that leads to like having to offload stuff to lower tiers of memory.

11:26

And I think, you know, n-gram is just like a natural extension of that because you're you're effectively getting like as Bryan was mentioning, you're getting layers for free in a way.

11:38

Since you don't have to build up the representation and you just have that offloaded on DRAM.

11:43

And you can, as Bryan was saying, you know, you can overlap that with other computations since you have the token ID at the beginning.

11:52

So yeah, I mean, this is like a general trend of just like Chinese labs are just trying to be more efficient because in a way it's a curse that the US labs have, you know, infinite compute and are hoarding HBM and stuff because we we don't have our hands aren't forced in that in that sense.

12:09

So yeah, it's it's been it's been quite interesting.

12:12

All the innovations feels like they're Okay. coming from China. yeah, interesting.

12:18

Now maybe you guys can talk about the application on the workload.

12:22

So obviously the reason to want to store a bunch of KV is driven by many concurrent requests or very large requests.

12:34

So either long context workloads or high throughput serving, both of which are in demand for anybody that's serving their model for coding agents right now, like a genetic coding.

12:47

So obviously, we had this big motivation to try to develop a benchmark that more accurately represents the workload pattern of agentic coding.

12:55

And we've applied this for all of the new models and released both the data set as well as the initial results. It's called the AgentX.

13:03

So, Cam, can you run us through what is AgentX and what's the implication on the change from like traditional AK1K to using AgentX and then how we were actually able to tease out some of these.

13:17

differences in efficiencies from the model architecture and clearly demonstrate them using this benchmark.

13:23

Yeah, yeah, good question.

13:23

So I think before we were doing AK one K, no prefix caching, just testing at a chip level. Now we move to AgentX.

13:30

Basically this is just, you know, agentic traffic, right?

13:34

So as you can see here in this image, right?

13:37

Each the weight, you know it When you code with an agentic harness, what's happening at a low level is you send a message and you know the the output gets appended and maybe some tool gets called and the output of that tool gets appended to your streaming of messages and then this is just keeps getting sent to the model because the model is stateless, right?

13:59

But I as we know, right, the KV cache is just growing like at an insane rate from turn to turn.

14:06

And, you know, if you know when you have million Context length models, right?

14:09

You need to store quite a bit of KV cache.

14:13

so the the idea with AgentX is we're we're no longer just testing like the chip level stuff.

14:19

We're testing the entire system how KV cache moves between chips and pre-filled decode disaggregation scenarios and how offloading works, right?

14:30

so yeah, actually if if you if you want to pull up InferenceX.

14:30

com, maybe, we can Mm-hmm.

14:37

Take a look at some results. Yeah, definitely.

14:42

actually, so in the in the the traces that we replay, right?

14:44

many people are are surprised that the cache hit rate, the theoretical cache hit rate is is very high, like above 98%, right?

14:52

So it's just about how much of that 98% can you actually realize in your system.

15:00

So maybe just click on like I don't know, B300 vLLM, right?

15:06

So the points that are have a halo around them, right?

15:12

this is with offloading enabled.

15:14

So you see that this is not very important for the higher interactivity points because really, you know, offloading to second year secondary tier of memory is slower, right?

15:22

And when you're at a higher concurrency and you want higher higher interactivity, you I'm sorry, I I misspoke.

15:29

When you want a higher interactivity at a lower concurrency, you don't have as much KV to store.

15:37

So offloading isn't particularly relevant.

15:39

But when we're at the you know 40 to 60, TPS range with extremely high throughput.

15:45

So this would be something like offload offline batch inference or even like reinforcement learning rollouts where you have like hundreds of concurrent agentic session.

15:56

Umloading is extremely relevant.

15:58

And you see like maybe like upwards of twenty to thirty percent cache hit rate just from DRAM instead of HBM.

16:05

So yeah, this is the idea with AgentX and we we we really just want to show like what a realistic workload is.

16:14

and yeah, I don't know, Bryan, maybe you want to talk about I think the coolest thing is the the profit calculator. Mm-hmm. As a result Yeah. of that.

16:29

Yeah, the cost of our benchmark is really quite cool in the fact that we can kind of like, because Simon and Racy's also have TCO models, right?

16:41

Please buy our TCO model.

16:43

Our TCO model actually calculates the total cost of ownership of of accelerator chips.

16:52

In other words, how much does it cost an hour to run them and basically like breaks down what this cost entails like your datacenter cost, your electrician cost, the chip cost, et cetera.

17:07

And we kind of give a preview for users of InferenceX, the final number, which comes down to the one shown right now.

17:17

with this cost as well as, so this is the cost basis.

17:20

And on open router, you can kind of obtain how much the inference providers are charging.

17:28

And from inference checks, we kind of know how many users, and how is the concurrency, the throughput.

17:33

And with all these numbers combined, we kind of calculate the profit one can make using different chips.

17:39

So for example, right now, with the best serving, is GB300, NVR72, I believe, you can make about a 40.

17:45

70 % margin, profit margin at those parameters for Kimi K3.

17:58

And yeah, with DeepSeek v4.

17:58

1 flash, you kind of see another trend where with certain parameters or certain settings, you kind of don't make a profit.

18:07

So I mean, it's really quite interesting seeing this because I mean, we do have, we are going to write up an article soon on the methodology of this, I believe.

18:16

Yeah, I mean i I it's quite simple, like it's it's I mean go go back to Kimi, Jordan.

18:23

I mean, the the greatest takeaway from this is like these are actual feasible profit margins that you can get from running open source inference.

18:33

Like these things are just money printers right now.

18:36

Until the cost of compute sort of normalizes, like To break this down, right, on the current GB300, like with Dynamo and vLLM, assuming a 60%, or sorry, what are we at?

18:47

Put it down to 60%, that's probably more.

18:49

So 60% utilization, right?

18:52

Meaning, yeah, obviously you don't have users fully saturating the system at all times, and maybe you have like some nodes down or whatever.

18:59

So that's 60% utilization with the 30% license fee that you have to give to Moonshot.

19:05

you can literally make a almost 50% profit margin serving this on GB300.

19:12

It's just a money printer, right?

19:13

So this is assuming like the cache hit rate that we get on AgentX and then looking at the open router data, but this is like very feasible, right? so if you have like $43. Yeah.

19:24

And and sorry, Sam, Cam, let's let's go back a minute to the AgentX you know data set and talk about that a little bit.

19:36

I I think we haven't necessarily introduced that clearly.

19:39

We we talked about how this approximates the use of Agenic coding and how this distribution of requests is interesting because like the P50 request is still w almost 100,000 input tokens.

19:50

It and you keep talking about cash hit ratio.

19:54

Bunch of people are probably like, how do they know that's real?

19:55

It's real because we're replaying our own data.

20:01

So can you talk about a little bit about how we collect that and why we think it's why we're pretty confident that like the replays of the data against these endpoints for the AgentX benchmark is actually approximating real world coding?

20:13

Yeah, I mean we we we have a proxy that intercepts all of the SemiAnalysis employees agentic coding sessions, and we just replayed them, right?

20:23

So we we just take them on then and do some post-processing to get rid of the bad requests, and then we replay them against actual servers running open source inference.

20:31

And yeah, I mean this isn't gonna be perfectly representative of all agentic coding traffic, but the fact of the matter is like The workload shapes are are quite simple, right?

20:43

You you all harnesses are kind of the same, similar, right?

20:46

You have some tasks that you want done.

20:49

You send that, you prompt the model, the model decides to use some tools, goes and gathers a bunch more context, writes that all to cache, and then does this self-agentic loop for however long it takes for the task to be done.

21:02

And During this time, the bottom line is most of the context is is a cache hit.

21:07

So with that, you need to be quite smart about how you store your KV cache, how you transfer your KV cache, how you route requests to different pre-fill servers.

21:17

And you as you can see at the top here, there's there's a lot of parameters you you can set.

21:21

We're using Nixl for the KV transfer engine, we're using vLLM for you know VLM for KV offloading to CPU, we're using DRAM offloading only, we're using the dynamo router.

21:33

And now this is like not necessarily trivial.

21:37

There are a lot of ways to do this, but yeah, to answer your question.

21:41

The data is representative because it's just like straight up just like coding sessions that we captured. so that's quite cool.

21:48

You can explore all this on inference. com. It's it's pretty neat. Yeah.

21:53

So when we're saying like, hey, back to what Bryan was saying, around the profit estimator per per gigawatt, like if we take our real world traffic, discount it by the utilization rate, and then divide it by real hardware cost numbers, if people are running the latest and greatest chips with the latest and greatest models, it's incredibly profitable right now.

22:12

And then Yes, if you're listening to this podcast and you have like a few million dollars, I would buy like a GB three hundred rack. No, I'm serious, dude.

22:19

Like you can R O it's just like probably make your money back in like a couple months. You sound serious, man. I believe you.

22:27

And it's probably better as well, for the inference providers because they have their own like private corners and we're just working with open source stuff.

22:35

Yeah, I mean, i even if you go back there, I mean you don't have to go back to page, even if you put like a 30 or 40 percent utilization rate, right?

22:41

Which I'm sure some of the providers in open router, no shade, but I'm sure they're getting around there, like you know, pretty shit utilization.

22:48

It still can be profitable, right?

22:50

So until the price of compute kind of re I I don't know, I'm not a finance guy, but I'm assuming this will normalize, right?

22:57

Like it's it's insanely profitable.

22:57

Well, the So the I think people question this a lot at the beginning.

23:02

And I want to bring in Alec here to talk a little bit about this.

23:05

So let me try and tee him up.

23:05

two things have been changing.

23:08

We talked initially at the start about how the models are getting more efficient and higher quality, which generally means people are able to charge a higher price for the new models because they're higher quality, but then make more money because there's it's a lower cost to serve them if they're improving in efficiency, especially with all these optimizations for n-grams and KV cache.

23:30

But then also the hardware is improving.

23:32

And so there's competition out there.

23:34

And specifically another article that you guys wrote was about introducing the TPU.

23:38

Probably the most exciting, looked forward to announcement here.

23:44

Here's a nice picture of the TPU racks.

23:46

A lot of people wonder like you can rent them in GCP why are they not on InferenceX?

23:50

Why is everybody talking about NVIDIA versus AMD?

23:54

But Alec talk talk us through the high level, man.

23:56

Like what's what's been up with TPUs? We got we got some data. Right, right.

23:59

So I think so we had we ended up testing out TPU V seven, which is Google's first real purchaseable accelerator through their own cloud.

24:12

And so we wanted to see like do a quick apples to apples comparison against B two hundred and B three hundred, you know, see how's the performance, is it like feasible and so on.

24:23

And as you can see from this Pareto curve, the blue is TPU V seven and it's like at the bottom it's That's the best part to be.

24:30

It's the cheapest cost per million tokens, you know, for everything except for like a single point. Which is fantastic.

24:37

and you know, that depends on that depends a lot on the on the TCO for the TPU accelerator.

24:47

And I think for this one it's specifically the external TP or external TCO, which is dollar twenty one cents, I believe.

24:56

I mean that's like super good.

25:00

It's like and like the performance they were able to drag out of it is fantastic and Yeah, and and and this is like this is early this is this is early days for TPU inference performance.

25:15

So it'll just get cheaper through software improvements. Right, right.

25:19

And like this is the open source version too. So like Yeah.

25:23

this isn't what the I guess what what they be running internally in big inference providers.

25:27

This is what the public's has access to.

25:34

Yeah, and talk about the trend maybe a little bit.

25:36

Like obviously TPU is a very different chip than B200, B three hundred, both in terms of like the silicon design, the programming model, how many people have access to it, how open source the community is.

25:52

But there's a lot of like you said apples to apples and then apples to bananas in the article in terms of comparing TPU V seven.

26:04

V seven is maybe the current generation, but V eight is coming pretty quick.

26:08

We're talking about some open source models that maybe aren't Google's highest priority to optimize for. TPU VLM is pretty new.

26:17

What's the trend for the future with TPU?

26:20

It's probably gonna get better from here, right? Right, right.

26:23

So I think as we saw recently actually we had we saw inf our infract and TPU had like the the mega kernels and like they got like a huge performance boost out of nowhere.

26:35

So like I'm I'm sure we can keep looking forward to seeing that kind of massive performance boost from them.

26:43

one of the things about the TPU for years is that Google would produce a ton of them, but then they would use them all for internal workloads or for really big customers where they get the whole support of the TPU engineering team.

26:52

You guys talked a little bit about VLM plus TPU and how this is in preview.

26:58

Obviously we're we're specifically reporting results from one model on one workload and comparing it to B two hundred, B three hundred.

27:05

It's a little bit contrived.

27:06

What's the broader trend here?

27:08

I think it seems like we're gonna have more TPU over time in the ecosystem. Is that fair? Yes. Yeah.

27:14

So I'll take this and then hand it off to Bryan.

27:18

You're gonna see a couple things.

27:21

You're gonna see mass externalization of TPU as we, you know, see more maturity of V7 and then V8 and so on and so forth, meaning that Google is going to stop allocating all of this to internal workloads.

27:33

They'll still allocate a ton, but they'll sell a bunch of it too.

27:37

With that, you'll see massive improvements to the open source ecosystem.

27:41

So for instance, right now there's no native torch backend for for TPU, right?

27:49

It's basically goes through this whole like different like intermediate representations and like I don't know, it it's really messed up, right?

27:57

So torch TPU is gonna come out quite quite soon.

28:00

and this is gonna be a native torch backend for for TPU, and this is gonna make things the programming model way easier, right, from a software perspective.

28:10

so you'll see immediate improvements in in open source inference.

28:14

And like this is just gonna be the trend for for quite a while, I think.

28:19

Bryan, you probably have some things to add.

28:23

Yeah, mean, TPU in the past, I may go a bit more deeper into this.

28:29

TPUs are built around systolic arrays, which prioritizes data movement.

28:35

And one problem with it, of course, was that the systolic array was of a fixed size.

28:42

So with smaller ones, have to do padding, which is kind of a waste of the compute.

28:46

And this kind of resulted.

28:46

And yeah, because TPUs were mainly used for internal stuff, A lot of the models were built around this limitation, whereas open source did not.

28:55

And this was one limitation, or one thing holding TPUs back for models on InferenceX.

29:04

So the full power of the TPU wasn't really used.

29:07

And a few years back, maybe just a year back, everyone was bullish on Google and DeepMind because they had the most amount of compute out of any lab. because of the TPUs.

29:21

I forgot what the numbers are.

29:24

Maybe something like 3 million TPUs being produced.

29:27

But yeah, everyone was very fooled because of the amount of TPUs Google had and the amount of research that they did.

29:36

I mean, if you went to any of the conferences, any of the ML conferences, you'll see a ton of Google DeepMind papers.

29:43

And the guys there are really quite correct.

29:46

There's a saying on X that everyone prominent in ML is either in DeepMind or has been in DeepMind at some point of their career.

29:55

Yeah, so the TPUs actually being externalized is quite a big deal, I believe.

30:04

And yeah, and it's quite some time since we've seen a real third competitor to AMD and Nvidia.

30:09

It'll be interesting to see what the next few years are like.

30:16

Yeah, I think the first thing I'm looking forward to is moving for the the support in InferenceX from private a private fork to the actual public repo and getting this stuff running on some persistent clusters there. That's really exciting.

30:29

The other thing is we we know that TPUs are being installed in Neo Clouds as well.

30:34

So for the first time there's gonna be hosting of them outside of GCP, which should open up the ecosystem to have some more customers actually using the TPUs beyond having a consume them through GCP and that's potentially gonna bring the price down yeah.

30:42

People will be People will as well.

30:42

be happier to not have to go through GCP man.

30:54

That that's like a turn off. yeah.

30:54

We don't need to make this a cluster max call, buddy.

30:54

Alright, couldn't help myself run.

31:00

Alec, t talk to us about the GCP console. What's your favorite?

31:02

Alec, give me your top three cloud consoles. What do you think?

31:07

My top three, well I can't say I've used that many of but I mean I j I just kinda leave it to to the command line to be honest.

31:17

I th I find the command Just the agents. line easier.

31:19

I don't have to navigate through all the crazy UI.

31:19

Okay, what do you what do think what do you think codecs or cloud code's favorite what do you think their top three CLIs are for the c for the clouds? GCP bro.

31:32

When you when you update GCP on the command line it it takes like fourteen prompts that you have to say yes to, dude.

31:41

It's like installs a bunch of shit. I think that's free.

31:45

I think software is free, man.

31:47

I think Claudio can Sure, sure.

31:47

do that whole install for you now. Or maybe your open claw. Yeah.

31:54

okay, so moving on from from TPU, some other new hardware hit the hit the wires on InferenceX as well, namely Vera Rubin.

32:02

let's talk a little bit about the difference between Vera Rubin and Grace Blackwell.

32:12

So Jensen's sandbagging again, huh? Certainly, yeah.

32:16

He he's definitely he was sandbagging.

32:19

I mean, look, it it's it's just an incremental improvement, right?

32:22

So I mean you see, in particular at the spot that people would act so this is DeepSeek V four pro, which I guess no one uses anymore, but in the you know, in the range of interactivity where providers would actually be serving, it's it's just like three X better.

32:37

So I mean, when you think about what that means is you can serve 3x number of users at the same interactivity.

32:46

So you can charge 3x the amount of users the same price, right?

32:50

But that's not how you'd price it.

32:51

But like, I mean, it's just more money, right? More profit.

32:53

So if you go to the profit estimator, you'll you'll see this reflected.

32:57

But yeah, I mean Bryan can speak more to the architectural improvements, but yeah, I mean just more flops, more memory bandwidth, like it's it's it's just better, man.

33:10

Okay, one thing I'll lead Bryan in.

33:12

One thing that I've talked about with Vera Rubin on this show before is the fact that it is a little bit more of a minor improvement going from GB three hundred to VR, NVL72 in terms of software support, in terms of like the physical deployment of the system.

33:27

Because let's say going from H one hundred or H two hundred, HGX to GB two hundred, GB300, you introduced the ARM CPU, you introduced scale up NVL switches for the first time.

33:38

He introduced a new 400 gig network or 800 for the for the GB300s.

33:44

you know, obviously black blackwell SM100 was a whole rev.

33:49

it had new FP4 data types that the previous one didn't.

33:53

Vera Rubin seems like more of an incremental change.

33:57

There's no big change to the scale up network, there's no new ARM CPU.

33:58

Everybody should have experience with liquid cooling already, deploying it physically.

34:03

So on one side, one hand, it might be easier to adopt.

34:07

But on the other hand, you might feel like, I'm not going to get as big of an performance improvement compared to GB300, and I'm maybe less motivated to go for the new stuff.

34:16

it seems like that's not necessarily the case initially.

34:22

So high level, like Bryan, maybe you can talk about the initial software experience of porting to Vera Rubin and then yeah, back up what Cam said, which is like there's definitely gonna be some performance benefits or cost benefits on a per token basis when comparing to GB300, right?

34:38

Yeah, I agree with you that the Vera-Rubin step up from the previous generation was not as big as Blackwell was from Hopper.

34:49

I mean, some interesting stuff regarding Vera-Rubin was that, of course, the display that we talked about earlier going to 8 high, the reduction in VRAM.

34:57

And of course, this may be the lab prioritizing memory bandwidth.

35:04

Because of the scale network, you are able to spread your model out.

35:08

and do stuff like YEP, which makes users of more GPUs memory bandwidth.

35:13

So your per GPU memory capacity isn't as important because your aggregate is what matters.

35:18

And another interesting new thing with Vera Rubin, of course, we can't talk too much about the intricacies of Vera Rubin, is that the new LUTB quantization that we mentioned in our previous article introducing. Ryan loves the P.

35:37

Yeah, it's very interesting. The first time we see. It is quite cool. Yeah.

35:42

So for those that did not read our article, Vera Rubin has a new hardware accelerated path called LUTB in their internal docs.

35:51

And this was confirmed already.

35:53

Like, their PTX, I forgot which version, I already talked about it.

35:57

And there are like open source Triton branches of Fox implementing LUTB.

36:05

From a high overview, all it does is a lookup table of quantization.

36:08

So use, if I'm not wrong, 3-bit keys to look up 8-bit lookup table values.

36:17

And this reduces down to about 3.

36:17

125 bits, if I remember correctly, weight.

36:25

And it's quite interesting that NVIDIA gave it a specific hardware accelerated path on Verubin.

36:34

And right now we don't see any applications of it yet.

36:37

But it'll be interesting to see what comes out of it.

36:39

Maybe it's some very specialized use case that the labs have that open source does not have yet.

36:45

But it'll very interesting to see what comes of it.

36:48

And thanks, Jordan, for showing what it is and how it compares to block scale quantizations like MXFP4 and MVFP4.

37:00

But it'll be interesting to see what.

37:03

is this is useful specifically and because you can actually emulate this for those curious you can actually emulate this in Blackwell to look at like the efficiency and how well it actually quantizes higher precision weights and you can run like evals on it and to check the perplexity law stuff like that.

37:27

So yeah a lot of work coming out of this yeah.

37:31

Verubin is quite interesting in some regards, although not really that much of a difference here.

37:39

But the performance does speak for itself, right?

37:40

We do see crazy increments in performance from InferenceX.

37:45

And this is quite early results, somehow.

37:48

And we are getting better results probably very soon.

37:52

I think before this poor car goes out, we'll probably have more updated results. Yeah.

37:55

I mean here's what jumps out to me, right? It's a curve.

38:00

And so obviously it depends which point you pick in the curve to compare to the baseline in order to see what the multiple is.

38:06

If you take some really, you know, contrived point here, it's like infinitely better than just not being able to serve people at two hundred tokens per second, because now VR and VL seventy two can do it.

38:15

But at a midpoint, hundred tokens per second to get three times better, and then on the high throughput cases to not be quite at the three times better, be more at One point four, so a forty percent improvement, you know, makes you question, okay, is this gonna be good for high throughput?

38:29

Interestingly, this reminds me actually.

38:29

Sorry, Jordan, but this reminds me of the CPX.

38:36

I mean, the noise has died down quite a lot on CPX, but Verubin introduces CPX for DSEC workloads and splitting up the pre-fuel and this decode.

38:44

And right now we don't have CPX on our InfantX belt.

38:50

Yeah, we look forward to getting CPX results soon maybe and check how CPX improves these aggregated workloads.

38:59

Yeah, it's it's just it motivates people to think about the flops versus bandwidth trade-off at the different ends of the curve and how you can optimize.

39:08

And so the obvious chart to me is just like the big step up for Rubin as a GPU is the memory bandwidth, which is like as you go from HBM three to HBM four, you get almost a three times improvement in memory bandwidth.

39:21

And that roughly seems to correspond with this, you know, three times improvement on the Big middle parts of the curve where inference is memory bandwidth constraint.

39:31

to your point, Bryan, the fact that this is an early result and these lookup table quantization schemes are not explored, and you know, some of the other hardware acceleration features in the GPU is not there yet.

39:41

It just seems like here's us demonstrating that roughly on a per dollar or per watt or per like adjusted dollar watt basis, Vera Rubin is still going to be more efficient at inference than Blackwell. It's just a better chip.

39:57

They're not pricing it super aggressively.

40:02

You know, aggressively, meaning they're not price gouging with this chip because they know people are gonna buy it.

40:10

Like if you're gonna get three times the performance in the chip doesn't cost two times more.

40:15

Like that's the that's a trade everybody's looking to make now going into next year.

40:19

And to go back to the profit estimator calculations from earlier.

40:25

It just seems kind of obvious that people are going to make this calculation going forward and say, Yeah, I want the latest and greatest. NVIDIA's done it again.

40:32

Jensen sandbagging performance again.

40:35

There's parts of the curve that are really memory bandwidth constrained, and you want to be able to serve that.

40:39

maybe last question for you guys.

40:43

You do a lot of work with these coding agents.

40:45

Can you talk about experience using fast tokens?

40:48

Like a lot of the implications of these performance curves seem to be related to like.

40:54

People are going to pay a premium for a hundred tus plus tokens per second.

40:58

we have some hot takes on the phone.

41:01

I I I won't share mine right now.

41:03

Can you give me your initial thoughts on like slash fast and slash ultra fast and if we're gonna see a lot more of this going forward for you guys using it?

41:12

OCam has spicy takes on this I believe.

41:16

Bryan, are a the start let you start?

41:17

Are you a super are you a a fast maxi or are you a a fast bear? What do you think? Nah, nah, don't do fast.

41:25

For my personal use cases, I have multiple agents running at the same time. I have multiple VSCodes.

41:31

In my personal use case, I use multiple VSCode windows and multiple agents on each VSCode terminal doing basically different tasks.

41:46

So it doesn't matter to me how fast one agent completes.

41:51

Because when I come back to that agent, it's probably already done, given how much time I spend on the rest.

41:55

But I know Cam has some differing views on ultrafast.

41:59

Yeah, I mean I really I'm quite quite bearish on it because I guess my thinking is that as tasks become more agentic, you you need to I mean really the the the power in these models is that they can call tools very very well, right?

42:19

So they can call tools to bring in more context and accomplish things.

42:22

But the tools take time to run.

42:27

And if you have a model running on like, you know, Cerebras or something that can serve like a thousand tokens per second, right?

42:35

But then the tool use time takes like 30 seconds, then that speed up in the generation of tokens just is not is not really realized because you're still bottlenecked by the tool tool use time.

42:48

So I I guess like if you're just like generating like HTML pages all day, like a hundred thousand lines of HTML, then yeah, I I I suppose it would be quite useful.

42:59

Well, that's a great cue for me to talk about my use of these models. yes. No, no, okay.

43:03

Alec, what's your take, man?

43:06

What No, I I I do have to agree with them, you know.

43:09

There's times where I where I need to do something really fast and or like get an answer real quick or something, gather context and you know, fast fast node's gonna be nice for that.

43:17

But for the vast majority of the time, I also have just multiple agents and I mean I've I can just do something else as they're running.

43:25

I d like I don't need the the code immediately.

43:30

I don't need it to be merged instantaneously in it.

43:33

You know, it can happen in an hour or two even.

43:36

So I mean, I just don't use it very much anymore.

43:41

I just I mean it's I mean so expensive, right?

43:41

Like I'm not sure if the ROI is there. So like I mean, yeah.

43:44

Especially when you're bottlenecked by by the tool use time.

43:49

It's just like really, really especially like the ultra fast stuff.

43:52

Like, you know, I'm not paying three hundred dollars per million output tokens or can drive right. It's insane.

43:56

Okay, so let me let me push this counterfactual.

44:01

I I I agree with you for my current usage, but I also didn't forecast myself using AI exactly the way I use it today a year ago or two years ago.

44:08

So I find it hard to predict what it's gonna look like a year from now and whether fast tokens aren't gonna be useful with any sort of certainty.

44:14

So let me push you guys to two extremes.

44:17

One is like, do you have the experience where a model is really big?

44:22

In other words, let's not say really big in this case, but like Or the network's really slow or the tool use is really slow or the model just slows down because the endpoint's busy.

44:31

And then you get really frustrated because something is just not returning something at all.

44:35

And so can you imagine a world where the models get 10 times bigger, 10 times less efficient, and we just need to speed them up for them to be even useful at all?

44:46

I mean, that's a very interesting Yeah. thing.

44:48

The scaling loss do show that the more you scale, the better results you'll get.

44:54

So kind of an interesting approach to just scale.

44:58

I mean, that's the bitter lesson, right?

44:59

Just to scale and let computing take over.

45:02

But I mean, I've been thinking about that quite recently, actually.

45:06

Like, why don't we just scale up and just go at it fast to reach a usable token output rate.

45:16

seeing stuff at what the industry is actually doing.

45:19

For example, Astra and rumors that 6.

45:19

1 sole are looped transformers, not confirmed.

45:27

It's really quite interesting to show where the industry is going.

45:33

It doesn't seem to be just scaling up is the way there's architectural changes, very interesting innovations. And yeah.

45:39

Sure, but even loop transformers and lots of reasoning is not free in terms of how much computation you need to do, right?

45:48

So Yeah, yeah, that's true.

45:48

Actually actually good point on the reasoning though.

45:52

Like if if you're on like ultra max, whatever the reasoning, that is obviously just purely to output speed, like you need to decode very quickly.

46:02

So that that's a good point.

46:02

I honestly, I think my prediction would be something along the lines of like you have a harness that can like route to a fast mode for like certain tasks and then like it can batch certain tasks that don't really require low latency, right?

46:20

Because when you think about it, this is basically what we're describing, where like if you're writing a you're bootstrapping a ra React app, then it would be really, really, really useful to have fast tokens.

46:31

But i if you're doing something like Bryan's saying, you know, running a performance benchmark and trying to hill climb overnight, i it it you know, the benchmarks take like fifty minutes.

46:42

So if it takes fifty minutes and then you spend forty dollars decoding a few thousand tokens and then just wait another fifty minutes, that that's just a waste of money.

46:50

So that that's my prediction, right?

46:52

My prediction is like it'll be very specialized, but It'll still be useful for some things. Yeah.

46:58

How about the the other extreme?

47:01

Let's say the models get really, really small, really, really efficient and still high enough quality.

47:08

Is there a scenario where you are just bootstrapping a React app or you're using your open claw slash muse slash instinct perplexity computer or whatever?

47:16

And you want have you had the experience of getting instantaneous responses and feeling like, wow, this is really cool?

47:24

I mean, that's one interesting direction. I don't know.

47:27

Personally, I don't really think there's a chance of us reaching this smaller model state.

47:34

Because it's been shown in the history of computing that if you make something more efficient, someone else is going to just lump more stuff into there.

47:45

CPUs, RAMs have gotten so much better.

47:47

But then websites have gotten so much more sluggish and so much more inefficient with the libraries.

47:53

Yeah, don't think, honestly I don't think a smaller model, like a world where smaller models are used will really come true but it'd be interesting if I'm proven wrong.

48:05

Better for the environment too as well. Interesting. Okay.

48:10

Appreciate you guys going on that little journey with me.

48:14

I have a suspicion that even when you guys use all of these different agents in parallel and you don't really care about the time it takes to return, that if it was ten times slower, you would still find a way to get upset.

48:25

If you went to bed, woke up in the morning and there still wasn't a response, that wouldn't be good.

48:28

So there's a a room for some sort of fast tokens in the future, maybe not ultra fast, Yes, there is, there is. right?

48:35

Also, Jordan, people in America can barely afford Big Macs, right?

48:38

So I don't know if they're gonna pay like, you know, a hundred dollars per million output tokens. So I don't know, man. But yes, I I agree, yes.

48:44

Like speed matters, right?

48:47

To some extent, the intel I always love when Jensen says this.

48:51

Jensen Jensen says this in every keynote, right?

48:53

Like faster tokens are smarter tokens, right?

48:56

Obviously he's inclined to say that, right? Because you know. He sells GPUs.

49:00

But yeah, I mean like to some extent that's true. Right.

49:03

I I mean, think about if you could, you know, basically have instantaneous output, right?

49:10

That would be extremely useful for for some cases, because you're just getting the answer right away. But I don't know, man. We'll see. Time will tell. Who knows? We will see guys. Okay. That's all.

49:22

dude, we should talk about TileRT. Dang, we're out of time. yeah, yeah, yeah.

49:27

That was a perfect lead in, man.

49:29

I was about to wrap this up. We we we can wrap it up.

49:31

It's i I don't think there's I mean No, talk about Tyler R RT because a lot of people when they saw the 6. 1 Astra or sorry, 6.

49:44

1 soul ultra fast mode get released at OpenAI Deviday, everybody was like, well, there's a Cerebras deal with OpenAI that's promising them 750 megawatts over the next few years.

49:54

Obviously, this is Cerebras, but Tyler RT is super capable.

49:59

And if our take on how big these open AI models are.

50:04

is correct and they're using some similar optimization to what we've got on screen here showing Tyler RT.

50:09

Well 200 tokens per second all the way approaching five hundred tokens per second on a decent sized model approaching trillion parameters like we got on screen here is totally cap possible just on NVIDIA GPUs, right? Yeah.

50:21

Yeah, I mean TileRT, for those of you who don't know, is basically just a latency optimiz inference engine optimized for ultra low latency, right?

50:32

So I think the takeaway here is that you don't necessarily need super like specialized SRAM architectures, whatever, to get like high, high interactivity.

50:46

So these are achievable to some extent, right?

50:49

Bound by the laws of physics, of course, by super optimization.

50:53

Optimized inference engines, right?

50:54

So this has been shown here.

50:54

There's there's also an AMD result.

50:59

And I think going forward, this will become a trend where inference engines will become more specialized for different use cases.

51:07

We're already kind of seeing this trend.

51:08

but yeah, TileRT is super cool because yeah, you don't have to have a new chip, just make better software.

51:16

Can you talk a little bit about how it works?

51:20

Obviously there's SRAM on the NVIDIA GPUs.

51:23

There's lots of different ways that you can configure the serving setup as well as write specialized kernels to actually serve the models.

51:32

Anything, you know, technical we can talk about on the overview? Yeah.

51:32

I think I I think I think Bryan can talk about like why it's good at high interactivity, but it's bad at high throughput, right?

51:45

Yeah, so basically like TileRT just you can see TileRT as one humongous mega kernel for the entire model.

51:53

like the problem at low interactivity is data flow.

51:58

So what companies like Cerebra, SambaNova have done is they build data flow machines where the hardware is built for data flow.

52:06

And for GPUs, this is a bit more difficult and in Taira RT you cannot see it.

52:14

only batch size 1 is possible right now and the model releases are quite slow.

52:20

This is because it takes a lot of software effort to actually get to this mega kernel because the GPUs are not that friendly for this kind of data flow stuff.

52:29

But it's really quite interesting that GPUs really are general, the fact that they can be used for this and kind of like replace the specialized hardware. And yeah.

52:38

But to what extent and what to what extent is this enabled by just having the AI like an AI agent hill climb one giant mega kernel that literally no one can actually read, right?

52:49

And then it's just like works and it it it's like good accuracy and high interactivity, right?

52:58

Yeah, but I mean, people have been doing this, that's for sure.

53:00

And of course, I personally do believe that kernel development is a very easy task for agents to do because it's very easily verifiable.

53:11

And we have seen a lot of agents on the leaderboards of GPU mode and other kernel competitions. And I don't know.

53:21

Right now, this is a very new space.

53:24

Recently, I've seen this kind of.

53:26

mega kernel ideas come out on MLX.

53:30

There's a report recently released on MLX.

53:32

As well as other developments, for example, the TPU mega kernel, which uses the same idea of this mega kernel for the entire model, which is very interesting. Yeah. We'll see where it goes.

53:46

Jordan, I have a question for you. Sorry?

53:50

Jordan, I a question for you.

53:50

If you're not to name any names, but if you're you know a chip startup and your basically entire thesis is like fast tokens, are you a little bit like concerned right now with the state of I mean all these improvements with megakernels like Bryan's talking about in jalapeno, for instance, like and just the general purpose GPUs just becoming faster at this region of of interactivity. What do you think?

54:20

I'm not concerned because I have a fundamental belief that demand completely outstrips supply and anybody who can produce chips is going to find a customer right now.

54:28

people just need a lot more chips of all sorts.

54:33

I think in the long term, all of these chip startups really make or break their business by their ability to deploy things that actually work and to be able to produce a bunch of chips.

54:44

Like you just If you're gonna have a really big business, you need to be able to manage your supply chain and actually produce a bunch of chips.

54:49

Like you need to be able to get good yield and you need to be able to build servers and just like build a datacenter and train people Yeah.

54:55

and all of that, which is you know, in some ways a much harder task than actually making something that performs really well.

55:07

I think the other thing that I've I've been clued into is the fact that High interactivity is high throughput if the interactivity is high enough.

55:16

In other words, you don't need a bunch of batching dynamics just to get a whole bunch of, you know, throughput, even if it's a one-to-one ratio between throughput and interactivity, because if your interactivity goes to a thousand, your throughput's at a thousand.

55:29

If it goes to two thousand, your throughput's at two thousand.

55:32

So I'd be driving interactivity as high as I possibly could with a lot of this custom silicon.

55:35

But it it's gonna be like expensive.

55:39

Like if if if you just have like batch size one and you have a thousand throughput at a thousand TPS then one user is paying for entire utilization.

55:48

You can do batch one, but if it's really, really fast batch one, you can just pipeline everything and still get high effective throughput, even though there's not a bunch of batching dynamics here.

55:57

So I I don't necessarily feel like fast tokens are by definition inefficient tokens.

56:07

I think that really depends on the hardware design and if you're as Myron says, using because eggs got expensive.

56:15

meaning HBM, you're going for caviar, which is the SRAM. so I I I don't Yeah.

56:19

think it's a per calorie thing.

56:24

It'll be it'll be interesting.

56:25

I mean, another axis that I'm skeptical of is like quite a few of the alternative chip startups, not to name names, do kind of rely on general purpose GPUs for prefill.

56:36

So like the heterogeneous stuff is cool and all, but if your argument is like, okay, well people are chip constrained and they need more chips, but like if the chip can't actually run efficient inference, like it's just kinda like you still need NVIDIA GPUs. I I don't I don't know.

56:49

A lot Yeah, the setup is the setup that makes sense to me is some pool of pre-fill decode whereby you oversize one of the pools to be able to serve a non pre fill decode disaggregated setup for some of the traffic.

57:04

In other words, if you're fixing your pre fill decode ratio ahead of time for the entire datacenter, I think that's totally bunk because it's just gonna your the workload pattern's just gonna change and you're gonna have some wasted, inefficient amount of silicon, right?

57:18

But if the workload pattern changes, you know, for more decode or less decode, and you can move those chips into a, you know, only like a homogenous broken up setup, I think that makes a lot more sense.

57:32

In all of these scenarios, if you depend on NVIDIA to be one of your GPU partners, and you're buying GPUs and stuff, yeah, the business decisions there make a bunch more impact on your success rather than the technical decisions because Look, NVIDIA acquiring Grok is just obvious like justification for all of these chip startups that they're onto something, technically.

57:57

Otherwise, you know, why waste twenty billion dollars acquiring somebody whose chip doesn't work?

58:03

It obviously works to do something NVIDIA GPUs can't at any price.

58:06

And so everybody else can do the same thing.

58:08

But the question is is not technically if this is going to be beneficial to anybody or if NVIDIA has the demand signals to wanna go and build this. It's business.

58:17

It's like are you going to be able to build enough of them and deploy enough of them and actually make them work?

58:23

And I think that's a quite a big question when like NVIDIA is controlling who gets shipments of GPUs and backstopping people who want to buy data centers and build them and stuff like that, right? that's a hard business.

58:38

Hard hard to hard to get into.

58:44

Yeah, I I and I think the other thing is like all of these chip startups have proven at this point, generally speaking, that they've made a chip that they can tape out.

58:50

I mean, and and get like it it works.

58:54

CPU training, everybody who's building custom silicon, you know, MTIA, even Maya, like the hyperscalers have proven you can build general purpose stuff too.

59:05

It doesn't have to be decode specific silicon.

59:07

So why not just build a second chip that's pre fill specific if you want to, you know, go and win.

59:13

build a pre fill and a decode chip, right?

59:15

Yeah, I think that's the all opinion.

59:15

Yeah, I mean I think that's kind of what Hall Pinu did, right?

59:15

Or that's what they are claiming at least. So See man.

59:22

I mean, yeah, like the the design, like the RTL is not free, but people have proven it's a lot less expensive than it used to be.

59:32

And the software mote, the quote unquote scooter software moat, is speed to develop these things as opposed to just like people ca there's something physically setting NVIDIA apart where others cannot reach a certain level of performance.

59:48

It's like you can if you spend the time and the effort and the money on doing so.

59:52

And so I expect chip startups to do that.

59:54

I expect them to build pre-fill silicon and decode silicon.

59:56

I expect them to realize the benefits of removing the margin of multiple vendors if it's NVIDIA or Broadcom or Marvel or whoever they're partnering with.

1:00:06

And then I expect them to have a really hard time deploying all of this stuff and producing enough of them for it to actually matter because supply chain and datacenter operations matter a lot, like I already said. So but I'm excited, man.

1:00:21

I'm like I know it kind of got buried in the middle of this conversation, but TPU on InferenceX is to me the most exciting thing here because it's just like a one way street, you know.

1:00:32

They don't start selling TPUs externally and then just say, we're gonna stop and keep it all for ourselves.

1:00:39

Like Yeah, no, no, it it's super exciting.

1:00:42

I think I mean Alec you've you've worked with Google.

1:00:47

I I think you probably can attest to to where you think this is gonna go.

1:00:52

And yeah, it it's awesome.

1:00:52

I mean we'll we're gonna have training them soon as well.

1:00:57

We're we're gonna have alternative alternative chip startups, we're gonna have, you know, everything.

1:01:03

That that's the goal, right?

1:01:03

So I I think going forward, like inference is just gonna be a consolidated place on the internet where you can literally get free information on the performance of chips on real world inference workloads and compare all of them.

1:01:21

So Infrarsex chat room soon, Kim? What do you think? On the website? InferenceX chat room.

1:01:24

I don't know, we'll see man. We'll see.

1:01:28

Gotta take that out the boss.

1:01:33

Dylan, what do you think about being a a Reddit moderator again, except it's on InferenceX dot com?

1:01:38

Alec, you want to be first moderator?

1:01:43

What'd your username on InferenceX socials be?

1:01:48

Man, I don't know if I ever want to be a moderator.

1:01:50

That sounds like such a pain. T T P U thing. Ha ha.

1:01:56

Aren't you kind of a moderator of the GitHub Actions queue for all of the runs that are going on right now, kinda? Yes, yeah.

1:02:05

Really kinda gatekeeping.

1:02:05

bro bro bro bro i i it's your codex, right?

1:02:08

You give me an API key and you say hey moderate all these drones.

1:02:11

Well actually that's a good idea.

1:02:14

We'll have an AI agent be the moderator.

1:02:16

hopefully there's no way.

1:02:16

Hopefully it's more of a Reddit and less of like a that'll be good. Yeah, yeah.

1:02:16

I w I w Yeah, see, they can be impartial.

1:02:22

I wonder if the AI agent will ban Dylan's posting on his own forum or if it'll let it slide.

1:02:31

Okay guys, nice nice way to end off.

1:02:34

Anything you think we missed here?

1:02:35

All right, little closing sentiment.

1:02:37

as we talked about this, it seems like there was six different articles that came out in between podcasts about InferenceX.

1:02:47

We talked about n-grams, AgentX, Rubin Performance, TPU performance, the DeepSeek V4. 1.

1:02:55

I guess we already talked about that, so that doesn't count.

1:02:56

Five things and TileRT is the fifth. good stuff.

1:03:00

Hopefully we can get you guys back on before.

1:03:05

the next five articles drop and do this a little bit sooner.

1:03:08

Appreciate you taking the time today and thanks everybody for listening. Thanks, Jordan. Yeah. Thanks for having us.