0:06
All right, we're here at GTC electrifying energy.
All right, we're here at GTC electrifying energy.
All day there's incredible announcements, people to meet, and all night what did you do all night? Waleed. Um, I stayed home.
I was whipping with three different Claude code sections. No, no, no, I'm kidding.
Of course, obviously a bunch of fun stuff happening.
Well, so Waleed, you're you're here.
Tell me about your company, what are you working on? Yeah. Um, so we're Macaw. Formerly Mako. Uh, now we're in Macaw.
But no, we're >> So so you you had to change the name, but you also changed what you're working on. A little bit about that.
>> is moving up the stack.
So what we had been doing, you know, previously was um, a lot of work on improving quality of kernel generation and doing a lot of work on validating kernels that you do generate.
Uh, these days you can actually get pretty nice kernels out of Claude code and Codex and everything.
And I think the entire Silicon Valley, like 80% of the startups here are starting to rethink everything now that coding has come this far.
I remember when Sonnet came out, like Sonnet 3, when it for when we first saw models that were actually really good.
And I remember people were talking about they're saying, "Look, all right, it does front end, it does this, but does it do this the complicated stuff? Does it do kernels? Does it do compilers?"
And it turns out the answer is yes.
I think the funnest thing about the Opus 46 announcement was like make a C compiler.
And then it spent like 20k and it made a C compiler.
[laughter] I was like, "Wow."
No, and you can like nitpick and be like, "Oh, well, there's this bug, there's that bug, there's this bug." But I mean, come on. Yeah.
>> Like you kind of just like told it what to do.
Um, so what we've been doing at Macaw is you know, it kind of forces you to get into like the product aspect.
Like what do the people whose job it is to, you know, make things run fast?
What does that person do throughout their day?
And how does it change now with Claude code?
And how can you build a platform that makes their life easier?
So so before stepping into what you're doing now, I actually want to ask a bit more about the code gen side, right?
Everyone's so excited about code gen, it's been something that people have been investing in for, you know, some time.
Um, there's a numerous startups now in the space.
Why are you moving out of it?
I I actually don't think we're we're not moving out of it is the thing.
So we still a core part of this is that we use LLMs to write high performance low level code.
Um, and what we've done is the platform that we're building, we want Claude to get better.
We want Codex to get better.
We should use that as a tailwind.
And then how do you build the rest of a platform around this kind of a capability that you can just keep assuming is going to get better.
The point is that you cannot um, have code generation itself be your product.
I think we're going to see a lot of basically giving the taking the code gen capability and then making these kind of end-to-end platforms that just do a lot of the work for you.
Um, and what what what goes into this end-to-end platform then?
Is it is it you're making the most of validation engines or what exactly do you do?
>> That's definitely part of it and that's a really important part cuz again, you have to have that feedback.
But the kernel code gen we started with input PyTorch get kernel.
I think the next step is going to be input a Hugging Face link, get full VLLM deployment with whatever kernels you need in there.
And then like let's think about InferenceX for a second.
Well, I think the obvious next step for what's probably happening with you guys, you see this Pareto curve where there's a bunch of different curves.
Now, the question for everybody else is like, "I want the good curve. How do I get that?
How do I get that for my model, my use case, which is maybe not the exact one that's on here?"
And so I think a big one is going to be "Here are the GPUs that I have or here is the cloud access I have.
Here is the model I want.
Here are the SLAs that my customers expect.
What's the best way to deploy this thing?"
And that's going to need um, you know, a lot of testing, a lot of infrastructure, a lot of like hardware to figure out what that Pareto curve is for that specific use case.
>> That's that's actually something that we're doing in InferenceX, right?
>> Right, cuz the natural next step, I would expect that, yeah.
At least uh, initially we were three models cuz we were trying to implement things like evals and implement things like YDP, disagregated prefill decode.
We're implementing all these features.
Uh, but over the last month roughly, we've added we've triple I think we've tripled number of models or doubled number of models.
Um, and those have been like day zero, we got the Pareto curve basically just by telling Claude to um, generate all the different configs, right?
Um, pretty simple and then cut all the configs that make zero sense to run. Right?
Like you're not going to run PP8.
Uh, you're going to pipeline it eight times or you're not going to do certain configs that just don't make any sense.
So then then getting it to first from first principles just eliminate uh, a lot of the configs.
And then there's there's configs that you actually don't know the answer because like, well, this one theoretically is better from a roofline basis, but maybe the kernels aren't optimized there yet or whatever.
And then and then and then we run those and then we have the Pareto curve.
Um, without using too much compute and able to implement a new model at least in the current leading VLLM actually laying.
It's still very manual engineer in the loop process, but I can I can see what you mean.
Like that that's going to be something that's much sooner to have. Yeah.
And I mean, I feel like every 2 weeks like the limit of what is being automated changes, right?
So I I totally think that, you know, everything is going to end up being automated at some point, right?
So what's next for InferenceX on carriers?
Are you guys going to actually do the dynamic submit a model and all runs?
Um, I mean, generally we accept PRs.
If someone wants to submit a PR for a model, great.
Um, the goal is to do day zero support of models now.
>> Now, what's the GPU spend on that?
Uh, >> Or just capacity on there?
Thankfully, we have great partners that give us free GPUs including CoreWeave, Microsoft, um, Oracle, Nebius, Crusoe, um, TensorWave, AMD, Nvidia, and I'm going to kill my There's one I'm missing. I know I'm missing one.
Uh, but we have great partners that give us compute for free because this is a free open source project.
We actually make money off of InferenceX.
Um, obviously there's a lot of learnings that we get out of We then turn into other products, but um, InferenceX, I mean, we're going to have day zero support for all the new models.
Um, we're adding more and more evals cuz it turns out people's kernels cheat. Yeah. >> Right? Um, >> Yeah, evals, man.
I'm telling you, validation engineering.
These things are faster, but it's like we only implemented GSM8K so far, but even in GSM8K you see a 10% swing even on different points on a curve for the same hardware. >> Really?
>> Because something they did on the curve, something they did on the kernels, something they did somewhere >> kernels.
I thought you meant like the full exact same, but just Yeah, it's it's it's a long the sweep, right?
So when you sweep everything like the kernels for this optimization point that's great at throughput or this one for latency actually has different performance. Yeah.
>> Um, which which is sort of like makes me wonder like you know, people are always like, "Oh, OpenAI downgraded GPT.
Oh, Anthropic downgraded GPT."
Often times it's in the in the they're like, "It's because they've gotten so much more traffic."
And it's like, "Well, actually perhaps there's some kernel screwery here." >> Yeah.
Do you guys use open source models? No.
But you use open source models.
Everything on InferenceX is open source.
We love open source models.
But like stop being poor, just pay Claude. Yeah.
Pay Anthropic, pay daddy.
No, we we've done a bunch of RL, um, on some of the open source We've done it on GPT-5, too.
Um, but yeah, the closed source models are a step ahead. Yeah. Exactly.
It's It's like 3 months difference in model timing or 6 months difference in model comp capabilities ends up being a >> Do you think that is true?
>> year or more in terms of like when you're able to release a project and then get the traction and get the customer feedback and, you know, this this iterative cycle.
So it's like you know, and this is why we spend everyone I encourage everyone to use fast mode even though it breaks the bank. It's like use fast mode.
Yeah, more tokens per second is more more intelligence.
>> Yeah, more intelligence per second.
More intelligent per second is more revenue eventually even though it's mostly just a cost.
>> Yeah, sounds like Jensen, man. Buy more to save more.
>> Well, actually he sounds like me. Oh.
Yeah, but he sounds like Jensen, of course.
Did you see the I'm just saying. I'm just saying.
I thought the Apple was good.
Dude, he spent 5 minutes on our slide, said my name.
He only said two people's names in the entire keynote, mine and the open cloud bro.
And now the open cloud bro deserves it.
I don't know why he said my name.
>> [laughter] >> Yeah, Jensen called me.
I told him not this time, maybe next. Yeah, yeah.
It it it it was it was I knew it was going to happen.
I was I was I was unbothered.
Didn't matter to me, you know. >> Yeah, it was cool. No, seriously though. Um, big things, man. Yeah, big things.
I'm excited to hear what you guys are going to do forward in the future.
>> Yeah, we'll announce the fund raising tokens and then you'll uh, you'll you'll you'll you'll see. Awesome.
We we encourage every new startup to do this, announce their fund raising Claude 4. 6 fast mode tokens.
Thank you so much for joining us. All right, man. Yeah.