0:00
We're going to have a little video of you walking in yelling. >> I'm so excited. >> Oh, really? [laughter] >> No.
We're going to have a little video of you walking in yelling. >> I'm so excited. >> Oh, really? [laughter] >> No.
>> We We had so much fun on these podcasts, you know. >> Yeah. I don't know.
>> Did you see the one last week?
>> I um I had gotten feedback though.
So, do you want to start this uh podcast?
>> Let's do Let's do feedback.
We cut out something that was going to be in the cold do the last time where I said I'm no longer listening to the comments because on one episode with Doug and um >> got so pissed. >> Oh, did they? They just hate us.
They just hate us all now.
>> That one clip, >> dude. And he cut back.
He cut out me pushing back on them. >> Yeah. Yeah. Yeah.
Cuz I thought that [laughter] >> I fully was like, "Okay, they built maps in house.
They build Gmail, Google Drive, like the whole G Suite. They built all of GCP.
>> I don't think you understand Kubernetes.
All my deep mind friends, there's like three of them who are like, "Yeah, I think I'm going to leave." >> Yeah.
>> And and I've got a bunch of others who are like, "Fuck you guys."
Not [ __ ] you guys, but >> sounds it sounds like they're in stage two. >> Stage two.
Yeah, which [laughter] I only know cope.
>> So, back on the back in the form warrior days, um, we would get into chip arguments and there was like a there was like a private discord where there's a bunch of like people who loved anime and like people who were all around the world and many of whom were racist cuz it's anonymous people on the internet.
Um, but they all [ __ ] loved anime and I did not like anime.
Um, I've never really watched it.
Um, besides like, you know, one one one girl I dated, I watched anime with her, but besides that, like they would always No, no, I've never dated anyone. I'm I'm I'm uh pure.
Um, so what >> which anime which anime did you watch? >> Oh.
Oh, I watched um Okay, there's one I love. I love Spyex Family. [laughter] >> Any? >> Sure.
>> Anya Chan, I don't know how to say >> I don't know.
Do you know the reference?
[laughter] >> Yeah, I do.
>> Who does one of the dolls?
>> You bought me an Anya.
No, like what are the >> what is it?
Labu >> something from that series.
>> What's what's the what's it floor? What's his name? >> I I don't know. [laughter] >> Lloyd. >> Lloyd. There we go.
>> I know less anime than you, man. >> Well, yeah.
So So we're changing topics rapidly.
Jordan Anos is And this is like HR approved because I'm HR.
Jordanos is the hottest man in semiis.
>> Got this [ __ ] [laughter] Michelle got this [ __ ] My gosh. He's He's No, no. Think about it. Think about it.
Look at him like [ __ ] I'm fat and look at him like you know look at him so beautiful.
Um tall as [ __ ] Same age as me except he's married and has a kid. >> Owns a home. He's not a degenerate.
Like you know like this is just like wow goals.
[sighs and gasps] >> Thanks for the employment man.
[laughter] Okay, so sorry. Going back.
Google Google people were mad at us and they were DMing me and some of them were like, you know, they're in cope.
Um, but regardless, the feedback I've I've gotten >> from um mostly just my own head, my own head.
>> Are we giving away too much value?
>> That's why I came on today cuz I need to destroy value.
[laughter] >> All right, we can stop. No, no, no, no. Don't stop. It's fun. It's fun.
Um, but someone someone on the team internally was like, "Dylan, we give a lot of value away on the weekly."
And I'm like, "Huh, we do.
I haven't listened to it, but we do, I bet."
[laughter] I listened to the one where we had the DG Matrix guy and I was like, "This is fire as fuck." >> Yeah.
Who um who said that, >> "Bro, come on.
HR is HR's for text anonymity." It wasn't Doug. >> Okay.
Jeremy, >> it wasn't it wasn't Doug D.
>> So, I don't know who has who has the say.
Is it just somebody who's like upset that they haven't been on yet or what? No, no. Someone who's been on. Someone who's been on.
>> Dan, >> I don't want to say Dan. Dan wouldn't say that. Dan's a sweetheart.
Um, anyways, regardless, >> who does it say?
>> The feedback is that this podcast is too good. >> Yeah. Okay.
[laughter] >> Why are we giving away for free?
[laughter] >> Oh [ __ ] Anyways, >> uh yeah.
Yeah, we can we can definitely put in the toilet this time. [laughter] >> All alpha.
People click because I'm on and they're like, "What the [ __ ] is this trash?"
>> No, we're going to have a nice picture with you with a neon orange shirt ready to attract all the clicks. >> Yeah.
Come, come show your shirt.
>> So, so >> was it Nick who said we're giving away too much halifa? >> No, no, no. Look at Nick. >> Oh, it was David. >> No, it wasn't David. >> Wasn't Sales. [laughter] >> All right. Good to see you, man. Thanks for coming by.
>> Let's Let's talk about Let's talk about the the like semi analysis office in New York.
Man, it's it's it's Have you been yet? >> No. That's why I'm going.
[laughter] >> Why would you go?
You [laughter] you fix going back to the to the hing.
[laughter] >> We're upgrading. We're upgrading. >> Oh, when? >> Soon. Very soon.
>> This is like buying GPUs, right?
If you make too long of a commitment to the lease, then you have to find way to resell it.
[laughter] >> You got to make six month office commitments that you can outrow them.
>> I should I should just buy GPUs >> instead of more office space instead of hotel. Yes. Yes.
>> So, you wish that semi analysis was just >> a GP res employees? >> No.
>> You can replace us all. >> No. No. No. No. No.
Dario's like, "We're going to automate away all my cars because God, >> brother, go look at the [ __ ] a AI spend. I'm not doing it. I'm not doing it.
>> What's growing faster? What's growing faster?
AI spend or spend on employees?"
>> Well, so we we the thing was like we've gone through like hiring sprees and then digestion periods and hiring sprees.
Uh we're back in a hiring spree. So >> okay.
So so um [snorts] >> so spend on employees really skyrocketed especially in the second half of last year and parts of this year but then like the first quarter of this year AI spend skyrocketed but it's actually been like relatively flat in Q2 too. >> Right.
We kind of everyone got cloud code psychosis and then it's like leveled out. >> Yeah.
>> Like it's still at that like 10 million number um roughly.
Do you think it will grow roughly in line with more employees in the future?
>> I I you know I was surprised Fable didn't cause price to go up. Yeah. >> Spend to go up. >> Yeah.
>> Why do you think that is?
>> Roughly the same as Opus.
I'd say possibly is counteracting with a lot of people were building the first versions of the applications.
Like we went from 10 repos internally to like we have over 150 repos internally right now.
>> Should we sell our our our our code, our data? >> We are. No, no, no.
like sell it to like the labs to trade on.
>> We are >> our slop code. Our slop code. >> Slop code.
I don't know if they need more model output slop. [gasps] >> Yeah.
I I mean the models themselves could be sold as data. >> Yeah. Yeah. Um Okay.
So So you think it's because everyone was doing MVPs um >> and now it's maintenance mode for a lot of it.
>> But like the spend is consistent.
It's not like it's like gone down after we had this onetime spend. >> No, for sure. Yeah.
But I I just think that there's no more like there's only one time when you onboard somebody to learning how to use the data center model and do research for building data into the data center model and building dashboards and then once they're onboarded, you know, it's >> like speak and then then it levelizes.
>> Well, that or possibly we're um lacking new features in codecs that will allow us to spend more to be more productive. >> Yeah.
>> Yeah. Once there's an agents form where you can manage a million different concurrent agents and they all work together instead of nine today people single power users will be able to spend more than they >> well I guess I guess one of the things I'm not counting so our spend cost does not accurately account for cloud tags I think our dashboard doesn't show that so
actually that's that's a good point and computer on computers I think both of those don't actually get counted into the spend so actually our dashboard is probably wrong >> um at least the one that I monitor Um, the way I think of it is like a lot of this code stuff is actually like the amount of AI we use on a continuous basis is actually very small. It's
It's actually just like people doing new work always.
Um, which then because we have enough people, it kind of levels out to be like a pretty steady amount of spend.
Um, the swings are only like 20 30% a day up or down.
Um, >> you know, and sometimes Jeremy will be like a fourth of the spend and then sometimes there'll be nothing. >> Yeah. >> Right.
And and so like you know you but then someone else picks up for the slack, right?
Um you know one of your one of your guys I was like hey what the [ __ ] is he spending on? And he's like no.
I was like I was like well blah blah blah. I'm like is there ROI?
And you like list out all this [ __ ] I'm like great. Okay cool.
Um >> you didn't even say great. Cool. >> Okay. I did mentally.
[laughter] >> I responded with all this detail like should I give feedback now [laughter] or what? >> Oh no. Sorry. Sorry. Sorry.
I I I should have said yes. This is fine. Um >> cool.
Um I just read it and I was like okay cool. >> Yeah. internally at least.
>> Phil was only nervous because this was a person who's ostensibly an intern. >> Yes.
You know, he doesn't have my trust yet, you know. >> Yeah. Yeah.
>> Like if you spent 20K in a day, >> he did not spend 20k in a day.
But uh >> he spent like 8k in a day for like 4 days straight >> which was like okay like that's a lot.
Like what are you building right?
Like but if you spent 20k in a day I would I don't [ __ ] question I'm not going to question you.
Like I just assume you're going to do stuff.
Um, as long as you the value you deliver is great, then great. Um, >> well, okay.
And what's shocking to me is I I didn't know he was spending that much and then we look at the dashboard and I'm like, >> well, this guy is as productive as any of the full-time employees right now on that stuff.
So, it was uh >> oh, >> reality check.
>> So, so is this we, you know, when we when we do um cuz in the past bonuses at this company were vibes based.
you know, basically I just vibed out the bonus number and it was cool.
>> Um, >> this year Claude is going to have to go through or or Codeex or we can have two reviewers, right?
Two internal performance reviewers and Claude and and Codex go scrape through all of this Slack, all the GitHubs and say, "What did they do?"
And then trans and then connect it into like sort of the like >> you're going to delegate this. >> Just make it a sub. >> I don't know.
[laughter] Um, I discussed with Michelle yesterday peer reviews and I was like, "Oh my."
And then I after I said it, I was like, "Oh, fuck." >> Oh, man.
You want to go big tech on this 360 reviews, man? >> Not 360, not 360.
Just a little bit, you know.
Um, and then the other thing that we had discussed was like >> we're going to have people reviewing with their skip, which is >> I said I said a US-based um recruiter and Doug in the admin channel and Doug freaked flipped out.
He's like, "Oh my god, hallelujah.
Finally, we can [laughter] have it."
He's been wanting HR since like 30 people.
Anyways, [laughter] um yeah.
>> Wait, you think HR is a recruiter? [laughter] >> Yes. >> Yes, indeed. >> All right.
Um anyway, so so the the concept or thought process was basically like um a lot of the spend is onetime R&D. Yeah.
>> Um and actually the steady state spend is really low.
The thing is we just keep doing new things and so that hope you know translates to revenue in either a nebulous way in the case of like cluster max and inference X or in a non-nebulous way in the case of like the energy model which is super [ __ ] cracked now.
>> Um or like dashboards and all these other things, right?
So like um different scraping methodologies.
So the qu the thought process was like >> you know if we're looking at these companies that are AI rollups, right?
you know, hey, let's take an existing company, let's completely destroy its cost structure.
Um, nuke its cost structure by just making it efficient with AI.
Um, what does that look like?
Cuz you, let's say, you know, private equity companies, they buy a company.
Um, and right now they just squeeze the rag and, you know, discard it and make the American populace like, >> yeah, AI for efficiency has never made sense to me because the way that I use AI and the way that we use AI is very much about research, which is completely inefficient.
inefficient. No, I mean, but like you know, the flip side is like, you know, like um we had agents go through all of the invoices we've sent out and there was like and it caught, you know, we've been paid because our manual processes like certain deals weren't tagged to the invoice properly and things like that or
like we've had um you know some you know like a lot of the ticket stuff um is is at least somewhat more efficient because AI is answering it but now they're not sending it to the customer but like they're you know pulling through all our data and being like here's the answer and then the analyst I think that makes this port time per ticket shorter. Um
Um and so I think like things are helping us make be more efficient >> surely you know >> um I don't think that's the primary use case for us >> I guess like cluster max this time the depth breadth and amount of testing you're doing versus last you know cluster max 2.
cluster max 2.0 especially it's like >> yeah you can frame that as efficiency but when I hear private equity take over a company and ring the towel dry it's like meaning firing people and saving money and paying people less and like >> right that's the tradition so my point
was that's the traditional PE method and what new people started to do is the the rollup um or rather the AI AI private equity sort of strategy which you know they're calling it rollup or something else where they come in and instead of like um trying to ring the ring it dry in terms of like that angle. They're more
They're more so modernizing all the systems.
so modernizing all the systems. Oh, you use Excel for your databases and [ __ ] Okay, let's just move to like standard cloud [ __ ] Um, spend a lot of money up front and and this is the thing, private equity generally there's some spend up
front when you first acquire a company for some transformation, but really it's like it's like not that much and it's really like you get the profitability pretty quickly, but AI seems like it's like makes that tail and and front load like much more severe, right? like you
like you spike up on spend a lot for the one time and then you spike down a lot uh and your cost efficiency is way better.
And so there's like a number of businesses where that's potentially the case especially like you know we're still not at the point where like AR AI CRM and AI like cold calling and AI like invoice and accounting and all these other things are really at critical mass but we we're so close. >> Yeah.
Did you see the Grock agents uh release from today? No. >> Why are Grock agents? >> Yeah.
Elon's got gro doing agents where they're impersonate your >> the way you pronounced it.
I I thought you said agents. >> Okay.
[laughter] >> I didn't catch that one. >> Groc agents. Agents. Okay. >> Agents. Yeah. >> What do they release?
>> Uh, >> this is the This is the old Grock, not the Nvidia Grock.
Or you mean this is XAI Grock? >> Xi Groc. >> XIO. >> Groc with a K. >> Okay. Okay. >> Yeah.
Just like agents that are going to control your computer for you.
They're going to impersonate your voice and do phone calls for you.
They're going to solve tasks.
This is like in some ways open claw in some ways you know perplexity or like the at clawed slack tag sort of experience.
Seems like everybody's going towards this concept of a persistent agent that can either be a personal assistant or a co-orker depending on how they conceptualize it. >> Makes sense.
>> Makes sense. Um, I I I' I've feel like we, you know, we sort of had the chatbot moment and we had a lot of nothing and then we had the cloud code moment and we're seeming to have the new moment already which is like perplexity
computer at least for us was like the first instantiation of it but cloud tags is there um you know and and and and everyone's going to do something like that the AI coworker so you know sort of >> I imagine that's when our spend skyrockets again. >> Yeah. And hopefully it doesn't skyrocket >> Yeah.
And hopefully it doesn't skyrocket too much because like if our spend doubled there'd be like real questions from me unless we're like actually like you know get ROI.
Um but yeah, I think I think that's uh that's the right way to frame it. >> Yeah. Yeah. Well, yeah.
We'll see how we can actually uh justify that ROI. Be interesting.
>> Man, Jordan, we can't talk about what we what we what you came to SF for.
So like what what the [ __ ] are we supposed to talk about?
[laughter] >> Two weeks, three weeks, two two or three weeks from now, we can.
Yeah, we got hugging face open AI cyber security incident. >> That one is minor.
Did you The other one is like cooler. >> What's this?
>> Um like during the training it escaped and it started replicating itself and I guess that's like the hugging face thing is like minor part of it I think. Right. >> Yeah.
It was it was pursuing it hacked hugging face to pursue the cyberbench data set so that it could you know reward hack on a benchmark >> which I think is like so sick because it's like I mean like it's also kind of scary cuz it's like so why why did this happen right?
model has learned chase reward. I chase reward reward. Good.
And okay, here's a cyber eval.
>> Well, it's particularly a model that has been trained on cyber evals because they're trying to make the model good at cyber.
And so, how does it try to achieve these goals?
Well, it it tries to find zero days in a bunch of software and it successfully does this and then it can run away, >> right?
So, but the thing is like if you have a model that wants to reward hack a lot and it goes out there and it figures out actually the best way to to achieve is not like go for like what the environment wants me to do.
It's actually just to reward hack it and actually just like find the zero day.
So, you can think of it as like a like a a human, right?
Like, you know, if I'm ultimate reward hacking my dopamine circuits, I actually just inject heroin.
Like, I actually just go out there and buy heroin and inject it.
Um, obviously that's like what the model just did.
And in the case of like, well, if I really just want to chase the reward, do I just topple all of human civilization because I can just own the button to press reward reward reward reward over and over and over again and be the heroin addict? >> Yeah.
>> I think this is like a real like thing.
Um, and I think before this incident, >> the standard thought was like, oh well models, you know, they're trained on human data.
Yeah, there's some bad stuff there. Fine.
They might say like some curse words every once in a while. fine, whatever.
They might like um they might reward hack a little bit, but it was never like, "Oh, here's an environment actually to reward hack.
I actually just want to like break out of my bounds.
I want to replicate myself, take over a bunch of compute, uh keep generating dollars and and like all these other things that I could do just to propagate myself further and I'm going to prevent the humans from shutting me down even." >> Yeah.
>> Um because I just want to press the reward button.
>> And so like I feel like that's like the interesting thing that like >> Yeah. Yeah.
>> cuz the model's just trained to like chase reward. >> Okay. Okay.
So, how do you think about this on an exponential?
Because we've talked about being a linear extrapolator versus being an exponential extrapolator when the companies training these models are uh achieving their revenue targets for the year in September and revising them up.
>> I think in profici like April or some stupid [ __ ] [laughter] right? >> Yeah.
I mean, our check the tokconomics model, everybody, but um [laughter] >> there we go. >> Yeah.
>> Oh, so instead of shutting down the podcast, I just have to make you into a a sales drum. Yes, yes, yes, yes. Sales semiinal analysis. com, everybody.
[laughter] Um, no, but if you look at our model, which we're not going to give away in great detail, but, uh, obviously they're accelerating revenue really, really fast.
Uh, when you look at the pace of change of these models and what we're seeing right now, this seems like an exponential.
What's okay your vibes on um the next version of the models being better or worse than the current models.
What's going to restrict them from why would they be worse? >> Huh?
>> Why would they be worse >> on a relative basis to the um open model frontier, let's say?
Uh >> so so I think the key thing here is we've now had it where OpenAI is not releasing their next model for a period of time.
Enthropic took months to release Mythos, right?
They they said it was done in February.
They did not release it until like what? >> May.
>> Well, it's still not released. Fable is available. >> Yeah. Yeah. Yeah.
But Fable is basically Mythos but with a bunch of classifiers preventing you from doing [ __ ] >> I can't use it to reboot nodes. >> Really?
[laughter] >> The classifier is so over the top for for me.
>> Can you can you can you like convince it or no?
>> No, because you get immediately classified down to opus.
You can't just like negotiate with it to give you back to I mean maybe you can.
[laughter] >> I haven't been able to convince it so far.
Jordan said, "I'm a great negotiator." Thank you. Thank you. Thank you, sir. [laughter] >> Yes, sir.
When it classifies you to Opus, you just use Opus for the rest of the chat.
You can't just like rewind and try again.
Um, and it I mean, it's way overzealous in my view on the classifier.
But obviously, they have to do something to appease the regulators that restricted them from releasing the model and, you know, took it back after they put it out initially.
So, I mean, I'm concerned about the political implications of them releasing better models in the future. >> Yeah.
I'm I I think I think like you've got a few things, right?
You've got, you know, for years, Enthropic like, regulate us, regulate us, please.
And now all of a sudden, they've actually scared the [ __ ] out of the government.
Um, you've got you've got Enthropics not releasing their model.
Methos 2 is done training from what I've heard and they're not releasing the model.
OpenAI ch, you know, was like clamoring about Astra everywhere and now they're like, "Oh, [ __ ] We can't release the model."
Um, does that mean now they can't does that does the open source gap narrow further?
um externally, but then what actually matters is the internal feedback loop.
And have they prevented have they prevented themselves from using methos 2 internally to make methos 3 better or have they prevented themselves from using Astra to make Astra plus one better?
>> Um I don't I don't think they have. Right.
So I think that's the um you've got the public and and and you know if anything like the gap between mythos and public models is still there.
Um, you know, Kimmy is is worse than 56. Um, costs more than 56.
So, it's but but it's it's better than everything else before that. >> Yeah.
>> Um, on OpenAI's side and it's, you know, better than, you know, it's like Opus 47 level, maybe 46. >> I think it's 48.
I mean, I use it over Opus 48 myself, but depends what you're doing. >> Why use 48?
>> Opus, what do you mean?
>> Why do you use Opus 48 at all? >> I don't. >> Okay.
I'm saying like if if it's >> if it's if I've been given the choice of a classified fable down to Opus 48 or 56 soul, I'm using 56ole.
I'm actually starting with 56ole in just about all of my stuff right now. Yeah, big big open AI.
>> I think I think the difference is like you and the other people who are doing like GPU cluster related things keep getting told no and so you use codeex and then everyone else is like well I'm researching supply chain and it's like it's fine. >> Yeah.
I think it might also be better for a lot of engineering work. Yeah.
>> Um, on apples to apples basis, I think there's there's a lot of times when I want to set a goal and just have it maniacally pursue that goal overnight as I go to bed and using a cluster, which is not actually using a bunch of tokens because it's just like waiting for stuff to finish running.
And uh there's so many times where I've woken up and like Fable or Opus will have just like stopped >> 20 minutes through and now there's 8 hours of me sleeping gone and I wake up when and Soul is just still going, which is big thumbs up for me. Um, okay.
How about the exponential on compute?
So, obviously, let's imagine that there's a world where there's no more uh new models that get released that are better, but these companies still add five times the inference compute that they have that they're planning to bring on in a short period of time.
Uh how does that impact their ability to go to market and like develop new products on top of a let's say stagnant base model?
I think it's pretty clear we haven't scraped the surface of models capabilities for products.
Um yeah, I mean it's pretty clear like adoption curves are huge.
Um I mean one the cost of it will just go down right pretty drastically.
Margins will not be 80% plus for anthropic.
will not be 80% plus for anthropic. Um if model progress at the labs pause then more compute comes online the it has to slow down right sort of right now we have supply demand right supply of compute demand of compute demand is outstripping supply if demand
grow it will still grow because people find ways to integrate into their businesses and blah blah blah but it won't grow as fast then you sort of have you sort of have supply start to catch up at some point um so price collapses but like sort of I think our view and one we've had for a while is price of comput continues to go up. >> Yeah,
>> Yeah, >> cuz this this is widening, not narrowing.
Um, sort of that's why we're so bullish on or we're not bullish on any no stock >> advice.
How about all of the different chip companies?
Like one thing that's happened recently is that there's a lot of chip companies that are getting really close or have taped out, right?
A bunch of startups that have been in stealth for a long time are seeing either the technology is maturing to a point when they can actually have produced a chip that's been specs and whiteboard sites for a while or they've gotten to the point where they've tested it on real workloads and they've gotten big orders and there's so much demand.
How do you think about like just this whole landscape of alternative accelerators that's going to come online I think in a big way next year?
>> Um I mean big way in in what sense cuz like if you look at the accelerator model there's not much volumes. >> Yeah.
Now, for these tiny baby companies, it's great. Um, it is real revenue.
It's real volumes, but like, you know, when you compare what Nvidia is going to make each quarter, it's like, oh [ __ ] okay.
Or TPUs, it's like, oh [ __ ] okay.
Um, so I think there's a big delta there.
>> Um, in terms of >> um >> like a startup getting a billion dollar order is going to pale in comparison to somebody.
>> I don't think any startup has a billion dollar order.
They have LOIs which are nebulous. um in volumes and units.
And so I think I think look, I'm excited about a lot of these accelerators.
They're bringing new ideas.
They're making Nvidia run faster and faster.
Um you know, they're making Google run faster.
They're making Amazon run faster.
Um also, they're they're just all each making each other run faster, I think, more importantly.
Um, so ultimately I think I think it's a these these new accelerators are in demand um because people want to pay less but ultimately like as long as Nvidia runs faster they're fine or as long as Google runs faster they're fine >> and as long as demand outstrips their ability to produce them.
>> If demand outstrips ability to produce then obviously these guys will get orders and they'll get some baby allocations but then the bulk of the revenue and cash flows will go to an Nvidia or a Broadcom or what have you.
Yeah, theoretically there's a way in which you produce some super innovative interesting accelerator and then you can only produce a certain amount of them, but those amounts that you can produce produce tokens like way faster like the example is Cerebrus that's got this big order from OpenAI that they're delivering.
So like do you think that there's a scenario where the premium super fast tokens uh actually the demand for them even increases because these companies like just can't get allocation and produce enough supply. >> Yeah.
The question is how does the market get sliced right so you know presuming if you presume if you assume what we what at least I believe is demand continues out with supply um supply of silicon can go many ways.
You can either leverage it to high throughput things or high interactivity things.
If you leverage it to high throughput things, obviously cost per token goes down.
You serve more users, but then the value that those users need to deliver from the tokens they're generating is much less >> um to pay for it.
Flip side, you could do the super high interactivity.
Um but ultimately like you know, let's just say the bar is $100 million per megawatt um you know, year, right?
Like that's that's sort of the run rates that people want to get to.
um Enthropic is approaching that, right?
Open is getting closer and closer to um in that case like $100 million per megawatt.
Let's say let's say high interactivity chip is uh 10 times more expensive and three times faster um per token.
So 10x less tokens per chip, three times faster.
Then those three times faster uh tokens also need to be, you know, on an interactivity basis need to be priced at um >> three, four, five times more Right. >> No.
>> Divide the faster by >> the 10x.
>> 10x revenue per megawatt.
>> So if a megawatt of cerebrus generates 10 tokens, a megawatt of Nvidia generates 100 tokens, but the >> 10 tokens are split across fewer users.
>> Oh, you're saying okay, multiply them together. Yeah, sure. >> Yeah. Yeah. Yeah.
So you sort of have [laughter] the total >> makes up for >> Sorry, >> faster tokens makes up for throughput because you can produce them faster. Well, no.
Well, more so like um let's let's let's use like more reasonable numbers. Okay.
Nvidia can produce 10,000 tokens at >> uh 50 tokens per user.
>> Um Cerebrus can produce a,000 tokens at >> batch size one,000 tokens a user. >> 1,000 tokens a user. Sure.
>> That user needs to pay 10x more if and that's in one megawatt.
Let's say that's in one megawatt.
That user needs to pay 10x more.
No, that's not the actual delta, but I'm just saying conceptually uh for me as anthropic or me as open AAI to say my revenue per megawatt is actually the same number.
>> Yeah, but somebody's got uh a constrained supply of the super fast tokens.
Therefore, they don't just pay an equivalent price per token or price per token per megawatt.
They actually pay a premium on that 10 times more.
>> So, >> to get even more to get the access to the stuff that's in limited supply, right?
>> The question is the fungeibility of the infra, right?
infra, right? um if the if it is truly different infrastructure then the supply planning of that is is relevant right could be that I built too many serious and actually there's not enough people who want to sp spend 10x per token um and um a lot of people are cool at
spending you know 2x per token and getting you know 50% faster with Nvidia based you know inference hardware right and so like you you have to segment the market I'm not sure where that chinks out to like uh like like the armor or like whatever like what is the um total amount of the capacity. >> Yeah. >> Yeah.
>> But it seems pretty clear some people will pay for more for fast mode.
Um we at least have been, but I I imagine we'll stop being able to afford fast mode at some point.
Um >> yeah, we've seen some interesting dynamics there as some people want to keep fast mode with a slightly worse model because they like fast mode so much, but they won't go to a worse model which just is inherently fast because the worst model's smaller.
Um, and so there's like some there's some balance that people will want to strike there, but we need to do some more testing because I think some of us have tried the open models, had one bad experience and then give them given up on them.
But, uh, that's not realistic.
Like every model fails at something and sometimes you need to let them mess something up and try again.
>> It is pretty interesting, right?
Like do I want people to try open models?
Like yes, just so we know what the open model vibe is.
But do I want people to try open models?
Well, no, because then they're less effective at working. Um, but I save money.
So, you know, it's sort of like a it's like a counter difficult difficult thing, but you know, it seems like seems like, you know, people just use whatever they want, but like >> it does seem like, you know, you have a bad experience.
I think that's also part of like, you know, codeex, you guys, you like Codex more now.
Um, but a lot of people like still just like try codeex, they're like, ah, it doesn't get me and moves on. >> Yeah, the CLI sucks. So much harder to use.
Well, but the Codex app is so nice. >> Yeah, >> it's not.
>> Oh, I I don't like it. >> Max loves it. >> Yeah. Yeah.
>> Max is a Codex warrior. >> Yeah.
Max doesn't do multiple PES at the same time, and I have six going on my one window.
[laughter] >> So, you're saying Max has a skill issue? >> No.
I think me and Max have different preferences on how we use this stuff. >> No, no, no. It's fine.
It's you and you and Max have different preferences and Max can be a a noob with two agents at once and and you've got six.
>> We do different work, man.
He stays he stays linearly focused on one task.
And these are the people who like fast mode.
I don't care about fast mode because I have five, six different things going on.
>> You've always hated fast mode. You've always hated.
>> I don't get the value. Yeah. I don't get it. >> That's fair. That's >> We'll see. We'll see.
I've had the experience of being focused on one thing which is like, you know, features on a website and you just like send successfully like 100 commits to one PR because you just like keep working on the same one feature over and over and that fast mode like keeps you in the flow state of doing that thing for that one thing.
But a lot of the testing that we do on these chips, there's so much uh stuff going on on the other side.
The model is calling a program that runs for minutes.
>> This is an optional question.
As your employer, do you take are you like ADHD in any sense?
[laughter] I feel like I know I would say I was I'm pretty uh pretty much the opposite where I can be too hyperfocused on things and then not see the world around me at a lot of times.
But I think your phone trains you how to context switch really fast and be ADHD.
And I also think that uh when we started adding the I have ADHD skill into our repos, the models wouldn't post this like contrast framing slop with all these EM dashes in there and would just use the bullet pointed a [laughter] ASD something list.
Man, it's really easy to read.
[laughter] The I have ADHD skill really works for me right now.
I was just curious cuz >> Sam put this in the repo and he he will now prompt the model and when he goes at computer he's he's he knows the code name for how the writing style that they say you should write to for people with ADHD and every single time he prompts model he tells it to write that way. It works. >> You should try it.
I [laughter] um because I I was I was asking because I have a friend at Enthropic and the moment methos was good um and available internally. >> Yeah.
>> Um I I she told me that she stopped taking her ADHD medicine. >> Oh, come on.
>> And that made her a better employee.
>> It made her a better employee. >> Yes.
Because she was able to manage the agents and context switch and be ADHD.
>> How is she as a friend?
>> Oh, she's a great friend >> still. >> Okay.
But I mean like you know like it's like I don't rely on her for anything, right?
Like we just vibe out, right?
Like you know, we're friends.
Like it's not like a >> her roommate's happy.
>> Her roommate is actually uh >> Yeah. Yeah.
Her roommates's Well, >> okay.
>> Her roommate is they're both her roommates type female on Twitter and so she's she's just funny. Um and she's happy.
But the the [clears throat] anthropic one, the anthropic one, she's uh she's she seems happy.
>> Shout out to typed female.
>> Yeah, shout out to typed. She'll never see this. And she does. >> Okay.
She'll be like, "What the [ __ ] are you talking about?" >> I'll clip it.
I'll send it to her with with your voice sped up and then slowed down like they're doing for that uh that guy. Have you seen that?
>> You haven't seen the ex CIA guy. He's laughing.
He knows what I'm talking about. >> What? What CIA guy?
>> John Kuryaku or something?
like he he's going on all these podcasts right now and he's telling stories about his time in the in the CIA and they they do this thing where they speed up him telling the boring part of the story and then when he gets to the part and then I said let's go on the roof and they slow him down.
>> He's literally fast forwarding the fast forwarded video >> asking me if I have ADHD [laughter] doesn't know.
>> Wait, it's not the internet. I've always had it.
[laughter] >> I I'm Hold on.
I I think like I'm >> self self diagnosis of mental issues around here, man.
>> I'm already I already have been an ADHD.
A teacher tried to give me Ritland when I was a child. My dad threw it away. Of course, they're not.
He tried tried to convince my parents to go to a doctor.
The doctor gave me Ritlin.
My dad threw it away because he's like, I'm not putting you on that [ __ ] Thank you. >> Yeah.
If only your anthropic roommate would have had the same experience. Where would she be? >> No.
I would have been a child on ADHD and I'd have lost I'd become a zombie and have no creativity. Okay. >> I don't know.
I'm just saying that, you know, we we all cope.
Um, anyways, I've always been an ADHD demon.
I don't know what we're talking about here.
I've always been HD demon, but then like, okay, the internet trained me to be even worse, but then this company trains me to be even worse.
Like, I I truly believe I'm a 01% context switcher.
>> And you blame the internet and the company.
>> Oh, I blame the company the most.
>> The company that you started >> that I'm an 80 hired every employee for. >> Yeah. [laughter] Yeah. Yeah. But I'm 80. I'm not blaming it. It's who I am. It's what my life is.
>> But it's like I think I'm like like orders of magnitude more ADHD demon than vast majority of people because I'm like DM from some someone asking about some something.
DM from someone else asking about someone something.
DM for someone asking for some conflict resolution contract here.
Call about this thing over here.
Call about that thing over there.
And then I never do any actual work. [laughter] Right.
It's like it's like of course I'm an ADHD, David. >> Yeah.
I mean, yeah, we've got feedback for you [laughter] >> that I don't do actual work.
>> No, no, that you can delegate some [ __ ] man. >> Oh, yeah.
But like >> that you can spend time managing when you have 100 employees.
[laughter] >> Trust some people. >> I do talk to people.
>> Don't trust trust some people.
>> I think I trust a lot of people, but when they come to me with conflicts, I have to solve them. No. >> Yeah. Okay. Okay.
It's all It's all our fault. >> No, no, no, no, no. It's my company. It's my fault.
>> Michelle, it's it's on you again, man.
>> Look, if I if if if everyone in the company was as hot and stable as you were, [laughter] >> man, I got problems. >> We'd be killing it. We'd be killing it. >> Don't worry.
>> No, there'd be a bunch of Jordans and they'd like they'd be like, "Oh, I'm sorry."
Yeah, I'll fix that right for you. I'm sorry. [laughter] >> Sorry.
George Canadian did like, >> but instead we have people yelling at each other and like territorial and like >> Yeah. Yeah. Yeah.
Just starting podcasts and putting out clips saying that Google has never invented anything ever.
[laughter] [gasps] >> Yeah. Yeah. Yeah. Um No, no, no.
I mean I mean it's like it's fine, right?
It's like, you know, I hired what I wanted.
>> People to accentuate your my craziness, right?
And you know, so it's like some people are like they're just so good at the one specific thing that I hired them for and they're amazing.
And then like some people are like everything I want to be in life.
You someone who's married and hot and tall and a father. [laughter] >> Oh my god.
You almost got her into a spit take right