0:05
Hello everyone.
Hello everyone.
Welcome back to Semi analysis weekly.
We're here with episode 20 with the tokconomics team. Things are moving fast.
We are recording this on Wednesday, July 15th.
I assume by the time this episode comes out, there's going to be a lot more releases of models, disclosures from these companies about their profitability, revenue, uh how many active users they've got.
Regardless, we're gonna record a point in time right now talking about token budgeting, uh, Meta's comput strategy, MSL futures, the release of Fable, Soul, and Anthropics profit margins. So, excited to dig in.
With me today, we've got Max. How you doing, buddy? >> Doing great. Thanks for having me. >> Welcome. We got Crystal. How's it going, Crystal? >> Hey, Jordan. >> And we got Joey. >> Hey, Jordan. >> Cool.
So, we're going to start with token budgeting. Crystal, over to you.
We got uh lots of conversations that we've been h having with enterprises over time and just asking them like what's going on.
Uh at some analysis, we're still token maxing, but others are moving into the time of austerity.
Can you take me through a little bit about what you found with this article?
>> So, there's a lot of like uh enterprises right now that are cracking down on their budgeting because they think their employees are spending too much.
Yeah, a lot of people are getting scared because they're like, "Oh, what if I don't have enough tokens to do my work?"
I don't think that's necessarily true then, and it applies to most people, right?
Because a lot of people aren't even getting close to that limit.
There's probably a few handful power users at each of these companies which will really be affected.
But other than that, I think it's been it hasn't been as big of an impact, I think, on most people's workflows as it has been portrayed on social media and stuff.
There's been a lot of different strategies that companies have been imposing whether it's like on a per person basis or like a monthly budget for the entire company from what we've seen.
And I think the per company basis obviously makes a lot more sense than like on a per person basis just given that some users will generate more like use from it than others.
>> And how do you see people actually enforcing this?
Like clearly you can burn through a budget faster if you're going with the super ultra premium max thinking model, but some in some cases when you're actually enforcing a budget, you take different approaches, right? >> Yeah. Yeah.
So it I've seen some people actually so I was talking to some people and they are like a smaller company and they have a much smaller budget.
So the way that they do it is they try to optimize how they're using the like more expensive tokens for like anthropic or open AI.
So they use like a cheaper model to like process a lot of it and like condense it down to like a smaller prompt and then they use the more expensive models or they use something that their company's not counting.
So there's a lot of like discrepancies in terms of what models actually count.
Some people or like some enterprises are counting their own models into the budget but some enterprises aren't.
So there's like if they aren't counting it, a lot of people will tend to like really push those to the limit, use those a lot and then use like the ones that actually cost money.
And you're seeing people maybe Max I can bring you in on this one.
You you've been seeing a lot more people burning tokens for coding rather than other use cases, right?
Like that seems to be the market that's growing the fastest.
>> Yeah, I think coding/ like software jury in general is just by far the most token hungry use case.
And I actually think this is why a lot of the token budgeting discourse is pretty like I think a lot of the budgets themselves are pretty uninformed.
Like I hear people saying that they want to make sure their sales guys aren't using Opus or Fable to write emails.
So they should definitely be using Sonnet instead.
And it's like dude like generating your email is basically free from a tokens perspective.
Like it does not matter if you use Opus or Sonnet to write this email.
Like I have no idea why this is the policy you're enforcing to try to like reduce your token spend.
I've also heard lots of stories of like certain companies where only the engineering team is allowed to use like cloud code or codeex and then you know everyone else only gets like codework or something like that.
I think that's pretty shortsighted and it's honestly you shouldn't have a cast system where your engineers are at the top and everyone else is just forced to use like demoris.
You should really be giving everyone the opportunity to try these tools and figure out how you can become more productive with them. >> Yeah.
Joy, are you using these models right now?
How would you feel if I got access to it and you didn't?
>> I would probably be I do use yeah some access not as much as like Max and other people on at semi analysis but that's probably just the nature of our work or my work right now and that's changing a little bit but yeah I'd be I'd be pretty pissed because I still use Opus 4.
8 a lot like to pull in to do more like aic research tasks that will pull in a lot of like different data sources do more like research type tasks.
Um yeah be I'd be pretty mad.
>> How do you think about the ROI question?
And I think a lot of people the the whole motivation behind token budgeting is clearly that somebody is not seeing a return on this investment.
Either that or they're just like not seeing it yet.
But I think in a lot of that research you're doing, you're seeing that there's a lot of benefits, right, to using >> Yeah, I think like with the market being so coding focused today, so probably 70% plus of AR lab AR right now on the API side is in coding.
And then two, like from what we've heard, it's very power user a power company driven.
So it's pretty widespread like coding.
There's no like major big company for anthropic.
Like meta is probably 3 to 5% of total anthropic ARR but it's heavily driven by power users.
Um if you're at a point where you're a power user at one of these companies I think the ramp data is like the top 1% you know the 99th percentile of companies they spend 100,000 you know per year on AI per employee.
And if you're if you've gotten to that point like where you're spending that much you're obviously getting pretty good ROI.
And I think like our conversations with a lot of enterprises were these top engineers could go, you know, well above, you know, their their limited bud, you know, whatever their budget limit was.
But if you were someone who was spending, you know, an insane amount of money, you know, you were getting budgeted already pretty quickly if you weren't seeing ROI.
Um, and I think, you know, there's always projects that you'll try that don't get you'll scrape that you'll scrap that and then move on.
But I think it was yeah it was pretty obvious that because the market is so coding focused like people are getting ROI they continue to spend.
I think we continue to hear then net new AR both anthropic and now open AAI uh with like codeex and in 5. 5 5.
6 like since um late March spend is broadening there's clearly a lot of like API token spent today.
Yeah, there's a lot of organizations like in the, you know, that aren't traditionally, you know, doing a ton of, you know, development type work or focused and they're, you know, they only allow their like middle back office workers to use, you know, Copilot 365 or SAT, you know, things like that.
Um, but I think the market will continue to broaden out.
Um, >> do do you guys did you guys have a take on the balance between API usage in applications versus like coding plans when it comes to people's individual usage?
Like do you ask them when you're talking to them if they are using a cloud code subscription versus API paying per token?
>> Well, I mean anybody who can use a subscription definitely should use a subscription just because it's so subsidized.
It's just you know OpenAI and Dropic are very wise to this and so when they have their enterprise plans like same analysis is on a you know a drop enterprise plan they just they don't let you use a subscription they force you to pay for a token because they know that's where the margin is but I think in general like anyone who can use subscription does use subscription and it's just the people who need more limits than that that pay for API.
>> I've even talked to like a few people at like startups and stuff and they're not on like the team or the enterprise plan at all.
they just buy five, however many like uh Claude or OpenAI subscriptions they need a month, charge it to the company card, and they just keep doing that.
If they hit the limit on however many accounts they already have, they just buy another one for the month because it's just more cost effective that way for them to operate than it is for them to actually use the API billing.
But yeah, if you're a startup that you're like young enough to where you don't care about like security or any of the other enterprise guarantees, like just absolutely buy, you know, 10 subscriptions per person, >> you know, paying for API credits.
Like I I think the recent uh, you know, study we did on this, it had a decently viral tweet said that >> for Codeex, I think $200 a month, you get like around 12K worth API credits and for anropic it was like 200 a month for around 8K.
Like obviously that's a no-brainer for one that is able to pick between those two.
>> Can you dig into that a little bit more, Max?
Like when you're talking about the margin profile of these businesses and then you're considering the subscription plan versus the API like what is that uh what does that look like?
I'm going to bring up the chart on screen that I think is the key one from uh that tweet thread. >> Yeah.
[clears throat] Um, yeah.
So, essentially, um, obviously when you hear that like a $200 a month plan can generate $8,000 worth of, uh, API credits, uh, the subscription is going to be much lower margin if the average user uh, has anywhere near 100% utilization.
Um, and so really the key question is sort of like what is the actual average utilization across our entire user base?
And this chart that Jordan showed us here is like kind of the break even percentage um for the various plans.
We can see for the Cloud Max 20x plan, it's like just 10%.
Um for the Pro and 5X, it's like a little better at 20%.
Uh OpenAI because they are sort of even more generous with their subscription limits is slightly lower at 11.
4% for the Plus and Pro 5X and just 5. 7% for the Pro 20X.
Um we're pretty confident that the average utilization especially for these like 20x plans is much higher than 10% and uh 5.
7% respectively for uh anthropic and openi uh which would mean that like forget being you know worse margin than API which we think is like likely 85% plus for enthropic at least uh it might just be like negative margin in general for these 20x plans and so of course like these businesses want to move as much volume as possible to API.
Um, and I think maybe Joey since she wrote the newsletter on Thropics business, this would be a good segue to talk about why we think Anthropic is potentially in a stronger position than OpenAI. >> Yeah.
And and I think too like once you move on to enterprise um subscription plans at Anthropic, there's no usage included in your subscription fee.
It's all it's all done at API pricing.
Yeah, I think at Enthropic what we found was really interesting is because the business you know today is so much more and this is this is changing but because it's you know 80% plus of ARR is on the API side and that is as Max showed is is really high margin
you know they've gotten to a point where they're profitable today um you know and some of that is on an on gap basis so excluding stockbased compensation but like operating profit um was was profitable in Q2, you know, could be profitable to the tune of 1 billion plus in Q3. Um, some of that's a function of
Um, some of that's a function of they've just grown so much I don't think they can hire enough or like plow enough back into training as as they'd like.
But because of that margin, that kind of gross profit advantage that they have today, um, they can plow as much back into training as they want um, and can kind of extend their model advantage and model capability advantage.
We think that's something they should obviously try to do.
So just because it's, you know, very profitable for them today, you doesn't mean they should go run at, you know, 30% ebid margins right away or as quick as possible, show that this is a very profitable business model, but continue to invest as much as you can into training as possible. >> Yeah.
Can you can you talk about the mix between consumer and enterprise?
Assuming that consumer is those paid plans and enterprise is pay peruse on the API like >> Yeah.
When you when you look at these things, clearly the bulk of like Twitter users are using the subscription plans, but then you get like one really, you know, big customer who's spending so much more money on the API tokens that it kind of drowns out everything else and that I guess makes it a good business. >> Yeah.
And because Anthropics really focused, they have been historically focused on enterprise.
It was 90, you know, it's 90% plus enterprise.
you know versus chat GBT has gone into you know started heavy in consumer up in you know back in Q1 60% of AR was in consumer and only 6% of the 950 million weekly active users end up actually paying for a plan and most people pay
for the $20 a month plan it's a much different business mix because then you know OpenAI has to subsidize and pay for you not subsidize but there's a cost to serving those other you know 900 million weekly active users that are free. So
So lowers their gross margins on a blended basis by about 20 points.
You know, that's changing.
So I think, you know, we think personally with some of the rise in OpenAI's API business uh since late March with Codeex and then 55 and 56 um that the 6040 consumer enterprise split, you know, has flipped now to 40, you know, 60 consumer versus enterprise here in Q2.
And it's, you know, from what we're hearing like in the channel, um, especially out of, you know, token as a service businesses like Bedrock and Foundry, um, is that, you know, Open AI has recognized this and it's starting to shift and even, you know, AR like net new AR is starting, you know, to look a little more even between anthropic and open AI. >> Yeah.
And it's kind of shocking like with the release of 5. 6 six soul.
The guys like Tibo on X who are tweeting about the daily actor user counts are saying that they're celebrating going to seven and then to 8 million active users.
But to be clear, they've got over 900 million weekly active users on the free tier of Chachu BT.
Like they've got a lot of room to grow to just get people using the number one paid product.
Maybe they got a few extra users using other paid products, but I would assume the bulk of people using the paid products are using codecs, right?
So that's not a great conversion rate at this point of in terms of who's actually using the full paid product from the, you know, base of 900 million. >> Yeah, it's funny.
You actually see it too in the paying rates for like even the consumer version of of Claude.
like cloud is like 50 like 9% of free users end up paying.
So it's a much more like focused user base.
I think you if we think of those 900 million weekly active users a lot of it's like Google search type replacement things like that.
And then you know obviously because of that consumer focus they've been working on advertising a lot consumers traditionally monetized you know people aren't willing to pay the retention curves are typically pretty poor.
um and how you know people how many people are are paying you know 3 months later, 6 months later, 12 months later.
Ads is a tough business model.
It's tough to tough to scale but you know maybe that's the eventual you they've put out some pretty crazy targets for 2030 you know advertising revenue but yeah it's consumer is a tough business enterprise is is you know obviously given the margins on tokens for API like enterprise is the place to be.
I think Anthropic had a pretty good lead because they're so lever decoding early, you know, the first six months of this year and it seems like it's shifting to more of a kind of equal two horse race uh here in mid July. >> Yeah.
Crystal, what are you seeing in that two- horse race?
Like are people uh that you're talking to using anthropic codecs both?
>> I feel like a few weeks ago when I was talking to people, people were pretty heavy on using anthropic cloud code specifically.
And if I asked them why, they would say that's just what I've been using, right?
And I'm used to using it now.
But now as people in like the consumer section start using like chatbt a lot more than like claw, that effect in like the consumer sector kind of starts trickling into like enterprise a little bit where like these people who were previously like claude code like power users didn't touch codecs at all.
They're kind of like okay, what if I like consider tinkering with codec?
So even though consumer is like a really small portion of the pie and the margins obviously aren't as straight as enterprise like and historically speaking open AI hasn't done as well in the consumer market as like Anthropic has but like they have the that traction in the consumer space that eventually like is going to trickle in and we're already kind of starting to see that in like just this past week already.
>> Max, what's your take? OpenAI versus anthropic.
Where do things stand for you?
Uh I think I remember last time I was on this podcast, we were talking about how it was it was right when 55 came out and we were talking about how things uh were looking really dire for open eye at the start of the year.
Um I think ARR growth was likely flat in like March and April which is the alarm bars are definitely ringing, you know, in Sam's head when that happened.
But then we said that 55 would like to be an inflection point.
Like this was actually a really good model, you know, probably opus 48 level.
And I think that prediction has played out even like sort of more quickly and more optimistically than we would have guessed.
Like some people here are saying that 56 is as good as Fable sort of just like straight up even though it's only half the price.
And we also likely think it's like a much smaller model which would be sort of a testament to OpenAI's guess like training abilities.
We think that like as Joey mentioned earlier, net new AR like month on month is like a comparable between open AI and anthropic now which is pretty shocking like people people were counting OpenAI out a little bit and I think Jordan might have been one of those people but [laughter] they're but they're back guys they're back.
>> I am nothing if not flexible.
It's always been about the model guys like everybody's like when are you gonna stop hating on OpenAI as soon as they ship a model that's good and it's good.
I mean, for the things that I want to do, yeah, I'm Yeah, basically default OpenAI at this point, despite the application not doing what I want it to do with multiple tabs and stuff like that.
>> Wait, are you using the Codeex app or like the CLI? >> CLI.
>> What's your form factor? CLI.
I feel like they don't want you to use the CLI, dude.
Like they they very actively want you to use the app.
What is uh what is pulling you back? >> Tabs.
I need like multiple tabs going. Yeah. Uh, okay.
I think I think the way you're supposed to do this because I agree like the first time I tried using codeex and I was like every single time I open a new chat, it's just like lost in the ether on like, you know, the lefth hand sidebar of all those tabs.
I think how you're supposed to do it is like you have one pin thread like per project you're working on and then you just like ask that thread to spin up subthreads anytime you have like a discrete ask.
And so you actually like never even interact with the vast majority of threads like in the left sidebar.
It's just like your couple pen threads that you use. >> Yeah, that sucks.
Second of all, >> second of all, I'm in remote SSHs all the time and their remote in the thing sucks.
So, yeah, both of those are I'm very against it. >> Makes sense.
>> I think they are working on this, but yeah, I'm I'm just I'm not reading the code to be clear.
>> Thanks for clarifying.
just yeah last time you were on I you know we had that discussion but you know still have to run some things still have to run some things from the terminal and see some output in the logs at some point and this is just like not part of the standard codeex thing but I'm not
the primary market like I understand that they are targeting vibe coders around the world you know >> we need to help people make more B2B SAS is like the core business model of this and that's I'm not the target market and that's cool I'll keep using the CLI. It's all good. It's all good.
>> Actually, I wonder if like internally most of their software engineers are using the app with CLI.
It's semi analysis for sure.
I think all of our real engineers are like CLI, you know, die cards, right?
Like they'll never give it up.
>> That would be a good question.
You're going to ask them that. >> Yeah. >> Yeah.
>> Well, next time we I see some open people, I'll ask. >> Sounds good. >> Yeah.
>> Okay, let's uh let's change gears and talk a little bit about who's going to come in third.
There's been some releases uh that we've covered a little bit. SpaceX AI is back.
They've got Grock with cursor bolted on and things are going well.
Meta has mus spark out in the world.
They've also stacked up a whole bunch of compute and they've teased launching a Neocloud, which of course M SpaceX was first to do.
So, what do you guys want to talk about first, the models or the compute for these guys? >> Let's do compute.
I think it's more interesting.
NASA is going to tell you to walk through the models.
I think yeah, >> I mean I think quip tlddr the models is like they're not bad but at this point I don't think it matters unless you can deliver like a true frontier model and they're obviously not true frontier so who cares.
>> Do you think anyone will >> max their API like product?
>> I don't think anyone's going to use the API until it's a frontier model. >> Okay. >> Yeah. Yeah.
>> So when they when they have an API, their token as a service business will be >> will be fire% margins too.
And >> I mean that's only if it's like better than whatever the best entropic and OPI model is at that time which I think easier said than done.
I think like Grock 45 spark 1.
1 they're only useful as like uh just like proving points along the scaling curve.
It's like okay these guys unlike Gemini they're like still like kind of in the game right now.
um you know they can train like adequate models.
It's not quite frontier yet, but there is a nonzero chance they can catch up I think is the way to view them.
I will personally be shifting a zero with my token usage to either Grock 45 or 1. 1. >> Okay.
But you said the other name which is shockingly in my view in fifth place right now.
Gemini, do you still that view? >> Oh 100%.
I think they're like clearly in fifth place.
[laughter] And like unless unless Gemini 3.
5 Pro is like better than the you know industry chatter that we're hearing, I think they're going to stay in fifth place maybe forever. >> Maybe forever. Okay.
Talk to me about compute.
So staying in fifth place forever would kind of be related to the quality of the model, but the quality of the model is kind of tied to how much compute these guys have. >> Yeah. Yeah. Yeah.
>> Yeah. Yeah. Yeah. And I think so I have to say that when I first saw the SpaceX entropic deal where Elvon rented like basically all of Colossus one to Atropic, I think I incorrectly viewed that as Evon when in reality like I think the really important question is like do you have the ability to claw back all of your compute assuming that you have proven that you can truly reach the frontier by scaling up your compute
and I think that's sort of like the view that uh Elon is taking with like uh cursor and XAI is essentially saying like I'm going to leave you guys with enough compute to sort of prove that you're capable of reaching the frontier and in the meantime I'm going to monetize all the additional compute by like renting it out on these like extremely high margin you know 3x maybe 4x market rate uh rental deals. I'm
I'm going to make sure there's this like 90-day clawback clause included in every single one of them such that if you do really well, I'm just going to claw back all the compute, give it to you so you can, you know, make digital costs.
And I think that is like a valid strategy for someone who still truly thinks they can like build RSI.
And I think you can really think of where it's like OpenAI and Enthropic are monetizing their compute by just like selling inference, Elon is uh doing these short-term comput deals instead. I think this is fine.
I think on the topic of meta compute like you know there are the rumors that they're going to become a neocloud.
I think as long as they do these like SpaceX style deals the clawback or they do like or they just like give it all to Rexus or they like do token as a service like any monetization method that allows them to in theory claw it all back if MSL proves they're doing well.
I think that's like an excellent smart decision on men's part.
Probably allows them to be more aggressive with buying compute in the short term and it still like gives them a chance to build RSI in the long term.
I think the issue with Google and why they're in fifth place aside from their model like Gemini 35 Spy Flash being bad is that they have signed all these long-term compute deals where they cannot claw back the TPUs that they're giving to right that's just like committed for the long term and I think it indicates like a lack of conviction
on their part that they think they can build RSI and I think you know Nom Shazir John Jumper like all these really high-profile guys at deep mind leaving is sort of because of this fact like they are hi pill their leadership is not hgi pill and they think that the only way they can do the only they can do in this situation is to go work somewhere else. >> Wow. Okay. Hot takes. So the two things >> Wow. Okay. Hot takes.
So the two things I want to dig in on on there.
The first is related to the clawback.
So this is a valid strategy but you can basically only do the clawback once and then nobody's ever going to buy from you and trust you again.
Or do you think you can actually between these two?
Well, I think in an ideal world, you only need to do it once because you only do it if you're confident that like I have I'm going to be able to breach the frontier by continuing like scaling, right?
Uh but I also think it's possible you could do it more than once.
I think in principle, Enthropic is like viewing this as a short-term deal.
And I think if Elon said like, "Hey, actually, I want it all back six months from now and then a year from now he's like, JK, uh we actually still can't train a frontier model."
If enthropic is so ripping at that point and it's in desperate need of like computer capacity, why would they say no?
>> Yeah, I guess financial motivations can make friends out of enemies.
So second thing is just on the topic of the business maybe Joey I can bring you in here.
What's a better business?
Selling tokens of your frontier model or close to frontier model or selling compute access on the short term at a premium over the market?
Selling your own tokens is probably the best model.
I mean, even like even selling someone else's tokens if you're AWS or or Azure is a pretty good business model as well.
If you can participate in some of the the upside revenue share, but if you can run above, you know, X market rates, it's not a bad business.
I think before like in like from the MetaMP compute perspective, it's been floated for a while.
You met compute has been floated since like Q3 of last year.
And it gives Zuck a backs stop to say, okay, if MSL isn't successful, we have means to essentially make decent ROI on our capex spend.
It allows them now to spend a ton more on capex in 27, you know, potentially 28 because there's a market and, you know, you've gotten these market signals that Anthropic can pay, you know, 3x market rates and, you know, you can do a small deal here and that's in, you know, the background.
investors can feel comfortable that you know if um if MSL is is deemed you know unsuccessful that you know Meta could have a pretty nice you know single client or multi you know few client like Neocloud business just serving the labs. >> Yeah.
But I mean I I thought this chart that you guys had in the MetaMP compute article was striking where you can see that the SpaceX plus Google deal they're monetizing this compute at 4x the market rate.
So, I think in the past everything you're saying makes perfect sense, but I would challenge a little bit that like this near-term SpaceX GB300 rad deals aren't a really good business.
Like they're clearly high margin in this, right? >> Yeah.
I mean, that's extremely high margin.
margin. I think like when you think about the you think about like the AR per per like gigawatt I think a lot of people be happy to get 48 billion of like actual token AR like of inference there are so I don't know that deal
doesn't seem very economical for for Google but you know >> telling a you think their meta compute is them telling a story to the market that they're going to be able to replicate this space XAI thing which is probably a one-time thing unless demand
keeps going so exponential that there's such a compute crunch that everybody's fighting for access to the GPU And while we have them and we can monetize them at like 4x the market rate from Neoclouds where Neo clouds I think we already I think our house view is that Neocloud is
a pretty good business on its own at like 12 billion per gigawatt approaching 50 billion per gigawatts like >> yeah and like >> and the labs have gotten so good at the AR per gawatt or gigawatt like number has gone up a ton. So token throughput's
So token throughput's been pretty impressive like just you know the models as well and then two how they've monetized them you the new models price higher.
So like if you look at like anthropic like revenue per megawatt was like 16 million per megawatt last year or 16 you know billion per gigawatt and now you know that's that's more than doubled in in three quarters and I think there's some people that think that can double again in two to three quarters.
So >> let's look on that trend.
I've got this chart on screen right now.
So, uh, total token as a service market, I two things jumped out at me.
One, how fast this thing is growing.
Like, I thought the market was pretty big in Q4 of last year.
And that looks puny [laughter] compared to our forecast that you guys have on this chart towards the end of this year.
Um, like way more than doubling as you said.
Uh but the other thing that strikes out that jumps out to me is just how bad Microsoft foundry is doing here.
Can you >> Yeah, and that'll be revised up with the recent OpenAI and API success.
But I think up until because open you know they were you know foundry was 90% OI and you know OAI was so consumer focused and and API didn't really this is not the the latest version of this chart.
Um yeah Foundry is you know doing a little better now. Maybe Vert.
Ex is doing a little worse given, you know, Matt's or Max's thoughts on on Gemini >> than that chart.
But >> you got to subscribe for the most recent tokconomics forecast.
These companies always we can't be giving this away for, you know, for free, dude.
>> You got to you got to pay for the tokconomics.
>> We were nice enough to leave the axis on that chart this time.
>> But yeah, I think we're like we're of the view and like when we talk to enterprises, you know, token as a service is a very popular, you know, choice for consuming tokens.
You know, AWS and and Azure have massive businesses with Fortune 500, global 2000 companies that, you know, spend nine figures plus, you know, on on cloud a year.
And if I can go to an existing vendor and have more model optionality, increase my ELA or my my credits um and then burn that down.
It's it's pretty pretty attractive um for both AWS and Azure.
And I think like you know if it was 5% indirect like sources like you know this token as a service for like anthropic 6 months ago like that's probably it's probably 20% now of your business.
>> Crystal what's your take on people using the big three hyperscalers token as a service stuff versus the smaller startups like a together base 10 fireworks anything they can get on open router.
I feel like as we see more of the bigger enterprises, so like I feel like a lot of the financial services industry still hasn't unlocked like and use AI to the level that they probably can and should be.
And that's where like Azure and Bedrock are going to be benefiting from because most of them probably already like buy their services, right?
And then once they do get like internal approval or whatever to be using these AI tools, like they'll just buy it through whatever existing channel they'll have already, whereas it's like probably going to be like that's where the bulk of the market pretty much is. >> Makes sense. Yeah. Max, how about you?
What do you think of token as a service from like the startups?
how fast together Fireworks base 10 they're growing compared to the hyperscalers that are kind of tied to the big labs but maybe not actually selling a lot of tokens to the startups of tomorrow.
>> Yeah, I mean I think uh together base and fireworks these are all super impressive businesses.
I do actually expect like opensource token volumes to grow slower than frontier token volumes, but that's only because I think frontier token volumes going to like absolutely explode.
I think open source token volumes are also going to explode just to a slightly smaller degree.
And I think like it's kind of funny.
I feel like if you talk to the VCs investing in like together fireworks and base 10 and you ask them for the explanation why oftent times it lead with like infrance is going to be like the largest market ever.
And so it doesn't even matter if like these companies can only capture like a super super tiny slice of this like extremely large market.
That's good enough for us.
And I think that thesis is like honestly more or less right.
I don't expect these guys to ever be doing like more volume than an open eye or anthropic or like nowhere near that.
I I think like this is in terms of like global token volume.
They might be like the highest percentage they've ever been like they ever will be rather today.
But they'll still be like good businesses in the future, I think. >> Yeah.
How about the rumors that they can hitch their wagon to some of the bigger labs?
So like SpaceX AI starts to win a little bit and then Fireworks grows because they're exposed to Cursor and actually have some I don't know exposure to that growth. >> Yeah.
Well, I don't think that's going to like last long term.
I think the only reason why fireworks got that exposure is because originally composer was post train on chem, right?
I think if you're a frontier lab that has made like a new model from scratch, there's no reason to believe why fireworks engineers would be better at optimizing that model than your own engineers.
So I don't think they can get like any share of that margin in the future. >> Okay.
So why do the hyperscalers get a share of that margin with the big three?
I mean because they just have like and I'm sure Joe can speak more on this uh but they just have like the enterprise distribution that the neoclouds do like Enthropic is not giving up this margin to bedrock so that Amazon engineers can
like make claw run faster and you know B300s or whatever or >> it's just so >> yeah sure trium but uh it's just so like uh all the existing customers that are reliant on bedrock uh can also use cloud malls >> sense yeah >> yeah I think the deal. I mean there's we
>> yeah I think the deal. I mean there's we hear more we don't hear anything but people tell us you know they think the deal could get reerruck um because it's obviously you know it was very beneficial for AWS I think we've written on AWS and their margins and how they
monetize you know versus just selling the bare you know compute for anthrop to run inference on you know to get that 30% you know or 20 to 30% revenue share you know that just falls down to the bottom line is pretty attractive AWS like Azure lesser extent GCP like obviously have massive customer bases
people's cloud estates are there the enterprises are very comfortable security and compliance and you know everything now with cloud even your more regulated industries are so it's a natural place for them to want to buy but to see that much of the economics for anthropic is obviously yeah I think
a lot of that was because when the deal was was struck and kind of the incentive to make tranium work on that you know they were able to to run with those terms and make it pretty favorable So yeah, we'll see what happens if those revenue share deals can continue. >> Okay, let me make one more attempt at
>> Okay, let me make one more attempt at the bullc case for these token as a service companies that aren't the hyperscalers for AWS and Google.
Obviously, a big portion of the benefit of getting token as a service from them is that their engineers are are working on tranium and TPU and that might be different than the experience that the lab has with the GPUs where they have more experience like running this themselves.
There's a whole class of chip startups that are coming to market right now and an obvious way in which they come to market is by partnering with these token as a service companies where hey the the biggest Frontier Labs aren't going to spend a whole bunch of time optimizing for you know the seventh best chip startup that's coming to market right now.
But if they do and that chip startup strikes something really nice for a given model that makes it a lot more cost- effective or a lot higher performance to run instead of Nvidia, then they've got a shot at doing something super unique.
Do you think that's a, you know, potential future in terms of where these kernel engineers who have learned a lot end up going and spending their time over the next few months or years?
This is a good point and it's reasonable in the short term, but if any of these new chip startups actually reach like sufficient scale, I think the labs will just dedicate teams to making their model like run really good on their chip.
I think there's just I think there's no world in which you have a new accelerator that's like meaningfully better than Nvidia and it's also being like uh sold at large volume that OpenAI is like relying on together to serve their model like on that chip instead of just like working directly with that company to you know develop the first party capabilities to run their model on that chip. >> Yeah.
You think it's going to go the way of Cerebrus where the chip companies to be successful are effectively going to have to become a neo cloud for themselves and exactly relationship with the frontier labs. >> Okay, makes sense.
Maybe we can finish by talking a little bit about MSL with Meta particularly Max.
I just loved the like crash course on RL that this article turned into.
not necessarily how it works from a technical perspective, but how it works from a business perspective like where people buy data, how they build these environments, what sort what the market looks like.
Can you kind of give a an over like previously in the world where everything's pre-training, it's just like whoever has a, you know, frontier class team and the most compute can just train the biggest model and win.
They follow the scaling laws and go there.
But now there's a scaling law related to RL.
How how is this playing out in your mind at a high level? >> Yeah. Yeah.
I think it's I think a it's like really important for everyone to understand that reinforcement learning or RL is like probably the most important scaling law for improving model capabilities today.
And there are a lot of people who believe that the only thing that's stopping the models from being able to do like literally anything a human can do on a computer is having sufficient RL environments.
This is sort of like data that lets the model like try to complete white collar tasks itself until and it can sort of like repeatedly try cleaning the task until it fully learns how to solve it.
And there's like an entire new industry slash like supply chain of these RO environment startups whose like entire job is to convert real world economically viable tasks into these RL environments that can then sell to the labs and allows use it to improve their models.
Um yeah, you see like most of the main players um uh on this chart Jordan has pulled up here.
Uh I actually think like some of these AR numbers might be slightly understated.
Uh I think it's like pretty consensus that uh the total sort of data budgets at the frontier labs this year.
So this is like um primarily OpenAI, Anthropic, Google, Meta and XAI, but also the longtail companies like uh you know Meta uh sorry Amazon, Microsoft like Thinking Machines.
Um I think all those sum together is going to be like well over 10 billion.
Um that's sort of like a 10x relative to last year.
It's very possible with 10xes again um in 27.
uh and this is just like a hugely important market for improving uh AI capabilities.
Uh this is probably like maybe the only market in the world where in customer demand I guess maybe compute was the other one but like in customer demand is just not even a question for all these startups like they will never have a contract turned down from the labs because it's too expensive.
It's just a question of like can you scale up you know creating high quality data fast enough and if the answer is yes we will pay any price for it.
So yeah this is like I guess maybe as one other like side note one of the reasons why Enthropics models are the best at coding today um or at least like they definitely were before 56 soul came out is that they were by far the most aggressive from buying coding data from all these oral environment startups.
Um I think some of the other labs are starting to like realize this and catch on.
Uh but it is definitely like a super important industry everyone should be.
>> Can you dig into a little bit of the process to create some of these tasks?
you kind of ran through this in the article by describing how Meta has moved thousands of engineers doing this work and uh and also dispelled the notion that this work is like meaningless soul crushing stuff and is actually you know pretty both economically valuable and I don't want to steal your thunder but you had a nice line on that one.
I think I said it was both potentially more economically valuable and intellectually stimulating than like your average big tech job.
So like yeah, I think a lot of people think like they hear the the phrase AI data and they still think like oh we have some random people in the Philippines who are like drawing bounty boxes or like uh labeling text as like NSFW or something.
And that's just like like dude, the models have already fully solved that.
Uh like your data is only valuable if it's uh doing something that the models don't already know how to do.
And so what this means in the case of software engineering is that in order to make a good like software engineering task today, it typically needs to be something that would take like a really good human engineer, maybe like a full day of work to solve in order to be something that like the model can't already run on itself today.
And so if you want to create this data like essentially you need like a really good human engineer to sit down and like think of an example problem that he would actually want to do.
You then need him to like create a verifier like usually a set of integration tests maybe along with the rubric that can check if the model actually susly completed this like dayong task or not.
Obviously this is easier said than done.
is easier said than done. And then you also need this engineer to like write a prompt for the model that like asked it to do this task but also in it sort of like needs to fulfill these like two competing factors where it needs to be like 100% fair and like unamiguous what
you want the model to do because you can't sort of like incorrectly fail the model during training because it like successfully uh did what your prompt asked for but your prompt was just not specific enough and so like the model didn't know it actually had to do some extra thing. At the same time, like you
At the same time, like you need your prompt to be realistic and natural sounding and representative of something like a human would actually ask an AI to do in the real world.
And it's like very difficult to get this balance right.
We've heard for like the highest quality coding tasks like the labs are willing to pay well over five figures for a single task.
And you know, this is obviously already entering into the realm of like how much you would pay a decent engineer for a full week of work.
And so I think that should sort of just like dispel any myths about this being like easy, you know, mind-numbing work.
I to all the listeners out there if you're looking for a new job, like if you are really good at creating RL like tasks, you can make seven figures plus annually at this point.
So maybe maybe consider that as a new job option.
>> Yeah, that maybe we have some listeners excited.
Can you maybe actually say like one click lower?
Do you have personal experience with striking that balance between easy for the AI to do and not impossible for the AI to do?
Like what sort of intuition could you give to a listener about what that means?
>> Honestly, it's like it's always changing.
So, so back in the day, like I actually did sell some environment data to the labs myself. Don't do it anymore.
uh but it was like much easier to create data even just like uh eight months ago when I was still doing it than it is today.
I say like today it often looks like you need to identify a specific failure mode that like you are aware of in the model and then you need to create like RL tasks that specifically target that failure mode.
And really the only way to know if it's like at the right difficulty for the model is like you have the model try solving it 10 times and you see how many times you're successful.
Um you sort of just like repeat that iteration.
So, it's like I I don't know if I can really provide any like blanket advice on how to find the right difficulty other than if this is like your first time trying to do this.
Your first thought is almost certainly too easy, try making it 10 times harder, like try that and then maybe maybe you'll be like in the right level. >> Interesting. Cool stuff. >> Yeah.
>> Okay guys, we got a whirlwind tour.
anything you think has uh that I've missed as we've gone through token budgeting, anthropics, profit margins, meta compute, MSL, what uh what have we missed, Crystal, what have we missed? >> I don't know. Nothing I can think of.
>> Just been too busy watching the World Cup.
Here we got we got Wednesday, July 15th, right when we're recording this.
We just got to watch England get knocked out as Argentina stormed from behind for a nice 2-1 victory. That was crazy, dude.
There's like a bunch of British people in the office right now and they're just like all depressed downstairs. It was crazy. >> No. Joey, how about you?
What's uh what's what's on your mind as we wrap up here? >> Nothing.
Um token token spend is up and to the right right now. So, it's good.
Um you know, I'm I'm excited for hyperscaler earnings starting next week.
We've got Google, you know, Amazon and and Microsoft the week after.
It should be pretty good, you know, on the top line.
>> Okay, let me go around the horn.
We'll close by getting a vibe check from everybody.
Joey, what's your vibe like on the market?
>> On the market or like I like I don't like as long as Lab AR is going up like and at a good pace like it doesn't decelerate.
I I think like I think things are fine. >> Yeah.
>> And like right now it's going up.
Open AAI is catching up like vibes vibes should be good.
That's not what you're seeing in semis you know the last few days.
But that's just summer you know >> summer momentum. It doesn't work. So it'll come back.
People come back to the office from vacations and semis will rip in a year and you know things will be good.
>> Yeah, we're going to see an acceleration after the summer pop.
Joey's going to shoot 84 and see Lab ARR go up and be happy.
Okay, Crystal, how's your vibe?
>> He's going I have to say I think it's going good. Same thing Joey said.
I need to leave San Francisco before the AI bubble pops.
So, very quickly moving away from San Francisco before it's too late.
>> Crystal, you think it's a bubble?
Next, we should we should publicly talk about our bet here. >> Yeah. >> Yeah. >> Right now. >> Yeah. We have a the loser.
So, Max chose the overunder of Anthropic ARR at 400 billion to to end um >> 2027. 2027. >> Yes. 2027.
And we're at like, you know, we're in probably >> You can't give them the actual number right now, Joey.
You got to subscribe to the tokconomics model for the actual but you know. >> Okay.
But just does somebody have who took the over? Who took the under? >> I took the under. >> I took the over.
>> Jeremy and Joey actually both took the under.
But we haven't we haven't decided what like the actual bet's going to be for yet. >> No this year. We'll figure it out.
>> You have to Did you guys see our pot at Rays last week?
I had >> no >> I pulled out a Canadian 20.
Rake pulled out 200 Singapore dollars, also some rupees.
[laughter] We had Dylan pull out some euros.
Like, we had five different currencies going on the table.
So, I whatever you guys are, >> I will throw in Canadian currency and uh and take the over with Max because we are >> Let's go.
>> We're exponential extrapolators here. >> Yes. Yes. >> Exactly. Exactly.
>> The loser The loser has to write a newsletter post of why they were wrong. >> Okay, that's good. That's good.
>> No, that's what it was.
I didn't realize we agreed to this.
Like I I must have missed that in the in the Slackard.
>> Are we doing like Frontier Lab fantasy here?
We need to pick a model for a given week and set up our team.
>> We'll bet on a we'll bet on Jeremy's spend. >> Jeremy's on Jeremy.
>> We bet on Jeremy's weekly spend and we'll have like an overunder and >> overunder.
>> Get everyone involved.
It's like isn't meta building an internal like Poly Market or Khi?
Like we'll do that for semi analysis.
Someone can buy code like Poly Market.
>> Dude, Meta needs to shut down that effort right now, dude.
Like what are they doing?
[laughter] >> Well, we'll have a we'll have a market for betting on Jeremy's tokens. >> There we go. >> I like it, guys.
There's no way that market could be >> Yeah. >> No way. Totally fair.
>> Hey, Jeremy, I've got 6,000 rupees riding on you.
Need you to hammer the data center model dashboard this week, buddy. Okay.
Well, guys, I appreciate you uh you taking the time.
Hopefully the listeners enjoyed it.
Nice uh devolved a little bit at the end.
Let's uh let's all get back to work.
Keep tracking those tokens. >> Yeah.