Ep. 036 - $200 Buys $12,000 of Opus Tokens, We Bought Every Plan (Tokenomics)

0:05

Hello, everyone.

0:05

Welcome back to SemiAnalysis Weekly.

0:07

This week I'm joined by Max and Andrew, the authors of the very controversial but very interesting “Anthropic Subscriptions Offer Five Times or More Value Than OpenAI.

0:15

” This was an article that we put out a couple of days ago, where the guys did limit testing for every AI subscription plan.

0:26

It's not just OpenAI with Codex and Anthropic Claude Code.

0:29

It also includes Meta, SpaceXAI, MiniMax, Moonshot, Z.

0:31

ai, Cursor, Cognition, anybody where you could be getting an all-in plan from.

0:40

And we're going to talk about the learnings that they've had.

0:42

Guys, welcome to the show.

0:45

Thank you for having us, Jordan.

0:45

I feel like this is a Okay, so first TV interview or something.

0:50

This, like, the setup is too polished now. I'm not used to it. Well, yeah. Okay.

0:57

Other people need me to like direct them with questions, but for you I think we can let you run it.

1:03

Max, the plans are highly subsidized, no?

1:08

Well, okay, it depends on what you actually—it depends on how you want to sort of define subsidization.

1:15

Relative to API pricing, the plans are obviously a better deal.

1:20

I mean we can throw up some graphics on the screen here or maybe people can read the newsletter if they haven't, but some headline results including you being able to pay like $200 for a Claude plan and get like $12,000 of Opus 5. 5 tokens at API pricing.

1:37

So that's obviously a good deal.

1:39

However, in terms of whether or not the subscriptions are like subsidized in an absolute sense, like, do you get—does Anthropic have negative gross margins when they're selling a subscription plan?

1:51

I think the answer is no.

1:51

The reason for this is that A, their margins at API pricing are incredible.

1:59

I think consensus is that it's like over 90% blended at this point.

2:02

And then the second thing is that most people don't use like 100% utilization.

2:08

And so you combine like a low sort of realistic average, which I think is around like say 10%, with this, you know, 90% plus margin at API pricing, you end up with subscriptions having like call it 50 to 60% like blended gross margins for Anthropic.

2:24

So that's obviously still actually a pretty good business for them, but compared to their like 90% plus margin API business, then subscriptions are still [unclear].

2:34

So maybe the takeaway is not that the subscription plans are subsidized, but that the API pricing for the current frontier models is like, you know, punitive, exploitative, price gouging, whatever word you want to make it. Also known as business. It's quite expensive.

2:46

They're selling to willing buyers at the fair market price, Jordan. This is capitalism. Yeah. Okay. Yeah. Understood. Okay.

3:01

So talk to me about the 100% utilization and how you actually did the testing to realize these like numbers, these results.

3:10

That's a that's an Andrew question.

3:12

Wait for the Yeah, Andrew.

3:15

Yeah, so I guess to figure out like the results for the subscription plans were pretty much just like we bought every single plan and ended up measuring every single token type.

3:26

And from that we're able to say what kind of workload we are running, in this case, generally agentic.

3:33

That's pretty much why everyone's using these subscription plans for.

3:37

And they'll tend to be around like 96-ish percent cache reads, a little bit of inputs and a little bit of cache write, around two point something percent and a little bit of output.

3:48

You blend that together, and then you can get the big numbers that we figured out.

3:52

Yeah, sorry, before you dive more into the details, I do want to emphasize that there have been people on Twitter before who did similar analysis, and they would essentially like just present one dollar value as like the equivalent, you know, API value of a subscription.

4:08

And that is sort of just a fundamentally oversimplified, incorrect analysis.

4:13

Because the way to think about your subscription plan is that by giving the company you know $200 a month whatever, they allocate you some number of like credits.

4:23

This is just some other like you know currency.

4:27

And you should think of each model token type combo as sort of costing a different number of credits and essentially drawing from your like allotted usage a different amount.

4:38

And then because sort of the ratio between different model token type like credit costs can be different from their API price ratio, the actual like API-equivalent value of your subscription plan can change dramatically depending on what model you're using and what workload you're running.

4:58

And what this means is that like you can't just make a blanket statement like, the two hundred dollar Anthropic plan gives you ten thousand dollars of value.

5:09

You have to say, like, the $200 Anthropic plan, when running a particular workload on a particular model gives you like some dollar amount.

5:17

And this is sort of the the nuance that we capture in our newsletter and something I really want to make sure like people understand. Yeah, makes sense.

5:23

Can you go even back to basics even more and just describe the different token types as you're referring to token types?

5:31

I'm not sure everybody might understand what that means.

5:34

Yeah, so when you're when you're in your Claude Code CLI or your Codex app, there's different token types that are being used there as you chat with the model.

5:43

The first one is the input tokens.

5:45

Input tokens are uncached, they're just fresh tokens that you send to the model.

5:52

When you're actually doing like an agentic multi-turn conversation, most of your input tokens are not from your direct chat, but are from things like web search, where they fetch something and then analyze it in like a small subagent in one shot.

6:07

In that case, it doesn't need caching.

6:09

And so that's where it won't use caching, and that'll be like a fresh input token cost.

6:13

The second one is a cache write.

6:21

Rather than reprocessing your entire prompt every single time when the when you call like the LLM API server, you can save the previous results into a cache and reuse them later for faster inference and also for cheap costs.

6:40

And so cache write tokens are previous tokens that you write to the cache and they'll they're billed once when they're written and then once you read them back, they're billed back as cache reads.

6:51

And the nice thing about cache reads is they're is that they're extremely cheap because they're you're essentially just serving from like the KV cache on the LLM server.

7:04

So yeah, so now you have inputs, cache writes, and cache reads.

7:07

Input is like generally, let's say, like 1x, a 1x price.

7:12

Cache writes are generally like 1. 25x of like inputs.

7:16

Cache reads are like probably around 0.

7:16

1x or less of input costs.

7:25

And then lastly you have output tokens.

7:26

So this is outputs is what the model actually generates.

7:29

This is generally the most expensive token type.

7:29

On Fable and Astra, these are like $50 per million tokens compared to like $10 on input.

7:38

Yeah, those are pretty much all of Yeah, and it it's pretty it's pretty straightforward to draw a line from like the token type Yeah.

7:44

To the actual processing that's happening on the accelerator that's producing the tokens.

7:48

In other words, reading KV cache from memory would be a cache read.

7:55

This is a lot cheaper in performance terms than outputting new tokens, which would be decode steps on the model, right? Yeah, correct.

8:03

And cache read is like probably the most important thing and like the way you do caching with LLMs, because, you know, LLMs are like autoregressive decoders.

8:13

If you don't do caching, you're like reprocessing the entire prompt each time, which is extremely time consuming and expensive.

8:20

And because most of your chat session is actually cache reads, because again you're reprocessing the entire thing every single time.

8:29

So if it's cached, all of that is essentially cache reads.

8:33

You want this cache read price to be as low as possible.

8:35

Because it almost dominates the price of your chat session. Okay.

8:35

So can you talk to me about the like the technical way Yeah, Yeah, well this is this is teasing our I just wanna say go back there.

8:35

This this is like teasing our our EndpointX testing a little bit.

8:48

But really, having like a high cache hit percentage is what separates like the good inference providers from the dogshit ones.

8:56

And this is like actually surprisingly, you know, difficult problem to solve, but it makes a huge difference in terms of the actual cost to like run the same workload, per sort of everything Andrew was saying about the different token prices.

9:09

Yeah, and even OpenRouter recently added like if you go to OpenRouter on any any given model, if you scroll down past the top level, you know, prices per provider, they have a nice effective pricing, which showcases like what the effective cost, blended token costs are, assuming like certain cache hit rates.

9:32

Of like the providers, which is pretty neat. Yeah. Okay.

9:35

So explain why cache hit rate or CHR is like a really important metric for people who work on inference serving today and why it's actually hard.

9:49

Like we've been saying, cache—cache is like conversation history, right?

9:52

So why is it hard to route a new request to where that conversation history is stored or to tier it off to lower performance but cheaper storage?

10:05

Yeah, if you're running a a real sort of production inference workload where you have like multiple different racks, multiple different deployments of like the same model on each of these racks, and then each one is processing like a batch that has tons of different user requests.

10:19

Essentially, like when I submit the like say the first turn of my user request, you have to be smart enough to like write all those tokens into the cache somewhere.

10:30

You know, it probably starts off and like staying in HBM if you're pretty confident I'm not going to sort of spit a follow-up that quickly.

10:37

But then if you know there's a large gap, you probably want to sort of offload it to a lower, cheaper memory tier is already some complexity there.

10:44

And then when I come back with sort of like the next turn of my request, obviously you as the inference provider have processed, I don't know, if you're big, maybe thousands, tens of thousands, like millions of other requests in that time delay.

10:56

But you have to like take my new request, probably send it to like the same deployment of like, you know, the model that you have, sort of find my previous tokens that you hopefully wrote down to cache somewhere, and then sort of route those to the same deployment and then sort of like reprocess the next turn.

11:14

And I think this this process generally of like sort of having to find the previous cache, like routing like the the subsequent turns of the same request to the same deployment.

11:28

Sort of balancing everything properly.

11:30

So you know none of your GPUs are like being overutilized and you don't have a ton of GPUs that are just like sitting there idle, burning money.

11:36

These are all very complicated systems engineering. Yeah. Yeah.

11:40

And there's also this TTL dynamic too, right?

11:44

Where like every new request is going to be tagged with a time to live for so it's going to sit in HBM, the highest tier, maybe for five minutes, and then the next one maybe for an hour or something.

11:53

But it's a punishment to the user if you let your old conversation die, so that it's no longer cached and now you want to keep it going.

12:04

You got to pay a whole fresh input cost or cache write at 10x the price or something to be able to rehydrate it.

12:12

So it's hard to Oftentimes more than 10x at this point. It's kind of crazy. Right.

12:12

Yeah, actually one of the one of the things like one thing that we called out, I think originally to our Tokenomics Model subscribers a few months ago at this point is when you looked at like you the original Anthropic API price ratios, it was like if cache reads were one, then input tokens were like 10, and then output tokens were like 50.

12:37

It was this like one, ten, and then five more ratio.

12:40

And one of the things we've like noticed recently.

12:43

With the price cuts that OpenAI and Anthropic have done, is they've been overwhelmingly concentrated on making the cache write the cache reads cheaper.

12:50

And why is this the case?

12:53

Well, like one, as we've established, this is like the most important lever to pull to actually decrease overall effective pricing.

12:58

But then two, and this is the thing that we pointed out to our Tokenomics Model subscribers a few months ago, is that cache reads started off at as like most likely by far just the highest margin token type amongst those three.

13:12

Like it's very possible that following this original price ratio, Anthropic was ripping like 99% margins on their cache reads.

13:21

And then now we're finally sort of seeing those prices get competed down again, you know, thanks to just competitive pressures.

13:28

That's one thing I wanted to point out.

13:30

Unfortunately, I don't think this has reached OpenAI yet because Astra is still $1 for cache reads as compared to Fable at 25 cents. Yeah, yeah, yeah.

13:38

Which is actually a big reason why you'll see on the plans.

13:42

Astra might have like higher dollar value on the Pro 200 plan versus Fable 5. 1. But actually Fable 5.

13:48

1 has more tokens and more usage because its cache write—its cache reads are 25 cents per million tokens versus $1 on Astra.

13:59

And that's besides the whole point that, you know, Fable 5.

14:04

1 is only 50% of your plan and you have another 50% to use. Okay, yeah.

14:08

So so it's take me through the exact like OpenAI versus Anthropic on the frontier model comparison in these plans.

14:16

Like this is the headline chart I think that a lot of people who have seen this article have seen, which is like, how can we compare Astra to Fable 5.

14:23

1 straight up on their different plans that they offer in terms of how much value you're getting on a per token basis?

14:37

For using these models, which I mean, the choice of model that you prefer is it seems like a very personal choice at this point.

14:44

It's not clear to me that anybody would say, like it's certainly not for my usage.

14:51

It's obvious to me that Astra is better than Fable or Fable is better than Astra.

14:55

They seem to just kind of trade off with different tasks that I'm trying to do.

14:58

But I yeah, tell me about like the actual usage.

15:04

Yeah, so funny enough, I think this was not the chart that was getting traction on Twitter.

15:10

We can talk about that one next.

15:12

This one I think people weren't complaining about because it shows OpenAI and Anthropic as like roughly comparable for you know the same, you know, two hundred or a hundred dollars a month or whatever.

15:21

The one thing to emphasize though, which people missed and which Andrew talked about earlier, is that you know, originally people were claiming that our other chart was like unfair. To GPT-6.

15:32

1 Sol because at API prices it is cheaper than Opus.

15:37

Well, for this like real headline chart where you're comparing Astra versus Fable, Fable is like effectively 40% cheaper at API prices than Astra or something, because their cache reads are 75% cheaper, and that's like 95% plus of your total tokens.

15:52

And so if you were to look at it from a token percentage, or for just a token point of view rather, you get like 1.

15:57

9 billion tokens of Astra on the $200-a-month plan, but you get like 3.

16:02

1 billion tokens of Fable on the same $200 month plan.

16:08

And not only that, but that 3.

16:08

1 billion only counts for 50% of your limit, not the full 100%, because Fable was restricted to only using 50% of your Anthropic limit.

16:17

So you can actually use like another million billions of tokens of Opus on top of that for the same $200.

16:23

So I think even this chart makes it clear that Anthropic offers a better value.

16:29

If we were to switch to the Sol versus Opus chart, I think it's an even more brutal Mog by Anthropic.

16:37

Maybe Andrew can go for this one.

16:40

Yeah, so this is a really controversial chart on Twitter where or X where everyone was kind of saying, this is so unfair because you know, 6.

16:51

1 Sol is so much more token efficient than Opus 5.

16:51

5, and everyone really liked to bring up like the Artificial Analysis Opus 5.

16:59

5 Max cost per task, and like the 6.

17:04

1 Sol Max cost per task, and it was like some like five dollars or something, and then 6. 1 Sol is like 72 cents.

17:13

And so people kind of use that to justify that this isn't a fair comparison.

17:18

There's like a few reasons why this is problematic, and like even if you were to assume that like you know, Sol is really that much more token efficient.

17:28

It's still—and if you do the math, like it still ends up being that like Opus 5.

17:30

5 is like two to three times like the value in terms of like tokens.

17:36

But yeah, the exact math here is so one, we don't agree that the Artificial Analysis Intelligence Index, yeah, is necessarily a good way to measure token efficiency.

17:47

But we'll we'll get to that point later.

17:48

Even if you assume that like sort of using these contrived, often broken, often saturated like benchmark tasks.

17:57

Are a good way to measure token efficiency. Opus 5.

17:59

5 on max is like the number one model on the index, and 6. 1 Sol is not.

18:05

And if you just use like Opus 5.

18:05

5 on medium, that outperforms like all the 6.

18:10

1 Sol results except for like Max.

18:16

So I think specifically 6.

18:16

1 Sol on xhigh and Opus 5.

18:16

5 on Medium score the same on the AA Intelligence Index.

18:24

And then if you look at sort of the cost per task for those two specific runs, it's like a 2 to 3x diff or something.

18:31

And then you're getting 5x more, you know, API-equivalent value according to this chart.

18:36

And so really you're still getting like 2x more value on Opus than Sol still.

18:42

And this is even if you sort of assume that this like broken benchmark index is a good proxy for token efficiency.

18:49

Yeah, the unfortunate part though is that like Opus 5.

18:52

5 is not really in the 6. 1 Sol class.

18:56

It's actually competing directly with Astra.

18:58

And like I don't know if I fully agree with this.

19:01

I think many people this is Anthropic fanboy Andrew's point of view. All right. This might be my point.

19:06

This this might be my point of view. Ha ha ha.

19:06

But I would say like Opus 5.

19:06

5, even on the you know, the broken benchmarks, is still like number one in the intelligence index.

19:14

A lot of people have been substituting like Astra versus Opus 5. 5.

19:19

Anthropic's been saying, you know, it's as good as, if not better than, Fable.

19:23

And so, like in that scenario, like you have Fable 5.

19:25

1 for, let's say, that's the Astra competitor.

19:29

You have that for half your plan.

19:31

And that's essentially equivalent to the $200 plan already on ChatGPT.

19:34

But then you also have 5.

19:34

5, which is a beast for the rest of your plan at like six thousand or five thousand dollars worth of value.

19:42

And when and when you do that comparison, it's like not even a question in terms of who's providing more value versus the ChatGPT, the ChatGPT versus Anthropic subscriptions. Because Opus 5.

19:51

5 is just like a step above 6.

19:51

1 Sol in most categories.

19:57

Yeah, for the audience, Andrew has a repo here at SemiAnalysis where he doesn't let—he has the AGENTS.

20:02

md that, like, you need to say which model was used to write this PR. And if it's not Opus 5. 5 or Fable 5.

20:07

1, he auto-rejects your PR. Yeah. Okay.

20:19

So this is kind of what I said earlier, which is that this is a very Mm-hmm.

20:25

This seems like a very personal decision at this point of which model you're going with and it's not scientific.

20:28

Well Andrew's actually enforcing this on everyone, right?

20:30

He's making all the InferenceX Astra fanboys switch to Opus 5.

20:33

5 if they want to contribute to his repo.

20:38

Hey, and they saw the light.

20:38

Everyone who everyone who started doing it, they converted Mm.

20:41

True, that's true, that's true. Yeah, so that works. They saw the light. Okay. Okay.

20:51

Can you talk to me a little bit more about the the methodology?

20:55

So like how did you go about actually measuring like per token type and how the limits fill up?

21:09

'Cause I I think this is also a way in which people get confused, Which is that like they launch one really big prompt and then their, you know, hourly rate limit or five hour limit or whatever it is just is gone.

21:25

And you know, you made the point at the beginning, Max, that like nobody is using these things one hundred percent of the time.

21:32

But what's the scenario that you guys actually pursued to use this thing to figure out how many tokens you could consume in these windows?

21:41

Yeah, so the rather than, you know, launching some task or benchmark and like seeing how many tokens it uses and you know the split for that.

21:50

We really wanted to like do like a kind of a scientific experiment to like extract like the exact value for the exact rate for input, cache write, cache read, and output token prices.

22:01

And so the way we do that is essentially kind of mocking out the way that all the popular CLIs call like the model APIs, essentially to mimic exactly how you would hit the inference server as a normal subscription user.

22:20

And then we use like essentially two types of prompts that let you maximize the token type that you care about and like minimize the token types that you don't care about.

22:31

So the for the inputs, cache writes and cache reads, we use like a small tag at the initial start of the prompt and then a whole block of War and Peace that just takes up a ton of tokens.

22:45

And then ask like a very, very simple question to the LLM to minimize the number of output tokens that it generates from the request.

22:52

Usually like four tokens, so just saying like history or fiction or something.

22:57

And like the way that the the reason this works is like let's say for input tokens, we can set a cache header to mark something for caching or not.

23:10

And we can also randomize the tag in front of the initial in in front of the prompt, which essentially guarantees that nothing will get cached and also that the API is not really trying to cache it because there's no cache header on there.

23:24

And this is present in Anthropic and OpenAI's APIs and other models.

23:29

And then we read back like the requests response to verify which token types we've actually consumed so that we know that yes, like this was all like for example.

23:40

For cache writes, it's like a similar story to input, except we do mark it for caching.

23:46

And so we see that the whole prompt gets cached besides the last maybe like model generated answer, which is like negligible tokens.

23:53

And then for cache read, it's a similar experiment as well.

23:57

And for that one, we just set the cache header.

23:59

One call is gonna be a cache write, and we just drop that from our accounting and then just keep sending the same request, and that gives us like cache reads on like a per request basis.

24:13

And then finally output tokens is just like write the super super super long like technical essay and then just have the models like rip out sixteen thousand output tokens.

24:22

And that kind of consumes that budget.

24:28

And we can like look back at the response and the meters and be like, okay, that was like two percent of our weekly limit right there or five percent of our weekly limits.

24:35

Output tokens are expensive.

24:37

They probably use like 5– 10% per one of these calls.

24:39

But yeah, by just like repeating these and just kind of graphing it out, like you can figure you can figure out like you can do the math and isolate exactly how much each of the token types costs.

24:52

And then we can use that to do any sort of analysis around like, okay, if you're running like an agentic workload, like under this token split, like this would be your like subscription value.

25:01

And then under like some chat some chat workflow, then this would be your subscription value, et cetera, et cetera.

25:08

Talk about how you get those ratios as well.

25:10

Like we're keeping good track of what a chat workload versus an agentic coding workload actually looks like, right?

25:17

In terms of ratios of input to output to cache read to cache write.

25:22

Yeah, the agentic one actually comes from our tokenomics dashboard, which has SemiAnalysis' own internal usage for like the month of—well, all the months, and then we just use like the month of September, one of the most recent.

25:34

And that's that was our own token split because all of what we do on the like enterprise accounts is just agentic work.

25:47

For chat, it's a little more difficult because the only way to really get a good token split for that is like the ChatGPT and have like the realistic workload.

25:55

But you can kind of estimate it based on the fact that you know there's a system prompt on every single chat message, or like chat thread.

26:02

And that system prompt is essentially cached for all users of Claude or ChatGPT or whatever.

26:11

And so that alone is gonna be like a big percentage of cache reads, and so that's how we kind of that's where we start.

26:17

And then you assume users are sending like shortish prompts.

26:21

Maybe they step out and they don't respond to the thread ever, or they come back and step in once or twice.

26:29

And so like under those assumptions, you kind of get like the split that we have in the article, which I believe is like is it like five percent ish input, some 20-ish percent cache writes.

26:42

Yeah, 13% cache writes, 75% cached inputs.

26:47

And then ten percent outputs. Makes sense makes sense.

26:51

Okay, how about some of the other coding plans?

26:57

Like one of the pieces in the subtitle or or header is that there's other all-in plans that aren't from Anthropic or OpenAI.

27:07

First of all, does this actually matter to anybody?

27:11

But yeah, let's assume it does. Ha ha ha.

27:16

What do you guys think of like Muse from Meta, MiniMax, the SuperGrok Heavy, Cursor Composer, GLM and Kimi and how they're actually measuring value compared to some of the other guys.

27:30

Yeah, so I actually kinda surprised that a lot of the Chinese guys, so Moonshot, Z.

27:34

ai and MiniMax, offered pretty good like API-equivalent value per dollar.

27:43

The reason why this somewhat surprised me is that they're they're all like so compute poor.

27:49

And I think if you look at their API pricing, it's even lower margin than like the Frontier Labs because it's a more competitive business.

27:58

That's like more like a commodity.

28:02

But some of them, I think I think Moonshot was the worst at like roughly 7x the API-equivalent value or something.

28:09

But some of them were, you know, in the same kind of 12 to 15 range of OpenAI.

28:15

And considering that OpenAI is sort of the exact opposite of compute poor and you know cheap API pricing, at least for Astra, that was a pretty surprise takeaway for But it's it's not maybe the takeaway for me on this was that it's not obvious that when you look at one of these charts,

28:32

that any of these plans are just like so much more compelling on a dollar-value basis that you would be like willing to switch It'd be like from OpenAI or Claude because like okay, I I get, you know, so much better value than on my Muse subscription. That I'm willing to take a slightly worse model because I just get like

28:55

That I'm willing to take a slightly worse model because I just get like get so many more tokens.

28:58

Like maybe it's two times or three times the open AI or Anthropic, you know, upper limit.

29:04

Or maybe it it's it's nice that it starts at a fifty dollar fee instead of a two hundred or five hundred dollar fee to like really get access to their top model, but nobody's doing a price war on the monthly subscriptions at this point, from what we can tell, right? Yeah, that's true.

29:21

And I think a big reason for that is nobody can touch Anthropic's margins at API prices.

29:28

And so if they offer a similar API-equivalent value per dollar, it probably means that their actual margins are like way worse than Anthropic.

29:37

And so this it really just underscores that in in the AI age, like being the person who actually owns, you know, the model everyone wants to use is a huge advantage.

29:50

Being a wrapper is especially bad.

29:52

We covered this at the end of the article.

29:54

But if you look at using like Anthropic models on Cursor, Cognition, the the limits are just like way, way lower than at first party.

30:05

And obviously, this is like, you know, when you if you just think about it for two seconds, like this obviously has to be true because Anthropic, you know, has 90% plus gross margins, and then Cursor and slash Cognition probably at best have like a 20% enterprise discount or something. Compared to API prices.

30:20

But it's like it was pretty cool to see this kind of claim validated by the data. Yeah, yeah, makes sense.

30:28

How about I do wanna small small shout out to Meta for being one of the highest value per dollar plans on like a fifty dollar subscription with Meta, you can get like two thousand five hundred dollars worth of like Muse Spark, and that's like, I believe it's like 140 or two hundred sixty million tokens per dollar or thirteen billion tokens a month, which is quite a lot.

30:55

Like if they keep the same rate and Muse Spark is like A little bit better.

31:04

Yeah, man, for just for that fifty dollar plan, that's like two hundred and fifty, maybe even like three hundred tennis reservations in San Francisco that you can have Muse Spark getting for you.

31:19

Tennis reservations per dollar? I'm not sure, man. I think so. Yeah, yeah.

31:26

Okay, let me let me quiz you about this one.

31:30

So we we've put out a chart similar to this before, Max, which was I thought was super interesting the first time, and it's like even more interesting this time.

31:38

We we have no visibility into what the utilization metric actually is on these plans.

31:45

Like There's obviously some loudmouth power users on X that are using 100% of all seven plans that they've got maxed out.

31:53

And there's other people that are a little more reasonable, maybe like myself, who goes and records podcasts for half the day and doesn't use his coding agent or it works in the background and stalls, and then I gotta go, you know, prod it to go do more. I take naps, I sleep.

32:06

You know, what do you think? Did you say naps?

32:11

Okay, I don't take naps, but Whatever. Yeah you have a wife. Jordan has a kid. Yeah. Fair enough. Yeah. My daughter takes naps. Yeah. Yeah.

32:26

So okay, what what do you think?

32:30

Like if you had to speculate, what do you think the gross margin percentage that these guys are on the coding plan actually is for the given models?

32:40

Like like where would you put these guys?

32:41

Is it the average user is using 10%, 100%, somewhere in the middle?

32:46

I think for Anthropic it's like it can't be much higher than ten.

32:50

Reason for this is that if you if you really squint at all of their leaked financials, it just does not tie out if subscriptions are like extremely negative margin business.

33:03

I think it basically needs to be like a decent, maybe call it somewhere between like 40% to like 7% margin business in order for just their leaked financials to all tie out.

33:15

We walk through all the details in our Tokenomics Model.

33:17

But that that's why I think the average utilization can't be much higher than like ten percent.

33:21

And I think it makes sense if you consider the fact that a lot of the subscribers are actually enterprise and slash on like business plans.

33:30

And I think if you're a if you're a business that's like buying two hundred dollar a month plans for all your employees and those people are only spending like eight hundred dollars a month on credits or something, thousand dollars a month on credits, like you're still you're still kinda happy with that deal and that corresponds to like Sub ten percent utilization. Fair enough.

33:50

In your discussions with people who have these plans, what's their initial reaction to you saying this?

33:57

Like I think initially anybody who's tried to pay per token who has a way to consume the the all in plans is kind of like I'd I'd never pay for tokens willingly because the all in plans are are a good deal.

34:10

But what's the what's the reaction, roughly speaking, when you've talked to people about the article?

34:17

Like they kind of agree, they disagree, like Do you think people are are around that ten percent, twenty percent utilization number or something like that?

34:27

Well, the guys on Twitter who are loud about this definitely aren't.

34:30

These are like the a hundred percent guys, you know, with five accounts recreating Minecraft every three days. Something like that. Yeah.

34:40

But I think their their like vibe of the plans kind of matches our raw data.

34:45

Or a lot of people have been saying for the past couple weeks that like Opus 5.

34:48

5 on a Claude $200-a-month plan in particular is like the most bang for buck thing you can get.

34:56

And also it's actually important to note that the $50 $200 a month the 50% cut that OpenAI made to the $200 a month plans limits that they announced at DevDay hasn't actually come into effect for anyone who purchased the plan before that announcement because they got sort of like a one-month grace period from OpenAI.

35:17

And so we were actually quite puzzled why Twitter was not like roasting OpenAI more.

35:23

For just dropping this 50% limit reduction and like you know, giving you a bunch of cope about how, like, because GPT-7 Luna is actually gonna be smarter than GPT-6 Astra, this is a win for you guys.

35:34

Like that was kind of dumb.

35:37

But no one like called them out about it.

35:39

I think that was maybe largely just because like A, they timed it well with a bunch of positive DevDay PR, and then B, they gave you this one month grace period.

35:49

And then perhaps like a month from now when the limits actually like go into effect for the majority of users, we'll see some more noise about it. Yeah.

35:59

That was good timing by them to announce the cut at like nine PM the day before DevDay and then just drown it out with like twenty awesome announcements the next day with all this other stuff. Yeah.

36:09

Dude, just imagine if Anthropic announced this.

36:14

Like they they would have been roasted alive.

36:16

They wouldn't have made it. Yeah. Makes sense.

36:19

Okay, what what haven't we covered here that we should be talking about?

36:31

That's basically everything in terms of the subscription limits.

36:33

I guess maybe maybe one thing to emphasize is that this newsletter is very much a point in time.

36:37

And as like new models come out, new subscription tiers come out, new things like ultra fast mode come out, or just you know, the providers decide like, hey, actually I want 8% margins on my subscriptions too, not just 50.

36:51

We expect all these numbers to change pretty dramatically.

36:54

And we will be sort of tracking that continuously and then you know reporting the updates as they happen.

37:02

And keeping in touch for model Yeah.

37:02

Quality and token efficiency, right?

37:06

Which is gonna gonna come into play in the Yeah. Future, I'm sure.

37:11

Dude, it's so hard to actually measure token efficiency well.

37:14

Like I wish there was a way to do it. Yeah.

37:17

'Cause you just you can't rely on these benchmark tasks too, but then without benchmark tasks it's like how do you actually verify the two models were, you know, similar quality? I don't know, man.

37:25

Yeah, I have that's Give two people at SemiAnalysis the exact same assignment, check back a week later, Yeah.

37:28

Measure the quality of that assignment and see how many tokens they used.

37:36

I think there's a huge, you know, confounding variable there, which is the person doing the chats. What do you think?

37:43

Yeah, what do you what do you think is has a bigger impact on token efficiency?

37:48

Same person using two different models, Astra versus Fable or something, or two different people using the same model to try to accomplish?

37:55

Andrew had strong opinions on this one.

37:59

I think I I came in, he looked at the SemiAnalysis token-spend leaderboard, and he has some strong opinions on this one. So if you Yeah.

38:10

Look at the leaderboard, people who use the top people on Astra and Fable are spending about similar amounts, but the Fable people are using more tokens.

38:28

This doesn't really tell us much about like how much they're getting done though, which is the difficult part of answering this question.

38:33

I'd say from like personal experience trying out like different models, it's you can definitely see token efficiency on OpenAI side.

38:43

But it's like what I like to think about as short term token efficiency in the sense that yes, I'll get like this immediate task done, but I feel like there's so many follow ups that I would have liked to just include the Opus or like Fable would have just done, and like 6.

38:57

1 Sol or Astra might have just not done it for the sake of completing the task faster.

39:03

And so like that sense, like in the long term, I feel like the Anthropic models builds you like the better like codebase and like a definitely more readable one.

39:14

If you read some of these like Astra tests, it's it's something else.

39:16

But and so that's like my opinion.

39:20

It's just like I think models matter quite a bit for token efficiency, over like which user is actually using them.

39:32

What about Andrew, what about the resellers of these things?

39:38

Like Cursor's whole business, Cognition, Perplexity.

39:41

Where do you where do you think about their influence on which model you might use?

39:51

Like is the concept of a router just still never showing up for you where you have to switch models or you you allow Perplexity Computer or something to use a given model for a given task because it thinks that's the model to go with and it has different taste than you do. Yeah.

40:15

I think people, I mean I don't use Cursor or Devin myself, but if I were, I probably would just select like my favorite model anyway.

40:25

I know, like, you know, Devin has like SWE-2 and Fusion, which is supposed to use like a cheaper model, alongside like a larger, better model to keep prices low.

40:37

And like this looks great on benchmarks.

40:40

I haven't personally tested it myself.

40:44

In the real world, but my general vibe is that if I can and like my company like is willing to pay, I would definitely just use the like an Opus cheaper model rather than like a SWE-2 model, for example. Or even like a Sonnet 5.

41:00

5 cheaper model, because I just like for me like I just trust those models to like do the right thing all the time.

41:06

Where I might not do that for like another like unknown model to me.

41:12

I think that like as you work with these models you kinda develop like what each model's good at and like where they might bite you.

41:19

And like that's actually relatively important when you're designing systems using Yeah.

41:24

How about the rise of open models?

41:26

Like I saw some claims Mm-hmm.

41:26

By David Friedberg on the All-In Podcast that there was a flip from Okay.

41:26

Open to closed—or closed to open models—from 80/20 to 20/80 in the last twelve weeks.

41:39

And you know, we wrote a whole article for our Tokenomics subscribers about how this is like not true because they're just looking at router data.

41:47

But what I learned through there is that there is a pretty meaningful move.

41:53

At least across the routers and at least across token share, but maybe not implied revenue quite as much, that people are using open models more.

42:03

So do you have firsthand experience with this or talking to other people that would lead you to believe that something about the Frontier Labs, either how they're messing with their prices or the quality of the models as they release new ones, is just like not far enough ahead from the open-model ecosystem that's leading people to Try out open models and then get stuck using them over time.

42:26

Think it's a question of how far ahead the closed source models are.

42:30

I actually think it's just a question of how good of a model do I need for my task.

42:35

And it is true that like a lot of software engineering, white-collar work in general, does not need like Fable 5.

42:44

1, Astra six level intelligence.

42:52

And so I think a lot of businesses, especially the ones who are like lower gross margin businesses, are very rationally making the decision to start offloading some of the easier tasks to these increasingly capable open source models.

43:04

I think this is going to continue like happening in the future.

43:07

And so then the question of like are the Frontier Labs a durable business model, like really just comes down to your belief on whether or not sort of the space or TAM Of new things that you can do with increasingly smarter frontier intelligence will outgrow the portion of tasks that are being offloaded to these cheaper models.

43:30

And I actually think this is kind of the biggest sort of difference in priors between people who are bullish, like Frontier Labs, versus people who are super bullish open source.

43:43

It's like if you were to sort of magically spawn.

43:48

100 million superintelligent PhDs who are experts in every field who also don't need to eat or sleep or anything.

43:55

Do you think the economy would be able to absorb that like relatively quickly and still get high ROI?

43:59

Or do you think fundamentally like the world's just not that intelligence-bottlenecked?

44:04

I think like reasonable people have very different priors on this.

44:08

But if you you know believe that the answer is yes to that question, you're bullish OpenAI and Anthropic.

44:14

The answer's no, you probably think we're gonna have like a recession, you know, relatively soon and you're really bullish like together Yeah. Interesting, man.

44:22

I'm I'm motivated to ask the question because I've been using so much Perplexity Computer in Slack recently, in particular using GLM 5.

44:28

3 because of all the refusals on anything cybersecurity related from both OpenAI and Anthropic were making it impossible to do my work.

44:37

And so I needed to use GLM 5. 3.

44:42

And I think that led to me discovering that yeah, it's pretty good for a lot of stuff beyond the you know. Yeah.

44:45

It's basically like a [unclear] level model on the screen.

44:52

I I mean on the cyber benchmarks it was it was crushing it and I I literally can't test it.

44:57

It to me it's a frontier cybersecurity model because I literally can't use the others for myself. So Yeah. True.

45:05

You know it's it's yeah, whatever.

45:07

I I can only look at benchmark model cards.

45:10

I can't actually use the other stuff, even though I'm in the CVP program to get verified to do all the cyber work.

45:17

I still can't get past the guardrails.

45:19

I need some more creative.

45:19

Creative prompting to get past the guardrails trained into the model, not the classifier outside of the model.

45:25

But anyway, waste of time to spend time trying to convince the model that my grandma's at risk and, you know, I'm gonna bribe it with a peanut butter jelly sandwich if it doesn't get this thing done, just may as well switch to GLM and it'll it'll do the work.

45:39

But what happened is like I'm in that same chat history.

45:42

It's got all the data inside of it.

45:43

And I didn't bother switching to the other stuff to just make a PR on a dashboard repo to put out some data that we got from that that work. And like, whatever.

45:50

For probably 10x less the token consumption in in terms of pricing, Yeah.

45:58

It put up the dashboard PR. Hmm.

46:00

I didn't actually No problem. Need Fable 5.

46:00

1 for you know 10 lines of code.

46:07

Although a lot of things do not need Fable 5. 1, for sure.

46:13

Yeah, but I also feel like I'm using this stuff day and night.

46:16

My wife would tell you I am up late using this stuff day and night, and I'm not cracking the Mm-hmm.

46:20

Top ten on the leaderboard.

46:20

So I still have no idea what the hell you guys are actually doing with all these tokens.

46:26

We're we're like consistently not at the top ten anymore either.

46:29

So really you gotta you gotta interrogate like, you know, Jeremy and Kyle and Andrew Wagner as another guy.

46:36

Yep, Jeremy will be back on soon.

46:36

After I'm finished asking him questions about datacenters, then I'll be interrogating him about what this fast mode is actually doing for him and how many subagents he's—he told me he just takes—one time he did admit that he he he takes like arbitrary odd numbers and says that's how many subagents to launch for a given task, which does not need subagents.

46:56

So just feel like launch 13 subagents.

46:56

It's like he's like, give me thirty-seven subagents to investigate all the permits for this one datacenter site.

47:03

It's a point of pride to be at the top of that topic.

47:06

And then he's he's like, shit, that was two thousand dollars. Alright. It's research guys. It's research. Okay. Okay.

47:13

Don't question my methods.

47:13

You know, I do wanna add one extra thing about GLM 5.

47:20

3, which is like after talking to some very prominent hackers recently in or actually in the world, it seems like even 5.

47:30

6 Sol, if you had the cyber unlocked, is a much, much, much better model than GLM 5.

47:36

3 on real world cyber tasks.

47:42

And so while like some of these models might achieve the scores on the benchmark, I think like 5.

47:48

6 Cyber, Astra 6 Cyber, Mythos 5.

47:51

1 actually unlocked might be a lot more capable than a lot of people think compared to these— Yeah, there is a Which is little scary.

48:02

You you can see why It is it is a re It is a real possibility that the people who are inside these frontier labs and and get to use these models before anybody else and understand their capabilities better than anybody else might not be completely lying and pursuing a convoluted scheme for complete regulatory capture and are actually concerned about the cyber capabilities of the model.

48:25

Do you guys think that's that's a possible scenario that's that's playing out right now?

48:30

I think it's very possible.

48:30

The question is whether or not they're going for regulatory capture.

48:33

I think he's meaning No, I'm just giving you a sarcastic like the researchers are genuine, like when they're like kind of scared of like Yeah, I'm giving you a sarcastic thing to say.

48:37

Maybe the researchers who have a Sorry, sorry, I was—OpenAI just released a new repo called Math, where I think they just like probably made more progress in the field in the past like 10 years combined or something.

48:49

So I got distracted scrolling that.

48:52

So I didn't hear your question.

48:53

I think they might have just solved another Millennium Prize problem. I think that's wow.

48:53

A better response to my question than what I was sarcastically Ha ha.

48:53

Asking, which is like, yeah, maybe the people who see these breakthroughs, you know, months before us and are really concerned about what's gonna happen next are being honest and genuine about their concerns and not pursuing some convoluted regulatory capture scheme. That's all I was saying.

49:16

In in more words and less That that's a ridiculous take, Jordan.

49:21

It's obviously regulatory capture. My god.

49:26

Anyway guys, I think we've derailed it enough and we need to go read an OpenAI blog post about some crazy math things. So Yeah, dude.

49:37

So the repo on GitHub is openai/math.

49:41

That's like that's high order.

49:42

And just a bunch of Lean.

49:42

Yeah, and then, and then, and then We're getting serious.

49:45

Like ten problems that I've never heard of but are probably like really important.

49:50

I'm about to learn some math tonight.

49:51

Well, guys, thanks for joining.

49:52

We'll talk to you about subscription limits and token efficiency another time. Sounds good.

49:59

Thanks for having us, See you guys. Jordan. See you there.