0:00
Feel the music.
Feel the music.
[Music] We came to feel [Music] [Music] Let me let me freak. [Music] team deathmatch.
Delta team, you are clear.
[Music] You're watching TVN. Today is July 10th. It's Thursday, 2025.
We are live from the TVN Ultra Dome.
the temple of technology, the fortress of finance, the capital of capital.
Today we are covering Gro 4 launched.
Uh we're going to break that down.
Uh the third browser war has begun.
Every artificial intelligence company is getting in the game, launching new browsers. A browser.
Uh the new Volkswagen electric bus is a flop according to the Wall Street Journal. Ouch.
Apparently, Linda Yakarino was not fired for the Grock dust up with the crazy hallucinations that were going on.
We have more details there.
And well, and I don't why would anyone think that she was considering that in the timeline?
That was for sure in the timeline.
People were people were talking about that like, oh, like this happened and she stepped down within like six hours. Oh, really?
My my read my my read on it was clearly Grock and XAI are not her domain. Yeah. No, totally. Obviously.
and and uh they they were Grock was saying some things about her that should never be said. Yes.
And my read on it was you know who who knows when when when Elon commented on her post and said thank you for your contributions.
It's like the boilerplate text and so I'm I'm I'm sure that their relationship is maybe not as good as it was day one. Yeah.
But I I almost thought it was maybe the, you know, Mecca was the the straw that broke the camel's back and and uh she basically said, "Look, like, you know, I'm I can no longer, you know, bet my career on this platform." Maybe. Yeah.
I mean, we we can debate it.
Uh the there's more reporting in the Wall Street Journal about what actually happened and the that story we which we'll get into is is kind of pointing to this idea that uh that it had been in the works for a while. Yeah.
and and and that was not the straw that broke the camel's back.
That was like like the the papers had been signed.
The that everything had been signed before that and and then and then and then the the dust up happened with Grock.
But the bigger news is that they actually got Grock 4 out and people are excited about it.
So, we'll talk about that.
And then the other Grock, GRQ, the CEO of which of that company we had on the on the show. Uh was it yesterday?
I'm losing track of time. It was very recently. No, it was Tuesday. Tuesday.
Um apparently they're out raising at six billion.
and we have some more details on that company. So, that's interesting.
Um, anyway, uh, let's tell you about ramp. com.
Time is money, save both, easy use corporate cards, bill payments, and a whole lot more all in one place.
They have a new agent launch today will be joining later in the show for him to break it down.
Um, so let's break down the Gro 4 launch.
Uh, DD Doss has a summary.
Insane that Elon Musk has pulled it off again.
Absolutely crushing the AI wars with Gro 4.
Um, and we can go into some of the meta crushing the benchmark wars for sure.
And there's a question about like are we postbenchmark? Does this matter?
What's the real question to be asking here?
But there's a bunch of interesting takes.
So just summarizing the core announcements uh posttraining RL spend was equal to pre-training spend for this uh for this release.
That's the first time it's ever been like that.
I think when you go back to the original RLHF stuff that Chatt was doing that kind of unlocked like oh wow this really really works.
Um, I'm pretty sure the pre-training spend was an order of magnitude or two orders of magnitude bigger.
Now we are truly in this uh reinforcement learning regime.
Um, $3 per million input is uh tokens.
15 uh dollars per million output tokens.
Uh 256,000 token context window priced 2x beyond 128k.
It's number one on humanity's last exam, which interestingly was a effectively like post-graduate PhD level problems, but across a bunch of different domains.
So everything from literature to physics. Yeah.
Kind of like the hardest SAT possible.
Interestingly, I I believe that benchmark was created by Scalei and and so Alex Wang is now at Meta trying to figure out how can we beat our own exam and El's just like I'm number one at your thing. Interesting dynamic. Yeah.
the the real test would be uh Elon, you know, t doing the same problem set himself and saying, "Look, well, yeah.
I mean, I was talking to Tyler about this before the show, like, you know, it's like humanity's last exam.
It's like really good at PhD level math, PhD level stuff, but like how often are you running into those types of problems?" Yeah.
I mean, that I think that's the whole thing about there's there's this concept of like spiky intelligence, right?
Where it's like, okay, it's really good at this very obscure problem that I I never deal with.
Yeah, but if I have a super long kind of like context window like or there's no kind of um like long term it it just completely loses its footing and then it's like useless.
Yeah, we're kind of in like less of the benchmark regime and more of the agentic like how long can the agent run?
So it's like we're in the 15minute AGI regime.
Maybe this is 15 minutes of like even better AGI, but we want to go to 30 minutes on Monday that this, you know, takes me back to him talking about continual learning being the next problem that we really need to solve because it's great if you have a PhD level expert in your pocket that can solve any problem in any domain almost instantly.
But if it can't learn and take feedback and improve on certain tasks, then it's basically like useless.
like if you had a if you had a PhD level, you know, uh uh you know, a PhD join your team to work on a specific problem, but it it was hard restarting at the beginning of every single task with no prior knowledge.
It would the it would be almost impossible for that person to succeed.
So yeah, humans still got it on that front.
humans still got it on that front. But at the same time like you know if you are trying to just really establish yourself as you know a at least a an API for tokens that that every business should check out y against anthropic or the the open AI
APIs just saying hey you know we're on the frontier or Gemini yeah um we're on the frontier is a good way and they certainly proved that with GPQA hard graduate math problems at 88% um the the really interesting news I mean It's worth calling out. It's worth calling It's worth calling out.
So, uh, Grock got number one on humanity's last exam at 44. 4%.
Number two is sitting at 26. 9%.
And then going down this list of all these different, uh, sort of challenges, uh, they are consistently well beyond the second place.
So, they are at the frontier now of all these different benchmarks. Yeah.
So, uh, Mike Nuke over at ArcGI says, "Zooming out on ARC progress, I'd say OpenAI's Oer progression on V1 is a bigger deal than Grock's progression on V2.
So far, the O series marked a critical frontier AI transition moment from scaling pre-training to scaling test time adaptation.
Um, and this was the the Oer progression.
If you remember that uh OpenAI was spending it was like thousands of dollars of reasoning tokens generated in the test time inference to actually get a good score on the V1 of Arc AGI.
And so it had to think a ton, but it was able to figure it out.
And at least it proved that that throwing a ton of tokens and a ton of inference at a problem and uh and letting the uh letting the letting it cook basically wound up uh producing progress there.
So that was kind of like a new uh just a new paradigm.
Um says whereas Gro 4 mostly takes existing ideas and just executes them extremely well.
uh in my opinion the notable thing is the speed at which XAI has reached the frontier and that is really like it it just can't be understated that uh it this is crazy.
You put a post from own in the in the chat. Um I'll pull it up here.
He says Elon Musk is such a beast.
I'm not even going to pure I'm not even a pure fanboy anymore.
How does he a lot of swearing in here got to keep the keep the timeline PG.
But how does he come out of nowhere with a cold start late to the game and ship Grock four and do it alongside everything else he's up to?
He's launching new political parties.
He's literally magnitudes above every founder. It's humbling.
So extremely everyone agrees that it's almost like he was a co-founder of OpenAI.
Yeah, I guess he returned.
You would have to you would have to, you know, be, you know, almost be a co-founder over there to to be able to do something like this.
Uh let me tell you about Graphite.
Uh code review for the age of AI.
Graphite helps teams on GitHub ship higher quality software faster.
You can get started for free at graphite. dev.
Um, if you want to ship like ramp, get on graphite.
Yeah, Chimath was was was saying the same thing.
Uh, somebody in his reply says, "Seriously, how does this guy produce what he produces?
Meta is buying talent at $200 million a year and Elon keeps his people at a fraction. It's mind-blowing.
Very deeply underappreciated edge for Elon, says Chimoth.
The retention of the best people happen when you can offer them a freewheeling culture of technical innovation.
No politics and few constraints and people in the comments are like no politics like what are you talking about?
Can get a little political over there but but but probably not within the engineering or at XAI, right?
Like it's probably just okay, how do we build the biggest thing? Cool.
Well, you can imagine the politics of like who gets the best spot for their tent in the office.
Tent, you know that there's there's a hierarchy, a tent hierarchy, you know, proximity to the bathroom.
I want to be directly under the air conditioning unit.
I want to be closer to my desk.
The windows can be nice, too, so you can, you know, pull down your tent a little bit and get a little view. Morning light.
I wonder what the political structure is of the tent city hierarchy.
Like, is there do they they is it dem is it democracy?
Do they vote for who runs the tent city?
I guess it's just a the XAI tent city.
It's probably just Elon at the top, but does he have a tent?
Something about San Francisco intents. Yeah, very funny.
Um but Swix is has been chiming in saying like we need community notes for LLM benchmark porn because um uh in the in the Gro 4 launch they highlight this AIM competition math problem and uh and and I mean it's and so Matt Schumer is basically saying AI AIM is saturated. Let that sink in. Uh Grock 4 got 100%.
It made no mistakes on on that benchmark which is obviously very impressive.
Um but there's this extra comment about the nature of AIM and so it's a cautionary tale about math benchmarks and data contamination.
Um apparently um you know like predictions was that that the models weren't smart enough to actually solve these.
But he says I used OpenAI's deep research to see if similar problems to those in AIM exist on the internet. And guess what?
An identical problem to Q1 question one of aime 2025 exists on Kora.
I thought maybe it was just coincidence.
maybe it was just coincidence. So I used deep research again on problem three and guess what a very similar qu question was on mathsack stack exchange skill still still skeptical I did problem five and a near identical problem appears on math stack exchange and so um like at a certain point if people you know put out
a benchmark then talk about it a lot online and then that gets baked into the training data you're just memorizing the result you're not necessarily actually learning everything it's still cool it's good it's good to have everything memorized but it it really it's not beating like the knowledge retrieval knowledge engine allegations and it's and we're not really in full intelligence. When Scott Woo was on the show earlier
When Scott Woo was on the show earlier this year, he was basically saying AI will win an IMO gold medal this year.
He felt very confident in that.
Y and I'd be interested to see how he thinks about um and I'm pretty sure new performance.
Yeah, I'm pretty sure the IMO gold medal questions are public once the IMO happens.
So every year they're developing new questions but then they go out there and then they get memorized and the solutions become discussed and you know there's all the context around that and so yeah it gets it gets kind of baked in.
baked in. know big question about how valuable are are these at the end of the day it's really just about like adoption and that's why you know we we were looking at the poly market uh for the best um the uh which company has the
best AI model at the end of July and XAI has has just surpassed Google which was sitting around 80% chance for a while and then started dropping earlier this week last week um started dropping and Now, XAI is sitting at 48%, Google sitting at 45%. Well, yeah, actually updating, it's
Well, yeah, actually updating, it's updating live. Google's back up at 49%.
Is Google planning to launch something new in July?
Because it feels like it feels like this market particularly is more driven by um Google's release schedule because Google might have something in the lab, but like they like to release things at specific times like they have it's a big company.
They don't just like who knows drop it.
Gemini team Logan over there might be fixated on this poly market. I need this. Yeah. Yeah. Yeah.
Oh, during during the wait he was like if if you need something to kill the time. Yeah. Google AI studio.
So I mean people people were were definitely memeing the production values on the Gro 4 launch because it it was supposed to start at 8 I think it went live at 8:45 or something like that maybe a little bit later at Pacific time.
Uh and I robot was saying yeah this this market is based on LLM LM Marina Marina specifically the text leaderboard.
So currently they haven't fully updated it so it's unclear. Right now Gemini 2.
5 Pro is still at the top but I think the expectation is once they get Gro up there it will be the top spot.
So we'll keep following this market.
There's over two million in volume already on it.
Yeah, it's so interesting that um Anthropic's not on this poly market at all because people talk about them as having like the best vibes, the best like big model smell, the best like you know interaction and Ella Marina is like
supposed to kind of like test that with these AB tests and yet like doesn't seem to be performing there but it almost doesn't matter because they're just focused on like the business at this point as opposed to like the benchmarks. So yeah, I don't know. It's all So yeah, I don't know. It's all changing.
Um we have a post here from Ben.
Hillel he says Elon Musk on AI.
So during uh the presentation a lot of people were critiquing the presentation saying that it it it was it didn't feel like super polished or whatever.
I I don't think that was the intent and and it was pretty fixated on the models themselves and and what went into them and and what they're good at.
But Elon did have this one quote in here where he says, "And at least if it turns out," so he's talking about uh you know what will uh you know what kind of impact AI will have on the world.
And he goes, "At least if it turns out to not be good, I'd at least like to be alive to see it happen."
It's like if we get the Terminator ending, I want to be around for that. Yeah. I want to experience it.
What does that say about his timelines?
Because it's like, is he expecting not to be alive?
I I I feel like most people that have been in the Doom category have been like the Doom's coming soon, not not the Doom's coming in 200 years.
I I didn't I I I read into it more like he he will find it interesting if that is the outcome and uh and and it'll be entertaining less so like will I be alive when it happens kind of thing. But who knows?
Uh there was another funny quote at the end of the art uh at the end of the presentation where uh Elon kind of looked around at the very end and he's like uh anyone else have anything to add and one of the engineers goes uh sir it's a good model sir and they cut it extremely online crew.
Yeah definitely definitely on brand.
Uh well Ben Hilac as you know he's been on the show he's a designer probably working in Figma all day.
Think big think bigger build faster.
Figma helps design and development teams build great products together.
You can get started for free at figma. com.
And we have our first product coming out very soon with Figma make cool that Tyler has been cooking on. I've been very excited.
He showed me he showed me it and I was like oh like someone built the thing that we were thinking about building like and he was like no like I I did this is in Figma.
I was like this is like an iframe on another website that like already exists because it looks like exactly what we want but it looks so good that like like it looks like he works on it.
He looks like it looks like he worked on it for like a few weeks.
No, it looked like someone else did it.
It looked like it was a professional product that like stole our idea basically.
I was like, "Oh, like someone else got to it."
That that was the vibe when I heard it. Yeah.
Well, how how has the how has the experience been?
Uh I don't know if you want to leak exactly what you're working on, but u Yeah.
I I I don't want to talk about it too um you know, closely, but But how many prompts did it take you to get where you showed me? Yeah, I mean maybe five. I can't see. It's so crazy.
This thing was so design is super it's really great. It's really good. Yeah.
The fact that it came out looking like basically like 90 like 90%. Yeah. Yeah. Yeah.
Uh and and I imagine that there's probably like the last 10% if we were really strict about like it's got to be on this exact style guide like that might be something where like you know Tyler winds up spending more time finalizing and customizing stuff.
But in terms of like just getting a functional prototype out oh man it was it was mindblowing. It was awesome.
I'm I'm I'm very excited about the the age of vibe coding.
Um this is an interesting chart from Tracy Aloway, been on the show.
Um the cost to rent an Nvidia H100 GPU hit a new low this week with annualized revenue at 95% utilization falling uh from 23,000 at the start of May to less than 19,000 today.
So that's not that big of a percentage drop, but it is.
But I mean it is a 20% drop. It's a consistent trend. It's a consistent trend.
I wonder how much of this driven is driven just by all of the frontier labs that are driving the most adoption are moving on from the H100 to the 200.
I don't know what else would be driving this because if if you can if you can still get like if you only take a 20% drop off of a full refresh of a new of a new of like a new hardware it's not the latest and greatest anymore.
pricing drop, not a utilization drop. Yeah.
Uh annualized revenue at 95% utilization.
So this is revenue per unit.
The util utilization is still very high.
It's the it's the price that you know these neoclouds are able to rent them for which is dropping. Yeah. Yeah. Yeah.
I mean the pro the like the the market's more competitive than ever.
There's more Neoclads spinning up and more people you know actually inferencing these things.
actually inferencing these things. And then I guess this is the question of like how how stuck will certain workloads get like if you if you have figured out a great use case for an LLM in your organization and it's something that's you know not
oneshotting your entire stack or whatever but it's just like you know we have data flowing through our systems and we are going to use you know LLMs are going to you know interact with every PDF that gets uploaded to our to our website or whatever and and so we're inferencing a lot. Like you might not
Like you might not need to put that on the latest hardware or update the hardware forever.
You might just like be like, "Yep, it's Llama 3. It works.
It's on H100s and it'll be on H100's forever."
And that piece of our business will just stay there.
Just like, you know, we have a Postgress database that, you know, works and we're not changing it every year.
We're not changing everything.
We're just like, we're just trying to cost optimize that and just hopefully the cost just comes down on that.
but like we've solved this particular problem, then we'll go solve new problems with new technology.
Um, so I I I think that I think that's probably what's going on here.
Um, but but it gets to the point of like the biggest question with Grock is that like the the model clearly is Frontier. It works.
It's it it it you know, like the the whole fine-tuning on the on the actual X account is is like a crazy final step of like system prompt and people were joking about that like, oh, they got to fix that.
It's like that's not what they're demoing today.
They're demoing like the underlying raw model which is clearly like just engineering focused as you saw in the in the in the demo the demo which was just like you know benchmarks.
Turns out turns out the secret ingredient to crushing every benchmark is to have the bunch of data from schizophrenic post.
No I I I actually think it's the design of the RLHF stuff and and the design of the the reinforcement learning pipeline. Tyler, you got anything? Um, yeah.
I mean, I I think just like so far what I've seen on X, like the overall response like vibe stuff. Yeah.
Is that people are saying uh maybe it was a little too kind of overfit on the RL like VR like verifiable rewards.
Like you kind of see this when um even in in the demo, I think it would it would sometimes respond in the answers with like uh in in like latte formatting. Oh, sure.
Which is like okay, that means obviously they've trained a ton on, you know, math questions, stuff like that, papers and stuff.
Um, maybe people are saying maybe it was kind of, you know, benchmaxed.
Uh, you see it like, you know, 100% on on Amy is like kind of crazy. It's like sauce.
It's like you don't want to get too too good. Yeah. Yeah. Yeah.
This is the thing about democracy.
Like if you win like 80% of the popular vote, it's like, okay, it was a blowout.
If you want 100% of the popular vote, like probably not a democracy. I don't know.
I mean, in theory, these things should be able to to do it, but uh I'm I'm interested to know more if we dig into ARGI, is there is there more stuff going on there? Are there any secrets?
Because it does seem like an kind of an outlier result.
You can see it from this Aaron Levy post.
Grock 4 looks very strong.
Importantly, it is made it has a mode where multiple agents do the same task in parallel, then compare their work to figure out the best answer.
In the future, the amount of intelligence you will get will just be based on how much compute you throw at it.
I was joking with Tyler about this that the the individual models are mixture of experts models.
So there's a whole bunch of uh of parameters, right?
And then the individual parameters like light up the different uh neurons based on an internal to the model router.
So there's kind of like the math section of the brain or the literature section of the brain.
And so this was like one of the this was one of the key breakthroughs in like GPT4, right? Was mixture of experts.
people think we're not super sure.
Yeah, we don't still we don't fully know.
But that's like an internal decision that happens within the model to be like let's go it's this feels like a math question.
Let's go down the math path in the model.
But then Gro 4 is doing multiple it's running the same model multiple times and then comparing the results.
And so now you have yeah you have multiple agents running mixture of X-ray models.
You have mixture of mixture of agents running mixture of experts models.
And the next thing is going to be like if you want the absolute best intelligence, you need a mixture of companies and you need like I send one prompt and it goes to Grock and Claude and GPT and a Gemini and a human.
Yeah, I wonder how open routers thinking about this stuff.
Um it is funny to think about the the the human version of that where you give five engineers on your team build you know the same feature and then kind of compare notes afterward.
wildly inefficient, but with with with software when you can do these things like very quickly, there's incremental cost, but you can, you know, have more confidence in in results.
And I mean, it's basically like having a brainstorming meeting with the whole team and just throwing up a question and being like, hey, like we have this hard problem that we need to solve. Here's my idea. What do you think? What does Tyler think? What does Ben think?
You kind of like go around the table.
Everyone kind of gives their input, their various expertise.
they kind of think through the problem in different ways and then you compare answers and everyone kind of coaleses around one strategy.
This is like how work happens in the real world with a meeting.
Um it's kind of the same thing but uh certainly expensive to do that.
So, it'll be interesting to see um where companies like ho how how eager are companies to jump over to Grock because it seems like it's been a big lever for Microsoft to have uh Grock in the ecosystem as kind of a stocking horse for all the other models because yeah, Satcha wants Azure to be very model independent, serve them all.
They have the I think they have exclusivity for chat GPT or GPT APIs or they have obviously like a great deal there with OpenAI.
Um, and so if they can if they can have Grock 4 as well, that's another, you know, tool in the tool chest to be like this top layer.
Satcha is in such a good position.
It's it's probably not discussed enough how much uh just by owning those end customer relationships and being able to vend in whatever model is hot at that moment and give people optionality and still get 20% of opening eyes revenue at least for now.
Yeah, he's also SOCK 2 compliant.
If you want to get sock 2 compliant, header of Vanta, automate compliance, manage risk, prove trust continuously.
Vanta's trust management platform takes the manual work out of your security and compliance process and replaces it with continuous automation.
Whether you're pursuing your first framework or managing a complex program um so yeah, Robot was uh was talking trash about the production values. I don't know about it.
They were just they were just noticing.
I didn't think it was that bad.
I think slides are worse than I'd create after getting into roped into a presentation with one hour notice.
You can tell the engineers made them themselves.
I think just this is just a reflection of the culture, right?
They're not they're Yeah.
very clearly is like screenshots dropped into a slide.
It's light mode screenshots on dark mode slides.
So like let's do black slides and then and then you come with your white with your white screenshots that are kind of like misaligned and not really evenly distributed.
like they didn't do like the the distribute evenly or whatever, distribute horizontally.
Um, still gets the point across and I think it's a reflection of their culture. Yeah.
And you know, it shows what they care about, what they don't care about.
They're not trying to be the most polished.
They're just trying to be the best. Yeah.
Uh, I robot kind of did like a whole like live tweet here. Yeah.
So, Elon was predicting the model will discover new physics within two years.
He said, "Let that sink in." Long silence.
One engineer laughs awkwardly.
Is that sooner or or later than his previous timeline?
Because he was he was talking about AI discovering new physics soon.
I don't remember if he was saying dating it two years or three years or one year before because this could be this could be that he's he's still excited about this.
He still thinks it's possible, but he thinks it's going to take longer than he said previously.
And that's kind of the more important update.
I don't remember what he said originally.
um see if Grock can find out.
But he was saying this at the Grock 3 launch that like that is the goal and and if you can get there like you've kind of you've kind of solved everything.
And Sam Alman was talking about that too.
Uh that if you can if you can create a super intelligence like that's probably the first thing that you'd want to do is like hey go discover all the new physics and like really help us figure out how the world works.
Um so you can solve um you know fusion all this other stuff. Um, I want to be clear.
I love all you guys at XAI and only want the best for you, but I'm going to continue to live post.
Uh, Elon attempts to give a speech on alignment involving a very small child, a child much smarter than you.
The monologue rambles with no conclusion. Uh, in sight, a pause. Yeah.
Will this be bad or good for humanity?
He says the, you know, at least if it turns out to not be good, I'd like to be alive to see it happen.
Uh, oh yeah, they had a polymarket integration.
Um, that was kind of interesting. Yeah, it's interesting.
um basically giving uh giving the model access to real time polyarket data so that it can help make predictions and sort of add context around the uh the market itself.
Yeah, that's interesting.
Um Elon asking the real questions.
You say that's a weird photo, but what is a weird photo?
I still don't understand why we're looking at weird photos of Maxai employees, but they were charming.
They're calling it super grock. Crazy features. 16-bit microprocessors.
What is I don't even understand what this is.
Um, oh they, yeah, they built like a game in Grock.
Uh, they had demo of a video game generated by Super Grock. It's a Doom clone.
Every time the PC shoots an enemy, floating text appears, reading Grockum.
Elon is fabricating timelines for product launches on the spot.
The engineering s the engineer sitting next to him is looking at the floor, face impassive, nodding. It's a good model, sir. For real, though. Congratul. It's a good model, sir.
I thought I I thought this post from the actual XAI engineer, Eric Zelikman, was funny.
It's like AI AI model version numbers over time. Did you see this? No.
So, it's this chart of the version numbers over time.
And you can see that Grock is versioning fastest because it's like at this point, what else are we measuring like the like at least they're iterating on the version number effectively as opposed and I guess this is a shot at OpenAI because they launched 4. 5 and then went to 4.
1 and they're kind of like, you know, there's this big question about like when will GPT5 come?
the expectations are so high for GPT5 and so they've uh they've obviously uh with the Grock teams that like hey at least every three months we release a new full number.
So, I wonder that the the five is a number that really no one has has has like gone for.
Um, and I wonder if Grock will do it first.
Like, if you draw the line on this, they certainly should do it, you know, in like three months.
They should have Grock 5.
And there's no reason that they shouldn't, but maybe there's And it's very possible that Colossus is the is the key. Yeah. Getting to five. Oh, the new data center. Yeah.
Uh, well, they'll need linear to plan that out.
Linear is a purpose-built tool for planning and building products.
Meet the system for modern software development, streamline issues, projects and product road mapaps. They linear. app need linear badly.
So hopefully they've gotten signed up.
Near said Grock on uh humanity's last exam gro 4.
Uh I'm not sure I buy even in the general case that there's a given humanity's last exam number which implies you discover useful new physics.
How would one make a benchmark of the proper shape for this?
You'd have to have a validation set of questions which are outside the scope of what we currently are able to do.
You could choose things on the edge of our knowledge distribution and then try and exclude.
Uh yeah, it is interesting like like if you like if you are able to memorize every hard math problem does that allow you to memor to discover new math like it's it's sort of a prerequisite because you have to I think where I I've imagined these discoveries coming from are having a single intelligence that has PhD level intelligence across like a single mind that has PhD level intelligence across every human domain, right?
And being able to combine ideas from different domains like historically a lot of innovation is just taking something from one field, bring it over here, making some combination of it.
I think Elon talks about the potential of discovering new physics, but again doesn't didn't spend a lot of time like breaking down how that would actually happen.
But um world is unpredictable. So we'll see. Yeah, it's interesting.
People are really pushing this idea of like okay like like we are accelerating like the the agi leaderboard is accelerating but I keep seeing this and and feeling deceleration like I am not feeling acceleration right now. Are you Tyler? Yeah, I don't know.
I I think generally I'm kind of like not that interested in a lot of these kinds of benchmarks.
Like I I think ARC AGI is more interesting, but just like the humanities last exam, the kind of general math physics knowledge, it's doesn't seem uh to be that like it doesn't seem to line up with like you see GBT 4.
5 kind of does very poorly on these things, but like writing it does really great. Mhm.
So like I I think I'm I'm more like if I were go to long short on like different benchmarks like the usefulness of them.
I think stuff like hle I'm kind of short long I'm like have you guys seen the uh Minecraft benchmark where build the two different Okay.
You basically two models build like a Minecraft. There's like a prompt.
It's like build a house then you can choose and then it's like their rank for the mines.
But but who who's who's grading that? The human.
It it's a human who picks between them and then it's kind of like a ELO. Oh, okay.
Um, but just like general kind of creative tasks. I think stuff like that. Aiden bench is good. Yeah.
Um, I think even in the Grock launch there was the vendor bench. Which one's Aiden bench?
Aiden bench is Aiden Mclofflin's benchmark.
It's just like it's it's kind of hard to describe how it works exactly, but it's just various like creative tasks.
Um, how like kind of novel its thinking is, the the like style of its text. Sure. Um, wait.
Is it just like it's just like whichever one he likes the most at the end of the day?
like he's the only greater.
No, no, there is like an objective like function that you can do. You can like run it.
It's not just like the idea that he's like open again. It it will be funny.
uh you know there come there there's a period of life where your SAT score like matters a lot and it says something about you and then a decade later it's you know what you can do what you have done starts to matter a lot more and so I do think we'll reach that point where it's like yes you can oneshot every hard exam question there is that you can throw at it but like what can you do for me? Yeah. Yeah. Totally.
And I think that's I think that's why like the bigger question is almost like you know chat GPT DAUs and like and like actual revenue revenue app installs and stuff. Yeah.
I mean the the revenue thing is interesting because you wind up in like B2B cloud world which is valuable but it's maybe less it's like it's more competitive because it's more commoditized and well yeah if uh you you don't have a lot of leverage in the enterprise if uh Azure is able to offer infinite models that are that are infinite frontier models open source models that are maybe just behind the frontier but great at certain tasks.
The the leverage isn't quite there, there will need to be another pretty significant leap.
Until then, you know, anthropic being really good at codegen, there's leverage there.
Y uh we we saw this yesterday with with Llama switching over to uh anthropic models internally and then you know just having a consumer app with a lot of users also very valuable. Yeah.
Yeah. The other interesting thing about the the foundation model layer commoditizing and it becoming like cloud and if you have a model uh you'll just be like vended in as an API to anything else uh like the token factory is that uh the the hyperscaler clouds are extremely
profitable like even though AWS GCP and Azure are all somewhat directly compet competitive and and they're somewhat perfect substitutes for each other they have not driven prices to zero such that in the way airlines are like deeply unprofitable like AWS and Google Cloud are both profitable. Yeah. Or you look at other commodity Yeah.
Or you look at other commodity sectors like oil and gas and I don't know if that's just because there's lock in.
there's lock in. I'm not exactly sure, but there's something about where, you know, maybe the maybe the counterintuitive take is that yes, they do commoditize and there are a few major foundation models that are frontier and they all are roughly the same price, but they all have decent lock in with their customers to the point where they're still able to extract some level of profit or they're just creating so much
value that even if they're taking like a small marginal slice on top of uh on top of the the cost to done that they're creating so much value that they still have 50% margins or something like that because like I mean this was the story
of AWS like no one knew how much money it was making and then and then they they had to break out the financials um in one of Amazon's uh earnings reports and it was like the AWS IPO as Ben Thompson put it. Um anyway before we get
Um anyway before we get to the next story uh let's tell you about Numeral HQ sales tax on autopilot spend less than five minutes per month on sales tax compliance.
Uh so the big news is that the third browser war has begun.
Um Google stock has dropped on the news that OpenAI is planning to launch a a Google Chrome competitor within just weeks.
And this is very interesting timing because it's time to browse. Yeah, time to browse.
Certainly makes sense to become deeper in more deeply integrated into the user's life. Makes a ton of sense.
There's a ton of benefits that come from having a web browser.
Um what was interesting is uh the we can go into what Google actually la or what OpenAI is talking about launching but this news this scoop leaked the same day that uh Arvin from Perplexity announced that they're finally releasing their next big product after launching Perplexity.
Comet the browser that's designed to be your thought partner and assistant for every aspect of your digital life work and personal.
And so Perplexity launched this on June 9th and then OpenAI the the scoop goes out via Reuters the same day.
And so this feels like very much like let's not let Perplexity get a bunch of attention and drive a bunch of people to to start daily driving Comet the browser because even though we're not ready to launch our competitor, we want I mean Arvin was on the show talking about Comet over a month ago.
He said it was really important to the business.
this was a big bet that they're making. Yeah.
Uh he uh and I'm sure both companies are racing to be the first to launch, but Dia the browser from the browser company uh also launched out of or they're still in beta, but they launched like a month ago or something like that.
So this is, you know, you're not going to be the first.
Oh, they launched a month ago with the DA browser.
That's interesting because I saw Riley Brown also posted the cursor for web browsering DIA browser and I thought Dia browser launched that same day, but I guess it had launched earlier. Um, yeah.
So, anybody that was an Arc user can download DIA today uh and chat with their tabs.
But, interestingly enough, Perplexity's brow browser and OpenAI's browser are both built on Chromium, the same open- source project that underpins uh Google Chrome and Microsoft Edge. Yeah.
So, it the cool thing here that means that they're compatible uh compatible with uh existing Chrome extensions. Oh, interesting. Interesting. Okay, that's cool.
Um, yeah, it's it it's I I I I want to talk to more people who were like active in tech during the earlier browser wars.
The first browser war was Netscape Navigator versus Microsoft Internet Explorer.
This is in the mid mid 90s, early 2000s.
Uh, Netscape was super dominant and everyone loved Netscape.
It was originally the Mosaic browser.
This is the Mark Andre project and then but Microsoft bundled Internet Explorer with Windows 95 and the distribution was so powerful that Internet Explorer actually wound up winning and became really really dominant.
But then there was this lawsuit and went back and forth but then uh basically in by the early 2000s Internet Explorer had over 90% market share but then they got kind of lazy and stagnant apparently and I mean I'm I'm not exactly sure ex what happened but they there was a lot more competition.
So Firefox, which was, I believe, like a spinout of Netscape or kind of like some of the same heritage there, um, began getting traction and then Google Chrome launched in 2008 and leaprogged everyone.
Uh, and Google Chrome was really focused on like speed.
It was the fastest browser.
Um, and they they did a whole bunch of work to optimize JavaScript so the pages would just load faster and run better on pretty much every computer that you had.
And so, uh, and then they had the open source project with Chromium.
And so they were able to kind of standardize the entire industry.
And so everyone's always been trying to draw uh analogies between like the browser wars and the LLM wars and like what's the role of open source in that?
Like is open source a strategy to wind up maintaining your your dominance?
How much does distribution matter?
Like Chrome was probably pretty easy to distribute because every single person was visiting Google just every day searching.
And so you just put this bar, hey, want to switch to the faster browser and people just do it because you have basically like, you know, billions of ad impressions on your product every day.
Will be interesting to see if chat GPT can get people to download their own browser on desktop.
I mean, I'm using ChatgPT on desktop in Chrome all the time.
I imagine which chat chat GPT model would you want to use as a default search engine?
That's the hard part because I always run into this problem where it defaults to 03 Pro, but that takes 10 minutes and so then I have to go to 40 and then if I'm in an 03 Pro flow and I'm talking to 03 Pro and I let it cook for 10 minutes.
It gave me a great answer, but then I want to just be like, "Okay, just like clean this up a little bit or summarize this or do some bullet points. I want 40 to do that."
So, I have to switch over. So, I don't know.
I I would imagine I'd go 40 as the default because I want speed, but even 40 could probably be faster before it truly replaced Google's very fast.
They've spent a very long time being fast. Yeah.
And I could imagine them doing a similar project to I believe it was like the V8 JavaScript engine.
They sent this team out to uh I want to say like Iceland or something.
Iceland or something. Uh they they basically sent like a bunch of engineers to like an offsite and they were like just go optimize JavaScript for like a month just go focus on this for like a month or months and come back when it's done like you have no other responsibilities than just like
optimizing this like compiler and they came out came back with the V8 JavaScript engine and created this whole like NodeJS boom people were running JavaScript on the server then and uh and I could see Google kind of doing something similar where they're like okay we have Gemini It's good at looking stuff up. It's a good knowledge
It's a good knowledge retrieval engine.
Go figure out how to make it load all the tokens for the full response in 100 milliseconds.
And that would be very very cool.
And I wonder if that's like a uniquely Google advantage.
Tyler, you look something up.
Yeah, it was in it was in Denmark. Denmark. Okay. I was close. I was close. Yeah.
I wasn't sure it was Finland or Iceland and Denmark. Yeah.
The interesting thing here, I'm realizing that tabs are definitely a light lock in to browser.
It's not just the default, but if you have six to 10 tabs that you've just had open for a really long time and they're like from a bunch of different things and you can't exactly remember what they were if you had to list them all off, but you know, you know, I I personally end up using tabs as like somewhat of a to-do list.
And so if you're spinning up a new browser and you don't have your tabs, it's like, oh, do I want to just like get rid of my my tab stack?
I have a bunch of tabs that just have stayed there for years and they're basically like it's basically like a mini operating system, right?
With like different apps that might be a Google sheet or something else. Yeah.
No, I know what you mean.
So there's very real lock in.
I could bring all those tabs over, but I have to then log in to a bunch of different services.
And so it's it's really really hard to actually uh win here.
I wonder if anyone's using, you know, in in Google Chrome, you can actually change the default search bar to, you know, when you type in the search bar and if you just type words, it just Google searches it.
You can change that to search chatgpt.
Yeah, like you can pass in a query parameter and it can just do that.
But I haven't heard of anyone actually doing that.
And I used to have I used to be such a power user of Chrome.
I used to have different code words basically.
So if I if I typed like I space and then a query, it would go to IMDb and search that specifically.
So you could you could have Chrome like route to any specific search any so you could press like Y space and it would search Yelp or you know anything else.
Um but I don't know if people are I don't know if people are doing that with Google with Chad GPT.
I think people mostly just like control command T and then hang out in Chad GPT.
Well, we'll have to ask uh Chris in 15 minutes about get an update on the browser wars because he was uh an early investor in I know one of those tabs that you have pinned right now. What's that? Adio, of course.
Customer relationship magic.
Adio is the AI native CRM that builds, scales, and grows your company to the next level.
You can get started for free.
I've had Adio open for thousands of hours in a row at this point. Yeah.
Uh so, Signal kind of breaks it down with the Open AI launching the web browser.
es this is the oldest play in tech.
Find product market fit with a single killer use case.
Then vertically integrate and horizontally expand until you control the interface layer itself. App platform.
Once you own the interface, you own the defaults.
Welcome to the next generation of browser wars.
I yeah what's interesting is there like Sam Alman at OpenAI and just the fact that OpenAI is a company like there is kind of a mandate to like vertically and horizontally integrate figure out code figure out research figure out devices but every company wants to do everything
but then sometimes they run up against barriers like there was a time when Google was like we want to win social networking and we want to beat Facebook and we're going to launch a direct Facebook competitor and they did and it didn't go well and then they shelved it
and then they wound up producing trillions of dollars in market cap just doing the thing that they do great and so the question is like the surface area of open AI they have to exp explore they have to experiment it's it would be stupid not to see if they could get a
browser and a device and a chip and a nuclear reactor and everything and sand get the get the sand get everything um but but there's no there's no guarantee that they will win the entire vertical stack and there will be the one company, right? I think my question is are these going
I think my question is are these going to be like is OpenAI's browser going to be an entirely new app other than their existing mobile app?
Is it is or their desktop app?
app? I yeah that is interesting because if they have to get people to reddownload a separate app then then that's then that's like an entirely you know they have a good fly you know they have a bunch of wouldn't just evolve the apps they already perplexity too I don't I don't
perplexity has uh is planning to to release this as like a new standalone app or it will be in the perplexity mobile app but yeah um yeah I mean I know I think comment's like its own thing because we were looking to download it and we need a code. Um, and you can't just get it if
Um, and you can't just get it if you're just on perplexity. Um, but I don't know.
All I know is that you should go to fin.
ai, the number one AI agent for customer service, number one in performance benchmarks, number one in per competitive bake offs, number one ranking on G2.
Um, so, uh, Arvin breaks down like his philosophy of, uh, of Comet, the browser that he's dropping from Perplexi.
says uh you can either keep waiting for connectors and MCP servers for bringing in context from third party apps or you can just download and use comet and let the agent take care of browsing your tabs and pulling relevant info.
It's a much cleaner way to make agents work. So that is interesting.
So I wonder how much like puppeteering will be in this because Chachi Chacht and OpenAI have operator that operates a Chromium front like a headless web browser basically but you can actually see it working and it's clicking things.
Um and so if they're like there's also the value of like the training data.
If you're getting people using all these websites, you have all this training data of like, okay, they clicked on the blue button, they clicked on the green button, they saw this, they they they entered, this is how they dealt with this form, this is how they dealt with that form.
And so that feels like very very valuable data if you can get it.
So it's probably worth duking it out even if it doesn't uh even if even if it takes a long time. Um, for sure.
I do wonder where where else they will um where they will plug in like Cle operates at like a higher level of abstraction with like the screen scraping and I wonder if we'll hear rumbles about either perplexity or open AI thinking about like moving up the stack to that level. I'm not exactly sure.
Um, anyway, uh, Dan Ivy's says, "We believe Apple needs to acquire perplexity for AI capabilities.
likely $30 billion range would be a no-brainer deal given treadmill AI approach in Certino.
Perplexity would be a gamecher on the AI front and rival cache given the scale and scope of Apple's ecosystem.
So people have been talking about this for a while.
It feels like there were talks and then they kind of stalled out and and arx revenue multiple. Wow.
I mean the product sense is good.
you use the product and and like there's like Apple hasn't been able to deliver on the product side.
They have the distribution, but they haven't been able to get things.
We talked about this before though.
The the most expensive acquisition Apple has ever made was Beats by Dre for $3 billion, which was a 3x revenue multiple. It' be a huge shift. Huge shift.
I don't know that I I think that Apple is embarrassed right now and and feels a lot of pressure to deliver.
I don't know if they're at the point where they would pay $30 billion just yet.
And even then, it's like hard to integrate and or even 14 or whatever their last private valuation was. Yeah.
And the big question for me was like perplexity is is built on a lot of different clouds, a lot of different tools, a lot of different models.
Is Apple cool with that stack? Yeah.
because if all of a sudden or do they want to just go direct to enthropic or open AI which they are in conversations with and every once in a while these these like um scoops pop up around perplexity and and Apple conversations and it's hard to read into that are these like is this like rumor mill like what's driving that rumor mill? Yeah.
Well, pull up the mag 7 um chart.
I want to see where Apple and Google are sitting today. Apple at 3. 2 two trillion. Google at 2.
1 trillion and Nvidia is holding strong at 4 trillion. Not bad. Yeah.
I mean, 1% of market cap, they're at 3. 2 trillion.
$30 billion acquisition to be, you know, to have a have an an AI product that clearly has a good roadmap. Isn't that crazy? I don't know.
Well, if you're making bets on any of the mag 7, do it on public. com.
Investing for those who take it seriously.
They have multiasset investing, industryleading yields, and they're trusted by millions, folks.
Um, so in other Apple news, they're preparing to launch the new version of the Apple Vision Pro.
They're just doing a slight iteration on the chip.
They're moving to the M4 chip, and they're launching a new strap, which was something people were complaining about because the weight, maybe it'll be better distributed.
People were switching out for the Pro Strap like earlier. It's so funny.
So, uh, last week, um, uh, Elon, uh, announced the America Party, or I guess it came out on Monday.
Stop stock dropped from $312 a share all the way down to $291 a share.
That is when Dave Portoi, I think, market bought.
He's he was saying he he's Davy Day Trader is back.
If that's not a top if that's not a top signal, I don't know what is.
But he was market buying like 10 million of of Tesla being like, I just think it's going to go back up to where it was and it's just been climbing since then.
It's back up to $38 a share.
Uh almost almost recovered. It's up uh 4. 2% today.
So on brand for Elon and basically gonna it looks like it'll just recover the price prior to the America party.
And Dave Fortnite, he literally was he basically his thesis was like, I think it's going to go back up to where it was in about two weeks and I'm going to make 10%.
And I'm going to make a million.
I mean, that was your thesis on Nvidia.
You were like, wait, it's like down because of Deepseek.
Like maybe it'll go back up. It did.
It was like the the most basic analysis and it worked perfectly.
It was it was fascinating.
May maybe that's broadly a top signal.
Just the idea of like simple analyses and like not necessarily needing deep insight to to call the market is good. I don't know. Who knows?
Um anyway, this Apple store is from Mark German, of course, the master of scoops and Bloomberg.
He's got He's on his fourth or fifth this week. He's on absolute tear.
I mean, this one is a little bit minor.
You know, they're going to include a faster processor and components that can better run AI stuff.
And so, um not that Apple has any crazy AI stuff that they really want to run in there.
I don't think that that's a major differentiator.
I've been thinking about how how like is AI a key unlock for VR and like I don't think so at all.
I think it's much more about the content and the use case entertainment.
I think it's a replacement for a TV for to to start and they need to just make it dead simple to to use as a I don't know.
We got we got a we got a demo of VR product a while back and it uh had some very cool native AI features. Yeah.
So there's something there.
But Apple's product doesn't feel like it's ready for just wearing while you're making dinner. Yeah.
So, that version, the one that significantly reduces the weight of the headset, they're planning to launch that redesign model for 2027, which feels so far away.
Uh I I know it's only a year and a half, probably the end of 2027, but so maybe we're talking two years, but uh that in in the in the AI race where we're like, yeah, AGI tomorrow, AGI next week, AGI next month, major news, ship a better we can't we can't slim down the headset and take off the screen and create a lighter materials like this this month. Like let's do it.
Um but hardware is hard and you know the stuff takes time. So good luck to them. I'm excited for it.
I'm I'm very excited for the next Quest.
Do you still have a Vision Pro? I don't. I had it for a month.
I took it back uh because um I just wasn't using it that much.
It was like heavy and I couldn't find and it had a bunch of things that like you had to do like these crazy workarounds.
Like I wanted just like an HDMI cable that I could plug into it and then just be like, "Okay, my PS5 is in VR now." And I couldn't do that.
It was like you had to like pull the pull the screen into the Mac and then screen share it in. There'd be latency. It was ridiculous.
The thing the use case that I still see is people using it on planes. Yeah.
But I just We got to check in with Tyler when uh give me your How many times have you thrown on the VR headset in the last week?
Did you play it last night? Break it down.
Have you turned is collecting dust?
No, it's I I I've been playing a lot of Call of Duty. It's in the VR headset. Yeah. It's a lot of fun. Hold your position. There's like no latency. I'm kind of surprised. No latenc.
You're doing in the cloud. Yeah. Okay. And it's online. It's multiplayer. Multiplayer.
So, you play multiplayer and you play like the latest and greatest Call of Duty basically. Yeah. Okay.
Black Ops like six, I think. Cool.
So, you have a controller and it's a big screen on the wall and you just chill there. But, walk me through it.
Is like 30 minutes a day. Yeah.
Probably like 30 45 minutes a day. You're fired.
No, this is not a gotcha.
This is This is true research. So, yeah.
I mean, I I honestly think that that that that the the the Quest Xbox uh the Meta Meta Xbox Quest or whatever.
I forget the name, but like that I think that's more like I think that's better to use than like a processor bump on the Vision Pro.
Just like deeper integration so that you don't have uh so you can just throw it on.
What's the actual time to, you know, if you want to turn it on, throw it on, start, get playing, get into a lobby, actually get your first kill, is that one minute? Uh, no. It's like 30 seconds. Maybe 30 seconds.
I mean, no, maybe like a minute.
It's not like noticeably slow. You're logged in.
You don't need passwords or anything like that.
It's not It's not like a hassle. Okay, that's cool.
And I I put the screen It's funny like I I have a TV in in my apartment, but I just put the screen right where the TV is cuz it's like the perfect spot on the couch.
A nice black black square. Yeah. Yeah. Yeah.
I'm going to have to get this back for you now. This sounds amazing.
Now, I don't have any time to do this, but um but but uh I I feel like the next what's on the feature roadmap that you would want to see like me uh Apple is bumping the uh the the neural engine and trying to upgrade the chip.
I'm not sure that that's the problem with the Vision Pro.
What would you like to see out of the Quest 4, I guess, is the next one that's coming.
Yeah, I think the main thing so I' I've tried the the Vision Pro and basically I mean the visuals are just like vastly superior.
It's it's it looks so much better than even the it's so it's it's the Quest Better Screen Meta Quest 3S Xbox Edition.
That's what I have and the screen is just way better.
I think that's I would say that's the main thing.
So, if they can just go find the supplier, if Meta can just go find the supplier for the Vision Pro screen and put it in the Quest 4, you'd buy it yourself.
Depends on how much it is.
I'm kind of broke, but I I would definitely be inclined to.
Not Not after you drop out and go full.
You would speedrun all of Halo in order to potentially win one. I would do that. You would do that?
You would do a very difficult challenge in order to potentially win one because you would you would want it. Yeah. Okay.
No, I think that's that's I would think that's definitely the main thing is are there any other are there any other nice to haves that you think might uh might shift people? Um I don't know.
I mean the it's it's very light.
It's it's way lighter than the the Vision Pro. Yeah.
Um but still I mean I I feel like light is very relative.
Like it's light to the point where you can do 30 minutes or an hour.
I you probably can't do like a full day or like or like four hours or any kind of like workout stuff.
I think I definitely would not do that. Totally. Totally.
But but what I'm saying is you're not training back enough.
I knew I knew guys I knew guys at UCLA who I I didn't go there, but like I friends that went there and they were so obsessed with Call of Duty that they would take a bunch of stimulants when the new Call of Duty came out and play it for 24 hours straight to get the max prestige because they were so addicted to Call of Duty that they would just stay up all night chugging energy drinks and uh and just to just to beat that.
And I I just don't think you could do that in VR.
I think after like two or three hours right now it's like too much and you have to take it off and get sweaty and tired.
Um but so so I feel like I feel like screen first then probably even a little bit lighter, a little bit more comfortable and then just drop the price as low as possible because if the next one was a hundred bucks, you'd probably buy it, right? Yeah.
And I I think stuff like like maybe I'd want another screen instead of just this monitor, but that's just an issue with the with like the visuals. It's the screen. Yeah.
It's it's got to be it's got to be competitively priced with a TV.
And the TVs are so cheap now that you got to just be like, "Yeah, I'm just picking one up."
Or the price of an of of AirPods or the price of, you know, it's it's got to be down in the low low hundreds of dollars to really ramp that up. But I don't know. It'll be interesting.
Anyway, our first guest is here.
Let's tell you about AdQuick really quickly.
Out of home advertising made easy and measurable.
Say goodbye to the headaches of out of home advertising.
Only ad combines technology.
Out of home expertise and data to enable efficient, seamless ad buying across the globe.
and we will welcome Chris Pike to the show. Welcome back, Chris.
Fantastic to have you on the show.
Thanks so much for taking the time. Great to see you.
Last time we got cut off, great to be back in the temple.
Last time we got cut off, we were we were having to jump and I was like, I wish we had another hour.
So, at least we have another 30 minutes here.
Yeah, it's great to get into it.
Um, what first off, what's top of mind for you?
Have you been tracking anything in the news that's that that that's kind of updated your thinking?
We were digging into Gro 4 and seeing is this an update to you know agent timelines.
It seems pretty great in the benchmark but is there anything else in the last week that's been like ah I can't get enough this story just in your world. It's a great question.
I feel like every week is a total blur. Yes.
Uh it seems like we're all waiting for not just these foundation models to come out but like the next open source models the big open source models to come out.
I think that that's super interesting to me.
Um the proprietary foundation models obviously are the frontier of research.
Uh but they're relatively inaccessible from a technology perspective because they're uh fundamentally rentseeking.
You can't run them on your own hardware.
You can't it's significantly less accessible.
Um and so I'm I'm kind of waiting for the next generation of open source models. Yeah.
Uh, one maybe underrated or underanalyzed Grock 4 thing that happened last night was I don't know if either of you saw this, but they they did this voice demo and they were like really pushing the accents really far.
I don't know if any of you saw this. Dyler, did you see this? Yeah.
And there's the the whispering the whispering. And so ASMR.
Wait, did you think it was uncanny valley, Tyler?
I felt very uncomfortable listening to it.
But at the same time, I think it I think it's a path where we're in the uncanny valley, just like we were with like sixfinger hands and stuff, and when they actually sort out the accents, the whispering, the intonation, the cadence, it's going to become a much more addictive companion potentially.
So, I want to I want to bridge to your piece and talk about um the the the different use cases that you see might people might kind of flow into with these like chat companions because you mapped out way more than just the the normal take in my opinion. Yeah.
Well, so I guess it's worth uh asking ourselves where where do we want other humans to exist and where will we accept like substitutes?
I think last time we were talking about it, you know, if um if you can imagine a situation where uh humans are getting in your way of doing something, then you kind of you hate them, right?
Imagine you're in in traffic, in gridlock traffic.
That's it's the most misanthropic you could possibly be.
You're like, if if none of you existed, I could just get where I wanted to go without you being here.
Um, but then there are just total other times where we will we would refuse to to accept anything other than humans as as that thing.
Um, I know that it's it's very popular to talk about like AI companionship.
Um, and I would never um I would never uh I would never say that people who get a lot of value from AI AI companions that that value isn't real.
Um but at the same time I think that when it comes to the allocation of let's call let's let's call like allocation of leisure hours we really care about other people.
Um whether it's like I I think last time I was talking about like going to fine dining or uh reality TV.
Um, I think I I mentioned it like we have uh the chess uh software that is way better than any human will ever be and it's not entertaining to us.
It's not entertaining to us because that's sort of like a there's no drama there. There's no emotion. Exactly. Yeah. Exactly. And so, um, sorry.
I I I I'm I'm just thinking about like the shape of companionship because in in your most recent piece you you call out like the imaginary friends that kids have, Calvin and Hobbs, Toy Story.
These stories resonate because they poignantly depict how colorful whimsical placeholders of our childhood slowly fade as society offers real alternatives.
Um, and and I'm just thinking about like like kids love imaginary friends, but they also in they also love like multiple IPs essentially.
Like they like Batman and then they also like Spider-Man.
And so I'm wondering like there's been this narrative for for you know a few years in AI of like don't build a GPT rapper because you're going to get you're going to get rolled.
you're going to get rolled. uh there's going to be immense concentration of value and there will actually be no middle class in this in this ecosystem and I'm wondering because there's this company Tolen that's you know kind of imaginary friend AIdriven and I'm wondering like how like we might
actually like is it possible that we're heading towards something where people are essentially developing new IP and yes there's still a power law in the companionship market, but they're but these models and these products are like much more opinionated to the point where they're there there actually isn't a one product to rule them all. And there's
And there's and there's a variety of products that that fit into different holes just for companionship, but also even just within like the imaginary friend hole, there's 25 different options.
And yeah, there's one that's popular, but then there's one that's onetenth as popular, one/100th as popular. But yeah, react to that.
Yeah, it's a it's a it's a really interesting Let me first go back to like the the this rapper concept, please.
Um, and I think it's really important to distinguish when rapper strategies work and when they don't work. Mhm.
Uh I would argue that the wrapper strategy works the best when the underlying infrastructure is purely commoditized.
Like you can choose across many different options.
Um where the wrapper strategy doesn't work is if you're being if you're if you're basically building on top of a monopoly and that underlying landlord is basically just going to be increasingly rent seeeking and squeeze you out of all margin.
Um so sort of embedded in the wrapper strategy is the assumption that over time you're going to be able to uh distribute your product on top of like increasingly commoditized infrastructure.
So for example like um like snowflake.
Snowflake launched uh actually just just on Amazon and then over time uh uh distributed its product across um Azure and Google and actually in doing so was able to expand the margin capture of its own product because it had a uh their vendors were competing to be their their underlying customer base.
Um going back to like this open source notion, this is actually why I am so interested in open source um the more that we have competitive fungeible models, the more that the application layer on top can really flourish um the more that we'll see distinct unique applications.
um and and kind of until we start to maybe S-curve uh near the top of of of the frontier.
I'm sure you guys are familiar with the um it's really popular essay, the the bitter lesson. Oh, yeah.
um which is basically the the you know no amount of fine-tuning or no amount of specific training is actually going to out compete just the fundamental uh advances when it comes to like uh more games with more compute. Yeah.
Um, but if we start scurving, if if we start if if these scaling laws break, which it seems like they are breaking on pre-training and test time compute and things like that, all of a sudden you have uh largely funible similar capabilities um at that commoditized layer and then we can really start to see the application layer flourish.
Yeah, I my interpretation of the bitter lesson right now is that the the impact of AI should be tracked less in benchmarks and less in individual tests of one model and more in the actual like volume of inference tokens being generated by humanity.
And it's fine that the that we don't have one central AI doing all of the work.
If we just give everyone an AI co-pilot for every single task, they all get better and and we'll build more data centers to inference more.
And eventually that will compound and compound and compound until the until the the the overall impact of AI is remarkable and and irmisticable in the same way the internet has.
But it won't be this like all of a sudden we unlocked this one incredible algorithm. Yeah. Maybe.
So, so I'm going to I'm going to try and m map a comparison that's probably wrong for any number of of reasons.
Like let's say let's say we ported the bidder lesson to the roll out of um uh PCs. Mhm. Right.
I think there's there have been a lot of comparisons of like AI feels like a new computer.
Um so like let's map the bitter lesson to PCs. Mhm.
Um maybe the the analog would be, hey, don't work on building software that optimizes for like a computer speed of like 50 megahertz because the the computer that comes out that's like 100 mehz or 200 meghertz or you know a a one gigahertz is is going to be you know is just going to blow out whatever software optimization you've you've achieved at that compute stop.
Um, and so I think about it a lot like that.
Now I think this really begs the question of okay, well what happens when the vast majority of our use cases are satisfied by the compute threshold.
you know, I I I feel like um you know, every everybody's like, who needs the nextgen version of this?
Because largely all of the applications that you can that you you use or you can run um work.
Uh and so I I it's very clear that we're in this like rising part of the Scurve.
Um but when when the S-curve starts to taper off, uh that's really when it comes in a question of okay, well um how do we think about um this the value delivery that sits on top of uh the underlying uh rent capture from from these foundational models.
Um you know uh You could you could think about it also um like video game consoles.
Um the the amount of the amount of creativity that uh video game developers are limited by is actually how advanced the video game consoles are and also how like how expand like what the install base of the video game consoles are.
And so I think one of the one of the challenges right now is like we have so few developers of AI applications. Yeah.
Like there there's like you can kind of like count them on maybe a few sets of hands, right? Which is insane. Really?
I feel like there's like thousands of startups that that count in like the AI software developer world.
I see market maps every single day. Um I'm sorry.
Uh there's like 10 in every B2B category like uh well okay so maybe like let's let's split the world between like enterprise use cases and consumer use cases. Oh sure sure.
So enterprise I would yes 100 a thousand% there's there's a lot of companies that are it's it's a it's a blue ocean sprint to uh vertical value delivery within in different sectors.
um because you're largely swapping um you are uh the the spend the addressable spend is um is opex which is insane like that's crazy your your revenue opportunity is just like headcount spend um and so that's for sure when tool spend like it's all the opex like to your point I mean I guess not like real estate or rent or something but basically everything else yes which by far the biggest cost center for any company.
Course anyone who's ever run payroll or anyone who's ever scaled a company is like, man, like humans are expensive. Yeah. Humans are so expensive. Yeah.
Um and I mean I I this is this is why I feel like uh everybody says that AI is the best candidate or the best argument that UBI is coming. Sure. Sure.
Uh yeah, Jordy, I think uh I think the picks and shovels meme became too dominant and there were so many over the last decade or so, there were just so many amazing outcomes of people building infrastructure and like very visible outcomes and it became cool to build infrastructure, right?
Like the Collison brothers made building infra cool and you have like Parker Conrad is like a fulk hero, right?
And like rippling it's like like and you have the like ramp is a good example of this.
Corporate spend management should not be cool.
They've built a you know really cool culture around it. interesting.
And there's this other there's this like weird kind of pervasive meme of like, you know, uh you have two years to escape like the permanent underclass.
And so I think people are like, well, I'm not going to just build something weird and fun.
I'm going to build enterprise SAS so I can make, you know, so I can escape the the permanent underclass.
And so I don't think there's been enough weird fun attempts from people like Tolen's the one you brought up earlier is cool.
not the most rational thing to say, you know, if you just want to build a big business to be like, I'm going to build a little alien AI friend, but like clearly there's demand for that.
And we were talking with Scott Bellski yesterday around just wanting like new fun weird consumer use cases.
And I feel like I think what you're getting at is like that whole area is like relatively under underexplored to date where like we've had a bunch of we had two browser announcements yesterday and they're both built on Chromium. Yeah.
And that's exciting and cool and we should we should talk about the potential for new browser wars.
But I think like yeah, the number of people that are saying I I'm actually yeah, you could call it a rapper, but I'm actually trying to create something entirely novel.
Like the example we gave yesterday is like a dating app based on like a digital twin that is just constantly dating other digital twins.
And you know, I haven't seen I'm sure somebody's working on that.
I haven't seen it uh yet, but there's like any pick any popular consumer app category and there's probably a way to entirely rethink it with this sort of uh LLM as a new computer at the core of that.
Consumer just seems so much more risky because it's like it's either a billion dollar outcome or a zero.
Whereas in in enterprise, it feels like well there's no way it's going to be a zero.
It's going to be a $10 million outcome or hundred million dollar outcome or a billion dollar outcome, but it's not going to be it's not gonna be.
So then you have 10 companies that each have 10 million where like they were trying to go for the hits driven business with the game didn't work and then they went into SAS and it worked. Anyway, sorry.
I think the additional challenge with consumer is um inference is not free right now. Yeah.
like somebody has to pay for the inference bill.
And so until you can run inference on device, it's still like every developer is doing the mental math in their head of like how do I like if any one of your users can can run you out of house and home if they abuse your your service? Yeah.
How do you build on top of that? Yeah.
One of the one of the one of the OpenAI researchers I think in Zurich that's going to Meta Yeah.
like posted like I didn't realize that I had this thing running.
It was like $150 a day just like every single day.
Luckily, I think he he'll be fine. He can take the hit.
Um I did I did before we fully leave like kind of consumer.
Uh what do you think of do you think that AI companions and LLMs broadly present a real threat to traditional social media?
like the idea of a companion.
It um a lot of people in our world are using these tools like very functionally like for doing research or getting answers or understanding topics, but a lot of people are using them as companions and that is somewhat of a social entertainment experience and and you see these charts ticking up of kind of user minutes in in in um so went from about five user minutes per day to over 30. There's around 30.
I think it's a 28 29 over like the last six months.
There's no blip on any of the other social networks.
They're not declining yet.
But personally, I'm finding that if I'm doing research on a topic, I used to go to YouTube.
I used to go to Instagram and just and people, you know, famously like search Tik Tok for answers to things.
And some of that is shifting over.
But but and if you think that an analog sort of anecdotal, you know, Yeah.
uh sort of experience for me is when do I use social media the least is when I'm with my family which is like companionship.
It's this social time or when I'm with friends like at dinner hanging out.
It's like rude to be you know using you know there's no point to go hang out with a friend and then use Instagram the whole time. Totally.
I I think uh I subscribe to this sort of cutting of our time as like we're we're either allocating labor hours or we're allocating leisure hours.
So we're either we're either trying to be productive or we're trying to um enjoy ourselves.
And so I would say that uh all all leisure allocation effectively competes against each other.
I think um maybe it was like the Netflix CEO that said that they weren't competing with HBO, they were competing with Fortnite. And that's largely true.
Like you know, you know, we only have 24 hours in a day.
You only have a finite amount of leisure hours.
So if I'm allocating one hour to this leisure activity, if I'm watching an episode of Love Island or whatever, uh that's that's time that I'm not going to be able to allocate to a different kind of leisure activity.
So to that end, I would absolutely agree that um AI companions almost certainly firmly in the bucket of leisure um more consumption of leisure equals less consumption of other leisure activities.
It's it's really zero sum.
The only thing that is going to make it nonzero sum is a fundamental advance on productivity that allows the leisure pie to be even larger. Yeah.
So, right, like maybe we have we we definitely have more leisure hours as as humanity now than we've ever had in the history of humanity.
Um, let's give it up for leisure hours, right?
Like like protohuman zero zero leisure hours like amazing amazing leisure hours.
amazing amazing leisure hours. um uh uh when it comes to labor hour allocation or like research or or utility um I would say like that doesn't necessarily encroach against um uh time on on social media and I think that all social media uh whether it's YouTube or Tik Tok or Twitter um they carry they they care
less about helping people get work done and they carry they care much much more about absorbing as much of um of attention as possible which is why like the the algorithms are so insidious of like serving you exactly the kind of saccharine thing that you want to consume next I wonder if there will be an incentive I mean there there will obviously be
incentive but I wonder how it will play out in the LLM chatbot interface because right now open AI and basically anyone one who has a a dominant consumer AI app um is probably seeing user minutes increase just naturally without putting in like growth hacks or retention loops or you know but you could imagine a world where to get to to get from 30
minutes to 60 minutes the LLM has to not just give you the response but surface hey would you like to follow up and learn more about this click these buttons that that's already kind of happening yeah it starts surfacing you stories that it knows you're interested in by what you've searched for a bunch of stuff. Let's
Let's give you a new breakdown kind of populating a deep research report.
You you're the it is funny that the push notification hasn't really quite hit it hasn't. That's right.
Chat apps and and it undoubtedly will.
Chat apps and and it undoubtedly will. I had this thesis that that push would be very important versus like pull like you have to go to chatpt and ask it for something and and someone is going to solve kind of like it's almost like an AIdriven newsletter or something where it it it it understands what you're
interested in and then generates the report before you can even ask it because it knows that you know if if you know Ferrari drops a new car I'm going to want a table of all of the details because I like consuming information that way in addition to just hearing commentary and watching the Doug Demiro video about it. Um, but but OpenAI could
Um, but but OpenAI could pre-populate that and just send that to me. But I don't know.
How do you how do you think that's going to evolve? Question for you guys.
Do you think we're going to pay for AI services forever?
I've been asking I've been asking that a lot. No.
I mean, I think I think I think that I think the better the more important question is will the will the will the average American pay for a an AI like an LLM and then will they pay for multiple like the the comp for this is in streaming where Americans I was about to say Netflix
Americans a lot of Americans pay for multiple streaming services but they're incredibly ruthless about canceling them on average like a lot of I'm sure a lot of people listening to this have had some streaming service billing them monthly, you know, for years that they haven't even watched. But but the
But but the average American is like, I'm not getting a lot of value out of HBO right now.
I'm going to cancel even though they might sign up again in like four months next time there's a hit show that they're going to watch.
And I think that um I think that uh there's not a right now there's this incredible demand for what's new and what's best, right?
what's best, right? like Grock will drive a lot of signups today because it is a a me the the Gro 4 heavy is like a meaningful advancement but I went and I haven't you know we've been busy this morning I haven't had a chance to sign up and and play around with it yet and I was getting plenty of value from Grock 3
like I was able to just search yeah I was asking Grock 3 about Grock 4 as well and it was actually doing a pretty good job which is and so and I'm not Yeah so I'm not I'm not I guess I probably get it through um X premium but so I have some I have some data here Netflix uh made uh for 39 billion last year. They're on track for 44 billion
They're on track for 44 billion this year. Uh 1.
8 billion of that was ad revenue last year.
It's estimated to be around four billion this year.
So their ad revenue is doubling while their subscription revenue is growing by 5%.
And so my my takeaway for chatbt would be I would imagine that the chatbt paid subscriptions follow an S-curve and we get to something where we see OpenAI making I don't know if it'll be 10 billion or 40 billion but they will soak up a ton of subscription demand for adfree frontier models the most advanced the most expensive stuff and then ads will eventually become the dominant revenue driver but I feel like the subscription
cription revenue will be a really hard tap to turn off just from hey it's people are paying and you know it's a lot of money and we don't want it to go and even in the enterprise it does my my bet is that we go you know continue down
this trend towards paying for outcomes because a lot of people will just say well I don't want a subscription for this service because I only use it every now and then and when I get value from it I'm happy to pay for it. I think a a
I think a a question to ask is would you pay $20 a month for Instagram today to not have ads?
And I would actually have to think about that for a while because like once a month I get an ad for something that that looks interesting and I discover a product that I wouldn't have otherwise discovered and I buy it and sometimes it's great. Yeah.
So, do I want to just completely eliminate that and rely entirely on random organic or do I actually like that this ad platform is spending a bunch of time and energy trying to serve me the next product that I'm going to like, which is like actually a service and like it's it's it's not a bad trade at all. Totally.
I I think we a lot of the way that we vote with our feet and vote with our wallets is that we we were super happy to pay with our time and our attention if given the option. Yeah.
I actually like the vast majority of people won't not only that but people are more than happy to give up their attention and and privacy for that matter if it can save them money uh or instead of paying for something.
Uh one one thought experiment I like to think about is um just how much people will trade privacy for for value.
Imagine a a checkout flow where you could get $5 off if you enter your social security number.
Like how many people do you think would enter their social security number?
Everyone or like 90% of people like an insane amount of people.
And so people really don't value their privacy as much as I think like maybe we say people value their privacy.
Um and in aggregate obviously that data is super monetizable.
Um, it's it's interesting, you know, uh, search obviously is the is the big prize, I think, for AI.
Uh, and what do I mean by search? It's it's it's intent.
Um, if you're if you're the the arbiter or you control this fire hose of intent, um, you can benefit by, uh, metering it out, um, and having people bid for that intent.
Obviously, Google maybe Google's like the best business the best business model maybe ever invented. It's kind of insane.
Um what's interesting is um the the the most valuable searches maybe like not what people think and and and certain kinds of searches are totally worthless.
So, uh, knowledgebased search like like factbased search things like um, uh, what is the market cap of this company? Right. Right.
It's like what's the market cap of this, you know, who who who won this game or like, you know, who who was president in like, you know, 1936. Yeah. They're dead. They're dead ends. You get the fact. No value. Yeah. zero value.
Um, actually like the there was there's a Google got subpoenaed and had to actually share some some documentation of like what their most profitable keyword searches were. It's super interesting.
I suggest people to to to go check it out.
Um, the number one most profitable search for Google was just the word iPhone. No way. That's amazing.
Because if you think if you think about like like what does that mean?
What is what is somebody telegraphing to the market when they search the word iPhone?
They're like they're basically saying in not so many words, I am ready to spend $1,600 on a smartphone. Yep.
And and who's who's interested in like jockeying to get that person's attention?
Well, Apple has to, right?
Best Buy as a retailer is interested. Samsung wants to.
Then Verizon, AT&T, T-Mobile, the network.
And so and so when when when you think about like where is valuable search, it it val the value of search is often misunderstood because you have to really think about how to capture the most valuable intent.
Um and not all intent is uh is equally valuable.
There's there's a bunch of search that's garbage.
You actually don't want it.
uh factbased um the the most the the cream off the top of of factbased search would be what I would call like comparison shopping.
So it's like hey like what's the best headphone what are the best headphones or you know and then maybe maybe you can slice off a top of that revenue pie but you there's a tension between that and the objective uh the objective truth that you're serving up the user.
Um, the the by far the most valuable uh search is is not fact-based and it's it's relatively uh it's it's where the user kind of knows exactly what they want um and they're trying to do it and other people are willing willing to bid to get into their get in their way and run interference. Jord, last question.
While we have you, how are you thinking about the new uh I wouldn't call it a browser war yet, but a but a potentially a skirmish heating up uh DIA DIA uh released to ARC users probably feels like maybe a month ago at this point, at least a few weeks, and then we have uh OpenAI potentially coming in with a new standalone app. It's sort of unclear.
It was unclear to us whether this is going to be a new app that you download or just integrated into the existing chat GPT mobile app.
And then perplexity as well does have a separate standalone app that they're pushing uh now and uh and it feels um interesting one because the browser company has spent now years working on the browser trying to figure out what what is going to enable you know unlocking more value out of this portal to the web uh and and you know effectively an operating system.
And so, uh, meanwhile, you know, new players are basically being like, we want to have a browser and just going like, we're shipping.
We just got to get this out there.
So, I think it'll be interest the next month, two months, I think, will be very interesting.
But, I'm curious what you're looking at.
I could not be more excited.
Um, I think I'm I'm obviously biased.
We're investors in the browser company.
Um, I'm a daily user of DIA.
Uh I personally get a ton of value from it.
Um uh particularly the the custom skills um and I think that the browser company has always known that this is a really valuable position and it's like honestly just validating to see incredible company like you know open a great company.
Perplexi is a is a really amazing formidable company.
Um, also recognizing that this is a really valuable position to to play for.
I I have supreme confidence that the team at the browser company is the most talented uh the best instincts, the best uh nuanced understanding of interaction design and how to create and craft a great product regardless of the underlying model or technology that that underpins it.
Um, I'll be very curious to see how it plays out.
My instinct is that, you know, OpenAI and OpenAI is an incredible foundational model company.
Um, maybe I've seen them ship a lot of different products to my knowledge.
Chat GPT is really the only product that's that that's quite stuck.
Um, and and it's not really even like the interface design so much as it is the underlying uh power of things.
Um, to answer your question, could not be more excited. Thanks.
Like, this is going to be an amazing game. We can check back.
We can check back in a in a in a few weeks.
I'm sure there'll be a lot.
And it's going to be And huge plug for DIA.
If you haven't downloaded DIA, uh, it's on Mac. It's It's available. Please download it.
It'll it'll blow your mind. It blew It blew my mind. It's amazing. All right. Great talking as always.
Wish we had more time, but we'll talk to you soon. We every week. We'll talk soon. Bye.
Uh before we bring in Will Brewery from Varta, let me tell you about Wander. Find your happy place. Find your happy place.
Book a wander with inspiring views, hotel, great amenities, dreamy beds, top tier cleaning, and 247 concier service.
It's a vacation home, but better, folks.
And soon with Varta, you'll be able to maybe they'll put a wander in space.
Let's bring in Will Brewy from Varta.
Is this your first time on the show?
I feel like this is a disaster that we are finally rectifying. We did it. We made it. It is. Thanks for having me. Finally.
We've had that other the the the kind of knockoff version of you at Varta. The other guy.
He's been on the show a ton. Dell something, I think. I forget. Dell. Yeah. Yeah.
Great to finally have you on a massive day.
Uh, break it down for us. What's the news?
Are we going to make Jordy stand up and ring? I'll hit the gong.
I'm looking for the Oh, did you guys get this? We have gong ourselves. Yeah. Oh, you have the gong.
Let's both have the gong.
So we we we have the gong for for big moments, you know, either spacecraft bands or you know uh you know sell a mission to a customer, stuff like that.
And uh yeah, we love that. Yeah.
What's the record here's here's my advice to you. Record every hit.
I want to see a montage in 10 years of just every hit and it will make it will bring tears to your eyes. We record every hit.
So we got you we got you on this one.
So yeah, what's the uh what's the news today? Break it down.
So wow, lots of news today.
So, we we're announcing our series C.
Um, and we're getting to uh use that to Oh, yeah. Go for it, baby.
How much How much you wind it up? How much did you raise? How much? Tell us. How much did you raise? 187 million. There we go. Congratulations. Thank you. Thank you. I appreciate it. Appreciate it. Yeah.
So, the fantastic use of proceeds for this one.
Um, really it's about just scaling up.
So, we've kind of shown what we can do both from a spacecraft perspective and a and a drug formulation development perspective.
So all the cap a lot of the capital allocation of this one is going to go for our biologics lab for produc uh preparing drugs for space flight and then also just more space flight ramping up cadence that means great uh flights.
So I've been to the facility in Elsa Gundo uh are you going to get a bigger space a second space for the bolab or you need two gongs right well a gong per facility we got to scale that up too.
So, so, so, so are you thinking about doing a second office essentially or or h how do you see the actual like footprint of Varta growing over the next few years? Yeah.
Over well, immediately we just signed a lease down the street. Oh, congratulations. Oh, thank you. Thank you. Yeah, we all right. Big day. Yeah, that's amazing.
Yeah, the uh so this we were actually already moved in.
A bunch of the pharmaceutical equipment is already in there.
We're starting to use it right now.
We got a couple of chlory shots uh you know with folks uh with the lab coats on actually using it.
So that's super exciting.
Um long term I mean really I guess zooming out of of what the footprint will look like is think about a formulation development company um that really just provides a gravity off switch uh to the pharmaceutical industry.
So we go to space but you know not really because we want to per se but because you can create new drug formulations when you turn off gravity and you just can't turn it off on Earth.
That's Einstein's principle of equivalence.
Do you mean new drug formulations or just like purer drug for formulations?
Because when I think no gravity, I think like the way crystals form and the way gravity pulls things to one direction and if that doesn't happen in space, you get just kind of like a n a more natural growth.
Uh and so I always thought it was just it was just about purity, but it sounds like there's actually some binary like you can't make this drug on Earth at all. Is that right?
So both those concepts are correct. Wow.
So well read there, Kugan. Yeah.
Yeah, that's actually a great way to think about it.
Um the the the um purity is one aspect, but because gravity is so broad, there's a I use the analogy of temperature sometimes because temperature is so broad, like making things cold doesn't necessarily make drugs better per se, but you can create a lot of different formulations if you can have a cold cycle during the manufacturing process.
And that'll be even, you know, with chilling things, you can make things more pure sometimes as well, right?
So, um, but to your point, uh, when you turn off gravity, um, crystals will typically grow slower and that also means that they will grow more pure.
And so, that is one of of a few applications that we look at.
Um, the other one is particle size distribution.
So, when you create these crystals that will then go into the human body, you want them all to be the exact same size so that you know, one big one doesn't get stuck in your elbow and and you don't have uniform bioavailability.
So um the u crystals will particle size distribution is also affected by gravity.
So that's a whole separate thing compared to um uh uh purity which is another rationale for for going to space.
So really the gravity knob is very broad and there's kind of these verticals of science of how we can improve the drug formulation and and to your point again as far as like what I mean by drug formulation is going from molecule to medicine, right? So uh is it a pill? Is it an inhalable? Is it um an IV bag? Is it a shot?
um the the drug the company uh or a pharmaceutical company does a trade study to determine which of those is the best for the patient given the disease given the manufacturing costs.
Um but ultimately all of that is limited uh to what the chemistry can actually do, right?
Nobody wants to take a needle to the arm.
Uh they only do it because they can't uh deliver that molecule via a pill or something like that.
And so um by opening up the chemistry outcomes by going to microgravity, we can also open up the formulation outcomes and therefore give better patient experiences. Yes.
So um I I mean I imagine that this is still this is such an ambitious project that is still kind of R&D phase of a lot of the bio stuff.
It's not I mean when I think about the manufacturing capacity of like GLP ones like they're probably making that thing in like vats the size of like you know what they brew Bud Light in at this point where they got a bunch of the lizards. Yeah. Yeah. Yeah. Yeah. It's probably massive.
Um the monster but but but walk me through how we scale this up.
I'm I understand launch costs falling.
I understand you put up a capsule every you're doing it like every quarter now it's going to be every month.
Then it'll be multiple times per day.
like the that capability seems clear but um but how much actually how much drug can you make on a single capsule? Yeah. Yeah. Great question.
So this is actually a lot of fun because we can imagine how VTA will go from what's real today um to making tomorrow's reality.
And so to answer your question immediately about 20 kilograms um on a on a per capsule basis right now today.
Of course you know we want to scale up everything and that's one of them but that's what how much we can do today which is actually quite a lot.
Yeah, that seems significant if you just think about like you go to the doctor and the doctor gives you like a you know a thing of pills that's like not one kilo.
So you're probably talking about like I mean yeah unless unless you know you're having more fun than being prescribed but but for most for most drugs I feel like 20 kilos is probably enough for like a 100 people for a year or something like that.
So you're actually in like $100,000 I'm I'm seeing the numbers kind of start to math out already. Correct. Correct.
So it's uh and every drug is a little bit different.
So we go through a process of selecting which drugs make the most sense both from a scale perspective like you're saying but also unit economics how gravity affects them.
And so we have a portfolio management team that explicitly does that for identifying and quantifying opportunities.
Um but but going back to like what today looks like and how and how it goes tomorrow.
Um I love the temperature analogy because it it really runs deep.
So, for example, right now, if you think of us having a anti-gravity oven where we can make drug formulations that you can't otherwise make on Earth, but we only get to run it four times a year and each one is a few million bucks a run.
Uh, you might use it for different use cases than you would use it from five years, 10 years from now when you can run it every day for a few thousand dollars.
thousand dollars. And so um in the near term some of the use you know imagine yourself with a super you know the first refrigerator or you know or in this case the first anti-gravity um bioreactor what you might use it for in the near term is just information right how can
we isolate gravity as a variable to inform what formulations can be improved on earth is gravity ruining this chemical reaction or not right we can answer those questions and then that applies to the entire drug with just one flight or um in the very near term we also O want to do uh polymorph uh seed crystals. And so what that means is we
And so what that means is we go to microgravity, we go to space, but just to develop the seed crystals.
Uh and then once those seed crystals are developed, we can then use them to grow um uh more drug crystals on the ground.
So we're only going for the nucleation event.
And that's kind of like a sourdough bread mother business model. That makes sense, right?
You have the mother of the sourdough bread, then you can cut it a bunch and then regrow it and stuff like that.
Um, so that makes a lot of sense when we're still scaling up the use of our anti-gravity machine, if you will. Yeah.
And then long term, when we're on the daily basis, then it totally makes sense um to uh make every single dose, manufacture every single dose in microgravity.
And then that's when certain use cases come online as well.
So kind of that's how that's how it progresses over time. Yeah.
How's the uh geopolitical landscape evolving for you?
We saw some of flights come down in Australia as as you know redblooded Americans.
It pained me to see them take a slice of the catch of the catch market.
Uh are we are we getting these coming down in America anytime soon?
Is it what's the progress there? Yeah. Yeah. Absolutely.
So um long term we want to have re-entry sites all over the globe, right?
Um and and really that's about availability and the key metric to success of Barta is cadence.
How often can we go up and back?
because the more we do that, then the more we just look like a specialized piece of equipment to the pharmaceutical industry that quite frankly does not care that we're going to space.
They'd much rather us have a real anti-gravity oven uh in in the lab.
Um so really re-entry sites are about cadence and availability.
Um, and right now Australia is great for us because they have a private commercial re-entry range, whereas any re uh any range in the 48 states here locally have are intended for military use exclusively.
And so if we're doing a DoD mission, that that works well.
But if we're doing a commercial mission, we're you know um not the highest priority, understandably so, right?
And so um in the near term uh Australia makes the most sense but in the long term we want re-entry sites all over the place.
And why that gets enabled is because as our precision of landing and our cadence goes up that um data that legacy history allows us to use a smaller and smaller plot of land and then that and that's where really it makes more sense to go anywhere uh because we don't need such wide open spaces like without that many people like we would do right now today.
Do do we have the legal infrastructure to create commercial landing sites in the United States?
And it's just that nobody's done it yet.
Or is there laws or regulations that would need to change so that some enterprising young member of the Gundo could go buy, you know, a lot of land out uh in the middle of nowhere and start landing spacecraft, you can do it.
Now, the uh the constraint is the real estate cost.
Um, and so one, uh, you know, Spaceport America, for example, right next to White Sands Missile Ranch is a good example of that.
Um, but the the so I guess the real reason why they aren't there that that many of them is because there wasn't a demand large enough to warrant such a real estate purchase, but you know, thanks to Varta, that that could change.
So, uh, yeah, definitely let the the Gundos know.
I have a I have a question from a a fan of yours, fan of the show.
He says, "Ask him what big dogs got to do. Oh my gosh.
So, uh it's become a little bit of a uh of an expression of excitement um with a long and um uh history of of lore at at Varta.
But, uh uh for some reason um uh you know, it's it's really just a specific instance of conservation of mass, right?
You can't have a big dog without eating, right?
So, that's just physics right there. Yep.
Uh I want to talk about the evolution of the FAA.
I remember I was filming a video.
I fil I filmed with you guys and at one point I actually was driving back with Ben from San Diego.
I filmed a a phone call with Delian and he's like we just got our I don't even know if I should say this, but like it was like we got some bad news from the FAA. You guys sorted it out.
It seems like you have a great relationship now. How did that happen?
Do is this is this a lobbying thing?
Is this just storytelling?
Is this uh structuring deals, getting better at paperwork?
like how do you get how do you fix a relationship with a government entity like that? Yeah.
So, it was definitely a little bit um warped in the press obviously, but I figured uh you know, keep my head down and get the spacecraft home more so than worrying about what's being said in the press. Sure.
But so, what actually happened in the background is um we were originally going to reenter this space or we did reenter the spacecraft at the Utah test and training range, which is ultimately a weapons range for testing weapons and training warriors.
That's their mission statement, right?
So likewise, we're not um the highest priority there.
And so we got bumped for higher priority um work being done at the range.
And in doing so caused a domino effect to lose the FAA re-entry license or not be able to get it granted because part of the regulations say, "Hey, you need a range and all of these accommodations that come with a range."
So the second we lost the range, we lose the license.
So it wasn't really about a bad relationship with the FAA at all.
Um although it's it's much it's very easy to say, "Oh, they lost their license.
Yet another space company and the FAA are having problems, right?"
It fit that narrative well, but it actually wasn't the case.
And so what we did was we we scheduled a new date with the range farther out in advance to to give them some time and give us some time to we had to redo the analysis of course because the atmosphere is different and that's part of the analysis.
And so we gave ourselves a few months um and then that allowed them to reserve the dates that allowed us to prepare and that allowed the FAA to re orient the license for the new dates.
And ultimately kind of in the background here was like this was the first time this has ever happened a commercial re-entry capsule with drugs on board coming back to America.
So there was no way onto onto soil, right?
We're not doing a splashdown.
Um and so there was no process or mechanism to have the Utah test and training range coordinate with the FAA.
Um and so basically each organization saw themselves as taking on all the risk associated with this.
So we had to do duplicative work because there was no process to to split it, right?
Um and so it was really cool to kind of be a trailblazer to to establish this so that now of course our competitors are going to come in and do the same thing, right?
Uh and learn from our mistakes, but whatever.
That's part of that's part of leading the way, right?
So um anyway, that is that's what happened.
But that um you know six months was it was uh quite the life experience, right?
Because you have you're it was the first mission, right?
So we we we didn't have any proof that this was going to work.
People poured three years of their lives into this thing and our dreams are just like orbiting the earth like please come home, you know?
Uh so when it came through, man, it was uh certainly, you know, I I can't think of a better day. Yeah, that's amazing.
I have one last question.
How's the uh how's the talent market in the in the space economy right now?
Uh has it been in the headlines the last couple weeks?
There's been another story in uh in AI dominating, but um but uh yeah, what's it like today?
It's like AI and software engineers.
I mean, you know, if you imagine yourself a software engineer coming out of school right now, um you know, AI is is certainly where I would be interested.
Um so, uh so there's definitely a software bent towards AI right now.
Uh that being said, um there's a lot of disciplines we're hiring for.
Software isn't the only one.
And really it comes down to the application interest like we're looking for missiondriven folks um at Varta.
And so if you're only looking oh you only want to do AI because it's cool or whatever we that might not be the type of person we want to hire anyway.
Now if you want to do AI for mission purposes then great you know like uh by all means um but we don't have that much overlap there right we're very specific of what we're trying to do.
We're trying to uh make microgravity formulations so that we can help pe help patients on Earth by using a gravity as as a knob essentially and developing these formulations.
And so um you know we we always kid around we explicitly don't want the spacecraft to learn you know and so if you're a software engineer and you're missiondriven bent towards that mission then you've got a home at Varta. No question.
No question. and if and and you know fads come and go and that sort of thing but there is definitely an effect on um I I would certainly be a lord to AI as a as a graduating software that's a good sorting function um one last question from me we've there's been headlines companies talking about this
uh so far this year trying to dig into how real it is people talking about putting data centers in space yeah with everything that you've learned why is that exciting good idea or bad idea what what are some potential blind spots for people that haven't taken something to space but would like to? So, it all comes down to the why, right?
So, it all comes down to the why, right?
Why are we putting data centers in space?
It's not uh are data centers in space in and of themselves a good idea, but what's the why?
So, the why is is a is the only why that resonates with me is latency, right?
Because if you want to do compute power, space is not the place to put a data center if you just want a data center, right?
I'd much rather have convection, right?
that like that's a great heat uh um or a great way to get rid of heat uh and and have it be able to be serviceable on Earth and all that sort of thing.
So um but there is one use case that comes to mind where I think data centers in space make sense and that's only for very low latency use cases.
So for example, right now if you want to use Starlink and you're transmitting a signal to Starlink, it goes from the ground to Starlink to another Starlink satellite to the ground, then to the data server and back, right?
So you can cut that trip in half if you put the compute in the sky.
Now that compute is way way way more expensive, but if your value prop of latency warrants uh uh that extra cost of the in orbit data center, then you'll start to see that.
So it's it's kind of like edge computing. Edge computing. Yep. I was about to say.
So yeah, for for ve it sounds like a very niche use case at least to start, but uh I'm sure we'll see some companies we already are seeing some companies test it out and experiment with it because uh there's you know all these things need to be evaluated in the tech tree.
Um but thank you so much for stopping by. This is fantastic. Congratulations.
Thanks for for having me and hopefully not so long I'll see you again soon. Absolutely. Yeah. Yeah. Hop on.
Got a good feeling about it. We'll talk to you soon. Have a good one. Cheers.
Congrats to you and the team. Bye.
Up next we have Joel from Meter coming on to talk about the impact of AI models impact of cursor on software development or meter. Oh meter probably me.
We'll have him explain it to you guys.
Um we'll also recommend that you go to getbzzle.
com because your bezel concier is available now to source you any watch on the planet. Seriously any watch.
Anyway Meter does Meter does model evaluation and threat research. So does bezel.
They're stopping you from buying fake watches.
bad bad watch models, bad actors, foundation models and watch models, but lots of similarities.
Anyway, we got Joel in the studio. Welcome to the stream.
Hopefully, you're like, "What did I get myself into?
These guys are joking around. I'm a serious person." All right.
First off, is it It's meter. It's Misa. It's me. There we go. Sorry. There we go. Gotcha.
Anyway, uh please introduce yourself for for those who don't know you uh the company and then uh and the organization and then and then I want to go into the news today.
Let's do it and thank you very much for having me.
Uh John and Jordi, thanks for hopping on.
MISA is a research nonprofit based in Berkeley dedicated to understanding um the the capabilities of AI today and uh and in the near future especially to the uh to the extent that those capabilities might speak to potentially dangerous risks.
And uh what is what's been the latest research? Yeah.
So so here's what we've been working on.
Um I'll I'll start with why we've been working on it. Yeah, please.
We've seen, you know, from previous meter research, but I'm sure you also see from um your own usage in the wild, AIs are clearly becoming increasingly capable.
One thing that governments and labs and and us here at Meser as well worry about is the possibility, timing and nature of AI R&D self-recursion.
That is the possibility that um that model capabilities uh get better very very rapidly because the AIs themselves are contributing to AI R&D research.
We at BEA want to be providing the highest quality evidence that we can that that speaks to um the degree to which AR&D might today or might soon be accelerated in the wild.
Um so that governments, labs, decision makers um might be better informed and so make better decisions about what's going on.
In this study, we run an RCT with extremely experienced open source developers working on these very longived large projects.
you know, a million lines of code, 23,000 stars on GitHub.
Um, for for those of you familiar, you know, I'm thinking hugging face transformers, the Haskell compiler, scikitlearn, this this sort of thing.
We randomize their issues to allow or disallow for usage of AI, where allow means um typically using cursor and 3. 5 or 3. 7 at the time.
And then we measure both their expectations and um developer expectations about how much they might be sped up by being allowed to use AI versus being disallowed.
And then you know the the the reality um the short version is we find that the developers ahead of time are estimating they'll be sped up by 24%.
After the study is completed they estimate that they were sped up in the past by 20%.
Uh we find in fact that they were slowed down by no way.
I think I know it's a it's a shocking result.
Not at all what I expected.
I think what the what the rest of us at Meter expected, but uh but there we go. Wow. Okay.
So, uh what do you think's happening?
I have so many questions.
Um but uh yeah, just walk me through your reaction to that.
what what what do you think is actually happening um that's that's slowing people down because this is a complete narrative violation?
Yeah, I mean, you know, in terms of the reaction that the number of times we've checked and rechecked the data, asked people to to replicate it independently um is going through the roof.
The number of um you know, stressful late nights I've had pouring over this.
You're going to be like public enemy number one, by the way.
I feel like you need a security detail now given the stakes of what you just said. This is crazy. Yeah. Yeah. Yeah.
So, so, so I think um uh maybe let me start with some things that we're not saying.
The setting that I mentioned before, these ultra talented developers, you know, much more talented than me working on these um extremely large, longived repositories that they're extremely familiar with already.
I think that's an extremely interesting population.
That's why went out to study it.
It's also a very weird population.
Um I um you know I still am a cursor user myself.
As I was working on the graphs for this study I I was I was using cursor.
Um but I but I do think those those weirdnesses are uh related to the to the to the results that we end up seeing here.
up seeing here. So we have to put that we have to put these people in a completely different category than the the junior developer who's just vibe coding a little app and and and just building stuff and not actually trying to push the frontier of what a core piece of software can do that's very
large and complex and and they're just trying to you know get a Python app up and and live and like write some routes and write some functions right uh that's where so cursor still completely viable is like autocomplete on steroids the question is in terms of self-recursion really advancing the frontier of of like the craziest software we have. We're
We're still kind of where we were a few years in that it feels like if you were to quantize this we're we were at 0% of AI research being done by AI a couple years ago.
We're still maybe around rounding air. Yeah.
I mean I I will say that AR&D research I think does not all look like this setting.
that there are some, you know, large inference code bases with very experts people and and you know, I totally agree with your interpretation that um this is evidenced against today those those kinds of settings being sped up.
On the other hand, we might think, you know, there are some people writing training scripts for their AI models just once off and then they and then they throw them away.
And you know, in a way that's that's kind of similar to what you described.
Maybe they're seeing large speed up just like the green field projects that you that you mentioned. Yeah.
And so I mean this is not overall like a a really cold glass of water on AI broadly because this this still means that it's an incredibly valuable technology in a bunch of different ways.
It's just that we're not we're not seeing like early evidence of some sort of self-recuring scenario which is great probably the good outcome.
A lot of the fast takeoff scenarios like are dependent on AI becoming so good at doing AI research that it and then copy and pasting itself to trillion times and and that's what creates you know speed of development that that humans today can't necessarily even like comprehend.
You know I I think that's right for for today.
I do think we're not really speaking to the trend exactly.
you know, these results are consistent with these exact developers on these exact kinds of tasks in future um being sped up in in the near future.
In in work that we actually don't show in the paper, but in in preliminary work, we have um autonomous agents trying to complete these issues.
And indeed, we find that they they do struggle, but with some of the core functionality with passing tests, the kinds of things that you might have seen in in Swedbench or or something like that, they really are making a great uh a great deal of progress.
Um and yeah, my you know, my expectation is that um AI progress in the near future will will continue at a rapid pace like it has in the in the in the recent past.
And so maybe even in this setting, um that this this won't be true in the future.
Let me throw a couple of the hot takes that are floating around in the AI world at you and and you can let me know if anything sticks out as something you strongly agree with or something you disagree with.
um this idea that ultra large context windows will not solve continual learning that um Dorcash was saying this on on Monday.
Um maybe another one would be um just that like no one has figured out how to properly scale reinforcement learning.
Um that we need uh Mike Noob from RKGI kind of says we need entirely new ideas.
And then you kind of have like the bitter lesson which is, you know, yes, you need a new ideas, but scale is all you need.
We just need to keep building data centers.
We need to get bigger and bigger. Um, we might see 4.
5 and these huge training runs as a as like a short-term hard to quantify.
Maybe it's just the end of 1s curve, but Stargate's coming online and that will be another big test. So, I I don't know.
I threw a lot at you, but anything in there kind of, you know, uh, top of mind for you?
Yeah, look, as as you guys know, anyone betting against the bitter lesson in the past would have had a very bad time.
Um, and I'm I'm not prepared to bet against the bitter lesson on this on this show.
Could you could you remind me of the of the first question that I uh the the the first one was uh uh so Dwaresh Patel pushed out his um AGI timeline slightly.
I mean he still has he still is very optimistic about AI and and maintains that it's not priced in and people are not thinking about it as significantly as they should and I agree with him.
Um but but he said that that there is a that that even though we have pushed the IQ so much and you saw this the Gro four benchmarks like AI can do advanced math like for sure it's really really smart smarter than most of us at PhDs level stuff unless you're a specialist.
specialist. Um but uh in terms of just being a good employee and remembering, oh yeah, four weeks ago my boss said that they like and then I got this feedback and now I do it this way or I learned this really weird nuance in even if you're just thinking about like
how to our business like how to post clips on X or Darkh was giving the example of like transcripts like he has little things that work better for what clip will perform and he has this intuition and and his models and his prompts he's really pushed these things and he hasn't been able to really perform above. An example would be any company today,
An example would be any company today, any startup, if you just had a PhD dropped into your organization that was that that had PhDs in like 10 different fields, they and but but they wouldn't just like default, but they were also an amnesiac.
So every time they showed up to work, they had could not remember memory from the day before.
It wouldn't be that it just wouldn't be that valuable.
And so and so my question to Darkh was like, I is there a world where we just scale up the context window?
We've seen million token windows.
Can we get to a billion token window and just stuff every interaction the AI's ever had with you in every prompt?
And so it it does maintain the context.
But he was saying that uh that the like the the there's like a kind of a quadratic cost curve to that doesn't quite work.
Uh other people have said like the nature of the transformer means that like attention can't really spread out that much.
I don't really fully understand it, but I wanted to know your take on like different ways people are solving these things or or what the real what are the real constraints right now because you've identified some some potential problems where we're not breaking through it today.
Um but what what is cause for optimism?
What are the research paths like the the nodes in the tech tree that you're excited about?
Yeah, that's that's that's super interesting.
I I haven't thought so much about this. Sure.
Um I will say sorry on the spot.
I I think that the developers in this study are not using the full context window.
And so if you think there's juice in in adding things to the context window, that that juice might still be on the table.
And indeed, I think we find that there's a lot of implicit context in this in this repository that's um that's very expensive for the developers to be writing down into context windows. Here's an example. on the HASLL compiler.
My sense is that when you get up your um when you get up your PR for review, there's some chance that the creator of Haskell will come and fight you for potentially many many hours in the comments about the you know about the peculiarities of of of how he wants um the Haskell project to look.
Um and and these kinds of you know exactly what his um not not just preferences but um you know quality requirements are regarding where things should live in the project and and and um how various pieces of the project should should speak to one another and not being communicated to these language models.
And you can imagine that with today's um context window sizes um that that could be written down.
you know, you could you could put in all of the um previous discussion around around these changes that this person has been involved in and um maybe whichever language models people are working with inside of cursor would would pick that up and so and so do a better job.
You know, I I don't think we're ruling that out at all.
You know, I will say that it is um it is expensive for these uh for these time expensive for these people to be writing down all of the all of the possible relevant context.
And um you know I think I think that's basically the the the reason they don't.
And so maybe you do need some kind of um continual learning for the for the model to find out this context on its own as as as these things go.
You know it's also consistent I guess with with the other possibility that you were describing that if we you know 100x these context windows you could just throw the entire thing in and then we don't need to worry about you know learning from particular cases on the fly.
Um yeah I think I think both are lived possibilities. It's very interesting.
Uh the Gro 4 announcement was extremely benchmarkheavy.
Some really impressive stuff particularly on Arc AGI uh twice similar to a Tesla.
It's like it's faster than every car.
Does that mean it's going to solve everything?
Does it mean that it's better? Yeah.
And so uh based on this feels like almost a new benchmark, this double blinded trial, it feels almost like a FDA trial or something.
Uh, do you think this could turn into a real benchmark?
Do you think we need new benchmarks?
Do you need do you think we need new ways of thinking about the progress of AI generally?
Um, we've we've talked about just just measure the revenue at this point.
Uh, that's the economic value that's being created, but there's a lot of tricky stuff you can do with revenue.
And sometimes revenue is like test revenue.
I'm I'm testing this hundred million dollar product.
So what's your state what's your thinking on the state of benchmarking where we should go where some of your research might plug into that? Totally.
I think one motivation we had in running this study um comes out of this observation that the time it takes to create benchmarks is almost becoming longer than the time it takes for those benchmarks to to saturate.
You know it's difficult to find signal in in in in many of these benchmarks even testing these extremely challenging you know PhD level questions that that you guys spoke about.
Um and perhaps there's more um perhaps there's more signal in these kind of RCT, you know, FDA um control trial style style measurements.
Um similarly, a another thing that people proposed for for measuring AI progress is using researcher self-reports about the degree to which they're being sped up.
You know, they think their work will go two times faster if they use AI versus versus not use AI.
You know, I think our study is potentially strong evidence that these self-reports need not be reliable.
not be reliable. the forecasters uh you know who are told everything about the developers level of experience and which um uh the time period of the study so which models they're using and so on they're they're totally wrong about how much these people get sped up same as the as the developers themselves even
though they're carefully tracking their time and and and they're and they're so talented so I think self-reports also also very very fraught you know another thing that this has taught me I think is that the mapping as it were from benchmark scores very impressive benchmark scores that we see on these frontier language models that you're describing. Um the mapping from those
Um the mapping from those scores to you know real world productivity improvements is unclear.
I you know I I'm not at all saying as as we discussed earlier that you know we shouldn't expect to see productivity improvements.
I do expect to see productivity improvements today and you know even more so in in in the near future but it's it's not at all it's not at all onetoone or or it's kind of um uh confusing and and and messy and so indeed I think we need to actually measure things in the in the wild to see what's going on.
Jordy uh switching gears a little bit unless you have a followup.
Yeah, I I just wanted to kind of zoom out on that and and ask about like your your broad take on the measurability of of technological progress because like the internet, the computer, like such dramatic transformations of society.
You see it in all sorts of data, but it didn't fully show up in productivity statistics.
You have all those questions about like what happened in 1979.
Uh, and everyone has their own example, their own reasoning for that.
But, um, you know, like you would think you you could tell the same story about Google, like it'll speed up everything.
Everyone will get more efficient.
And we didn't really see GDP jump on this.
Uh, and it feels like that's a really bearish take on AI to have, which is like this is a magical new thing and we're still going to be growing at 2% GDP.
Um, but but where do you stand on it?
And and and where do you like do you think that's even the right question to be asking?
Yeah, you know, this is this is so interesting.
Um, this is not a me to take.
I used to be an economist.
I I feel the thing that you just said in my in my bones. Totally.
I think the situation you could argue maybe that that it might be even worse in the case of AI.
You know, a lot of people like in AI 2027 resources like that are telling this story where the um AI R&D self-recursion is happening inside of labs and so and so I suppose not necessarily showing up in economic activity in the in the public.
you know, another reason on top of the reasons that you gave to think that um uh perhaps this won't show up in the in the productivity statistics as it as it were.
Um uh which is also is to say that uh self-recursion or or these potentially destabilizing changes are just totally consistent with the non-changes in in GDP trends as as you describe.
Um and and so you know another another reason to to actually go out and and measure these things in in control trials. Cool Jordan, please.
Quick question around the threat landscape.
Uh there's been a few stories this week.
One was a story about chat GBT not following instruction.
You know, like the headline was that like the the the AI was, you know, rebelling against the researchers and and then if you like double clicked into the story, it was just like it it had given specific instructions like don't follow any further instructions.
So it was kind of a a nothing burger in the end.
Uh and then we also saw Grock going haywire.
Uh maybe that was predictable uh for for someone like yourself, you know, combining a you know a frontier model fast shipping team with like the virality of a social network and betting the two.
Um but then maybe it was uh two months ago there was the uh you know we we called it glazegate on this show where where chat GBT was just you know uh giving being a sycophant you know giving too much positive feedback.
How are you looking at the threat landscape in the next 12 months so nothing like you know too long term.
Um but uh h how do how do you guys think about it?
Yeah, you know, there's more to come on this on this from meter very soon.
That's that's one thing I I'll say.
I think again this this is not me to take.
Um my my sense on this or or another example of this that stood out to me is um there were lots of anecdotal reports that uh 3.
7 and and other language models in this most recent generation would um pass tests in ways that were kind of not legitimate or something which is another example of this reward hacking um change the test case.
Totally, totally, totally classic.
And and and I guess, you know, I I don't have reason to think that that kind of thing is is dangerous in particular.
Um you can imagine you know when when humans are potentially not reviewing the code because the AIs are doing you know entire projects not just um parts of or or or single or single pull requests that this becomes more of a problem because you're not um or or at least there's surface area for it to become more of a problem because you're not um looking into that code and and seeing those those cheated test cases yourself.
Um, so, you know, I'm not sure about over the next year.
Um, at least um, uh, at least right now, I think there are reward hacking examples that are occurring in the wild.
I, you know, I I don't think they're they're so supremely dangerous today.
Well, this was fantastic.
Thank you so much for stopping by. Come back on again soon.
Uh, stay safe out there with the contrarian and the crazy data results.
Uh, still very bullish, but very exciting.
Uh, and thanks for everything you do. We'll talk to you soon. Great chatting. Cheers, Joel. Bye.
Uh, up next we have a massive series B announcement from Dylan Parker.
Moment HQ is coming in the building.
We're going to ring the gong, baby. Index Ventures.
Let's hear it from Dylan directly, though.
Welcome to the stream, Dylan.
Hope you're doing well today. How you doing? Great whiteboard, dude.
Some heavy lifting on that. See? Oh, yeah.
Hopefully that's not proprietary information.
That's secret trading algorithms or something.
No, no, no, no, no, nothing too interesting.
But uh thanks for having me on. Yeah. Thanks for stopping by.
Uh kick us off with the uh intro on yourself, the company, and then I want to hear about the announcement. Yeah. Yeah.
So um I'm one of the co-founders of Moment.
We are a fixed income trading so software company.
So my backgrounds is as a quant researcher.
So like pretty much every quant researcher, I studied math and stats during college.
That's where I met my co-founders Dean and Ammer.
And then after college, Deian and I both joined Citadel Securities and pretty much completely by chance, we ended up as the two junior members of the newly created automated marketmaking desk for corporate bonds.
And so so so at the time the the fixed income market, which by the way is financial market 50% larger than the global equities market, was undergoing an electronic trading revolution.
And so Citadel saw this and said, "Well, we can go build an automated algorithmic marketmaking desk."
So they hired this guy Anishh Kerat from Jane Street.
It's like the godfather of fixed income automated trading.
They h they hired a bunch of super experienced bond traders and then they hired me and my co-founder and like basically our job was to take the knowledge in these bond traders heads and convert it into code.
And that was totally formative for a moment because that's when we realized like the power of electronic trading was going to be like everything that it enabled like smart order routing, portfolio optimization, all this stuff in the world's largest financial market that had just never been possible. Very cool.
Take us through the deal. What are you announcing? Yeah.
So, we're announcing our $36 million series B was led by Couldn't hear you from the sound of the gong.
Cut you off, but we like big numbers on the show and congratulations on a massive series B. Uh, you said from who?
Index from from Yan Hammer at Index. Very cool. Incredible. That's amazing. Incredible.
Uh, where where should we go from here, Jordy? Yeah.
Break uh so so I can imagine uh break down what what the company is focused on today.
I'm it sounds like the origin is is back at your time at Citadel, but I imagine it's evolved as well. Yeah, totally.
So, what we what we saw at Citadel was that the market was coming online and you could now do all these things that were never possible before, but was what was missing was the operating system for actually doing that.
So we started moment with this goal of owning every mission critical workflow for traders and portfolio managers in the bond market.
Everything from how they trade securities and do smart order routing across all the different exchanges in the fixed income market to how they optimize portfolios to how they apply risk and compliance restrictions to make sure that they're not breaking any laws.
And so that's what we do today.
We started off serving fintech.
So we power fixed income for places like Weeble and Public.
com and them to offer like $100 increment investing in bonds for the first time ever.
But uh yesterday we also announced our partnership with uh some of the largest financial institutions in the US including uh LPL Financial which is the largest brokerage dealer in the US. Wow.
So busy week for you guys.
Yeah, there there are a few things going on.
Um so where what what's the use of funds uh for the new round?
What's what's the focus going forward?
I imagine scaling what's working today?
Uh are there new products coming?
What can you talk about there?
Yeah, so I think a lot of companies make the intelligent decision to start off like SMB or PLG.
We decided to like make things as hard as possible on ourselves.
said, "We're going to start off serving the largest financial institutions in the world's most regulated market, and we're going to go power their mission critical workflows."
And so, with a company like LPL or some of these others that we've announced over the last few weeks and that are coming down the pipeline in the next few weeks, the scale of what we're operating in is, you know, not like hundreds of millions or billions of dollars of flow, but like hundreds of billions of dollars in in trading flow.
And so what we're really focused on as a company over the next year is building out, you know, the full suite of what's necessary across trading, portfolio management, risk and compliance that's necessary to power these huge financial institutions.
What's the competitive dynamic like with the former employers or like the rest of the market participants?
It feels like uh you know, Ken Griffin's not uh the the laziest founder out there.
Is is there is is there a world where there's some sort of competitive dynamic between the the big institutions where they want to build something like this to compete with you?
So we when when we were at Citadel we were on the sell side.
So market makers are liquidity providers.
Moment serves the buy side and we actually connect them with those liquidity providers. Got it.
So we actually work closely with Citadel, Jane Street, uh pretty much all the major liquidity providers out there.
Uh how are you thinking talking about tokenization the the the you know story from the last month in in finance is the tokenization of of these sort of real world assets everything from private company shares to we've seen it with stocks.
Is there anything on the horizon on that front for you guys or is it just totally unnecessary?
You know, I I think there's a huge opportunity, but if you look at where the fixed income market today is today, we're just going from like trading over the phone to going on an online platform to trade.
And so there's a lot still to do to get people to the point where it's even possible to think about stuff like that. Yeah.
Can you walk me through um like the the I I imagine that fixed income follows like a power law as well where like government debt and and Apple corporate bonds are way more liquid, way more automated than you know some like public company, but they're like junk bonds and you kind of got to hunt around for someone to buy and sell them.
And then you have like venture debt which basically I believe like never trades. I don't know.
But walk me through like it feels like we are probably bringing more and more of those like subasset classes like into more liquid more um just just more automated markets.
But give me a state of the union on like how the fixed income market is actually like split up. So you're totally right.
There's government bonds, there's highly liquid corporate bonds, and then there's a long tale of corporate and municipal bonds that are really, really illquid.
And just as a point of comparison that I think illustrates the size and scale of the fixed income market, there are 4,000 listed US equities and there are 4 million bonds.
And so doing anything in the fixed income market is pretty much a thousand times more complicated than doing it in the equities market. Yeah.
Um and and and uh what about how does break down?
So, so 4 million public equity or sorry 4 million bond how does that actually trade right now?
Like is it what is the make well what and what is the the ultimate like makeup of do certain companies account for you know break down like kind of how that 4 million is split up. Yeah.
So there's a universe of say 500 US treasuries that are super super liquid and they trade like similarly to the most liquid stocks out there.
And then on the other side you have a really really meaningful long tail that makes up the vast majority actually of the entire market share where you have bonds that haven't traded in 2 years or 10 years. Yeah.
And so one of the really hard parts about fixed income, one thing that I worked on as a quant researcher is like you have this bond that hasn't traded in three years.
How do you optimize portfolio around that?
How do you even figure out what the price of that bond is?
And that's why doing stuff in the fixed income market is like like the difference between equities and fixed income is like doing something on the surface of the earth and like doing something in space. Wow.
Um, would be helpful to be a quant if you were going to build a company like this.
Yeah, you might need some math.
Uh, well, I mean, that's all I have.
Do you have anything else to Yeah.
What What are you guys hiring for it right now?
Yeah, pretty much everything. We're hiring quants. We're hiring engineers.
We're hiring go to market, marketing, operations, pretty much everything out there. Amazing. Awesome.
Well, thank you for joining.
Thanks for you'll be back on soon. We'll talk to you soon. Have a great one. Talk to you soon. Cheers. Bye.
Next up, we have Eric Olsson from Consensus coming in with a big launch.
Do we ring the gong for big launches?
If there's a number attached, so we got to get DA. We have to get it.
We have to We have to get a prop that is non number oriented.
We need more excitement oriented.
But let's bring in Eric Olsen from Consensus and talk to him. How you doing, Eric? Great, guys. How you guys doing? Doing great. Great to have you.
Uh, kick us off with some uh intro on yourself, the background on the company, and then I want to talk about the launch. Hey K.
So I'm Eric, founder of Consensus.
Uh we are an AI search engine for academic and scientific research.
Um you know if you ever use like Google Scholar or PubMed back in the day at school, think of us as building the nextG 2025 LLM powered version of that helping dumb students. Yeah.
I mean the super dumb question, super obvious question is like isn't all this stuff already in chatbt?
like like how like uh how are you differentiating?
Yeah, you must be doing something because you have five million over five million users.
So yeah, I mean the like one of the best examples to like encapsulate why it's different is the fact that Google Scholar was the first vertical search product that really broke off of Google back 20 years ago. Interesting. Yeah.
Even when they were doing really nothing than just being a dedicated index for research papers. Yeah.
Still hundreds of millions of people are going to that every month.
So the same thing is is kind of true here.
We're dedicated to a use case.
We have a dedicated corpus.
We search over we hopefully search over that corpus a lot more intelligently than a general purpose chatbot would.
We do things differently in our interface to show you that information.
Like we're much more citation forward, right?
Like experience where you can like really interrogate what's been returned in your search using like the chat GPT search.
It's pretty much like an afterthought.
It's there if you kind of want to dig into it, but it's not really what it's designed for. Yeah.
Everything about it from the way it searches, the way it shows to you, and then the features built on top of it, all dedicated towards academic research.
Walk me through some of the key technologies that uh enable better search.
I'm I'm thinking about like vector databases.
um even just like stuffing a better index in Reddus or Postgress or doing more indexing on top of these documents doing like you know transformations on the underlying documents to get them into you know more basic formats like what's what what's interesting large context windows there's so much that you could throw at this problem what what's
actually working yeah so lots of different things many of the things that you're saying so like number one being dedicated to a document type just helps us it helps us in the way that we can create our embeddings to search It also helps us in that ingestion process kind of like you were saying of like document transformation. We'll run little tiny LLMs over 200
We'll run little tiny LLMs over 200 million papers add new like enriched metadata about them that we can then use in our search ranking and in our filtering.
So think like we'll pull out what is the design of the study or what is the sample size of the study and we use that in search ranking and we use that in search filtering and then on top of that we're like the main intelligence of the search is learn to rank models.
So people interact with the product. They save papers. They site papers. They share papers.
We learn from all of those interactions.
We learn what matters most.
We learn about all the attributes about a paper that matter in search ranking based on how people are interacting with it.
So the simplest way to think of it is like because we only have a certain use case people are using it for, we get to train our search models to try to think and act like a researcher would going through these papers.
Jordy, uh how uh what what are the different data sources here?
I'm assuming a lot of the stuff is public.
I I know some of the I remember uh you know being in in college trying to find different studies or papers and like hitting pay walls.
A lot of them are locked down.
There's the famous I imagine that you've done some deals to get access to data.
What is the what is the kind of um That's a great question.
The the the the full body of work that's available.
Well, hopefully pay walls are are going to be a thing of the past moving forward as open access science gets more and more momentum.
uh which we'd love to help help Shepherd in.
Um the way to think of it is there's like three different layers of access levels of access you can have in consensus.
So there's one there's fully open access science that's all publicly available.
We're able to ingest the full text, show it freely, let you download it. All is well and good.
Then the rest of the bucket is paywalled content, but there's two levels of access we can have within it.
So there's the buckets where we have deals with publishers trying to get as many done as possible where we're able to use the full text in our search and in our analysis.
We're just not able to display it to a user.
Benefit to the publisher is we're helping them drive traffic, get people to see that.
Hopefully this like snippet search ranking is engaging.
Then they go into it and drive a purchase.
And then there's this third bucket which is we just don't have a deal with the publisher yet.
Fully behind the payw wall.
We're using what is publicly available.
So that's like the abstract and the metadata of the paper, which goes longer forward than you think.
Like the abstract is specifically designed to be this perfect like nice summary of the paper.
It's like can go a pretty far distance in search ranking and even some analysis using abstracts, but obviously nothing nothing more brutal than than being a college student and almost getting like the information that you need from an abstract and realizing like, do I really have to pay like $50 for this? Like single fact.
The craziest is that even if you wrote the paper, you still have to pay for it.
Hand it off to a publisher, they publish it.
So, I could literally have published this paper and if I come across on the internet, I still have to pay for That's wild.
Uh Jordy had this question earlier about the nature of scientific discovery.
Elon Musk at the Gro 4 launch was talking about uh his timeline for discovering new physics is two years now based on the progress that he's seeing at XAI.
Uh and Jordy was making the point that a lot of scientific discoveries come from mapping different disciplines together um or just invent you know inventions invention generally is is apply you know the the mind of a computer scientist to a biology problem or vice versa.
Um are you seeing users do those type of searches?
Is is this product useful for that type of scientific discovery?
There's been this like lingering question in artificial intelligence about if you were a person that had read every single research paper, you would probably make yeah you would make discoveries and connections across things and yet that hasn't happened.
Maybe it's some fundamental limitation of LLMs or AI at this moment.
But what are you seeing and what's your take on that concept of like cross functional pollination? Yeah.
I mean, everything basically that humans have invented new comes from pattern matching across disciplines and like that's how we create new ideas. Yeah. Yeah.
I mean, I'm not an AI researcher, so like I don't have the single most informed take, but also nobody knows what the heck they're talking about in in this world.
Um, you know, I think it is probably a fundamental limitation of LLM given what we've seen.
like a measure of I'm going to parrot this from from Francis Sho in his YC talk the other day, but that a measure of intelligence is the efficiency by which you process information and apply it in different domains and that just like isn't what LM are really doing great right now despite the fact that they've processed so much information.
So our take at consensus would be more get people to the edge of what is known and then let them do the inherently human part of science which is create these new insights and new discoveries.
Like every science experiment that's ever been done starts with a review of the literature.
Like think about it as like you're getting the foundation of knowledge underneath your feet.
If we can, our goal at consensus would be speed up that part as much as humanly possible and let us do the thing that humans are better at than machines right now, which is that pattern matching, which is that coming up with new ideas.
And if we can make that loop move faster, like heck, that's a freaking valuable and powerful thing.
Uh, switching gears completely, I know you were at DraftKings prior to this.
uh on the sort of research and and analytics side, what is your thesis around the um ultimate um collision between uh sort of betting activities and AI last night.
Grock announced a partnership with uh Poly Market to try to bring in prediction markets to to try to basically help make uh the model itself smarter.
Uh ho how are the how are the big players like DraftKings even thinking about uh AI?
I'm sure a bunch of people have like chat GBT rappers specifically focused on sports betting and things like that.
But um how do you think the big players are thinking about it?
Yeah, I mean well I left draft games in 2021 so I can't say that I was there when people were think worrying too much about AI models.
And also the natural question I always get is how the heck did you go from sports betting to science?
And the answer is uh my parents and my grandparents and my sister are all teachers and scientists.
I was athletic growing up and I loved applying numbers to sports.
Um but I I actually have something kind of interesting to say here.
So my job at DraftKings was I was building models to find the professional gamblers on the site.
So like you'd look at all previous betting history and demographic data.
You try to make predictions on is this person actually have an edge over the market?
Um, and I would have to imagine that with better and smarter and more powerful models, like people's ability to themselves have an edge on the market would increase in the short term and then the markets obviously catch up and figure out how to bake all that in.
And I mean that is the beauty of markets, right?
Like whatever technology the people have on the side of betting into a market, so does the provider who is putting up that market and they get the information from the people they know that have the best models.
So, I think it's going to be an interesting cat and mouse game moving forward as it's always been with sports betting.
Just instead of, you know, Johnny Two Shoes in New York getting inside information about injuries, now it's somebody with a super powerful AI algorithm that's predicting games above market. That's fascinating.
I I have to imagine that in the in the AI era, the insider knowledge about injuries is is even more valuable. 100%.
But but but to your point, you could probably number one way to know if you need to limit somebody is if they are ahead of an injury because it means they're somewhat connective.
They're doing the sophisticated they're on the inside.
Yeah, that makes a ton of sense. Insider trading. That's fascinating.
I I didn't think about that in the context of uh sports betting.
Uh well, thank you so much for stopping by.
This is fantastic and congratulations on the launch. Appreciate it.
Check us out at consensus. app.
Deep search launch today. Thanks, guys. Awesome. We'll talk. Thanks for coming on.
Uh up next, our lightning round continues with Rita from Zero Entropy, uh a YC graduate, uh doing automated retrieval and announcing a seed round of $4. 1 million.
So, let's seed round alert. Seed round alert. 4 million.
That used to be a series A couple years ago.
Now, just keeps ticking up.
Congratulations on the round. Welcome. How you doing? Thank you so much.
Super excited to be here. Thanks for joining. Um introduce yourself. Introduce the company.
uh how'd you get started? What do you do? Yeah, so I'm Rita.
I'm one of the co-founders of Zero Entropy.
Uh a little bit about myself.
So uh my background is in applied mathematics.
Um I have two masters in the field, one from Ekarpi technique, one from Berkeley.
Um I guess I started more into the computer vision side of things and then I discovered GPT2 and GPT3 and I was like, "Oh my god, this is this is huge."
And I started thinking about, you know, personal assistance and stateful AI systems.
And I guess that's what led eventually to zero entropy and building retrieval systems and bringing context into LLMs.
Um, and so that's what we do.
We we build search for RAG and AI agents. Okay.
It it feels like a crowded space.
I know a few founders that are working on rag.
There's also rag implementations at the hyperscalers and the clouds.
Um, how are you differentiating? What's the key insight?
What's the pitch to companies to come over and use your service as opposed to the other options out there for retrieval? Yeah, absolutely.
I think it's about having the right abstraction.
Um, so we solely focus on the retrieval side.
We don't do the entire rag end to end because we believe that developers need to have their own prompts into generating the answer.
They need to use zero entropy as a search tool for their own AI agents.
Um, we're also developing our own uh models.
So we just released um a reranker yesterday um which was pretty exciting.
Um and I guess the winning solution needs to be extremely accurate but also extremely fast and just be production ready and easy to implement um for various use cases.
What's your take on benchmarks currently?
It feels like solving a really hard math problem and and you know retrieving the right document at the right time are somewhat unrelated.
Uh, and so how do you evaluate if your system's getting better?
Yeah, that's a great question.
Actually, um, the evaluation side of things is is is very messy.
Um, almost everyone that I talked to, they basically rely on manual inspection, um, to make sure that their retrieval is working correctly.
So, um, we've been looking into the evaluation side a lot.
Actually the very first thing that we did is release our own benchmark that was on legal documents and that really evaluated just the retrieval step um of rag meaning from a question was I able to pull all of the documents and only the documents that I needed because the problem is that if you feed your LM too many tokens that it doesn't need it's just going to hallucinate.
So the precision and the recall side of things are are extremely important and we're we're rolling out our own um evaluation solution um in the next few weeks that we've been using internally so far.
What does the rest of the stack look like?
I know you said you were kind of rag provider agnostic.
Are you also model agnostic, cloud agnostic, database agnostic?
Like where have you actually made bets?
What uh what pieces of the stack are you particularly aligned with?
Yeah, I think um you know building context uh context engineering um is going to be a new you know class of of products that needs like the data layer but also needs small LLMs inside the retrieval pipeline.
We see many teams either feeding everything into the context of the LLM entire knowledge basis because they they weren't able to make retrieval work properly um and we see teams having a very simple pipeline.
I think the winning solution needs to be somewhat in the middle and basically orchestrating LMS to rewrite the question properly, summarize the documents and creating more metadata associated with each of the documents that are indexed.
Um, and so that's what we're doing.
Um, and building this this this solution that works really well and almost gets to the precision um, and the accuracy of a large LLM while still being pretty fast um, and pretty optimized.
What's the appetite been like for this product in the enterprise versus new companies that are building new AI products from scratch?
It feels like they might be uh just the the the AI agent infrastructure companies.
There's a lot of them and it feels like they're selling to a new crop of companies and that's where the revenue is accelerating most aggressively.
But what are you seeing in the market?
Yeah, I think the adoption for products like this usually comes from like bottomup type of approach where developers are experimenting with new approaches and new techniques uh and then larger enterprises um catch up.
Um so that's what we've been what we've been seeing.
Um in terms of experimenting with models, I think large enterprises also do that pretty easily.
Um so for things like the reranker that we just released um there's also appetite from larger companies in integrating that into their current systems.
current systems. Is there a uh is there a case study that you are have your eye on amongst the big tech companies like like we think that our software could improve Netflix or YouTube recommendations or something like if the deals could just magically happen where's the lowest hanging fruit like
for me you know if I could do anything in I would just get whisper into Siri and so when I dictate a text message it's just perfect and it's much better than what they're currently using um what's on your wish list for you know consumer tech company or or big tech company that everyone knows and they're not taking advantage of something like this? Honestly, for me it's it's Slack. Um I
Honestly, for me it's it's Slack.
Um I always struggle um you know I I can never find anything on Slack.
Um and um something that we've been doing is annotating our own conversations like appending keywords to our own threads to be able to find information.
But uh we have a lot of our internal research and a lot of things going on on Slack and we find it pretty difficult um to find the right stuff.
So I think companies like that could benefit and it would provide a much better user experience if you could, you know, just magically find all of the information that you have in there.
Yeah, I've been noticing that with Gmail, like the the amount of email has just grown so much and the and the and the amount of text in each email has grown because of all the trackers and cookies and stuff behind the scenes.
And so when I search for something, it just pulls up completely random emails like every time.
And it doesn't understand the hierarchy of in an email, I care a lot more about what's in the subject line than what's in the footer.
And so if I'm searching for, you know, artificial intelligence or something and someone has that in their footer that, hey, I run an artificial intelligence company, that's not what I'm looking for.
I'm I'm looking for the thread that I was talking to somebody, a close friend about AI and I want to pull that up first. Yeah.
I think that's also why, you know, basic semantic search is just not enough because it basically will pull all of the similar information but not the most relevant or the most helpful.
Keyword search is is the same. It's not very smart.
Um, and I think it's just it's just such a waste because there's a lot of information that you could have access to and it would make your work so much faster and you're just spending time like rewriting your question and trying to make the system understand what actually you're looking for.
So I think that you know the query side of you know the user intent um query re rewriting is also super important. Yeah.
IM message search absolute disaster. It is absolute disaster.
It's like I I know I'm in a text message with Jordy and someone else and so pull that up and it's like here's six. It never works.
Um I also noticed fix it all.
There's the there's going to need to be a shift in the way people search.
a shift in the way people search. I remember hearing the story about Google where there was some Google engineer who was running a test on like uh you know how many it was like how what's the world record for you know the the the marathon or something like that and they were they were using the typical keyword
boolean search and they weren't getting good results and then they sent it to a user and the user just asked the question in natural language and it just hit it the first time and so I feel like people still I at least I have been a you know a email user for a long time when I go to my email search I often I'm searching in like the keyword world instead of just natural language. But
But Google has I mean they're they're experimenting with the AI search thing.
They have a 50 like 50word limit right now.
You can't just type a whole prompt in in Google search.
Like they need to kind of reimagine what that search box is.
And then there also needs to be a a consumer change in how and how consumers interact with that particular uh like UI element basically.
Um but thank you so much for stopping by. Congratulations. Good luck to you.
We will talk to you soon. Have a good one.
Up next, we have Elliot Hersburg from Amplify Partners coming in the studio.
They are in Data Dog, Chain Guard, Runway. I love Data Dog.
That's maybe the Golden Retriever and me, but I love Data Dog so much.
It's the greatest company name ever.
Uh right up there with a. com. Get a pod five Ultra.
I'm back on back on my game.
I'm still uh still behind where I was relative to a month ago, but back up into the 75 range.
I'm going for 90 tonight. Good luck. Good luck.
They're calling the Pod Five the first fully immersive sleep system that works with any bed.
Pod 5 actively adjusts your temperature, elevates your body, and plays integrated soundscapes to improve your sleep.
I have some new copy today. Simply too good.
Welcome to the stream, Elliot. Hopefully, you're here. How you doing? What's going on?
Hey, what's going on, guys?
It's a pleasure to be here. Thanks so much. And I love the suit.
You are dressed fantastically. Sign of great respect.
I feel like there was a, you know, there was like a period where people would actually sort of match the vibes and have the suit for the technology brothers.
And I feel like it's it's dropped off.
So I want to bring it back, you know. I appreciate it. I appreciate it.
We we we you're in the boardroom.
You know, it is uh it's I'm glad to be here. Yeah, it looks great. Yeah. Proper uniform.
I I I wanted to get a a state of the union uh from you on a few things, but why don't you just kick us off with an introduction on you? Yeah, for sure.
Uh my name is Elliot Hersburg.
Um I started my career as a experimental biologist.
So I was in the lab trying to make new treatments for cancer.
Um got super frustrated, decided to retrain as a computer scientist.
Uh so I became a computational biologist and was obsessed with that as sort of a practitioner for about a decade and then uh got really obsessed with writing about it.
So I was um started writing a newsletter called the century of biology writing about companies in the space uh data on the frontier and then that sort of was a rabbit hole into investing with a friend of the network with uh none other than Pachy McCormack.
Um so spent some time at not boring where I was uh writing and investing and then recently joined Amplify.
Uh we just closed 900 million in new capital including 200 million for a dedicated um Did you just Is this Are you announcing the new the new capital the new uh the new fund today or was that a little bit ago?
Uh it was a little bit ago 900 million including 200 million specifically for bio that I'm helping to build. Okay.
So yeah, you guys were not loud enough about that. No.
Well, that's part of the thing.
I feel like it's a it's um a really quiet fund where um for three of the four funds for Amplify, they've been in the top 5% of venture returns. Wow.
Um not top desile, but like top 5% like really good at what they do.
And I feel like it's just um not thought of as much and just like very quiet and stealthily doing really phenomenal work and so um yeah, excited to talk more about it. Okay.
State of the Union on bio.
I want to know about uh where we are uh in artificial intelligence and technology helping advance bio.
Uh we've seen alpha fold, we've seen kind of tools and amazing breakthroughs.
Uh I think everyone has a really concrete idea of the impact AI is having on software engineering whether it's like amazing autocomplete, you have cursor, now you have agents.
Um, where are we in the deployment of AI tooling in bio?
A lot of the narrative just jumps straight to we're going to oneshot cancer. And I love that. I'm optimistic. It happens eventually.
It doesn't feel like we're there.
But where are we actually in terms of the impact on productivity in AI with bio? Yeah.
So, you guys know like the Gartner hype cycle, right?
hype cycle, right? you have these like huge swings uh for new technology where um I've been working in this field for like a decade and there's been a bunch of companies you know Amplify invested in recursion which was one of the early leaders of this right like there's been this general sentiment that you can make
a huge amount of progress with new data new tools new technology and life sciences and it just turns out that it's like a really hard problem right and it takes time it takes some time for like the market to ingest that actually figure out the right business models and strategies and so there was like a
really like strong wave of early adopters and then there was some disillusionment and disappointment where it's like oh it actually turns out that it's really hard to oneshot a cure for cancer and then you know as the as that sort of happened there was just a bunch of breakthroughs in the technology right so it became consensus that this
is making a huge impact on hard problems in biology where we had the Nobel prize for alphafold right so like uh a Nobel prize going to an AI lab to uh epidemis in part to David Baker at the University of Washington because it's actually starting to make real uh meaningful impacts on hard problems in biology. And
And so that's true for sort of molecular machine learning where you're thinking about um designing new molecules and proteins.
The uh virtual cell research where you're trying to actually model how cells behave.
Um and so you're seeing this sort of step change now where like there are a couple sort of uh you know actually faster than Moore's law curves.
um DNA sequencing is is is decreasing in cost faster than than Moore's law.
And so you get this huge data tailwind plus the tailwinds in machine learning and modeling.
People are actually starting to scale these models and it's just like getting pretty impressive pretty fast.
Where where are we on uh new drug discovery companies?
We're targeting a single thing.
We're going to build a drug to solve a problem versus we're going to start a company that's SAS. It's tool.
It's going to help with all different drug companies. what's what's working?
What's more overhyped, underhyped?
Where what's your take on like picks and shovels versus drugs basically? Yeah.
So, like short answer is we do both.
Um, so you guys had Jake on the other day at Centivax, like absolutely insane founder.
Uh, was one of the early computational immunologists at Stanford and like he's making a medicine that you just couldn't otherwise make without these technologies, right?
where it's like all these impacts within biotech and within modeling to make these just really incredible drugs you couldn't otherwise make.
And so that's people just like making these singular things that are like really phenomenal.
Like there's also a real platform opportunity there for other things beyond the universal flu vaccine.
And then I think like if we take like a little trip down like history lane and think about like how hard it was to actually sell software in the life sciences.
Uh one of the early companies in the space is Schroinger and they've been around for about 25 years.
um they're a molecular dynamics company and it just took an extraordinary amount of of time and effort to actually saturate and get people to adopt the technology and get people to pay for it.
And you know they're a phenomenal business.
They're a public company.
Um they're actually vertically integrating into making their own drugs.
But the thing that we're hearing consistently like from CEOs of top pharma companies um is just like there's huge demand for new infrastructure.
people realize this technology is here now and that they need to adopt it.
They're hearing this from their shareholders.
They're hearing this from the scientists at their companies.
And so there's just a very different moment um post AlphaFold and post even chat GPT where like they're using it, their kids are using these models and they're just like, "Oh, I really actually need to adopt this."
And so, uh, I think opportunity for both like fundamentally new picks and shovels where you sort of replace experiments with compute, um, and then also just fundamentally new drug products. Jordy, last question.
No, I mean, uh, I think we should have you back on as as new news hits because, uh, yeah, I think there we'll do the TVPN, uh, bio drop in, you know, it feels like there's, uh, within the traditional labs, bio has been used as like a as a as almost like marketing, right, of being or or even Elon yesterday saying like, we're going to discover new physics.
Um and uh I don't think that that's obviously not where like a lot of the true innovation is is happening.
So um yeah, let's make it a regular thing. Yeah, this is great. Thank you so much.
When you think about that, right, like there's just been a couple things like I had I had an investor who made a joke that like there's just a couple things that have consistently delivered venture returns and that's like software and also drugs.
And so like as far as a physical prediction goes for it being something that's super valuable, um if you can make your inference be a billion dollar uh drug product, that's a pretty good spot to be. We're excited about it. So see you guys soon.
There there's uh who I forget who we had on when when Trump had the executive order around drug prices.
We talked to a few different people that were saying biotech has like on average been a terrible asset class be and there's some amazing outliers. Weber maybe. Yeah.
Like b b b b b b b b b b b b b b b b b b b b basically there's all these amazing outliers that that do deliver returns but if you just index the market you were going to underperform dramatically underperforms specifically. Yeah.
in the venture category feels like this could be a massive shift where suddenly you know the next five years become the golden age of of venture uh bio investing.
bio investing. I mean there have been some massive companies before right so you have like breakouts like the Janentex of the world huge companies there actually is you guys you guys should have Bruce Booth on who's an OG biotech investor at Atlas Ventures he's done a bunch of analysis in the fact that like there's actually some
interesting just return data for biotech versus tech where it's like not as gloomy as you would as you would think but I think in general like biotech's hurting right now uh we want to make a world where it's actually like engineering right and that you're actually just getting these like really scalable able amazing medicines and like I think that's where we're headed. Incredible. Incredible.
Well, thank you for joining. Great to have you on. Yeah. Awesome. Thanks. Talk to you soon. Have a good one, Ellie. All right. Bye. See you.
And next up we have Kareem from RAMP.
The man himself coming in to talk about the launch. Ramp's new agent launch.
Uh is he in the waiting room?
We will bring in Kareem from RAMP to chat.
Second time on the show he hopped on at uh Hillen Valley. That's right.
First time as a remote guest. Great to see you. How you doing, Kareem? Hello.
It's great to see you guys. Can you hear me? Okay. Yes. Loud and clear.
I don't think you need much of an introduction, so why don't you just kick it off with uh the announcement and break down the launch today and then we'll have a bunch of questions. Yeah, of course.
I mean, it's been a a very exciting day for us at at RAMP.
We uh finally announced our uh guess our first agent.
We're going to be announcing a lot more agents soon, so it's hard to keep track sometimes.
Um we've been playing with a lot of tech internally.
internally. We think we're in a very uh interesting space where maybe maybe the thing about it is a lot of people from the outside look at ramp and think of maybe visualize the card they think about the um the fintech aspects but at at the end of the day
like what we're really trying to do is um help reduce the drag uh on companies uh that happens when there's just a lot of work and between teams and a lot of like papers being passed around, questions being asked, uh, the things that really get in the way of of of doing work. And that first agent that
And that first agent that we're building is is is really just that like it operates in the me the messy middle between finance teams and every other team trying to spend to move the business forward.
Um, and yeah, that's that's basically what we launched today.
So, it's a um an agent for controllers.
Um it knows a lot more uh about the expense policy of of a company, the rules that are in place that govern spend uh than any single employee and it knows uh a lot more about every single transaction than any single person on the finance team.
So it can operate in the middle and automate all the little decisions and the extra work that needs to get done to figure out uh what's in policy or not.
And it's immediately available.
Like that's part of the power, right?
Is that if you're if you're if like in in the in the sense of like if an employee wants to decide whether or not they can buy something or or something's in policy, you no longer have to be slacking somebody, you know, it could be in the middle of the night or something like that or off hours where there's creating that drag, that delay, right? 100%.
Uh that that's certainly one part of it.
like you can you can ask questions about your your your policy and ask questions about specific transactions live to figure out whether they would be in or out of policy.
But more interestingly once you make a transaction it's already doing work to go and figure out like well that transaction that you made at that restaurant um it looks big but if that was a dinner with 10 people maybe it's not as bad as initially thought and that's actually in policy.
So well that information is in your calendar, it's in your email, it's in um sometimes outside of just the immediate context of a transaction.
So the the the agent will go out on the internet in some cases uh contact uh um vendors or pull data from APIs on your behalf to really gather all that context and make better decisions um on behalf of of the company.
I I want to talk about like this like the the word agent and the decisions to like like how how agents are fitting in the different stack of a tech company these days because like there is kind of always there's kind of always been an agent behind the scenes working you we think of these as like cron jobs before
it's like there's a there is a longunning process that when a receipt comes in it gets tagged and there has been for I think years I don't want to share anything you can't but like there's been an LLM interacting with receipt data for a long time, but it's been fully agentic in the sense that it was behind the scenes. And so I' I've
And so I' I've been thinking about this in the context of like meta and like some of the value that Zuck is going to be getting from having a frontier AI model.
It's like there's so many workloads inside a business that has billions of of users that just happens behind the scenes and and these are agents, but they're like almost internal agents.
And so I'm wondering about your decision to Yeah.
position position an agent as like this is a userfacing agent versus something that we're just going to have a process that's running behind the scenes entirely.
Well, there's a bit of a difference, right?
Because when you think about the these processes running behind the scenes for the most part, like the code is pretty deterministic. The tools are the same.
It's built for accuracy and auditability and you have a high confidence.
it could trace back the the path that uh the the old school agent, let's call it that, went went through. Exactly.
And in this case, like it's less deterministic.
You give the agent a set of tools.
You could tell the agent or you can essentially give it access to um let's say ability to call, ability to email, and you could be like go figure out a way to get the receipt.
That's what you know about the restaurant.
and it will browse the web and figure out that that's the phone number of the restaurant and then try to call the restaurant and if that doesn't work then it will try to email the restaurant until it achieved that goal of of getting you the receipt or it fails and you can then interact with it.
Um in in this case like the instructions that we are giving the agent as we're building it are very high level.
you're just giving it high level instruction um and access to tools and that's very different from like the the old way of building these these processes this these processes in which you had to be like very specific about all these paths.
So it would take a lot longer to build to build these systems to debug them to update them um etc.
We lost you your your zoom background is turning like like a ghost. It's very funny.
It's a super a super intelligence.
I think you just need a little bit more light light on your face.
We I think we actually I I I actually lost power.
But wait, you lost power. There we go. I'm back. There we go. That's wild. Much better.
Um I want to talk to you about the data walls that are going up and some of the battles that are playing out in like the enterprise uh world because uh when when I when I read stories about, you know, companies that want to do like enterprise search, you can see that well, you know, maybe Google doesn't want you to be taking maybe they want that for themselves.
RAM's in a very different position, but at the same time like there's just evolving policies about, you know, how friendly this is a classic with like Amazon not sending the itemized receipts to Gmail because they just didn't want to give Google the the data.
Um, but as a Ramp customer, I want the Amazon details pulled in through via Gmail via the ramp integration.
So talk to me about like how's the broader trend playing out and then how do you go to big companies and say hey like you know work with us our clients want to be able to pull data from your service and we're not going to build a delivery network Amazon so you're not a you're not a we're not a competitor for you. Yeah. No for sure.
I mean, mo most of the data that that we need at the end of the day is is like data that is quoteunquote owned by our users, the businesses that are on ramp, their employees. Yeah.
Um I think it's a little bit easier to to to operate in the B2B space because like those uh I guess what what what what governs who owns the data and whose data it is is a lot clearer than in in the uh uh a lot of consumer applications.
So like in in our case it's like what data do we really need to know to in in in the case of the agents that we just launched to figure out whether something's in policy or not.
It's metadata about the transaction, right?
like um what's in the receipt at the end of the day like stores owe you a receipt that's your receipt right we we get that information you have information that we get through the networks through Visa the the the metadata about the trans the geographical location of the transaction
maybe whether it was an inerson transaction or not there's data that's in your inbox in your email which again like that information is also owned by the company um we we haven't really encountered heard um a lot of of of push back and challenges. I I found most of
I I found most of the challenges in in getting the data to be more like technical.
How do you make sure you get it quickly clean it up and get it accurately as opposed to ones where there are third parties that are trying to make it harder and harder for us to access the data. Mhm.
U been in that in the previous company that that was kind of the story of the previous company. Yeah, of course.
We we had a lot of these problem.
I mean there was there were lots of funny moments at at Parabus or previous companies where uh we were uh I mean we're really building an agent for consumers to help them save money on their online shopping right and we're trying to log in on their behalf to Amazon accounts and Walmart accounts etc.
And of course they'll put blocks, they'll put captures.
And today those captures seem like a joke.
I think uh any version of any half recreent version of of Chad GPT or or Claude is able to solve those captures very easily.
Well, that's one of the ways the internet's getting worse right now is the captions are actually getting so hard and annoying. Yeah.
When we go to the gym in the morning, Jordy has to log in.
Takes him like 2 minutes to get through the captures for this gym.
It has like the most like military grade security to get to a gym login and it's just it just gives you a barcode that you just scan.
barcode that you just scan. But but but but on that I'm I'm actually interested because I can imagine you know Ramp has tens of thousands of customers like highv value business customers and other people that are building agents I'm sure would love to actually be able to make
actions on the ramp platform but at the same time you guys are trusted to handle the finance you know basically the finance the finance back office for these companies and you don't want like an agent like hallucinating like saying, you know, based on and taking actions on the ramp platform. So, I'm curious
So, I'm curious um how you see that that dynamic playing out because I'm sure you've been approached by a lot of companies saying, "Hey, we're building this agent to do this thing.
We'd love to be able to, you know, get authorization." Of course.
I mean, we're we're thinking through that um a lot right now.
through that um a lot right now. I think there are good ways of exposing the right information to the right agent as long as um our customers are very aware of what what they're exposing and there are a lot lots of interesting
applications for us to work on like in in the case of uh any large purchase at a company there are multiple uh parties within that company that need to review it or approve like you want to review a certain vendor and and and look at their uh data protection policies. You want to
You want to look at the the legal agreements.
In some cases, you want to negotiate the price.
And you can imagine a a day in the future where uh a lot of our customers have an agent tool that they trust or agents that they trust for legal work, agents that they trust for IT work, etc.
And we're uh very interested in in in uh actually uh working with with some of these companies, but we got to figure out on our end how we expose the the right interface uh so that we're we're ensuring really like the the security of the data of our of our customers.
So uh we it's it's an ongoing uh uh discussion work stream. Yeah. Uh last question for me.
Um the the the Gro 4 launch was very benchmarkheavy.
Gro 4 launch was very benchmarkheavy. Um it seems like you know the consensus is that uh it's a good model and so as soon as I see that it's now about cost per token and and so I want to hear from your perspective what drives decision-
making how big of a line item roughly like or how much time is spent thinking about LLM inference optimization at your scale like like roughly you know like how big of a deal is it and then and then what what is the workflow to decide can we use a cheaper model? How do we do
How do we do you have internal benchmarks?
Are you just checking these things?
Like how are you making decisions about which model to use for what problem?
Yeah, that's a great question.
I mean, I'm a lot more paranoid uh about being too slow to try uh the newest model than and and uh the latest and greatest tools than I am by uh maybe over spending a little bit in in in in one area. Sure.
I mean, the amount of time and money wasted at companies doing BS work is just insane that uh uh if we're debating whether you can make something faster by spending extra dollar or half dollar like the value that we're able to create is so big that I don't worry about it too much.
But we do have internally um uh somewhat imprecise like stack ranking of the different places where we need to make inference calls.
And in some cases, they're very simple, high volume, um kind of low risk, right?
Like you're you're trying to um normalize or clean up some like merchant data to figure out the appropriate spelling and maybe the right like photo to use.
It's not the end of the world if it's not not like perfect.
Uh we're doing it at high volume. It better be cheap.
So we have a kind of stack wrecking of like this is something high volume where we need to be cheap.
this is something that's low volume and high stakes where you need to be accurate.
And we'll generally try the the newest and greatest models in in the places where we think will make the biggest difference.
And over time like we'll break up some workflows and some parts of it will become uh cheaper, more repeatable with uh smaller versions or cheaper versions of the model and and and it will just evolve.
Um, I mean we come from I mean I remember like micro optimizing every single thing on our AWS account back in in in 2014, right?
Like we were it was it was a lot harder back then like um I I think we we also pride ourselves in in being the the time and money company.
So we do care a lot about making sure that we don't waste our own money and our and our own time.
But I would say that the TLDDR is like our time and engineering time is the most valuable thing here and I'm a lot more focused on that than than anything else. Yeah.
On on the time issue, uh what do you think about uh the various latency tradeoffs?
I'm sure if if if a if an employee wants to know is this in policy and you hit 03 Pro and it waits 10 minutes like they're probably just going to slack their manager and ask them.
Um but you're going to get a really accurate answer that's really detailed.
And so how do you think about those trade-offs in latency?
Yeah, I mean it really depends on like where where in the workflow are we making that inference call, right?
Like if it's live in the interface and the user expects a quick answer, we'll be using some of the faster models. Sure.
But the reality is like a lot of these agentic workflows that are being kicked off at ramp like happen behind the scenes, right?
Like you make a transaction, you maybe get a very quick question from Ramps AI to gather a little bit more context.
little bit more context. like that's enough and then from there it'll kick up another kick up kick off another task that can be a little bit slower that'll happen in the background and by the time it reaches uh a bottleneck uh or it'll reach a place where it needs like additional feedback uh it'll be in someone else's like notifications or on someone else's Slack like you could take
a little bit of time when like the work is going from one person to another person but less when it's like the same person interacting with the interface Um that's yeah that's generally the the thinking but cool yeah I have I I tried some of the the the the newer brow the newest browsers today and and like I I tried Comet today I tried a couple weeks ago and I think what they're trying to do is incredibly cool. Uh, but like I I often find myself
Uh, but like I I often find myself thinking like, damn, like I wish this was a little bit faster.
And I know I know it's coming, but I think unlike some of the the the browser agentic calls like you want it to be really fast. Yep. Yeah.
I was thinking about that in the context of the of the OpenAI browser.
And unless they figure out something that makes it basically 10 times as fast, I'm still going to default to to Chrome if I have both of them open just because I'm like, well, I just really need a fast answer here.
That was the Chrome innovation, right? Chrome won on speed.
Like they just went and optimized the code and and they and they nailed speed and it was enough to leapfrog.
Uh and and so you could see I mean that's like the bullcase for like Apple coming from behind is like yeah they like it feels like if if XAI and Enthropic and and Open AI are all kind of like r and Gemini are like roughly at the frontier.
If you can just get something that's at that frontier, not any new innovations, but hyper optimized and it runs locally on your phone and and it's put spitting out like tons of tokens every second like you have a product that would be very very it would be very rapidly adopted. Uh it's exciting. It matters a lot.
I mean, I think one of the like weirdest UX patterns on on on Shad GPT now is that I have to do the work to figure out whether to use uh 03 or 03 or 40 every time.
Do I have 10 minutes or do I want the and and 40 is always so good that I usually don't need to, but then I'm just like, well, I want the best of course and like I'll come back to it.
And it's such a weird paradigm.
It's going to be something that dates us.
And I just know our kids are going to be like, what did you have to do back then?
You had to rewind the VCR tape.
You had to put the disc in the Xbox.
You had to pick which model to use. This is insane.
It's so legacy and it's going away.
But we're just in this weird like we we don't have a model router solved.
And it feels like the easiest thing is like which model should we use for this? I don't know. We'll see. Yeah.
And if if I mean I I don't know if you guys I I I grew up in Lebanon.
I still remember the the days of dialup where Yeah.
You would have to uh uh kick everyone else off. Select Well, exactly.
Well, select the phone line.
In our case, like, okay, which phone line am I going to use? Like, I don't know.
Can't you tell me which one is free and like pick it for me?
And like, no, I still have to, right? Yeah.
It seems like the easiest thing to do.
do. And also I mean well this is just uh you know complaining about the app that I use 30 minutes a day at least uh chatbt but um I I almost wish I could just define it in the prompt and just say hey you use 03 pro and then here's the prompt as opposed to needing to
click the UI change it switch it and then pick instead of just being able to go back and forth and I don't know I mean it's it's a good sign because like like people are using the stove so much that they're frustrated by these like niche UI things. So, you know, it's an
So, you know, it's an exciting time.
There's a lot of I forgot who who it was who who posted on on X.
I think it was like a couple weeks ago that like every company is like one great UX breakthrough away from something amazing.
And I think that will be true for a long time.
Like there's a lot of alpha right now and just great UX and and good patterns.
Uh we we haven't figured it out.
We're still in the maybe terminal phase of of personal computers, right?
Like when is the mouse going to come out?
when are the right gooies going to come out?
Like there there's a lot of that happening right now and yeah, it's a fun fun time to be building.
One one last question for me.
Uh on Monday, uh Dwaresh released an article and then came on the show kind of talking about his timelines around, you know, when an AI agent would be able to do his taxes, right?
Sort of like aentically basically like fully agentic experience being like I want to do my 2025 taxes and then it just sort of autonomously runs.
runs. How do like big how are like you know uh Fortune 500 CFO like what what are their timelines around um maybe maybe you just tell them what the timelines are like okay by by 20 2028 you know we're going to be able to do this for you um but but h how is the the the sort of
um finance arm of the seauite kind of anticipating uh like the rate of advancement obviously like the agent today is a step towards that future, but you'll obviously need a variety of different agents or or Well, I I think in terms of capabilities of LLMs, we we we're there. We we have
We we have the capabilities like the the bottleneck on on on being able to to do this today is like having the right context, right?
It's it's uh well some of that context is in my head.
So the AI needs to know to like ask me the right questions efficiently so I can answer those.
But like even when I'm working with my accountant like pick the best accountant in the world for your personal taxes.
If you just tell them like file my taxes, they can't do anything.
Maybe you tell them file my taxes and here's access to my email.
They could do a little bit more but they can't get it fully.
Just tell them like, "File my taxes.
Here's access to to my email.
You can call my the wife as much as you want.
Uh you can look through my drawers and you give it more and more of these things."
Like maybe it could do it, but it's going to get lost.
It's going to take forever.
And and and really what we we we need to do even even for businesses is like what are the right like patterns for us to extract context that's in in people's head organize it um get them comfortable with uh connecting different tools like your inbox and and and things of that nature.
And I think in terms of tech and capabilities we're we're there.
We're not we're not really missing anything.
So, there's a lot of UX and buffs I like.
Yeah, we almost agent that can email me a question and put it in my inbox, which is effectively my to-do list.
And that's what my accountant does when that taxes happen.
They email me and say, "Well, that's why this is cool.
You can just you can take a picture of a product and and ask it if I buy this, you know." Yeah. Is it in policy? Yeah. 100%. Yeah.
Uh well, thank you so much for stopping by. This is great.
Well, we'll definitely see you soon, Kareem. This is great. To the whole team. Talk soon. Talk to you soon. Bye.
Um, and that is um the rest of our guests. We are through that.
In other news, uh, Periodic Labs, there's this scoop from Natasha Muscerinus.
Uh, the startup being co-founded by Liam Fetus and Eric Dogus Kubuk, great names, is in talks to raise hundreds of millions of dollars in funding at above a1 billion valuation.
And the two-month old startup is looking to apply AI to physical science, starting with discovering novel materials. Wow.
Let's give it up to the two-month old unicorn.
We got to have these guys on the show. That is extremely fast.
Uh I also like this post from David Pell.
We're we're getting him back on the show ASAP.
We had a lot of fun talking to him a couple months ago.
Uh he said, "I'm touring apartments in New York and just about every new build has the same soulless aesthetic.
Flat walls, white paint, no cornises, no ornamentation, just a room in a box.
only one uh one real estate agent said to me, "If you want something with character, you're going to have to stick to pre-war buildings."
Look, I'm all for some efficiency gains, but we've created a world where new things are soulless things, and that's how a society as modern as ours, and that's not how how a society as modern as ours should function.
Intuitively, you'd think that a wealthier society would build more would build more beautiful things, but not ours. And I completely agree.
agree. What's crazy is that this isn't I mean I don't these apartments look nice um but this continues all the way to $20 million houses that are still bland and I think it's I think it's mostly because maybe time and and and uh and all the
difficulties with permitting because if you are if you're if you even have the resources to build something from scratch creating okay I want these ornaments and I want this and I want like something that's really expressive of my personality. Well, now you if you want
Well, now you if you want that, no one else wants that.
So, you have to build it and you have to and you have to, you know, underwrite it and you're going to be underwritten to code.
You make sure it's to code and then get it built and and then and then the secondary market value is going to be less because not everyone wants Hurst Castle.
Whereas, if you build if everyone builds the exact same thing, they're perfectly it's perfectly liquid market because every every apartment is interchangeable with every other. Yeah. So good point.
It it's kind of it's kind of a function of just like uh modernity, but it's more a function of uh people not, you know, just risking it ever on building a disaster project, making their forever home.
People learn the lesson of uh William Randph Hurst too much.
They should have just like never learned that lesson.
Just ripped it and just send it and just build something that no one else will want to buy and will take decades to build. That's always the best.
Well, I have a good place to end it.
Uh Rob Petroso says the original Hermes Birkin bag prototype just sold for $10 million at SE.
There was a two-minute standing ovation.
He says bull market confirmed.
We love We love a bull market. The original prototype. Fascinating. That's wild. Makes sense.
It's uh incredible lore and uh I wouldn't be excited for a bull market in and alternative assets such as Birkens.
be great and uh you should be too.
But uh that's a great show folks.
We will be back tomorrow. I cannot wait.
We will talk to you tomorrow. Have a good day. Cheers. Bye.