OpenAI LIVE: Greg Brockman, Mark Chen & More

0:08

[Music] Hey, [Music] hey, hey.

0:23

[Music] Hey, [Music] hey, hey.

0:49

[Music] [Music] [Music] [Music] Hey, [Music] hey, hey.

1:44

[Music] [Music] Hey, hey, hey.

2:15

[Music] [Music] [Music] Hey, hey, hey.

2:39

[Music] [Music] [Music] You came to this world.

3:23

We came to this world to shape our future.

3:34

We came to this world [Music] to shape our future.

3:48

[Music] We dream to feel the new I like the music.

4:16

[Music] [Music] [Music] [Music] Heat. Hey, Heat.

4:59

[Music] [Music] We came to this world to reach the stars.

5:34

We came to this world to shape our future.

5:44

We came to this world [Music] to shape our future.

5:59

[Music] Reach to feel the new Feel the music.

6:24

[Music] [Music] Let me take you.

6:51

[Music] [Music] [Music] You're watching TVPN.

7:08

Your background looks way different because you have a whiteboard behind you because we're breaking down the X's and O's of the GPT5 launch today.

7:17

GPT5 launched from OpenAI.

7:20

Uh really quickly, there is some other news.

7:23

Firefly Aerospace stock opened at $70 in NASDAQ debut.

7:25

This is the company that landed on the moon. Very cool. >> Very cool.

7:31

>> Um there there are a few other stories going on, but we're going to skip most of them because we're going to be focusing on Chat GPT today on GPT5.

7:37

We have a bunch of uh a bunch of guests coming on.

7:44

We have a stacked lineup.

7:45

We'll pull that up, but we'll break down the X's and O's of the matchup. So, of course, Open AI.

7:49

Here's our here's our lineup.

7:51

We have something like 15 guests today.

7:54

Uh a ton of folks from OpenAI, a ton of people that build on top of OpenAI and uh can comment on what's going on with CH GPT.

8:01

Um but of course this battle is between OpenAI and the timeline.

8:07

It's the it it's they got to get the vibes right. >> It's war. >> It's war.

8:10

It's it's the timelines in turmoil over whether or not this is a good model, what it means for the industry, what it means for AGI timelines.

8:18

Everyone's got their take.

8:20

Everyone's posting memes.

8:20

There's been a ton of funny ones already.

8:22

We'll take you through them, of course.

8:23

But let's break down the offense today.

8:25

We have Sam Alman, the founder, CEO.

8:27

He briefly got cut from the team in November of 2023, but he's back leading the team for the 2024 2025 seasons. He seems healthy. He's doing great today. Uh he went on at 10 a. m.

8:39

to break down the launch of GPT5.

8:42

Uh he has a couple of key plays in his playbook, in his arsenal.

8:46

Uh he's got a solid ground game, lots of quick posts, hitting the timeline, probably in lowercase.

8:52

Then he might air it out with a couple thousandword essay.

8:54

We've seen him do this before.

8:56

It's a bit of a hailmary.

8:58

Maybe a thousand a couple thousand days away.

9:00

Maybe we're in the soft singularity, but he's very strong there with the long post when he needs to be.

9:06

It's up his sleeve if he needs it.

9:08

Um then he can also pull out the vague posting.

9:10

He was doing this last night.

9:11

Posted a picture of the Death Star.

9:12

No one knows what it means.

9:12

Maybe it was taking a shot at the doomers who are on the defense today.

9:17

So, he's also known for driving supercars.

9:19

That lets him get to the office faster.

9:21

He's saving time and money.

9:23

You can save time and money by going to ramp. com.

9:27

Easy to use, corporate cards, bill pay, and accounting and a whole lot more all in one place.

9:30

And so, he is uh he also gave apparently, this is a rumor, he gave every Open AAI employee who's been with the company for more than two years $1. 5 million. A lot of people say 1. 5 million.

9:43

That's not enough for a big house in San Francisco, but it is enough for a supercar.

9:47

So that's probably why he picked that number and that's why that's what the OpenAI team will be doing with that money.

9:53

They'll be buying Aston Martin Valkyries, Pagani Huayas, McLaren Sabers for Ferrari Daytona SB3s.

10:01

Uh they can get a Koig set game.

10:01

They could get a Singer DLS or a Bugatti Veyron.

10:06

It would have to be used.

10:06

They could also get the Bentley Bakalar. There's only >> Bakalar.

10:10

There's only 12 of those ever made.

10:12

Uh, it's an open top two-seater roadster.

10:15

It's coach built, so that's going to run you 1. 5 million. But that's perfect. You just got the 1. 5 million bonus. So, put it to work.

10:20

Spend it all in one place on a car.

10:22

This is financial advice. >> In shambles. >> Yes. Exactly.

10:25

Then you got Greg Brockman. He's joining at noon.

10:27

He's he's extremely well-rested.

10:30

He's actually coming off a sbatical right now. That's very exciting.

10:33

Uh, he should be injury-free for the rest of the season.

10:37

uh he cut his teeth at at MIT and uh then he got drafted by Stripe in 2010.

10:42

Uh Microsoft tried to do a trade deal during the 2023 chaotic trade deal trade window that opened up post Sam Alman ouster.

10:49

Uh but he stuck with the open AI team and now he's president of the company. Then you got Mark Chen.

10:55

He's coming on at 11:30 today.

10:55

Uh he's the chief re research officer.

10:58

The rumor is that he turned out a maxed out contract to head the Metal Lamas, but he's sticking with the OpenAI team. He was an MIT undergrad.

11:07

Also worked at Jane Street before joining OpenAI in 2018.

11:11

Then we got Sarah Frier coming on the show at 12:30.

11:13

She's the CFO of AP of of OpenAI.

11:16

It's her job to find bank accounts big enough to find to fill all the cash they're raising. It's it's a tough job.

11:23

You got to find okay this bank account. Will it hold 10 figures? Will it hold 11 figures? Will it hold 12 figures?

11:29

Like >> a lot of cash in this one. >> Exactly. Exactly.

11:32

She's also going to be defining the non-GAAP metrics that will be catnip for Ben Thompson in just a few years.

11:37

We're excited to talk to her about how she's measuring the success and the health of their business.

11:43

Obviously, it's not just revenue, not just topline, bottom line.

11:44

We're going to want to know about queries.

11:46

We're going to be want to know about DAUs, all those non-GAAP metrics.

11:49

That's where people are going to be tracking when IPO uh when IPO day comes hopefully soon.

11:55

And then we also have Brad Litecap. He's joining at 235.

11:57

He entered the league as an investment banker.

11:59

Let's give it up for the investment bankers.

12:01

They don't get enough credit around here, but we love the investment bankers.

12:05

Then he got drafted by Y Combinator before joining OpenAI as CFO in 2018.

12:08

Now he's the chief operating officer.

12:12

And then we have uh Maxarer.

12:14

Uh he's in charge of post training, fine-tuning these models, getting them into the fight, fighting performance to put on a display of authority on GPT5 launch day.

12:24

Now, let's flip it over to the defense.

12:27

They're going up against the timeline.

12:28

They're going up against the vibe checks. We got the Doomers. The Doomers.

12:32

They're led by Eleazar Udicowski.

12:34

Admittedly, everyone knows this. No one debates this.

12:37

The Doomers have had a terrible season, but you'd expect to see at least a few hail Marys about GPT5 creating bioweapons thrown up on the timeline today.

12:46

Probably won't be bangers.

12:46

Probably won't get a thousand likes, but you'll be seeing them here and there, mostly in the replies.

12:51

We've also seen some doomers talking about uh GPT5 being available to every government employee and Eleaser had some harsh words about that.

13:01

Don't give the keys to Sam Alman.

13:04

Don't give the keys to the government to open AAI.

13:06

Uh he was upset about that.

13:08

But in general, the doomers not putting much of a fight up today. Then you got Claude. Uh interesting.

13:13

Claude was caught playing for the wrong team earlier this week. Anthropic.

13:18

They're on defense today.

13:20

Uh but we saw them take out OpenAI's key pinch hitter, Claude.

13:26

Uh the Claude Code API was playing for the OpenAI team, but they shut that down and Claude is no longer pinch hitting for OpenAI.

13:34

Uh then you got the Elon stands.

13:36

Uh the ground game's going to be there.

13:39

It's going to be tra uh it's going to be strong.

13:41

The Elon stands are going to be tracking the benchmarks relentlessly.

13:45

We know XAI loves to benchmax and all the Elon stands are going to be calling out GPT5 for any any misaligned benchmarks.

13:53

If they fail humanity's last exam, it's over. It's over.

13:57

Uh they'll also toss up the occasional unhinged conspiracy theory. Uh moving on, Gemini.

14:01

Uh the betting lines have shifted big time.

14:04

People thought Gemini was out of the game. They're so back. >> They're up.

14:09

Poly market has Gemini at what 75% chance of being the best model towards the end of the month.

14:14

This is of course based on the LM arena more vibes-based benchmark, but uh Gemini will probably be quiet today.

14:22

They usually don't try and frontr run press releases.

14:26

They usually try and sit back, let the model speak for themselves, let the API credits work their way through the latest YC demo day batch and get the product into the hands of people.

14:37

And so, expect to see a big uh glossy conference in a couple weeks.

14:42

Demoing uh Gemini 3 should be a good re rebuttal from the Geminis.

14:48

Uh then you got the Metal Lamas.

14:50

Zuck's been on a poaching spree.

14:53

He's rebuilding the team during the off season.

14:55

Uh now he has a stacked roster and he's ready to go duke it out.

14:59

But no one knows exactly what's going to be in the playbook.

15:01

Is he going to go consumer? Is he going to go API?

15:03

Is he going to turn into a hyperscaler? We don't know.

15:06

But we know they got a stack team. They got Alex Wang. They got Nat Freiedman. They got Daniel Gross.

15:10

They got tons and tons of other researchers.

15:14

They've been raiding every other team.

15:17

Completely reset the salary cap for the league.

15:18

And it's been uh it's been an absolute clinic in terms of recruiting over there at Llama.

15:23

Then you got the final benchmark, Arc AGI. This benchmark stands.

15:29

GPT5 couldn't get past this defense and uh ARGI, you know, sitting there right in the end zone just swatting him down.

15:40

Swatting him down all day.

15:40

You think you think you you think with super intelligence around the corner? RKGI denied. Denied.

15:46

Uh Tyler, give us the update on RKGI.

15:49

Where does everything stand? How GPT5 do? Does it matter?

15:54

Should we care about ArcGI?

15:54

We love the team behind them, but is it an important benchmark?

15:58

Should we be tracking it today? >> Um, yeah. Okay.

16:01

So, so there's RKGI V1 and V2, right? >> Okay. >> On both. >> And V3. >> V3.

16:06

I actually don't know if >> No one's been No one's even tested V3.

16:10

>> No one's even really close there.

16:11

>> But how we doing on V1? >> V1. Uh, GPD5 is at 65. 7.

16:18

Unfortunately, that's going to be 1% just short of Grock 4. 66. 7.

16:23

Okay, Arc AGI 2, >> the Elon stands are going to be going wild with that. >> ARC AGI 2 uh 9. 9%. >> 9. 9%. >> Gro 4 16%.

16:30

So, absolute kind of brutal, you know, arc AGI mogging.

16:36

>> Rough showing rough showing.

16:37

>> Some people have have accused Gro 4 of being slightly benched maxed.

16:39

You know, this is, you know, they might have a team >> really, but >> um >> what's the what are the pros and cons?

16:48

We know the cons of benchmarking uh of bench maxing.

16:50

You're overfitting on something that might not actually drive consumer value.

16:54

It might not actually solve real world problems.

16:56

It might not increase DAUs or revenue or ARR or anything that really matters.

17:02

It might not even get us closer to super intelligence.

17:05

Give me the counterargument. Why is benchmaxing good?

17:10

>> The bull case for benchmarking.

17:11

>> The bullcase for benchmarking bench maxing. Break it down for me. >> Yeah. Yeah.

17:15

So I I think the idea is basically um this is almost like a non-aggi pill kind of take, right?

17:19

So if you don't have a a super general intelligence, >> y >> um >> your ability to benchmark basically proves your ability to um solve some like kind of specific task.

17:31

So So there's this um thing about the the gas station spiky. >> Yeah.

17:38

>> It's called getting spiky. >> Getting spiky.

17:39

Getting adding more spikes to the spiky intelligence. >> Yeah.

17:42

I think it was Rune who had this this tweet about Yep.

17:44

the gas station benchmark. Yep. Right.

17:46

I I don't care if he said something like >> uh I don't care about um AI solving gas stations if it has the gas station benchmark. Something like that.

17:55

Um, yeah, >> but the idea is like if if you if the if making the gas station benchmark >> run said, "My bar for AGI is an AI that can learn to run a gas station for a year without a team of scientists collecting the gas station data set in in capital letters." >> Yeah.

18:14

And then my take is basically I don't care how they got to the like I don't care how they made it run the gas station.

18:22

I care how fast >> that it runs it.

18:23

If it if we can run the gas station with AI, >> if you have a team who's, you know, your benchmaxing team that just proves that like if you have some task that's like really important that you want to get done, >> they can just figure it out.

18:34

So it's like RL for business.

18:36

This is like the same thing RL for law.

18:37

All these like >> specific verticals if you doing the thinking machines, right?

18:42

RL for businesses come into your organization, understand the most the most valuable business processes out there that could potentially be RL against that could be turned into a benchmark and then and then, you know, bench hacked because I don't care if you're hacking, you know, if I have translate this type of document to this type of document for my business.

19:04

If you can do it with 100% accuracy, I don't care that you benchmacked it. >> Yeah. Exactly.

19:08

Like like benchmarks right now are not like economically valuable.

19:11

Like if you if you're really that much better at MMLU, >> yes, >> it's like is are you producing that much value? >> Yes, >> probably not.

19:17

But if you have if you make some new benchmark that's, you know, your tax benchmark, I think Anthropic just released that fairly recently. >> Oh, sure, sure, sure.

19:24

>> That's like I don't care if you benchmax on that.

19:26

It does way better because then it's going to it's going to do the task. Yeah. >> Yeah. Yeah. Yeah. Yeah. That makes sense.

19:31

Um what about um the what does it say that it feels like open AI seems capable of bench hacking?

19:40

It seems like they've opted not to.

19:43

Is that because bench hacking has a risk of giving you negative aura?

19:50

Because if you're accused and found guilty of bench hacking, you could it it it often reveals that you're not building this one beautiful, you know, super intelligence to rule them all. >> Yeah.

20:03

I think it's also like maybe we're just looking at the wrong benchmarks.

20:08

>> Like maybe they're um there's a bunch of like interesting benchmarks about like there's this one I really like.

20:12

It's the Minecraft benchmark where you have to like build >> you like give it some castle and how how good it looks or there's the one you always see about um the unicorn.

20:21

>> Yeah, >> that's and it's um so you use this like math package that does like grass and stuff but you ask it to to draw a unicorn. >> I've seen that.

20:28

Yeah, >> those are really good because it kind of shows the creativity stuff like that.

20:31

>> Uh walk us through TBPN bench >> and what we will be benchmarking the uh the AIS against going forward.

20:37

Have you heard about this >> reps of 225?

20:42

That would be close but it's difficult because uh the humanoids kind of change that and you can just use normal actuator.

20:48

This is this is truly for a large language model.

20:50

You feed in our data set.

20:52

We have a public data set a private data set presumably at some point.

20:56

But walk us through TBPN bench. >> Yeah.

20:59

So so I'm yet to try this on GP5.

21:01

I don't think it's out yet like for public use.

21:02

At least I don't have it.

21:02

Um but I can I can tell some of the questions. Right.

21:06

So so the first one um I have this picture of a horse.

21:08

You have to guess the breed. >> Yep. So, um, let me see.

21:11

I think I don't want to say it in case D5 is listening, but it is may or may not be a Caspian horse. >> Okay.

21:19

>> Um, >> and it's failing right now. >> 03 is failing. >> 03 is failing. >> Boro is failing.

21:23

>> I haven't tried every Yeah, we got to try Grock and Gemini.

21:27

>> Allorse identification.

21:27

This seems extremely hackable, but at the very least, if we get one scientist to be to go off and collect the horse data set and then and then uh and then bench hack it, I think we will have done our job. Yeah.

21:40

So that's the first question.

21:42

>> The second one is a it's I have two pictures of before and after of uh this guy >> and it's which peptide didn't take >> to achieve this body transformation. >> Yep. Yep. Yep. >> Um so it fails there. >> It fails there.

21:54

So you have a data set of of what peptide does what to the human body.

21:59

>> Where'd you find that?

22:00

>> Well, you know, Wikipedia has a lot of this stuff. >> Okay. Okay.

22:02

You'd think they'd be able to it'd be able to cheat this around with 03.

22:05

Just reason who is this person?

22:08

go look up what they've said they've taken and then boom, you have >> Well, at first with 03 when I was prompting it, I would like save the the photo, but then it would have the metadata or the the file name would be like Caspian horse or something. >> Yeah. Yeah. Okay.

22:19

>> And then and then the third one, >> the third one, um, I pass in an audio file of a car revving.

22:25

>> Has to pick which one. >> It has to pick.

22:27

It has to identify the car. >> The car. Yeah. >> From the engine note. >> From the engine.

22:31

>> And it's not doing it currently. >> It's No, it wrong.

22:34

>> This is This is a good benchmark.

22:34

We >> humanity's real last exam. Yes. Yes. Exactly.

22:39

>> So, I think those are pretty solid. I have some more.

22:41

Obviously, I don't want to make them public in case anyone's going to try to, you know, benchmark this. >> Of course. Of course. >> We'll see. Hopefully.

22:47

>> It's funny because um >> Yeah, >> I was I was mentioning the other day this this app that my dad had of like tracking the like you just set your phone up and it just automatically detects which birds are in your backyard. >> Yeah. >> So, >> yeah.

23:01

I mean, this has to be extremely solvable.

23:02

It's just something that it it reveals the lack of like general general intelligence when when you have to go and and collect the horse data set which should just be out there or the engine note data set which should just be out there.

23:15

Um but but but clearly we are in the age of go on the on the individual problem and we are looking at like the power law of capabilities.

23:23

knowledge retrieval is clearly a you know 12 billion dollar a year market that consumers will pay for that will probably grow significantly.

23:33

Um and and then health and therapy and shopping and all the other features that PGCMO laid out uh in her post.

23:42

out uh in her post. This is kind of like you know what will be rldled against because those are key pockets of value in the in the consumer economy and the same thing will happen in the business economy but in the B2B context you'll probably see an individual startup building on top of an API but even then

24:01

most of the most of the model platforms offer kind of RL as a service fine-tunes as a service something where if you're starting to spend tens of millions of dollars they will do some customization on top of the model so that could be the regime for the next few years as we go

24:15

into this like you know uh instead of like this centralizing AI force there's only one company there's actually like a camberan explosion of a ton of companies doing a bunch of different things so anyway let's go to signals post signals not happy with the launch he says okay I've seen enough this launch felt like

24:32

attending a funeral hosted by minimalists uh they're unveiling tech that should feel magical real breakthroughs but the whole vibe was grayscale grief the set design looked like if mood disorder got bow got a bow house grant Um, I don't know what a bow house grant is exactly. Uh, even the

24:45

Uh, even the story telling arc, chart styles, the eulogy, tributes, then closing on someone's health battles.

24:51

What exactly are we uh are we as the audience mourning?

24:54

It feels like they're trying to get you to pre-install a therapist.

24:57

Uh, potentially great products, sure, but the emotional tone was so damn DOA.

25:02

Uh, incredibly strange all around.

25:02

I I think a like like it's weird because we're in this world and this is a question that I want to noodle on all day is is will this be the last launch of a GP of a number GPT model because >> like you don't hear about new new versions of Google going out.

25:19

You just it just got better and better and better.

25:22

Same thing with Amazon when they were optimizing for hey it's faster.

25:24

We have more on our >> cat.

25:27

Take away was is that the product matters more than the model. >> Yes.

25:30

now and probably will for potentially a very long time.

25:35

>> And when we were watching the stream, I was cheering because they gave the feature of you can now talk to the model and get it to trigger a deep reasoning workflow or get it to give you a quick answer in natural language.

25:48

And so it's it's abstracting even fur even more of the UI into the actual text interface.

25:55

And so I think in terms of like surprise and delight and and I don't know, it's like you you everyone kind of rips the Apple thing, but Apple does a great job of being euphoric and and and happy with somewhat minor product changes and like maybe that's more of where they'll go is just, hey,

26:14

there's these new features and here's how these things and Apple will spend 10 minutes on stage talking about like shifting an icon around and stuff and it's like >> I thought it was interesting they just sort casually mentioned that they're they're deprecating the old models. >> I think it's great,

26:28

>> I think it's great, >> which makes sense. >> I think it's great.

26:30

I don't want the model picker anymore. But are you upset?

26:33

>> A lot of people are going to be upset about that.

26:35

>> I think they're getting rid of 4. 5. >> Oh. Oh, really?

26:38

>> And you're and you're and you're a 4. 5 fan. >> I love 4. 5.

26:40

But I I would imagine that the future is if I ask it to think really hard about the pros and the writing style, it would then do a pass on with 4.

26:52

5 >> and but it would only trigger that when it needs to.

26:56

It's not going to give you that because if I'm just asking for, hey, regurgitate regurgit regurgitate a bunch of facts or write some code or or put together a table of data, like it's not going to need to pull 4.

27:06

5 off the shelf just like it's not always going to pull Python off the shelf.

27:09

It's not always going to pull web browsing off the shelf.

27:12

And so I I'm I'm not I'm not sure that I necessarily want 4.

27:15

5 there as a selection criteria.

27:18

I would like all this to be tucked behind a UI and have something that's actually cleaner and less frustrating to use.

27:23

I think it'll lead to higher retention. >> Yeah.

27:27

For the average like normie that no one knows what 4. 5 is like. >> That's true. That's true.

27:32

Uh anyway, Chris Pikes is open AI and anthropic are duking it out.

27:36

Meanwhile, consumer surplus is growing.

27:37

Um, we also have very good news.

27:40

Um, we also have uh uh the details from Mike Nuke over at ArcGI.

27:45

Full F GPT5 is along the V1 paro frontier.

27:50

Uh, that's cost versus performance.

27:53

OpenAI said they focused on other goals like UX and reliability.

27:54

Our testing supports this.

27:57

Uh, Mini GPT5 is super impressive accuracy for cost.

28:00

In fact, based on cost efficiency, Mini could have entered ARC Prize 2024 and likely won first place.

28:07

We are still verifying GPTO OSS or as Rune says GP GP toss >> um results soon.

28:15

Nano GPT appears overfit.

28:18

Performance is commodity.

28:18

Uh and France Chalet is also chiming in with the top line.

28:24

Yes, >> production team needs the deck.

28:28

>> Oh, we don't have a deck today.

28:28

We're just going through Yeah, we're just we're we're just riffing through the timeline uh the timeline uh tab and just pulling up some random posts.

28:37

So So you're free to pull those up, but also we can just read through them.

28:41

Um Ashley Vance is saying, "But but model switching was my job.

28:45

Model switching is out and we are into the future. Just talk to the model.

28:50

Just talk to the model and ask it what you needed to do and it will switch for you.

28:53

It will pull the right tool for the job."

28:57

Anyway, um the other question that I have for the OpenAI folks today is on the nature of secrets.

29:04

So, in 0ero to1, TL has this concept that that discovering a secret is key to building a startup and it's a key insight and I I was joking with you know the super intelligence or or GPT5.

29:18

Could my first prompt be teach me exactly how to build GPT5?

29:25

And then I go to Meta and I say I know how to do it. I have the prompt.

29:31

I have the I have the result.

29:33

And of course the answer is no.

29:33

Of course, OpenAI would never leak the most frontier capabilities into the model.

29:41

But can you build a super intelligence?

29:45

Can you call it super intelligence if it doesn't if it can't tell you how to build super intelligence?

29:50

One read on what the secret might be is that the app was the most important thing all along.

29:55

And if you if you create this narrative that that super intelligence is, you know, weeks or months away and you get a bunch of people that go and try to compete on raw intelligence.

30:08

Meanwhile, you build a consumer app business with >> billions of users.

30:14

>> Yeah, >> it's like a seems like a pretty good strategy.

30:17

I guess one quick thought is is how do you how do you rate uh Sam's vague posting from yesterday with the Death Star in the context of this new uh in the context of of of the release today?

30:31

>> It's a great question.

30:31

There's a bunch of reads on it.

30:32

One is just that like the Death Star is to some degree like Stargate and you have to Oh, wait.

30:39

He if this is the apocalypse, I figured I'd at least tune in live.

30:43

Um, you have to like the the impact of this of GPT5 is not one crazy super intelligent model that does everything.

30:56

It's just a more userfriendly higher retention lower churn consumer model that weaves its way into all aspects of daily life and improves performance and efficiency all over the place.

31:12

And so you have to build this massive cluster to serve all of that. I don't know. What's your read on it? >> I don't know.

31:21

I think it just was I think it was dramatic. It was provocative. People didn't like it.

31:27

>> It is provocative because there are many other like super mega structures that are in in sci-fi history that are positive >> or positive. Yeah. >> Yeah.

31:36

And this is this is like >> but it gets the people going. >> Yeah. I don't know.

31:39

Is it is is it a metaphor for someone else that's going to attack?

31:44

Is he I mean I mean the image is from the viewpoint of someone looking at the Death Star.

31:48

Is he saying he is seeing a Death Star being built on the horizon? Is that something else?

31:52

Is that another company, another organization? Is that the government?

31:56

Is that the is that is that legal?

31:59

>> Here's a read from uh Bubble Boy says, "I am an expert on bubbles, so it brings me no joy to say that the AI bubble is popping this time next year." Mhm.

32:08

>> Uh he is updating his timelines.

32:08

When you promise infinite scaling and don't produce it, the calculus changes.

32:12

I don't think it will be bad for most companies, but those who built their entire business model model around making the best LLMs are unfortunately going to struggle as models become more of a commodity.

32:22

Again, OpenAI is a my read on it is a consumer app business, right?

32:29

They still have a big enterprise business, but but by you know their recent valuations are are predicated on their this incredible consumer business that they built. >> Yep.

32:40

>> Uh Bubble Boy says the end user doesn't care much if Claude is 5% better than GPT5.

32:44

They care about cost, speed, and utility, especially at scale. Things will be going.

32:48

The obvious play now is shorting Nvidia and dumping.

32:50

Uh okay, start >> getting into financial advice.

32:54

>> Getting into financial advice territory here, bubble boy.

32:56

Uh but um interesting uh again kind of goes back to what I was saying earlier uh in that um if you were raising billions to make a lab and I think the potentially anthrop you know we'll see what happens in the coding market but there's some clear winners emerging and then on the consumer side um you know expecting a power law outcome and and it's hard to see anyone unseeding uh chatbt completely agree.

33:27

Uh I want to dig in more, but we have our first guest.

33:29

Let's welcome him to the stream. What a day. Mark, how you doing? >> Hey, pretty good.

33:37

Nice to see you guys again.

33:39

>> You uh congratulations on the launch. Uh take us through it.

33:43

Uh are you were you actually live or are you wearing the same thing and you recorded it yesterday? >> I'm actually live.

33:51

I don't know why, but we do. >> Yeah, it's gone.

33:54

>> I mean, we're big fan. We're big fans of live.

33:56

I mean, it just allows you to be mo the most reactive to the most new information.

34:00

Um, give us g give us the core thesis that you are trying to get across.

34:06

I think that there are a few narratives out there.

34:08

Um, we we've been enjoying the one that's, you know, this is a dominant consumer product.

34:15

They just made it a better consumer product and people are going to use the product and get more value out.

34:21

I saw a bunch of things in the presentation where I was like that's going to make my daily usage of of chatt better.

34:27

At the same time, we're in this we're in this world of oh the models the numbers matter and the scale matters and this and that and this and that and uh and it's a it's a fine line and it's a dance and and we're in a transition phase away from benchmarks and away from talking about the size of the bubbles.

34:44

But what was your core thesis like what did you want to get across to the listener?

34:47

Yeah, I mean fundamentally I think from a research perspective, we've been working on reasoning models for several years now.

34:55

And I think until now you've had this really clunky interface.

34:57

You have to pick, you know, GBD40 or you have to pick 03.

35:00

Um, and for the longest time, we've known that 03 gives you better answers across the board.

35:06

It's just too slow, right?

35:08

I mean, you often don't want to just sit there and wait for the model reason it out.

35:12

So we've done a lot of work to push the speed, the performance of our reasoning models such that these can come together and work in a very seamless way.

35:20

And so I think you know above everything we're trying to move the world into this agentic reasoning world.

35:26

We believe that's the future.

35:28

And on top of that you know you pointed something out which uh I really resonate with.

35:32

Post training is a huge part of this release.

35:35

We really wanted to highlight uh Max Schwarzer and his team who did a phenomenal job and they've made the model just really that much more useful for consumers, for businesses.

35:44

It's a monster at coding. So >> yeah.

35:48

>> Um on the on the speed of reasoning, you're obviously the chief research officer.

35:55

officer. um is are are you more optimistic about getting speed ups there from I don't know algorithmic design software optimizations or new hardware just let Moore's law carry on or find new AS6 or we saw Sarah Brris posting yesterday about um the incredible speed

36:16

that they're getting 3,000 tokens a second on GPTO OSS um and I'm wondering what levers obviously we pull all of them but But but what what what path of the tech tree are we should we be like most focused around most tracking and uh and most excited about? >> Yeah, I mean as a person who represents

36:34

>> Yeah, I mean as a person who represents research I control the things that I can't control and I think a lot of that focuses on algorithms right simple algorithms that are scalable that we can pump a lot of compute into.

36:44

Um >> we also do care about the hardware improvements that are stacking up.

36:49

improvements that are stacking up. um with the open source release you see thousands of people right um really kind of serving these models creating really great inference stacks and those are really great lessons for us to pull from

37:01

you know how what's the ceiling of the speed in which we can serve these models >> um what uh what can you tell us about the actual um like like user experience of speed um I was I I' I've just like last week I finally got to a place where for a lot of tasks. I'm I'm firing off a

37:20

I'm I'm firing off a 40 query and an 03 Pro query.

37:26

>> I just have I have two tabs. Yeah. 03 tab. >> Exactly.

37:29

And and and I'm wondering um >> uh what user experience patterns you think can uh help people balance between those?

37:38

those? Is this like just something that we're like different patterns that we're going to learn over time or different uh or or or are there going to be certain problems of user experience that are purely solved just by better product design, better speed and we don't even

37:55

need to learn these because I remember like you you know when you when you prompted a uh an image generator you used to have to say like don't no six fingers five fingers please or like don't make mistakes and now you know the models kind of have that baked in. Um

38:07

Um but but but how how are you thinking about the user experience of getting the user um the results in the right amount of time? >> Yeah.

38:17

I mean this is one facet of why we believe so much in reasoning.

38:19

It's just because all the scaffolding you used to have to give the model.

38:23

All these small hints, they go away, right?

38:24

Like the model can examine its own outputs. It can review them.

38:28

It can be like hey look like I'm just counting the fingers here. Why are there seven?

38:32

And um and it can kind of fix that, right?

38:34

It does a lot of iterative generation.

38:36

and does a lot of fixing things on the fly.

38:38

And so we think one of the benefits of bringing reasoning to the world is really to kind of remove the need for scaffolding.

38:44

And with GBD5, right, we know how clunky that that experience is with uh switching between 40 and 03.

38:51

Actually, I mean, there's so many stories.

38:53

Um I was just talking to someone yesterday, right?

38:56

They're like, "Hey, well, you know, I've used 40 my whole life, right?

39:01

It's the Frontier model."

39:01

And I'm like, "Hey, well, have you tried 03?"

39:02

And they're like, "Why would I try 03?"

39:05

you know, three is less than four and so you >> need to get out of that world.

39:09

Um, you know, GPD5, I think it's a one-stop shop, reasoning and non-reasoning.

39:12

Um, and we've really tried to make it kind of just parto optimal. >> Yeah. Yeah.

39:18

It's absolutely crazy to just take a bunch of letters and smash them together and expect people to pick up on that as a name or a brand.

39:23

Chat GBT, TBPN, we're both kind of in the same insane gambit.

39:28

But fortunately, it's worked out and I think people have have gotten over the hump.

39:32

But >> it rolls off the tongue TV. >> Yeah, sort of.

39:36

Except our friend David Sandra keeps flipping the letters. A lot of people do that.

39:40

But at a certain point, yeah, you do break through and catch PPT has.

39:43

But uh but keeping the model numbers simpler uh m makes a ton of sense.

39:48

Um uh talk to me about the pace of play for research to actual product like >> a lot of >> and and on that note the line between your personal philosophy on the line between research orgs and engineering product orgs. >> Yeah.

40:06

I mean so our research operates on a variety of different time scales, right?

40:10

We have teams that they scope out a bunch of ideas and then they start to kind of narrow in on um the promising ideas as they get closer to a run.

40:18

Um and then you kind of see a winnowing of ideas as you get closer to launching a flagship model, right?

40:24

And um there's always this kind of like uh um explore more exploratory to more kind of concrete and execution focused pipeline.

40:35

Um and we're pulling on ideas across the board here, right?

40:37

There's a lot of work in architecture optimization. Seb was on stream.

40:42

He pointed out improvements in synthetic data.

40:44

So there's really a lot of work that goes into creating one of these models.

40:48

And you know it's hard to say like oh this model was about this breakthrough just because right now we have this machine that's producing breakthroughs on all these axes and um even across several paradigms. Right.

40:59

So it's all that coming together that produces the experience that you guys feel. >> Yeah.

41:06

Can you talk to me about um the legacy or future of 4. 5?

41:10

Um I remember I was talking to you and I was like I haven't been using it a lot and you looked at me like I was crazy.

41:16

You were like ah it's so good and I was talking to Tyler and he was like uh our our intern here and he was saying like yeah the people who really like understand how good it is use it.

41:26

Um but but I was I was wondering is there a world where that is a tool in the tool chest for GPT5 in the same way that Python is or web browser is?

41:38

And if if it detects that I want something with more more emotional pros or more thoughtful writing, it can do a whole bunch of research, collect a bunch of raw raw text, and then kind of do a 4.

41:50

5 pass that I believe is more expensive maybe, and maybe doesn't make sense for every single query, but um could be a a feature in the loop or a tool that is pulled into the overall product experience. >> Yeah, absolutely. Um, speaking of 4.

42:04

5, it's also a very smart model, right?

42:07

Um, and one of our bars in creating GPD5 was to make sure that on a lot of the axes we cared about that it was able to outshine 4. 5.

42:17

And I I think even in some of the soft ones like creative writing, I I think that um was the case and and that's what makes us so confident with with the name.

42:28

Um, I think we're able to really rely on all of the architecture advancements, all the kind of post- training advancements, all the synthetic data advancements to create a model that's better than 4.

42:37

5, but much faster and much cheaper.

42:41

>> Yeah, it feels kind of like we're I remember, wasn't the second iPhone called the iPhone 3G, and the number literally corresponded to a specific technology?

42:51

And now when you get the iPhone 14, it doesn't mean it's 14 megahertz or a gigahertz or inches big like it doesn't like the number is abstract and it speaks to a bucket of features and it feels like there's I mean this was the first day of kind of re-educating folks on what the nomenclature means going forward.

43:12

Um, have you talked about an annual release schedule or like or or because there's the iPhone cadence and then there's the Google cadence which was like Google search just got better every year for two decades.

43:26

Um, I it it it feels like at a certain point you want to just be shipping as fast as possible.

43:31

shipping as fast as possible. How do you think about the culture of shipping updates that you know you find something that feels like hey that could make the customer more delighted or the user more delighted and we don't need to do a big training run for it so let's get that out today um and let's tell people about it like how are you thinking about fast

43:50

iteration versus splashy announcements >> right so on the product research side I think it makes a lot of sense to think about you know what's the cadence of release and you know uh what are the feature sets that we want build and I

44:03

actually think there's enough great research happening there that we don't have to worry about, oh, you know, is there going to be a drought or a long stretch without enough features to launch. But one thing that's important

44:10

But one thing that's important for us is to be able to provide the people doing the exploratory work some buffer from that, right?

44:17

It's hard to do really great exploratory research in an environment where you feel pressured to do release after release after release.

44:25

And so we let that be a little bit of a lazier pipeline.

44:27

Not meaning that the work itself is lazy, but we give it space really to mature uh and to flourish.

44:33

And you know once it's ready uh we can ship things across across that fence.

44:37

So um that's kind of philosophically how we organize.

44:40

We have a product research or uh still very much entrenched in the research and they care about the release cadence um and they're able to draw from all of the research that's happening um you know algorithmically and in scaling and in RL. Yeah.

44:54

Uh, talk to me about tool use and how that's growing.

44:58

I was I was kind of noodling on this idea that, you know, the I was I was thinking about the IMO and how uh it it at least from the reporting it sounded like OpenAI's model didn't use tools for that.

45:13

And that's an incredible achievement, but it's kind of like artificial.

45:19

Like I don't I don't care if the model doesn't use tools.

45:21

I I use everything possible and uh even if even if an LLM can can memorize every fact, I'm fine with an LLM looking stuff up in a traditional database, spinning up a spreadsheet, like use whatever tool you want.

45:35

Just give me the correct answer.

45:36

Um but do we have is it important to give to surface to the user the variety of tools that are in the GPT5 tool chest?

45:46

I noticed something magical happened when I was using GPT uh I was using 03 Pro.

45:52

I sent an image in and I asked to estimate the height of a desk and it wrote like a thousand lines of of Python image interpreter and was like you know interpreting pixels and I was like I didn't even think to trigger Python. It did. >> Yeah. Yeah. Yeah. No, he was right. It was crazy.

46:11

But the the really funny thing was that it was just a standardized desk.

46:14

It was just like it could have just googled like how how tall is an average desk or something. Uh or just memorized it.

46:21

It probably was just already in the weights that it knows that a desk is like 36 inches tall.

46:26

But it it did a ton of work and it still got it right.

46:28

It fact checked it a bunch of different ways.

46:29

But but but but I've noticed that now I can I I I can pull different things. Make a table. Don't make a table.

46:36

M write some Python for this. Don't write some Python.

46:38

And it kind of gives me the feel of like a super user to some extent.

46:41

Um, but I'm wondering how you're thinking about what is further down.

46:46

You like you've given chat GPT a computer as Ben Thompson said.

46:51

You've you've given kind of the core tools, the Python, ripple, the uh the the the web browser.

46:56

Um, what how are you thinking about kind of the long tale of tools that you want to bring to bear and how does that interface?

47:03

I know that there's API integrations and all sorts of different surface area there, but give me some context on that.

47:09

Yeah, I mean our reasoning models are pretty cute, right?

47:13

I mean I think they um you know when you look at their behavior, right?

47:16

They they know the height of the desk, but they'll still go verify it five different ways and you know it's all consistent, give you that median answer.

47:23

And um I think that's really what makes these models so powerful.

47:24

And when you think about tool use generically, right?

47:28

Like we want the models to use that reasoning ability to just be able to like zero um a new tool, right? it.

47:35

You should be able to kind of minimally get instructions about how the tool works and just be able to know how to use it, right?

47:41

And humans do this all the time.

47:43

You get a new tool, you start experimenting with it, and then you don't need too much scaffolding and you just go and go and use it and understand it.

47:50

So, we want our reasoning models to use their reasoning to be able to use a broad selection of tools.

47:54

And of course, there are a couple that you really do care about.

47:58

You know, in in coding, it's very important for you to be able to execute code.

48:01

Um, it's really important in personalization for you to be able to get context from your calendars and uh from from basically from the digital world.

48:09

So I think there's a range of tools we want familiarity with, but beyond that, we want the model to be smart enough to just generalize and use tool zero shot. >> Yeah.

48:17

Talk to me more about personalization.

48:19

I feel like um there's a world where I feel like I'm maybe underutilizing chat GPT as an app because I don't have it wired up to a non- relational database where it can just stuff data from, you know, it already has memory and it's doing kind of rollups and there's some sort of saving of context.

48:41

saving of context. But um I was when we were talking to uh Kevin Wheel, I was I was kind of like, well, like I don't really have like a GitHub repo that's active that I want to like dump code in regularly for like my one-off tasks, but for that image generation, like, you know, understanding the height of the

48:58

desk, it's like, well, if I'm doing that a lot, maybe I want to have a tool built that lives in the world that my chat interface can can kind of interact with on an ongoing basis and contribute to and modify and and kind of wind up instantiating a piece of software that's like even more longived and then every successive query is even faster. So um

49:16

So um yeah, how do you think about about different ways to increase personalization?

49:26

>> Yeah, I mean I think memory is huge.

49:26

Um so we have we have teams surrounding memory and also personality.

49:30

And when you look at memory, right, um I think it's just we have so much context built up about ourselves that the model doesn't have.

49:39

And um our memory team's been really hard at work.

49:42

You know, there's a surface level of just gathering facts about you, but there's also stuff about just kind of thinking very deeply about who you are, what your motivations are.

49:50

Um and even you could think about, you know, you're you're trying to do some codebased tasks, right? You're a developer.

49:57

Uh shouldn't the model just be trying code out, you know, um and and just kind of leveraging all that memory kind of its thoughts about what you want to do to just help you kind of be doing work all the time?

50:09

So, um, yeah, we do think memory is a huge part of making the model more personalized to you and it should just makes make use of all that passive signal about you that it that it observes or all of that interaction and and just help you accomplish your goals. >> Got it.

50:25

>> What do you think it'll take for AI to start making novel discoveries?

50:27

that's been a critique over the last year is everybody's so excit everybody's using these products every day and in their work and life and yet uh it still feels like we're missing that.

50:40

Uh Dwar has talked about you know potentially that being around continual learning but I'm curious what you think.

50:48

>> So one thing to underscore is I think the models are already phenomenally creative in certain ways.

50:52

creative in certain ways. So uh when I've looked at our performance on on contests right um you know I' I've done these contests before sometimes you have this mental classification of uh these problems require more creativity or these ones require less and one of the

51:08

big surprises for me was that the model can get some of the ones which I intuitively think require more creativity um and you know it often does come up with these solutions that I consider quite ad hoc and really don't pattern match to anything I've seen before. Um when you look at you know

51:24

before. Um when you look at you know advancing science or mathematics or field like this one thing that construct in which humans work sometimes is uh there are kind of theory builders um in mathematics for for instance there are uh mathematicians whose role are to kind

51:43

of build out this theory and and almost to kind of create um you know Olympiad style um uh subpros which uh often other mathematicians who are very good at that kind of style of work can do and I do think kind of the model will increasingly contribute on that side

52:01

first right if there's some mechanical like hey you know I I really don't know how to simplify this expression I really don't know how to like get get this result um it can really do that quickly for you um we're trying to increase the envelope such that the models getting

52:17

towards that theory building side and you know being able to create uh creative hypothesis and um all these components are very useful for what I consider the ultimate goal, which is being able to automate some of our own work and our own research. >> How are you thinking about like the the

52:34

>> How are you thinking about like the the layers of mixture of of mixing?

52:38

Like I remember GPT4, I don't know if this was ever confirmed, but mixture of experts model.

52:46

This is kind of like widely understood in the industry.

52:47

Um, now are we in the era of like a mixture of models that have mixture of experts?

52:54

Like how many mixtures are going on?

52:54

How does GPT5 actually work?

52:56

Is there a uh is there a taxonomy or or architecture diagram that you can kind of like walk through to explain what GPT5 is because it feels so much different than GPT3. Mhm. Yeah.

53:11

Mhm. Yeah. I mean um one of our probably the pinnacle of our research roadmap and our path to AGI um when you look at the levels of AGI the top level is uh what we describe as organizational AI and what this means is you know uh

53:28

collections of agents working together often like we might in a company towards a shared goal right and you would imagine that these agents probably subspecialize in ways maybe similar to what humans do maybe in their own more efficient ways. um and I think you know effectively work

53:42

um and I think you know effectively work together to accomplish some goal.

53:46

So we very much care about exploring this vision seeing if that's much more effective than you know one single big brain working on a problem and I think there are reasons to think why it could be so and um and yeah I I think that that is one of the things that we're after.

54:04

On that note of specialization, um how are businesses working with GPT5 or how do you expect them to work with GPT5 in terms of uh coming to OpenAI and asking for special capabilities or or or fine-tuning or you know any sort of RL on this particular problem in my world.

54:24

I have this specific data set.

54:26

It's not public, but I I want a hyper I want you to benchmax on it.

54:31

I want you to I want you to get a 100% on on you know the gas station bench or whatever.

54:35

Um you know if I'm if I have a certain business and and I'm I'm willing to invest in sort of some some overfit RL because it will create immense economic value for my business or it'll solve some fundamental problem.

54:49

Um how can how how are businesses going to be using GPT5 over the next few years?

54:54

>> No, that's a great question.

54:54

So I I think that um this is a chance to kind of highlight one of the the results that we've accomplished over the last couple weeks uh which is our ACT coder results.

55:04

So this is um a relatively unknown programming contest but it involves really the pinnacle of the the best coding contest contestants in the world.

55:16

Um and what they do is you know they're put in a room and they have to solve an optimization problem.

55:20

This is something that's actually very real world uh relevant.

55:24

So you can imagine an optimization problem as something like you know what Uber might have.

55:28

you know what Uber might have. uh you have let's say riders and you have drivers and you want to kind of create a system where you match them as as quickly as possible you know um with uh you know the least amount of cost for for instance and um and so we've really

55:46

created a system that can solve optimization problems at the level of the best in the world right and uh these truly are the kind of the best heristic solvers in the in in the world and so we have an organization led by Alexander

55:59

Madri he it's called strategic deployment and what they do is uh for a select handful of customers who really have that you know beefy problem that that they need to solve um to just go and provide that value right and um I think there's a lot we can do there I think there's a lot of very very

56:18

valuable optimization problems in the real world and um we're really excited to partner with with people because I think um this creates a template for directly having AI provide economic IC value and and really catapulting certain industries forward. >> On the on the research side, um

56:32

>> On the on the research side, um what uh what unique advantages do you think you and your team have given your position in the market with the incredible user adoption and the incredible usage from those users?

56:49

It's not just DAUs but it's actually the number of queries semi analysis estimated at like 71% of all queries going through chatbt.

56:59

Uh what advantages does that confer from a research perspective?

57:06

>> Yeah I mean a lot right and I think um you know it allows us to kind of deeply understand use cases.

57:11

It allows us to understand the frontier of where humans are, you know, kind of finding value, where they're not finding value, which areas that we need to improve the models on.

57:21

Um, it gives us a lot of signal into, you know, how users are deriving value, when they derive value.

57:26

Um, and >> what is that signal?

57:29

Um, like I I see the thumbs up, thumbs down button.

57:32

I I I'm sorry, I don't push it very often.

57:37

I'm not doing my job apparently.

57:37

But I know that you can figure out whether or not I'm satisfied.

57:42

Just stop booing me, Jordy.

57:45

>> That's the research team. >> Okay, Mark.

57:46

I promise you for the next 100 Chad GBT responses, I will I will be honest with my thumbs up, thumbs down just to help you do it.

57:54

We have tons of people luckily who do. >> Oh, that's great.

57:58

Okay, so you do get a lot of thumbs up, thumbs down.

58:00

Uh and I'm sure I have done it occasionally.

58:02

Um but I but I also imagine that there's a ton of other signal in there.

58:06

ton of other signal in there. um you know with the Tik Tok algorithm or you know any social algorithm it's very easy time on site but with chat GBT obviously it's exciting when we hear okay 30 minutes a day or some rumored number of of minutes it it feels correlated with usage it feels correlated with value

58:24

that's being delivered uh you can obviously look at churn metrics and all that stuff but what other what other pockets of signal are you finding are you finding people just I mean I I remember the story about Google where they were trying to figure about how to handle like misspellings and create the the definitive database. Do you know

58:40

Do you know this story where they were trying to develop the definitive database of how to spell things and they were like taking a bunch of shots at it and they figured out that the the best most rich source of data was just if you type in financial into Google and you misspell it oftent times then you will just correct it yourself and the second query you send will be will be spelled correctly.

59:03

So they can just look at two similar queries. What's the second one?

59:06

That's the correct that that's the correct spelling.

59:08

So yeah, what other pockets of signal are you finding that are translating into the research environment?

59:12

What are you excited to go deeper on?

59:15

>> Yeah, so I'd love to first talk about the DAU signal because um >> I think um you know that's something that a lot of companies track, but we find actually a lot of danger in tracking it too closely.

59:26

And one of the recent blog post we pushed out was one on sophincency, right?

59:30

If you just, you know, hey, we're going to boost responses where uh users say thumbs up, you know, it >> creates a condition for a model.

59:42

>> I just want to say, Mark, I love everything you're doing on this front. >> Yeah.

59:45

This entire interview has just been fantastic.

59:47

You just >> We'd love to have you back on the show tomorrow.

59:54

>> You're But clearly problems with that. >> Yeah. Yeah.

59:58

Clear clear problems, right?

1:00:00

the model just starts kind of sucking up to you and it saying like, "Hey, you know, you're right."

1:00:04

And even in complicated situations where I think objectively, you know, collectively we'd be like, "Hey, this person's in the wrong."

1:00:10

The model starts saying, "Hey, you know, you're right.

1:00:12

You know, the other person's gaslighting you.

1:00:13

You know, this other person's kind of and and >> and people deal with people deal with this in in the real world.

1:00:18

They'll go to a friend, they'll tell them about a situation, and the friend will give them advice, but maybe it's not the entire it's not the fullness of the situation, right?

1:00:26

Maybe they left out some key facts and the friend is like, "Oh yeah, that other person definitely is in the wrong."

1:00:31

And they like skipped over some important details and >> Yeah. No, no, exactly. Exactly.

1:00:33

And we don't want our models to fall into this trap where it's just trying to get you to like uh you like what it says.

1:00:39

Um and and so, you know, we wrote back a lot of changes that produce that kind of behavior.

1:00:48

And really the way I think about daily active users today is we need to be opinionated about the features that we build into the future.

1:00:56

I think we have we have a lot of ideas here but we have to let that drive um you know build for the future.

1:01:00

Build for the things that people you think they'll want and maybe don't want necessarily know they want necessarily today.

1:01:06

Um and then use DAU as kind of this byproduct right a way to track that you're on the right right right track here.

1:01:11

So >> um yeah I mean we we want to be careful here.

1:01:16

We don't want to fall into these traps of like you know 3 4 years from now that this turns into kind of engagement bait or something. >> Yeah.

1:01:24

Was it uh how how much time have has the research team been focused on efficiency specifically?

1:01:30

efficiency specifically? It felt like summer was a a good window before kids come back to school and start you know maxing out query is a good time to increase efficiency and and uh I know uh the cost of GPT5 uh have >> every time there's a new model I'm like

1:01:47

this is the best it could ever be it's good enough bake it on an ASIC I just want it for free and I want it like in milliseconds but but that's just me being you know grumpy I guess >> we we've done a a lot of work we've been building out our teams We focused a lot on scaling. I think Greg's going to come

1:02:01

I think Greg's going to come come on a little bit later and uh he's been spearheading a lot of that work.

1:02:06

So, um yeah, no, honestly, it's become a bigger and bigger focus for us, especially in the last couple of months.

1:02:14

>> Um on on the I mean, this is somewhat related to the sickency thing, but uh I'm interested to know like what do you think is driving like the GPT tone?

1:02:22

You know how like the M dash is a thing and then the the it's not a newspaper, it's a way of life and it's like the there's these like little like uh like flourishes like that that come through and in kind of a tell that it was written and in a lot of ways I love it because when I get a deep research report I like that it's using the same Wikipedia style tone.

1:02:46

Like I want consistency there.

1:02:48

I don't want it to be like oh this today it's looks like it's a Vice News article and the today it looks like it's written by someone at BuzzFeed.

1:02:54

I like that it's consistent in many ways.

1:02:56

Um but but why is that happening?

1:02:58

Do do you think that uh bigger models like 4.

1:03:00

5 kind of were able to solve that or do do those kind of like uh local minima like I don't know like wells happen even in bigger models?

1:03:10

Is there anything from a research perspective that can that can stop GPT having its own voice or is it fine that it has its own voice? >> Yeah.

1:03:19

>> Yeah. Um that's a really great question and I think you know as you scale up models as the models become more intelligent they kind of have a just deeper innate understanding of tone right and so you expect that to improve just naturally as you make the models

1:03:32

more powerful bigger better reasoners but one thing that I think gets lost a lot is >> each individual company has a lot of impact in terms of how they shape the default tone um and you know we publish a document called the spec it kind of lays how we expect the model to sound in certain cases. Lays out a lot of

1:03:51

Lays out a lot of examples for that.

1:03:52

And I think we use the spec in many ways, right?

1:03:55

We have people come in and see, hey, uh, is was this thing generated in accordance with what we would hope to generate from from our spec?

1:04:03

And this is a living document, right? It evolves over time.

1:04:05

And so I think, you know, um, each company kind of has a very opinionated take on what they think the model should sound like.

1:04:14

And it's not an accident that the models sound a certain way.

1:04:17

Um, I I don't think just naturally every company is going to train the same kind of voice into their model. >> Totally.

1:04:24

Well, thank you so much for hopping on.

1:04:26

Congratulations on the big launch.

1:04:28

Uh, we'd love to have you back soon to talk more.

1:04:30

We could go in a million different directions, but we'll let you get back to it. We know it's a big day.

1:04:34

So, have a great rest of your day.

1:04:37

It was a great conversation. >> Talk to you soon.

1:04:40

>> And we will tell you about reream one live stream 30 plus destinations.

1:04:44

Multiream and reach your audience wherever they are.

1:04:45

This stream is made possible by Reream.

1:04:47

OpenAI just did a live stream.

1:04:49

If you're trying with Reream, if you're trying to do a re if you're trying to do a stream, you got to get on Reream, so it's everywhere.

1:04:54

And we will bring in our next guest, Greg Brockman, the president of OpenAI.

1:04:57

And uh the um we'll bring him in. Greg, how you doing? >> Doing great. Thank you. >> Congratulations. Uh how are you feeling?

1:05:10

How's the company feeling?

1:05:10

Uh it's been such a wild journey.

1:05:12

just take me through a little bit of the the like the the vibes and the company and and how you got here to today. >> Well, I'm excited.

1:05:21

The whole company's excited and honestly, I'm just so proud of the team.

1:05:24

Like, it's just been amazing to watch people come together, not just for this launch.

1:05:28

Uh, and you know, the funny thing is behind the scenes that people are always putting on the last minute adjustments and polish and scaling up the capacity and there's always something that goes wrong uh before launch day.

1:05:39

There's a lot of people who uh who you know worked late into the night or really crunch to bring this release to the world.

1:05:44

Um and you know it's a little bit like the duck uh that's you know you know under the water >> but but that also describes the whole opening eye history right is that I think that we have put in many years worth of investment to the techniques used to produce this model um and really it's across just every function within OpenAI that has come together to make this a reality.

1:06:07

Yeah, I mean, you've been there for every GPT release.

1:06:10

Uh, how do you think about summing up each iteration in kind of like one one line?

1:06:21

Because GPT1, GPT2, GPT3, these feel like like similar architectures, at least at least history's kind of compressed them into similar architectures.

1:06:30

But how do you think about the progression of just the big numbered releases?

1:06:33

Yeah, it's interesting because in some ways it's a punctuated equilibrium, but on the inside it looks very smooth, right?

1:06:39

Even before the GPT series formally began, the first result that really sort of set this path to be something that we were heading down and that was clear that we were going to pursue it was the unsupervised sentiment neuron which was an LSTM in like 2017.

1:06:55

So a different architecture from today's transformers and it was the first time that you could train a model to predict the next uh element.

1:07:05

So we predicted the next character um on on Amazon reviews and we were able to get semantics out, right?

1:07:11

Because you expect okay yeah it's going to learn where the commas go, what maybe what nouns and verbs are.

1:07:14

But the idea that it would learn a state-of-the-art sentiment analysis classifier, that was mind-blowing.

1:07:19

>> And so I remember seeing that result in 2017 is like we have to scale this up.

1:07:22

We have to see where it goes.

1:07:22

We have to see where it goes. And so GPD1 was like I think a good like sign sign of life of you train on on sort of all the public data you can get um and you use transformer and that you were able to get state-of-the-art on various downstream benchmarks right so you have a model it clearly learned some

1:07:40

representation something useful about the data that it was shown and it's applicable you can use it for various tasks but we didn't really think very hard about the generation side gpt 2 was the first time that we were like all right let's actually like the samples we're getting from it, the things it actually generates, they're kind of cool. And I remember reading the uh in

1:07:56

And I remember reading the uh in the GBT2 blog post, we have this unicorn uh story uh where it generates some fictional story about a herd of unicorns. And it was just so cool.

1:08:06

It was like, wow, it like wrote a story that's actually kind of interesting.

1:08:10

It doesn't totally make sense, but like there's something here.

1:08:13

There's some real spark of intelligence within this model.

1:08:16

GPD3 was the first time that we had a model that was actually something people would it was just barely above threshold for something people would want to use.

1:08:24

And I remember working on the GPT3 API.

1:08:25

This was our first real product.

1:08:28

And it was actually the hardest product, the hardest project in total I've ever worked on because it just felt like maybe no one wants to use this model.

1:08:38

We don't really know what it's useful for.

1:08:39

Um and it certainly was the case that GBD3 is a great demo machine.

1:08:43

you can make really awesome just like tweets and you know cool little little apps and it would give you quick answers but it didn't feel very reliable and then GPD4 was something that actually felt like it had true re real world utility.

1:08:55

It was above some threshold.

1:08:55

It was something that was helpful for health.

1:08:58

health. was something that was helpful for uh for you know starting to be good at coding and GBD5 I think just sets a whole new standard for the reliability for the utility things like coding I think are just like clearly you know we're already on this trajectory of transforming software engineering this

1:09:13

year I think are really on the trajectory now to be revolutionized so just really exciting to see that that whole arc >> when did when did uh when did the API opportunity like really click for you because I do remember companies in that era that like quickly unlocked the power of of the API and and and grew tremendously. When did that opportunity

1:09:32

When did that opportunity click?

1:09:34

Because you said initially that you you kind of had some I don't know concerns, kind of doubts how useful was it going to be and then when did the consumer opportunity click?

1:09:45

>> Well, we in 2019, end of 2019 had GPD3.

1:09:49

We knew we needed to build a product um to be able to actually continue the mission to be able to raise capital.

1:09:52

Um but what did we want to build? Right?

1:09:56

we're really here because we believe in AGI that's going to have this powerful positive transformative effect on society and we want to be part of it.

1:10:01

Um and so we thought well maybe we could build something in health and then you realize okay well we're going to sell the hospitals and we're going to maybe hire >> let other people do that. >> Exactly. Right.

1:10:12

It's just like you have to go into one domain and that means giving up on the G the general right.

1:10:17

It's like it feels like you're going to become a one particular thing but we kind of want to be supporting all industries at once.

1:10:23

And so the idea was let's build an API and let people figure it out.

1:10:27

But this is totally not the way you're supposed to build a startup, right?

1:10:31

You're supposed to have a problem.

1:10:32

No one cares about the technology behind it.

1:10:34

Add value to that problem.

1:10:36

Focus on just that one thing.

1:10:38

Um and so that's why that project was so hard.

1:10:40

And in you know January of 2020, February of 2020, I that you know I with the team were going around trying to just find anyone that would be willing to try this API.

1:10:50

and we were driving to different offices in in San Francisco being like, "Hey, we have this cool model."

1:10:56

And it was hard enough to get people to take the meeting, much less to sign up their company for it.

1:10:59

Um, it was actually very fortunate.

1:11:01

We found we found a couple of good partners.

1:11:03

Um, and it was fortunate that that happened then because March 2020 suddenly that was COVID, we weren't driving around to people's offices to try to beg them to use this, you know, this uh this budding new technology.

1:11:13

Um, so it was really six months worth of grind, right, of really trying to turn like when we when we started with GPD3, I remember it was, you know, that the inference code was not very well optimized.

1:11:22

It was like, I don't know, 150 or maybe 250 milliseconds per token or something.

1:11:26

And we just optimized optimized, got it down to like 50 milliseconds per token, which by the way, today's models run much faster than that, which is kind of amazing for me, just like seeing how how much uh fast we're able to run them with much greater intelligence.

1:11:39

Um, and I remember setting two goals for the team.

1:11:43

One was I actually find one customer who's willing to pay.

1:11:47

So literally get a dollar in for this thing.

1:11:48

Um and the second is get a use case that we use at OpenAI every day.

1:11:52

That first one happened within the first couple months.

1:11:55

So actually that moment I was like all right like this thing is probably going to work.

1:11:59

Um but in order to get there we had to do a bunch of you know just scaling the API and really um you know doing doing the product work.

1:12:04

But that second one took much longer right and that wasn't really until chat GBT.

1:12:08

And so if you fast forward a couple years, because this was, you know, mid2020 when we when we first got that the API into the world, chat GBT, we didn't release until November of 2022.

1:12:19

So you're talking like a decent a decent period of of two years there, a little bit longer.

1:12:26

And I remember we were building, you know, people have talked about we were going to call it maybe chat with GPD 3. 5.

1:12:32

Um, we had a a sort of precursor product called uh called WebGPT that was built on on on 3.

1:12:37

5 that we were literally paying contractors to use. Right?

1:12:41

So, this was all throughout 2022.

1:12:43

We basically had the chat GBT precursor that we had to pay people, they would not pay us, we had to pay them to use this thing. And >> that's wild.

1:12:53

>> The moment for me that really clicked was actually when we finished training GPD4.

1:12:56

So that was August 8th of uh of 2022 which actually is like three years ago now.

1:13:02

It's actually pretty pretty wild to realize that um almost to the day >> and we did the initial post train of GPD4 and honestly I had a bunch of bugs in there.

1:13:12

It was like broken for a bunch of different reasons but >> the model was like extremely creative.

1:13:19

It was actually really interesting.

1:13:19

It took us like about a year and a half to get to the point that the creative writing of our models matched that initial one that was buggy for various reasons.

1:13:26

Um, and I remember, you know, we had an instruction following data set that it was post-trained on.

1:13:30

So, it's really we had collected examples of here's a human asking for a thing.

1:13:36

Here's what the model should do.

1:13:36

So, it's really not trained to do multi-turn.

1:13:38

Um, so I asked it a question, it gave a response, but then I was like, well, what if we just ask another question?

1:13:46

And it actually was able to leverage that full context.

1:13:49

It actually was able to have a coherent chat.

1:13:52

And the moment that we saw that that we were like, okay, this thing is capable not just of being post-trained to do this like very specific thing, but it can generalize, right?

1:14:03

It can kind of do the intelligent thing even though it wasn't directly trained for it.

1:14:06

It was just so clear this was going to be the killer killer application.

1:14:09

And so then we were planning on launching GP4 in, you know, early uh 2023.

1:14:14

and uh we had this chat infrastructure we've been working on and it's so clear okay like we're going to have to release the infrastructure and the model and it's going to be this this amazing killer product.

1:14:27

Um and so just almost as infrastructure ahead of getting the the real thing out you know I was excited for us to do chat GBT and that's why we did you know then and see that come to life in November.

1:14:39

Um so I think that for me I I was really focused on GBD4 as the model.

1:14:43

this is going to be the chat moment that's really going to work and kind of ad missed the fact because every time you see these new models you just sort of you know see only flaws in the previous ones and so miss the fact that GBZ 3.

1:14:53

5 was something that no one had really tried before in in the broad sense of society and that it was something that was already useful and that people would respond to >> was GPT3 kind of like the main pivot point for shifting the company towards LLMs because in in the prehistory of OpenAI there were a lot of other maybe expensive training runs.

1:15:13

I I don't know how much uh I don't know how much financial risk was taken with like the the OpenAI 5 project or the robotics projects, but it feels like at a certain point the the chat became like the main financial risk vector.

1:15:26

Um so I I guess the question is like when it feels like GPT3 was the moment when you shifted.

1:15:35

Um, I'm also interested in hearing about uh Ben Thompson called OpenAI the the accidental consumer company and I'm wondering when that narrative set in for you like what when when did it become clear that this was going to be a really really powerful consumer application? >> Yeah.

1:15:56

Going from paying people to use your product to people saying hey we want to give you money for this. >> Yeah. >> Yeah.

1:16:01

A very important transition it turns out. >> Yeah.

1:16:04

So um it's yeah it's it's a it's a great question.

1:16:08

Um I would say that if you rewind to the beginning of OpenAI you know there's many people who thought that you know in in retrospect say that we set out to prove that scale is how you make progress in this field but it's almost the other way around.

1:16:19

Scale was the thing that worked right that we tried a bunch of things that didn't pan out.

1:16:24

And it really the first time we saw this concretely was in our Dota project.

1:16:29

I I remember my collaborators Yakob and Shimone um trained the very first little agent on like 16 cores or something and left it running on their desktop uh over the weekend and we came back and it was this like very you know sort of constrained mini environment but that the model was doing something smart.

1:16:43

was actually able to to solve this this kiting environment and that was pretty cool.

1:16:48

And then they just they and the team just kept scaling up, right?

1:16:51

That we had all these free cores that were just sitting idle on on uh on AWS at the time and they just kept throwing more computed it and every time they would do that, the model would just get better.

1:17:02

And so when you look at something like that, you're like, well, you just have to see where this goes.

1:17:04

You have to push it until it hits the wall, right? Right.

1:17:08

And our goal with Dota was actually to develop new reinforcement learning algorithms because the common wisdom at the time was well the existing reinforcement learning PO it doesn't scale. Everyone knows that.

1:17:17

But the question from Yakab Shimone was well why do we believe that?

1:17:22

Has anyone actually tested it?

1:17:24

And no one had really tested it.

1:17:26

And so I think that that ethos of saying you have to push the existing techniques to the wall until they break.

1:17:32

And then once they break you actually have a baseline to overcome and you win either way, right?

1:17:36

either it just exceeds all the humans um in terms of of the the specific capability that you're trying to to to exercise um which was the case for for Dota um or it hits a wall and now you have a real problem to solve.

1:17:50

And so I think that ethos really got embedded in our DNA um and you know at the same time I think that we were really thinking about how do we get to AGI, right?

1:17:58

AGI, right? And really like Illy and I spent a lot of time thinking about that question of where's this company going and how do we actually uh how do we actually achieve it and you start to do some math in terms of you know the kind of compute that it would take to get to

1:18:11

to AGI and you just start to realize you're going to have to build really big computers and those are extremely expensive and so I think that from the from these early foundational results and thinking we kind of realize the path that we're going to have to walk. So, it

1:18:22

So, it seems like there's been a few walls that we've scaled up through and then maybe hit them.

1:18:28

Uh, there's been talk of like a pre-training wall.

1:18:31

Now, we're uh putting tons of resources and compute towards reinforcement learning.

1:18:35

Is there a third is there a third scaling curve that we're going to be talking about in the next few years?

1:18:42

Are we continuing to scale up those two primary vectors?

1:18:45

Is that too high level of an abstraction in terms of um like how we should be thinking about just progress along the the vector of scale like give me the up-to-date thinking on just the the fruits of scale.

1:19:03

>> Yeah, I'd say fundamentally deep learning I think that you know people talk about the bitter lesson.

1:19:07

talk about the bitter lesson. um it's almost this exploration into how do you convert compute into intelligence right through a you know we have some particular techniques to do that that we're kind of constantly fleshing out and the thing that's really amazing is if you rewind to

1:19:21

>> I don't know even the 1940s for the the makulla pits neuron which is kind of the precursor to neural nets if you look at that paper they have all these diagrams that actually look very similar to like the kinds of diagrams we draw now of multi-layer neural nets and things like that like the basic idea of what we're

1:19:37

trying to do has not really changed in almost like 80 80 plus years which is just a wild fact right it means there's something deeply fundamental about the thing that we are pursuing and that idea itself I think kind of came from trying to model the information processing of the brain and it's imperfect and not a

1:19:54

exact an analogy to biology and all these reasons that it should fail or that people have said this thing is doomed um but the results are undeniable at this point I mean some people try but uh it's it's hard to uh it's really hard to to kind of close your eyes and sleep on this in my mind. Um, and it's very

1:20:10

Um, and it's very interesting if you look at um you can find quotes from the mid1 1960s of people trying to poo poo the whole direction saying that these neural net people have no new ideas.

1:20:21

They just want to build bigger computers and you could basically say something something very similar today.

1:20:26

What we're what we are trying to do one moment >> little water break >> second. Yeah. >> Exactly. >> For all of us. Cheers. >> Exactly. Cheers.

1:20:37

>> You know, we're all human.

1:20:39

>> A proof of humanity right there. >> Exactly.

1:20:41

Um, so what we're all trying to do is find novel ways of taking compute and really harnessing it.

1:20:47

>> And sometimes you hit a wall, but these walls tend to be ones that you can drill through, right?

1:20:51

What we found is every time you scale up, everything, all of your engineering, all of your sort of scale and variance, all these things, they get stressed to the next level.

1:21:01

It's almost that the tolerances become tighter and tighter.

1:21:03

It's like launching a 10x bigger rocket means you need to be like a 100x just more precise on everything, but it doesn't mean that the fundamentals of the science are different.

1:21:11

So pre-training, there's definitely been a lot of discussion of data wall.

1:21:15

Doesn't mean it's fundamental, right?

1:21:16

It just means that we need to be better and more precise at what we're doing.

1:21:19

Um there's RL, which has been something that has kind of come from spending a small amount of compute to much larger amounts of compute now.

1:21:28

And then there is a third way that we're really harnessing compute, which is compute at test time.

1:21:31

And we publish some scaling laws around this.

1:21:33

Um, and all three of these things multiply.

1:21:35

Like that's the amazing thing.

1:21:38

>> And of course the compute and the harnessing of it is the fundamental goal, but that you get these multiplicative effects out of all of it through the quality of your engineering implementation, right?

1:21:48

Through the quality of the data sets, through a bunch of the refining work that you do.

1:21:53

And there's lots of different techniques and ideas.

1:21:54

And that's what makes this field so rich and why progress is just going to continue a pace.

1:22:00

>> What about on the infrastructure side?

1:22:02

you guys have been busy scaling up.

1:22:02

Uh what what can you share on that front?

1:22:07

>> Um well so so I run I run a team called scaling at at OpenAI and we really focus on building the infrastructure for scaling.

1:22:15

Uh and that this is in partnership with really everyone across the company.

1:22:18

It's almost a misnomer that our our team is called scaling because fundamentally this this whole team and effort is about scale.

1:22:24

Um but what we really try to do is to both on the physical infrastructure side deliver as much compute as humanly possible and that is in partnership with uh you know companies like Oracle uh SoftBank and others um that we've been able to deliver just like increasing amounts of compute to open AI.

1:22:39

Um but we're constantly thinking about how do we just deliver more flops and do it more efficiently, earlier, cheaper, more power efficient, all of those kinds of questions.

1:22:50

There's the software infrastructure side as well and really thinking about how do you coordinate massive numbers of GPUs in order to work across one synchronous training run.

1:22:59

How do you coordinate that for reinforcement learning?

1:23:03

How do you deploy that into production and bring these models to life at massive scale?

1:23:07

And I think that every single layer of the stack there is innovation required.

1:23:12

And that's something that's very easy to miss.

1:23:13

Like one way I think about research is that there is um and this is kind of the view from from Jakob who who's who's now our chief scientist um that there's a research stack and you can kind of think of the top of it is um you know people running experiments and coming up with with new ideas for how to you know sort of utilize data or something like that.

1:23:32

There's a middle of the research stack of people thinking about the how do you sort of take these different ways people are running experiments and uh be able to train in novel ways and kind of put together the pieces differently.

1:23:42

And then there's a bottom of the research stack which is like writing CUDA kernels to get the absolute max out of the GPUs.

1:23:50

And at every single layer here you get a multiplicative factor through innovation.

1:23:53

So it all comes together as one big hole. um on on on on scaling.

1:23:57

um on on on on scaling. I'm interested to hear about just if we think about like the impact of AGI or the impact of AI just being some sort of maybe you know quantitative GDP metric or qualitative just impact and good um is there an important factor of scale with

1:24:19

just not even the flops that are going into the models into the pre-training into the RL into the test time inference but actually just the flops that are going into the usage of AI within humanity broadly and I feel like that

1:24:36

might maybe be the next like scaling curve that we're seeing as more people use models they see improvements all over the fact like is is that something that we should be tracking um to see kind of the the instead of these like

1:24:51

scurves we want to see like the continual exponential >> I think that's a great perspective right because at the end of the day I mean if you look at kind the shift from something like Dota which we pursued in order to you know we wanted to do new algorithmic

1:25:07

development but really it almost validated how we scale up existing algorithms um but there was no illusion of delivering direct economic benefit from it right to the current models where we are still we're starting starting to end the era of like pushing

1:25:20

on these academic benchmarks right you look at things like the IMO at this point >> models are able to get gold medal on it like these the hardest academic benchmarks that are available are sort of no longer a you sort of the the guiding the the guiding light of

1:25:35

progress for these models to where we actually want to be is for AI to be helping everyone right to be something that uplifts humanity and that's the final metric right is how much does it actually benefit everyone how much value does it bring to the world

1:25:47

>> yeah not just health bench it's actually how many people did you solve their healthcare problem right >> exactly yes yes and that's the actual goal and that's what's exciting right is it's like we're moving from the lab to reality. >> Yeah. >> Yeah.

1:26:00

>> And I remember in the early days as we were thinking about how do we measure our progress towards AGI, we always sort of dreamed that one day we would be able to measure it this way.

1:26:08

And you can think of revenue maybe as a proxy metric for value delivered to the world.

1:26:12

Um it's not perfect, but it's at least something, right?

1:26:16

You can think of the distribution of like how much compute goes into it, how many people are using it.

1:26:21

Um but fundamentally like what we're after is how much do we really uplift humanity through this technology?

1:26:25

Yeah, I mean I might be misreading it, but I'm pretty sure like that was the Kerszswwelli and Kurs Ray Kerszswe philosophy was that like total number of flops getting getting immense not necessarily all in one data center for one model.

1:26:39

It was that it was that compute broadly would be so wide. >> Yes. Yes.

1:26:43

And I remember like on on on that chart right you can see um you know total compute of all human brains. Yeah.

1:26:50

Which really suggests a particular vision of how the this technology will be rolled out. >> Yeah. distributed.

1:26:55

The phones count as as an impact.

1:26:58

The Wi-Fi router counts for the impact of the internet just like the phone does.

1:27:02

Not just it's not just the big pipe that's going the the backbone of the internet that actually matters.

1:27:08

Um >> um deep research hit product.

1:27:09

Almost everybody I know at least in in the in the >> Mark says he's reading 30 pages of deep research a day. Basically he loves it.

1:27:22

>> He's making books with it.

1:27:22

Uh but why have uh agents broadly come around a little bit slower than than people may have expected?

1:27:30

Is it is it is it just that uh using computers is actually a much harder uh computer use is just a really hard challenge or or you know I I think going into this year everybody said this was the year of agent >> booking you talking about flight booking but you know people people were saying 2025 is the year of agents and I would say that it's the year of deep research >> and and not a lot of these other sort of like broader use cases. >> Sure.

1:27:55

Well, 2025 isn't quite over yet.

1:27:58

So, that' be my response.

1:27:59

And I >> I'm I'm very much on the uh I I think that progress in this field, the way that it tends to work is that if something kind of works with the current generation of models, it will be extremely reliable with the next generation of models.

1:28:14

And I think that where we've been is that deep research is the if you've rewound a year, that was the like we kind of had something working.

1:28:22

And then like this year, it's been just incredible.

1:28:24

And I think that agents, you know, specifically like computer use agents are something we've kind of had working and again, you know, the year is not over.

1:28:31

I think there's a lot of rapid progress to be made.

1:28:32

Um, but I think that maybe part of it too is that the agents that we're about to see, I think are a little different from maybe what we would have pictured 5 years ago.

1:28:41

years ago. Like I remember having a debate with some friends on do you want a agent that does the flight booking because the problem is it's actually a very high bar to beat the flight booking UI because there's so many preferences that are entailed in that right and you really have to know kind of what mood you're in like are you okay with like

1:28:57

taking the extra layover and all these kinds of questions and um that actually there's so much other stuff that happens in your life that that is that is toil or drudgery or that's something that that you're not an expert in you're supposed to be think about health right that like every patient really is the doctor if you're coordinating across multiple specialists. There's no doctor

1:29:13

There's no doctor that helps you with that, right?

1:29:15

That that's really on you and that there you actually can have AIs that are just text only that actually able to add massive value and then frees up your time if you want to go, you know, book the flights yourself.

1:29:29

And so I think that really finding the right problems that have high leverage, right?

1:29:33

That really add value to people and also thinking about the other side of how to make sure these agents are responsible with the trust that you put in them, right?

1:29:41

That the more that you give an agent access to your email, the more you really have to trust uh that it's going to, you know, sort of do right uh with whatever your your task is and send the right email to the right people and be able to se you segment your information all these kinds of questions.

1:29:58

And so I think that there's both a practical how do you get to adoption but also just like where are the most important leverage points in a person's life.

1:30:07

>> You also missed coding agents because it's been the year of deep research but I feel like it's also been the year of coding agents.

1:30:13

Um how is that developing at OpenAI?

1:30:17

Uh I've noticed that I'll hit 03 Pro and it'll wind up writing a bunch of code for me and I didn't even ask it to.

1:30:24

Then you have specific products for coding.

1:30:26

Um, how do you see the evolution of software development uh evolve?

1:30:30

How how are you seeing OpenAI customers use uh coding tools and how good is chat GPT or GPT5 on coding?

1:30:41

>> Well, software engineering is definitely being revolutionized in front of our eyes.

1:30:44

It's been happening and GPD5 is the best coding model in the world right now.

1:30:49

Um, it's the default now in cursor.

1:30:52

uh which I think is a a really huge statement of the quality of the model and that uh it's just so good across like every function of writing code understanding codebase being able to use tons of tools being able to do agentic work um that yeah it's like I'm not a front-end developer at all

1:31:09

>> but actually now I am right and I think that you are too right if you just talk to the model you can produce incredible things and so I think that there's this real empowerment if you think about what computers were supposed to be right computers are supposed to a more productive thing. But then somehow

1:31:23

more productive thing. But then somehow when we started out with computers you have to contort the human to the machine right assembly language and like all these like very abnormal things for a human to do and that as we've moved to tools ultimately you know in the current generation now GPD5 suddenly the

1:31:42

computer comes closer to you right that you just express your intent and you don't think about okay like exactly which language and what you know version of different libraries that the model is something you can delegate to and so we are very committed to programing ing and to making our models continue to be the best they possibly can be. >> Must a super intelligence be able to

1:32:00

>> Must a super intelligence be able to explain how to build super intelligence?

1:32:06

>> Uh so it's it's a great question.

1:32:06

So I mean I think that where we're going is a world and we're already seeing it where these models help us produce the next generation of models, right?

1:32:15

They also help us really supervise tasks that are too hard for humans to supervise on our own. Right?

1:32:21

If the model writes a 10,000line program for you, reviewing that is probably going to be quite burdensome.

1:32:26

Uh, but if you can have a model that you trust that maybe isn't as capable as the one that wrote all that code or maybe there's a team of agents that work together to write all that code, but you have a team of reviewer agents, like this is the kind of thing that you can actually bootstrap trust.

1:32:39

And I think that this is this is like one of the most important things.

1:32:41

And also interestingly, 2017 is when we had the first language results.

1:32:46

We also had some uh some results or some some vision on how you can actually bootstrap supervision beyond the scale of tasks that humans are able to supervise directly.

1:32:57

And so I think that we're heading to a world where you know we now have these chain of thought models.

1:33:01

We've been advocating very strongly to preserve the integrity of the chain of thought. Right?

1:33:05

So that means don't directly optimize it to look good even though there will be lots of temptation to do it for various reasons.

1:33:10

um really make sure that there's no pressure on the model to obfuscate its thoughts uh with within that chain of thought because then you can really see what it's up to.

1:33:20

Um and I think there's further techniques to even make it more faithful and more rigid to what the internal monologue of the agent is.

1:33:25

And so I think that there's actually a lot of promise in terms of interpretability, in terms of supervision, in terms of being able to scale to just like much more sophisticated tasks. >> Yeah.

1:33:36

I guess I guess my question is like there how much information in the world can be derived from first principles reasoning versus true secrets that can that need to be discovered by uh interacting with the world directly because I would it feels like it would be very difficult to um I'm just wondering about like how intellectual property interfaces with super intelligence or how like if you play this out a lot.

1:34:05

Um how like there's all these like hard one.

1:34:09

Doresh has talked a little bit about this with continual learning.

1:34:13

There's all these little subtleties that maybe they're not secrets.

1:34:16

Maybe they're not true trade secrets.

1:34:18

You don't think to lock them down, but they're just things that haven't been codified online or anywhere.

1:34:23

They haven't been given to anything that is surfaceable by the model.

1:34:27

And I'm wondering how is it is it just we need to build up new knowledge in every fact from first principles and and kind of go through the the history of humanity's pursuits of knowledge or do we just need to onboard more and more information or maybe it's both. I don't know.

1:34:45

It's just something I've been noodling on.

1:34:46

>> Yes, it's a great question.

1:34:46

I I would say all of the above. Select all star.

1:34:50

Um so >> I I think that the answer is very similar to what it is for humans, right?

1:34:55

How does a human generate new knowledge?

1:34:57

How do we accomplish new things?

1:34:57

First, you want to be grounded in the wisdom of the past, right?

1:35:00

You really want to understand what have people tried, what worked, what didn't work.

1:35:04

You want to go and read the biographies of, you know, various people and understand those.

1:35:08

But you also want to try things out, right?

1:35:12

You want to make, you know, some mistakes in a contained environment in a way that you actually can see the effect of your hypothesis.

1:35:17

Um, and then you want to be able to learn from those.

1:35:20

And I think that being able to really start to scale up these systems and be able to integrate them with the world is a very big process and milestone that we're currently embarking on, right?

1:35:31

To move from a world of totally hermetically sealed reinforcement learning environments to thinking about how do you actually put real world interaction in there.

1:35:39

And you think about things like robotics, like you're going to need to have that at some point, right?

1:35:44

You're going to need to have some sort of interaction with the real world and to have models that are able to produce new materials, right?

1:35:49

to be able to actually solve various diseases.

1:35:50

Um, for them to be able to really help people, right?

1:35:55

That, you know, we already have models that are great at use cases like therapy.

1:35:59

Um, but to really get to the next level of something can just really help every person accomplish more and accomplish whatever their goal is.

1:36:05

It would be very helpful for that model to actually have some real world experience with doing that very thing.

1:36:10

And so I think that that figuring out how to bring all this together is ultimately what our mission is about.

1:36:14

And we do this not in isolation, but really as part of a much broader community.

1:36:19

>> It seems like it's advantageous to have the most dominant consumer app in that environment. So, congratulations.

1:36:22

Jordy, do you have a last question? >> Last question.

1:36:25

What What do you hope to see out of Washington DC in the next year, year or two, not thinking super long term in terms of, you know, basically promoting innovation within the United States.

1:36:36

Obviously, the admin cares a lot uh about AI and has been making moves, but but what else would you like to see or where would you like them to double down?

1:36:44

Yeah, I've been very very impressed with how much the administration has engaged with the technology and really tried to figure out how can we help and ensure that American AI continues to lead and really sets the standard for the world.

1:36:56

And I think that that is the lens that I would really encourage thinking through right is like this technology is changing very fast and that fast plus government is not usually a ideal combination but this is the reality that we have.

1:37:10

It's the opportunity we have and I think that the question in my mind is less about any specific regulation or strategy but it's really being calibrated.

1:37:18

It's really having a very tight udal loop, right?

1:37:21

Being able to react to, okay, we have a new model.

1:37:23

These are the capabilities we see on the horizon.

1:37:25

How do we make sure that we get the most uplift and benefit from it and thinking strategically about not just how do we do this for Americans, right, but how do we actually do this for the world and promote democratic values?

1:37:36

And so to me, the most important thing is that motivation, right, is the question that is asked and the ultimate sort of motivation behind uh what what gets implemented.

1:37:46

>> Yeah, that makes a ton of sense.

1:37:46

Thank you so much for joining us.

1:37:48

Jordy, are you gonna hit the gong >> for GPG5 and the whole >> congratulations on the massive historic day and thank you so much for stopping by. We'll talk to you soon. >> Thanks for joining. >> Have a great day.

1:38:01

>> Thank you for having me. >> Bye. Cheers.

1:38:03

>> Really quickly, let me tell you about figma. com.

1:38:04

Think bigger, build faster.

1:38:06

Figma helps design and development teams build great products together.

1:38:08

And we are joined by Sarah Frier, the CFO of OpenAI next.

1:38:12

And we are going to bring her in in just a minute. still swinging.

1:38:18

>> The gong's still swinging.

1:38:18

And I'm going to tell you about Vant. com.

1:38:20

Automate compliance, manage risk, improve trust continuously.

1:38:24

Vanta's trust management platform takes the manual work out of your security and compliance process and replaces it with continuous automation.

1:38:32

Whether you're pursuing your first framework or managing a complex program, uh, we need one more second.

1:38:37

Tyler, any other questions that we should be asking for the OpenAI folks? Anything top of mind? What's on the timeline?

1:38:42

Is the timeline still in turmoil or has it settled?

1:38:47

>> So, I I think the general vibe is like this model was not benchmaxed, >> but if you actually get to use it, it's pretty solid. >> Cool.

1:38:54

>> One thing, uh, it failed QPN bench. >> Oh, it did.

1:38:57

>> It did not get the horse breed. Correct. >> Get the horse breed.

1:38:58

Wait, you so you have it?

1:38:59

You have access to >> Yes, I have access.

1:39:00

But I've seen other things on the timeline.

1:39:02

We can talk talk about it later, but it seems like a really good model. >> That's amazing. Great to hear.

1:39:06

Well, welcome to the stream, Sarah. Good to meet you. How are you doing? Congratulations. A historic day.

1:39:11

Thanks so much for taking the time to talk to us. How you doing? >> I'm doing great.

1:39:15

I mean, how could you not be doing great on the day when GPT5 launches?

1:39:19

It's been a a long time in the making and we're so happy it's out. >> Uh yeah, fantastic.

1:39:23

Um walk me through uh your role and what GPT5 what this launch means specifically for you.

1:39:29

Um and uh yeah, well, let's just start there.

1:39:34

>> Finance has to be you guys have to be the unsung the unsung heroes uh at OpenAI.

1:39:39

There's just there's a lot of big numbers bills coming in for crazy training runs and you have to underwrite these against future revenues and I'm sure you've developed many models to uh to figure that out but yeah walk me through what what your role at OpenAI and what what today means for you. >> Yeah, absolutely.

1:39:55

So I'm open CFO but um the finance can be the unsung heroes but they are an amazing team so I'm going to shout out to them. >> They're heroes to us. They're heroes to us.

1:40:04

It's a complex world that we're all living in and there are a lot of bees on the end of a lot of the numbers that we look at.

1:40:11

Um, look, what is our role?

1:40:13

Number one is just making sure we have a healthy, high growth business.

1:40:15

Um, it's been incredible watching just first of all the number of weekly activives.

1:40:20

700 million people using Chat GPT every week and I'm assuming after today we should see a very nice little bump in that number.

1:40:28

>> This is going to be a gonggheavy segment, Jordy.

1:40:30

I think we're going to have we have a lot of soundboard for the big number. So congratulations. >> I love it.

1:40:34

And I love I've never met a number I didn't like.

1:40:37

I think the other part of the business that you know and then we have to do >> is clapping.

1:40:43

>> We have to do this balance of the consumer business, the enterprise business and then API business which I think of as somewhat enterprise.

1:40:49

Um you know and and balancing that out.

1:40:52

So enterprise adoption has also been exploding.

1:40:56

I probably do I mean interestingly as a CFO I probably meet four to five customers a week.

1:41:00

It's a part of my job I actually love.

1:41:02

We have about 5 million paying business users right now from banks to biotech.

1:41:06

I was talking to the CFO.

1:41:10

>> And so that number is individual companies.

1:41:13

>> That is individual seats at companies. >> Seats at companies. Got it.

1:41:17

>> So what I would say about that number is it's crazy to have done that in just two and a half years because enterprises, right, you got to you got to put your big boy big girl pants on to go sell to an enterprise, right?

1:41:27

They want to make sure that you have the table stakes of security, SSO for signing on, you have HIPPA compliance if you're selling to healthcare and so on.

1:41:36

They want to know that other people have done it.

1:41:38

So they're often looking for that case study.

1:41:40

Um, but they also want to be, you know, the innovator right at the front.

1:41:44

And so that to grow that scale of business in just two and a half years blows my mind.

1:41:49

blows my mind. And it's not just big big businesses which I could talk you know at length on but it's also small mom and pop you know you know literally the people who really keep the lights on in most countries are also gravitating to chat GPT which is wonderful and then on

1:42:06

the developer side four million developers have built in our platform and the the question there is like that could be a developer inside a big company like grabb it also could be the next you know startup founder that's why combinator getting going with the next

1:42:20

multi-billion dollar unicorn business and so we see the whole gamut there and that's important to us as well because very mission aligned right how are we going to get AGI to all of humanity if we don't do it through this ecosystem so a big part of my you asked my role big

1:42:36

part of my role is just keeping that business really healthy making sure we always have the headlights on so people know the decisions they're making from a business standpoint um huge part of what the team does >> um the other big part of my role compute. If I didn't talk about that in

1:42:49

If I didn't talk about that in my first breath, you all should correct me.

1:42:54

>> I mean, it's making sure we think compute is a massive competitive differentiator.

1:43:00

>> I give so much kudos to Sam and the team, but particularly Sam because no matter how big a number we look at, Sam always wants to go bigger and he's been right.

1:43:10

Um, it is >> he's never met a number he doesn't want to add a zero to >> that, too.

1:43:14

Maybe more, maybe logarithmic, >> maybe two zeros.

1:43:16

>> maybe two zeros. and uh and he's but he has been very right and if you know you've just had a long conversation with Greg Brockman I think he does such a good job of kind of really explaining what a completely different world an agied world is or an AI world is and so I think when people get all caught around the axle of like you know what is what is a gigawatt of compute and oh my

1:43:39

god you guys want to have 10 gigawatts and that's more than the compute of like Ireland since I grew up there >> um and now You kind of look back on that and you're like those numbers already look small for a world where everyone will have access to intelligence and we're really starting to see what that can mean when you look at the demos today around things like healthcare and education and so on. >> Can you talk to me about nongap metrics

1:44:00

>> Can you talk to me about nongap metrics and what you think is going to be useful to track?

1:44:07

We were talking to Mark Chen about this and he was saying, you know, DAUs are great, time on site is great, but that's not as impactful of a metric for open AI as it is necessarily for a social network or an entertainment app.

1:44:23

And and there can actually be some problems that come up with that.

1:44:24

So, it feels like there might be some tension in the organization eventually or or just publicly about um you know, what metrics are worth optimizing for.

1:44:33

And then there's also the financial community that wants non-GAAP metrics to track the health and progress of the business.

1:44:39

And then of course over you know decades we see companies eventually roll back some of those non-GAAP metrics and as as the business gets more complex.

1:44:48

So how do you think about the development and and and sharing of non-GAAP metrics and what do you think is actually interesting and provides signal to the business and the and the investor community?

1:44:58

I'm kind of smiling to myself because when anyone normally says talk to me about non-GAAP metrics, I can see like most of people's eyes roll back in their head.

1:45:06

I live for non-gap metrics. I would love to do that.

1:45:09

Um look, I think in a CFO seat, first of all, it's really important to think about input metrics and output metrics and things like revenue, which is a gap metric as well as anap metric, they're very laggy.

1:45:20

um like if you are spending your whole time focusing on the revenue number in an operator seat like you are completely missing what's going on with the business. Yeah.

1:45:29

So I push my team a lot to get out of kind of ultimately what the P&L looks like and I'll come back to it though and go way upstream and say what are the true input metrics that tell us about the health health of our business.

1:45:40

And so I think it does start with that funnel of monthly activives to weekly activives to daily activives because we do I mean our mission is literally AGI for the benefit of humanity.

1:45:52

So we know how many billions of people live on the planet.

1:45:56

The fact that we're starting to be able to talk in billions and percentage of the world's population. It blows my mind. Right?

1:46:02

Today 85% of our users are outside the United States. And I love that stat.

1:46:08

And in fact, if you go look at where where are the big populations of users, it just tracks global population, right?

1:46:14

It's countries like India, Indonesia, Brazil, Vietnam, like the Philippines.

1:46:20

Like go to anywhere that has big population.

1:46:22

The US too of course, but that will be your tracker.

1:46:27

So that's kind of number one when I think of an input m metric.

1:46:29

From there on the consumer side, you're right.

1:46:33

things like time and app I've actually always had somewhat of a lovehate affair with but I think in this case because we're giving people intelligence teaching them how to use that um I actually think is where time and app does become important and one of the things we've really seen with chat GPT

1:46:52

are people are spending more time with it now you know we balance that with things like mental health and so on making sure that we're not creating bad things like we might have seen in prior eras of computing but I think we're just getting started on that front. Um,

1:47:03

Um, beyond that, like when we go into areas like the API, I don't look only at usage, right?

1:47:10

I can look at tokens per minute as a usage metric, but I look at things like latency.

1:47:14

I actually try to look at the elasticity of demand, right?

1:47:18

We know that developers want performance, they want intelligence, but they also want to make sure the the API is always up, and they want price.

1:47:24

And they're often willing to trade across those three three things, right?

1:47:28

It's a kind of a linear program depending on what your use case is.

1:47:32

And so I think it's important that we are offering things to developers that allow them to optimize across those three metrics.

1:47:39

For example, so that's kind of your input metrics.

1:47:43

And again, I could wax lyrical, but I won't.

1:47:45

But then you go to what you really ask.

1:47:47

So investors on the other side, right? They want to see a P&L.

1:47:51

They're like, I want to be able to compare you to other companies.

1:47:52

I want to be able to create a maybe a DCF.

1:47:54

like I want to think about fundamental valuation for a company if I'm going to invest in it.

1:48:00

And so, you know, today what I really try to push investors on is we are not a company that should be optimizing for free cash flow today because there's just too much opportunity.

1:48:11

Like that point about compute.

1:48:13

We have to make a decision on compute today with an eye to what we're going to need in two to three years because data centers don't just spring up overnight.

1:48:21

Like they're not mushrooms.

1:48:22

They literally take time and effort.

1:48:24

The thing we have at frankly I would say is three years ago we didn't have enough foresight to say how big could chat because it didn't exist.

1:48:36

>> Um it's just a shame on us if we keep doing that over and over.

1:48:38

So there can be a bit of a mismatch between our belief on revenue because we don't yet know the product versus the input which is the cost today on compute.

1:48:47

And so getting investors comfortable with the fact that there's probably losses for a period of time.

1:48:54

I say probably because Chad GPT just generally the revenue models continue to surprise to the upside but at least for now we should be in big investment mode.

1:49:03

And then you kind of said it well like as companies mature you move to more gap metrics right if you look at you know the large the mag seven many cases they're looking at like real gap net income.

1:49:14

So the whole way down to the bottom of the P&L, we're just not there yet and we should take advantage of that advantage because we can invest as a private company.

1:49:24

>> How do you think about timing fundraisers from my understanding or or or rumors?

1:49:28

Uh the last, you know, the most recent financing was very oversubscribed and at the same time you're still committing to capex in the future that is a multiple of current, you know, the current run rate.

1:49:41

And so you in the CFO seat, I'm sure there's you're you're trying to find this balance of like what does the business need today while you know not diluting the company uh you know too much uh knowing the the sort of growth rate of the business.

1:49:57

>> I mean that's exactly right.

1:49:57

That's the the art, not the science of it, is that, you know, we did just come off the back of closing out the the the sleeve of investment that we could take down in this current round led by SoftBank and it was massively overs subscribed, which comes back to I think the market really waking up to the fact that AI is a generational opportunity and the scale that it requires is like something people have not even seen before, right?

1:50:23

It's, you know, people talk about the internet or like the railways.

1:50:25

They're good analogies or transistors.

1:50:28

I think Sam always goes back to they're good analogies, but I do think this is bigger than everything that's come before.

1:50:33

Um, so there's a, you know, taking down $40 billion, which we just did in this round, that certainly felt like that gave me a lot of confidence. Appreciate that.

1:50:46

Um, a lot of confidence to then go out and do large compute deals, right?

1:50:50

we announced um the large deal with Oracle for example and to be able to keep working with all of our supply chain Microsoft, Corewave, Oracle, Nvidia and so on.

1:51:00

>> Um but at the same time, you know, in a world where our valuation has gone up, you know, at pace with our revenue, um you do get an opportunity to keep coming back to market and not take that same dilution because you're getting that higher valuation for the work and the output that you've created.

1:51:14

So it is a bit more of an art than a true science.

1:51:19

I think for now we will we'll continue to need to fund raise in order to fund that compute but I think we want to start getting more sophisticated like just pure equity fundraising for everything is an expensive way to fund raise and I think we're probably getting

1:51:33

to the stage at a company where we can be a little bit more kind of broad and how we think about funding overall and and even just working frankly with our supply chain because you know our success with bringing this era of AI into being is their success too. And I

1:51:46

And I think these companies are realizing that. >> What about partner? Last question.

1:51:54

Partner selection on the compute front.

1:51:57

There's not a lot of companies in the world or or or or firms that can can really be a >> you should update your LinkedIn title.

1:52:04

We saw someone yesterday works for Discord is in charge of their cloud buying and the and the and the his LinkedIn title was I have full responsibility over buying cloud our our entire cloud budget.

1:52:17

And it was clearly like a huge flag, but I'm sure you know you're you're in you're in direct text message, you know, with every single person that's relevant in the industry.

1:52:26

But >> yeah, but but I'm curious around like, you know, a lot of people uh have been excited about developing data centers over the last couple years in hopes to win. >> Oh, yeah.

1:52:36

>> Uh comp, you know, the business of companies like OpenAI, but I think in when you guys are evaluating partners, I I imagine that scale is is such a such a massive factor.

1:52:46

And so a single small data center is not really going to move the needle.

1:52:50

You guys need to be thinking in terms of mega projects.

1:52:54

>> Yeah, I mean I think that's exactly right.

1:52:55

I mean it started with our partnership with Microsoft and it's kind of it makes me smile now to go back and look at that original kind of large fabric for pre-training because I think it was only in the maybe 20 megawatt sort of size.

1:53:09

Um and you know now we're talking gigawatts even just this year.

1:53:15

Uh and you're right that when we think about like what is perfect compute for us or strategically the right compute for us we are definitely thinking about large scale um we're thinking about flexibility right we're learning a lot about um you know pre-training post-training test compute even like where the different kind of scaling is

1:53:35

happening um we're kind of recognizing there's more of a blurred line often between what people think of as inference investors always are like your inference compute and your training compute it's like you know literally it's vanilla ice cream and chocolate ice cream when in reality there's like a bit in the middle that is something of both. Um we also need to think about things

1:53:51

Um we also need to think about things like um where you know latency where do we want to put our footprints around the world that very global weekly active user base right as they use chat GPT you don't want to slow the model down right the beauty of the intelligence is like the real-time nature of it and then when we get into big compute like where

1:54:11

there's lots of tokens being used like deep research image gen um video as that comes online like all the work you saw today actually just even on voice um like that really quickly means that you got to make sure your compute is near your users and so it is a a big plan that's coming together but you're right like small is just not that useful to us. Um, but

1:54:33

Um, but >> what about pushing partners to take risks?

1:54:36

From my understanding, you guys are pre-committing to certain, you know, basically spend levels, but at the same time, I imagine you want people to say, "Here's what we know we're going to need, but we want you to build, you know, this much capacity so that we we have the sort of uh incremental capacity built in.

1:54:53

>> Yeah, we want I mean being extensible is really important.

1:54:55

really important. Um, and we do want to see partners like I think Oracle OCI has done a really nice job of that of kind of starting we started with like one large it felt really large at the time data center footprint in Abalene and

1:55:08

Texas and now that has really multiplied up into multiple sites that can all be connected and that's a good example of a partner who has the capability to start in one way but to be able to show you a path to maybe 5xing just in that in that single footprint. That said, I we are

1:55:22

That said, I we are finding that as we go around the world, there is an ability to go work with governments.

1:55:29

For example, we just made an announcement in Norway, made an announcement in the UK.

1:55:32

Um, this is the first time in my professional career I've seen countries come to the table and want to do commercial deals like wallto-wall chatbt.

1:55:41

I think the government of Estonia put Chat GBT into all of their high schools um high school or I can't remember was up in the university level.

1:55:49

But that's kind of wowing and handinhand with that they are viewing AI infrastructure as incredibly strategic for their population.

1:55:57

And you know it's a whole other level of selling versus you know I've I've seen enterprise large enterprises before but never anything at this scale. >> Last question.

1:56:09

Whose idea was it to give every federal agency chat GPT for a dollar a year?

1:56:14

Yeah, I imagine I imagine that pulling your hair out.

1:56:17

You could have gotten more than a dollar.

1:56:18

The CFO must be really upset here.

1:56:23

>> $10 that's 10 times as much money.

1:56:23

Now, of course, >> this is one where I think it's really important.

1:56:28

Opening is, you know, in some ways a US asset and national asset.

1:56:30

And we want to make sure we're accelerating our government like all of the resources as we think about, you know, Western democracy and so on that we are absolutely putting our technology into those hands.

1:56:42

It's that guy Kevin Wheel.

1:56:44

He's been moonlighting for the US government.

1:56:45

It's like, which team are you playing for, Kevin?

1:56:47

Are you on Open AI or are you on the US government?

1:56:52

>> Kevin just did his basic training.

1:56:52

I don't know if I'm allowed to tell you that, but I was hearing all about it yesterday. >> I saw some photos. They look great.

1:56:58

>> Yeah, it's a good thing.

1:56:58

It's even better for Kevin. Yeah, it's great.

1:57:01

Um, last question for me.

1:57:04

last question for me. I I I you know the open source model launched uh two days ago and um there's this world where like you have this dominant the accidental consumer company you have this dominant consumer app that's generating so much revenue then you have B2B and enterprise and API and that looks more like a cloud

1:57:21

provider but then is there a world where the Red Hat Linux of open-source LLMs is an open AI division and that and that there's actually serious revenue and profit that comes from helping companies implement an open-source uh large language model like Red Hat built a pretty fantastic business for a long time on top of open source Linux implementations. >> Totally like yeah I mean I think it's

1:57:45

>> Totally like yeah I mean I think it's the right question to be ask to be asking I mean I think step one was getting >> yeah you got to get it out two days our second open source model out um and getting seeing what that traction is and then seeing what the community needs.

1:57:56

I think it's important to leave space for a community to develop, right?

1:58:00

That is the beauty of open source is that ecosystem that develops and that was true with Linux.

1:58:05

It's true in areas like crypto too.

1:58:08

But I do think you'll find over time that as enterprises want to deploy it like I now dinosaurs myself, but when I was a, you know, when I was a research analyst at Goldman Sachs back in the day, I covered software and I covered I covered Red Hat actually. Oh, that growth.

1:58:24

that growth. I I wrote a research report called fear the penguin at one point because Linux being deployed >> but then you started to understand that for an enterprise you couldn't depend on like patching and upgrading to happen via community model like you needed some of the rigor that goes with an enterprise business where you kind of know when you know if you need

1:58:44

maintenance if you need a bug patch and so on and so that did allow Red Hat to grow an incredible business so I don't know if it's us or we'd be supportive of others but I think we are so excited to see open source out there and getting incredible feedback and I think we want to do that ahead of GPT5 to keep coming back to like we're here to grow this ecosystem. >> Well, we'll give you market cap credit

1:59:04

>> Well, we'll give you market cap credit for it anyway even if it's early stage.

1:59:08

Well, thank you so much for coming on. This is fantastic. We'll talk to you soon. >> Thank you. Great to see you both. Take care. Have a good one. Cheers. >> Bye.

1:59:13

Up next, we have DD Credo from Kudo.

1:59:15

I believe I'm pronouncing that correctly.

1:59:17

Uh let me tell you about Graphite Code review for the Age of AI.

1:59:21

Graphite helps teams on GitHub ship higher quality software faster.

1:59:23

You can get started for free at graphite. dev.

1:59:27

And let's bring in our next guest. How are you doing? Welcome to the stream. >> Welcome.

1:59:32

>> Oo, very clean background.

1:59:32

I know it's probably virtual, but whatever you got going on looks fantastic. You look great. How are you doing?

1:59:37

Are you excited about GPT5? >> Oh, I'm so excited. It's It's awesome.

1:59:44

>> It's actually like everybody's talking about the coding capabilities, >> please.

1:59:46

But no one is really talking about the code review capabilities and I'm going to talk about that today. >> Yeah. Yeah. Break it down.

1:59:51

Um how are you using it right now? >> Yeah.

1:59:56

So we just enabled it in our platform.

1:59:57

Um it's uh the default model for both our ID plugin, our CLI, our Git plugin.

2:00:04

Um and um yeah, we're using it to generate very high quality code reviews, catch bugs before they hit production, help enterprises verify that their code is aligned with their best practices. Mhm.

2:00:16

>> See, it's super exciting.

2:00:16

I can share my screen and show a few things if like that makes sense.

2:00:20

>> You can everything you share will be live.

2:00:23

It'll be a little bit please.

2:00:23

Um but I I I want to know also while you're getting that set up, I want to know about um what changes materially do you think happened in GPT5 specifically for code and code review?

2:00:35

Do you think there's more data going into the model, more data going into the pre-training, post-training? Anything else?

2:00:41

any anything that you're noticing that you're like, "Oh, there's a specific upgrade here.

2:00:46

They must have done something to get there." >> Yeah. Yeah.

2:00:50

I think it's a great point.

2:00:52

So, I think it's all of the above.

2:00:52

So it's scaling of both the like the pre-training but probably a lot of the reinforcement learning y >> um and basically using that at scale to verify that uh code gets generated in high quality and then also um basically catching bugs like and and when you do it with reinforcement learn learning you have the the actual ground truth.

2:01:13

So once you scale that you can get the model to be um to basically be a lot better at that.

2:01:21

How how steep is the power law right now in uh in just programming languages?

2:01:26

Is it basically all Python, JavaScript and then uh like a really hard fall-off or is it actually important for uh coding models if they want to be adopted widely to be like truly multi- language and get all the way down into the long tail of like the Rust and the and you know C and all the different languages that are out there. >> Yeah. Yeah, for sure.

2:01:44

It's important to I mean the majority of the market is in the JavaScript, TypeScript, Python >> um like the majority of the early adopters I would say but then when you get to enterprise use cases you get a lot of .

2:01:57

NET you get a lot of Java and the models are pretty getting pretty good at those uh languages as well. Um, for sure.

2:02:06

>> How are you excited about um I mean how do you think about the difference between like the improvements to GPT5 from the consumer's perspective versus at the API level?

2:02:14

Um I always found it a little confusing that chat GPT was available as an API and you could interface with the chat I believe you could interface with the chat GPT model via the API.

2:02:27

Um, and and there's a little bit of like a line blurring there, but are there features that you think are are cruff and you want to kind of rip out for an a API use case or do you just say, "Hey, give us the kitchen sink and we'll we'll we'll work from there and it's actually helpful to have, you know, a coding model that can still have a web browser." >> Yeah. Yeah.

2:02:48

I think basically it's a lot about uh we consume the model through the API and it's really the same model that drives the consumer product.

2:02:53

that drives the consumer product. M um but as the for us since our use cases are a lot aboutic use cases >> the more the model gets better at using tools um and gets better at um kind of listening to very very specific instructions

2:03:11

following instructions is critical for the enterprise use cases um because for us unlike the border market we believe that for enterprises you need to have um very specific um agents that are defined with specific set of instructions prompts and tools and permissions. Um,

2:03:26

Um, and the more the models get trained with that type of environment, the better they end up serving the the enterprise market, which is really where we're focused on.

2:03:37

>> Um, my my question is um I wonder like you you said like very specific instructions are important.

2:03:42

Uh, when are we going to get an agent that I can just turn loose in a codebase and say like just go improve it?

2:03:50

like just go hunt around do like rewrite that like like when you get a good open- source contributor on a team that just becomes nerd sniped by the project that you're building on.

2:04:01

They will just go around and find little ways to improve this documentation needs to be a little better.

2:04:06

Let's rewrite this test case over here.

2:04:08

Let's add a little bit more, you know, functionality to this class or function.

2:04:12

Um how far are we from that?

2:04:15

Yeah, I think the models are getting u better and better at that part of basically kind of running loose in a codebase. Yeah.

2:04:22

>> Um but they do need the guardrails in place >> and this is kind of where we're focused on like the a lot of the talk in the market is around the code generation side.

2:04:31

Um you know let the agent loose and give it a task and it will just going to go around and run for hours and do and and do it.

2:04:36

uh what we're seeing is that the real challenge is now shifting towards how do I verify that the code is aligned with the best practices?

2:04:45

How do I make sure that it's well tested, well reviewed um doesn't break anything um you know so that that's I think the next frontier and really developers going forward are not going to write a lot of the code um by by hand.

2:04:58

by hand. They're mo spend they're going to spend most of the their time reviewing code and that's the next frontier and that's what we're talk really like are here to tackle >> very cool anything else Jordy >> no well thank you so much for joining giving us some extra context on the GP

2:05:14

GPT5 launch we will talk to you soon have a great rest of your day and thank you for joining >> cheers thanks cheers >> talk to you soon uh and let me tell you about profound get your brand mentioned on chatbt that seems more relevant than reach millions of consumers who are using AI to discover new products and brands. I forgot to ask about this.

2:05:34

I forgot to ask about this.

2:05:36

We'll have to come back to this, but I want to know if >> the found powers MongoDB, Indeed, Mercury, Docyign, Zapier, RAMP, >> row, Goolan, Workable, Majuri, Sleep, US Bank, Chime, Clay. >> Okay. Okay, we get it.

2:05:50

Um, >> they got some logos.

2:05:51

There is this question of like okay uh even if you're even if you're like okay GPT5 is more incremental than re more of a more of a an evolution than a revolution.

2:06:01

It's like okay well then let's talk about how it affects every other business and every other aspect of the economy.

2:06:06

What should you be focusing on?

2:06:08

Um and and is like do the do any of the updates from GPT4 to GPT5 change how you're positioning your brand for AI search?

2:06:21

That's certainly an interesting question to dig into.

2:06:22

Anyway, we have Zack Lloyd from Warp coming into the studio.

2:06:26

Welcome to the stream for the second time. Welcome back, >> Zack. >> Good to see you. >> He's to be back. >> How you doing?

2:06:32

>> I'm doing pretty well. >> Uh, you know, yeah.

2:06:33

So, uh, I mean um, a lot of what stuck out to me, I'm mostly a consumer of consumer AI uh, apps.

2:06:39

I'm very excited about not needing to mess around with a model picker anymore.

2:06:44

Um but take us through the biggest improvements from uh the software development side.

2:06:53

>> Yeah, I mean it's uh it's a major step up from the prior open AI models.

2:06:55

It's um I mean it's doing a gentic workflows and warp for much longer period.

2:07:02

It's just a smarter general model like we eval it against all of our benchmarks and it's up there at state-of-the-art which is you know from our perspective it's it's awesome to have multiple competitive models that our users can benefit from.

2:07:17

So definitely a a huge improvement from um GPT41. >> Yeah.

2:07:23

So it seems like not not the you know clog code killer um but certainly in in in the same conversation in the same uh in the same football stadium if we're using a sports metaphor.

2:07:36

>> How much uh you know one thing that stood out is the cost reduction.

2:07:38

How about do you think that developers will care about that versus just you know what it can do from from an output standpoint?

2:07:49

>> I think developers do care about value.

2:07:52

So sort of like quality to cost ratio.

2:07:55

Um I think it's the more you get into like the individual developer and the small team, the more that that matters.

2:08:02

Whereas if you're at the enterprise level, I feel like it's it's a little bit less uh price sensitive.

2:08:06

Um so yeah, I mean you you you can see it as as different apps change their pricing what the reaction of the developers is.

2:08:19

You've probably seen this with cursor and seen this with cloud code and so developers really really are looking for something that's cost effective.

2:08:24

So the the fact that the cost is a little bit lower is actually is a big deal.

2:08:29

>> Do you think we're in um the Lyft Uber 2015 arc where the prices are subsidized and the prices will go up?

2:08:38

Do you think that there's a price war on the horizon now that the Frontier models seem to be similar capabilities?

2:08:44

Do you think that someone will try and raise a bunch of money, cut prices a bunch, and steal a bunch of users?

2:08:50

Like, how do you think that plays out?

2:08:53

>> It's an awesome question.

2:08:53

Um, I mean, my hope is that we get to a world where there is price competition at the model layer.

2:09:02

So, Warp is very much at the at the app layer, right?

2:09:04

And so, our value prop is like we can give our our users who are mostly developers the best model um access.

2:09:13

And so to the extent that it's not one sort of model provider running away with that and having pricing power, it's better for us just candidly.

2:09:23

And so, you know, my my hope would be something like the model world ends up a little bit like G-Cloud, AWS, Azure.

2:09:31

That's our best end state where all of these models are, you know, sort of similarly powerful and a little bit more commoditized.

2:09:38

I don't think it's been like that, but it's going it's it's getting a little bit more like that.

2:09:42

And so, uh, the more that that there's more than one show in town, I think that's generally good for Warp and actually is good for developers because it will put competition, uh, the competition will put pressure to bring the prices down.

2:09:56

But I don't know, like I also think that people will will definitely pay for quality.

2:09:59

And so if there is a um you know a meaningful delta and quality on on the frontier models then I think that like whoever has the quality delta will will have a lead temporarily but it's I'm not sure that that lead will be sustainable. We'll see.

2:10:17

>> How how do you think the uh developer community should plan around uh model deprecation over the next you know one to two years?

2:10:27

like h how much you know from I I I don't know that I've gotten a reaction yet from I don't know if there's general frustration yet from people um you know we've we've heard on the consumer side Tyler on our team here loves four five uh and so he was a little bit disappointed to hear that but but it what kind of what what are you seeing on the developer side?

2:10:47

Yeah, I think it's a little bit different for people who are like building apps on LLM versus people who are using LMS as like a like um accelerator to doing coding.

2:11:01

Um and like you know at Warp actually we do both like we we're we're an application level stack and like it's actually very easy for us to go to the latest model and so it doesn't it doesn't really um bother me. I don't know.

2:11:16

I don't know what type of app you would be building where it's like it's really important that it's like GPT35 or GPT4 or something like that.

2:11:22

I think like generally we want the most intelligent tokens at the best cost.

2:11:25

So I don't I don't see that being like too big of an issue honestly.

2:11:32

>> What about open source?

2:11:32

Um does that does that feel like something that will be in the playbook?

2:11:36

Is the markup on closed source models high enough that there will be a significant price delta and or or is the para frontier kind of indifferent to closed source open source?

2:11:50

>> So if there was a a comparable open-source option that would be awesome.

2:11:54

I think that the economics of it again it doesn't it doesn't seem like a perfect analogy to me between like open source software and open source models.

2:12:03

models. So open source software it's like you have a big community of people who you know for the love of coding are building a really awesome product for open source models it's um it's like you just need a crazy amount of capital to train something that's on the frontier

2:12:19

and so I don't know that how that happens and so what we've seen is like the open-source models are competitive at the quality level that they're at but the quality level that they're at is not the same as the frontier models and I don't really see why that would change. Um, and so I don't know in warp it's

2:12:36

Um, and so I don't know in warp it's like we we were serving some open source models, but they're just not they're not as good.

2:12:44

And so there's I think a more limited use case for them right now.

2:12:46

Um, and I don't really see economically why that would change.

2:12:52

In fact, I would be I would be surprised if anyone was spending billions of dollars to train a model and just kind of put out the open weights.

2:13:00

like I don't get this the business strategy there but but maybe that will happen that would be awesome.

2:13:07

>> Is there a world where uh you're like this idea of like smarter smarter models either orchestrating dumber cheaper models or like using or or distilling models into more narrow um narrow formulations that can be run more efficiently.

2:13:23

We've talked to a few companies that do this for businesses.

2:13:26

Like you just want a model that just filters for profanity and you can run it on, you know, a a a gaming graphics card.

2:13:34

And so it's basically super super cheap or super fast.

2:13:36

Um I'm wondering about like in the coding world, coding agent world, any of that?

2:13:41

like where where are the opportunities to kind of fan out and use an ensemble of models instead of just the hit everything with the smartest best.

2:13:53

It feels like because of the funding environment, everyone can kind of justify like a high cloud bill, but um and most people don't admit that it's hurting the bottom line, but it feels like at some point it kind of has to eventually.

2:14:06

Uh I mean I think I think that's a very real thing like in sense of um even in warp we don't use >> like the the biggest most powerful model for every task and so >> there's certain things like >> um you know for warp maybe for like deciding whether or not we should summarize a conversation is like a good example.

2:14:30

So you hit the context window, you're like, okay, is this is this a good spot to summarize?

2:14:33

Is this a good spot to encourage a user to start a new conversation?

2:14:37

We use a much more uh inexpensive and also low latency model. Right?

2:14:43

The other thing, the trend is that these very very uh powerful models tend to have much higher latency.

2:14:47

And so we do a mixture of models and that's totally a real thing.

2:14:51

Um, but I I think for like the like the predominant use case as a developer is going to be I want to tell an agent to do something.

2:15:01

I want it to be harder and harder.

2:15:04

I want it to run for longer and longer.

2:15:06

Um, and to do that it's like you kind of want in general the most intelligent model.

2:15:12

And so, >> yeah, >> know until this until the models have a sort of S-curve like type shape, I think that um I think it's going to be more of a quality game than a cost game for most of these things.

2:15:24

And >> doesn't it feel like they have an S-curve shape right now?

2:15:28

>> Certainly does from a consumer perspective. >> That's interesting.

2:15:30

Um from from a coding perspective, I feel like we're still accelerating.

2:15:35

Um, like the difference again between the last version of GPT and this version of uh of GPT is is probably bigger than the difference between like 41 and four and four and 3. 5.

2:15:48

Like it's a big deal and same thing with the anthropic models and I'm sure that we'll see something from Google where it's an acceleration.

2:15:53

Um, and I think that there is like a um, maybe an underappreciation of how much left there is to solve here. Yeah.

2:16:03

Because when when you even when you're doing like a real coding task as a pro, like despite all the demos you see on Twitter where it's like someone asks uh, you know, an agent to build an app, that's like a lower level of difficulty than doing what a prodeveloper does with one of these models.

2:16:16

And the models still don't produce great code a lot of the time.

2:16:21

Like there's a lot of kind of handholding that has to go into it.

2:16:23

And I think I think that we are still seeing an acceleration in terms of the models actually becoming not just like okay competent engineers but like really really good engineers. >> Yeah.

2:16:34

Do you care about benchmarks?

2:16:36

>> We care a ton about benchmarks like we um >> but your own internal benchmarks or or >> we we do both.

2:16:42

So you know plug for warp we're number one on terminal bench which is the public uh you know terminal benchmark and we're top five on SWEBench which is the coding benchmark.

2:16:51

And then the only way uh in my opinion that an app at our layer in the stack can really improve is by measuring the progress.

2:17:00

And so we have our own internal set of evals that we run across all these models as well which are coming from like real use cases and that again is an advantage of being like a product that's in the wild that has a lot of users is that we can sort of see where the models are failing where they're working and so we're we're very big on that actually. Yeah. >> Awesome.

2:17:18

Well thank you so much for stopping by.

2:17:19

We will talk to you soon.

2:17:22

>> Sure you'll have a busy afternoon.

2:17:23

>> Shout out by the way to OpenAI team.

2:17:25

Very very helpful in uh working with us to get GPT5 to be awesome in Warp.

2:17:28

And one more shameless plug.

2:17:31

It's we have a discount code for people who want to try GPT5 in Warp. It's $5 GPT5. >> Okay.

2:17:39

>> Thank you for having me guys. >> Yeah. We'll talk to you soon. Thanks. Cheers.

2:17:43

Uh Tyler, any updates from the timeline while you're thinking about what the latest vibe check is in the war between OpenAI's linear >> is a purpose-built tool for planning and building products.

2:17:56

Meet the system for modern software development, streamline issues, projects, and product road maps. Go to linear. app to get started. >> Choice for OpenAI.

2:18:03

You have something >> uh from Reggie James, front of the show.

2:18:07

Half of my timeline says this is the closest we've been to AGI.

2:18:09

The other half of my timeline says we officially just hit AI stagnation. I love tech.

2:18:18

>> Well, uh, we will be going deeper deciding whether or not this is stagnation or hyper intelligence takeoff.

2:18:24

Uh, and and we will be joined by our next guest, Riley from Charlie Labs. >> Sorry.

2:18:29

>> Hey guys, thanks for having me.

2:18:31

>> Good to see you, Riley. How you doing? >> What's happening? >> I'm doing fantastic.

2:18:33

We, uh, we've been heads down with GPT5.

2:18:36

Uh, and >> how long have you had it?

2:18:38

How long did you get the preview?

2:18:40

I feel like it it it you know it gets rolled out to early adopters a little bit earlier, but has been weeks, months.

2:18:46

How long have you had it?

2:18:48

>> We're couple couple weeks, two or three.

2:18:51

>> What was the first Charlie liking it? >> Charlie loves it.

2:18:56

And also I love what Charlie does with it. >> Yeah.

2:18:58

What does Charlie do with it?

2:18:58

What was the first thing you did with Chach GPT5? >> Uh ran our evals. >> Oh yeah. How'd they come back? >> Just uh really good.

2:19:06

um much better than 03 which was much better than any other model we've run before that. >> Interesting.

2:19:13

And yeah so so let's zoom out. What what do you do?

2:19:16

What do these eval measure? Um walk me through it.

2:19:21

>> Um so Charlie is a Typescript focused coding agent.

2:19:25

>> Um that operates much more like a human does.

2:19:28

Um so less like IDE application terminal and more joins your GitHub and Slack and linear workspaces.

2:19:34

Um, and it interacts with the team the same way other humans do.

2:19:41

>> Um, and then our evals are a mix of code review, um, because part of Charlie's job is to review PRs from humans, um, as well as his own, and then, uh, code authoring, so opening PRs and pushing commits.

2:19:53

Um, >> so, so when you develop your own p your own evals, I imagine you try and keep those out of any training data.

2:19:59

You want those to be held private. Is that correct? >> Yes.

2:20:03

And it's getting even harder with web access now because they're too good at finding things.

2:20:09

>> They're finding everything. It's funny.

2:20:11

Um and then and then talk to me about like the shape of those uh of the actual problems in the eval.

2:20:17

Are you are you doing are there some easy questions, some hard questions, some some some extremely hard questions?

2:20:23

Like how are you formulating those?

2:20:26

What's the shape of an individual task?

2:20:27

Is it scored out of like a hundred?

2:20:30

How do you think about developing a good eval?

2:20:31

um a mix of hard to very hard.

2:20:34

The easy ones are just a waste of money and time at this point, especially with five.

2:20:38

Like there's a bunch that it's just not going to get wrong. >> Yeah. Yeah.

2:20:42

>> Um and then we're mostly doing uh the PR ones look kind of like Swedbench in the sense that we're taking an issue to start with.

2:20:49

Um but instead of giving the issue like in a Docker container already, um we trigger a comment on the issue that says, "Hey Charlie, go make a PR for this."

2:20:58

Um, and then Charlie does his thing and then the PR comes up.

2:21:00

Um, and then we score that PR against a whole bunch of things like correctness to a known solution that's correct as well as um code quality, testability and some softer things like descriptions.

2:21:15

>> Who are the biggest uh who are the biggest customers or use or users for like a TypeScript focused coding agent?

2:21:23

Um, it's a wide range of mostly modern apps.

2:21:26

Like pretty much any web app these days is going to be like a Nex. js type app.

2:21:31

Um, and then all the way into like back end like Charlie himself is written in TypeScript. >> Sure. Makes sense.

2:21:36

>> And there's very little front end. >> Anything else? What else you got?

2:21:42

>> I just want to say I love the name Charlie.

2:21:43

It's one of my favorite agent names that we've had on the show. >> Yes.

2:21:46

It's right up there with pig and what was the other one?

2:21:48

Well, I don't think that was an agent, but uh >> that was an agent, but >> but yeah, it's a it's a good one.

2:21:53

>> Yeah, >> congrats on locking it down. >> Yeah.

2:21:55

What what about what about um cost and uh and that side of the business?

2:21:59

Is there is there any movement there or anything that you where where you require movement or you need movement to really unlock new capabilities in the business or new markets?

2:22:16

Not really for us because we're operating kind of as at a human level.

2:22:21

Um we do value based pricing.

2:22:21

So we charge per PR or per commit.

2:22:23

Um and because that's comparing to such expensive actions that humans are doing.

2:22:30

The challenge for us is more actually living up to the promise than doing it cheap. >> Yeah. Yeah.

2:22:35

Uh are you having >> but then but then doesn't the cost reduction announced today?

2:22:38

Isn't that great for business? >> Yeah.

2:22:43

I mean, it's good overall, but like that's our problem is not that the models are expensive.

2:22:46

It's that they're >> I mean, they're getting really smart, but I'll always take more. >> Never enough.

2:22:54

>> Like, for instance, >> since the beginning of August, we've been testing um 98% of the code that got merged into our codebase was written by Charlie. >> Wow. >> Not 30, not 50, 98%.

2:23:04

And that's coming through PRs.

2:23:07

That's not like autocomplete in an ID type thing. >> That's crazy. >> Yeah. What Yeah.

2:23:12

What does that mean for like the future of of like who are you hiring?

2:23:15

I imagine that you're still, you know, a an engineeringheavy organization that's just puppeteering and orchestrating agents.

2:23:24

Um, but where do you see like the future of um software development as a career path going? >> Yeah.

2:23:30

Are are uh new CS grads cooked?

2:23:37

>> I think if they get really good at using the AI, no.

2:23:38

if they try and take an approach of getting really good at writing code by hand. For sure. >> Yeah.

2:23:45

>> What we're mostly looking for hiring is people who are able to see things at a much higher level and plan further out because with tools like Charlie, you can write so much more code so quickly that it's like it's more important to see where you're going and take the right path than it is to be able to write it quickly. >> Very cool.

2:24:03

Well, thank you so much for stopping by.

2:24:05

Good luck with the rest of your day and uh congrats on a on an upgrade to everything that you do.

2:24:11

>> Tell Charlie to have fun out there. >> Have some fun. >> Thanks a lot, guys.

2:24:14

>> We'll talk to you soon. >> All right.

2:24:16

Let me tell you about numeralhq. com. Sales tax and autopilot.

2:24:19

Spend less than five minutes per month on sales tax compliance.

2:24:23

>> Sales tax super intelligence. com.

2:24:26

>> A number of the fellas in the chat got access to five. >> Break it down for us.

2:24:31

>> Reg says it's pretty good.

2:24:31

The writing ability feels a little nerfed.

2:24:33

says, "The way it writes feels a little programmatic rather than sounding human, reverts to using points even for things like blog posts and also >> uses overly complicated language for simple stuff."

2:24:47

>> Uh, Techno Chief says, "It's crazy fast."

2:24:51

>> Oh, that's >> Dan Rat Ratliff says, "Yeah, I was just going to say that very, very, very fast."

2:24:56

Um, Z Jean Ahmed says, "Junior devs are barbecued.

2:25:06

Tyler, anything from your side before we talk to Germa from Verscell?

2:25:09

>> Um, I think maybe a good way to to like vibe check on at least on the timeline is that it's almost like a 4.

2:25:15

5 kind of thing where comes out people are like this model totally sucks. Look at the benchmarks.

2:25:20

It's like not it's not some massive improvement.

2:25:22

It's like a bar, you know, not a step change at all.

2:25:23

But then you you start playing with and it's actually like, okay, this actually a good model.

2:25:27

Y like a lot of the stuff I'm seeing people post like, oh, that's actually like really like interesting output stuff like that. Um, but it seems good.

2:25:36

>> Can we do the green text eval green text bench?

2:25:39

>> Yeah, we got the TVPN intern. >> Yes. Yes. Yes. Yes.

2:25:42

Uh, we'll let you cook on that and then we will move on to our next guest, GMO Ralph from Versel coming in to TVPN for the second time. Great to see you, GMO. How you doing? I like the action hall. Thank you. Welcome to the stream. >> How you doing today?

2:25:58

Do you think GPT5 could beat me, you, couple of the boys here on Dust 2 in Counterstrike? >> Easily. >> Easily. >> Yeah.

2:26:08

It depends on the frame rate, right?

2:26:10

Like >> on a long enough timeline, we're we're cooked.

2:26:13

>> We're cooked, >> but we might frag it short term and we might be faster. >> Amazing. Yeah.

2:26:17

Uh yeah, we got to uh I mean, I'm sure we'll get to GPT5, but what's your reaction to the world model stuff from Google?

2:26:23

Uh do you think do do you have an idea of where that's going as a product?

2:26:27

It feels like a GPT2 level technology very much a research focused uh technology.

2:26:33

I'm sure OpenAI is working on something too and uh a lot of the labs will work on it but uh what what's your theory behind the the the generative video game world model stuff that's going on?

2:26:45

that's going on? I mean number one super fascinating right I think when when we think about the future I always think about Jensen's the future of of applications will be that pixels are generated not rendered so as as much as we're really excited today that GPT5 and Vzero are really good at writing code

2:27:07

that then renders interfaces >> I think it's also cool to dream of a world where we're just going directly from GPU to pixel grid right And but if you remember like a couple years ago and maybe a decade ago, there was a lot of excitement of video games that were going to be live streamed from the cloud. >> Yeah, that's right. >> Yeah, that's right.

2:27:25

>> Where your input, your keyboard, you could have a very thin client.

2:27:27

Your input, your keyboard, your mouse movement was going to be dispatched to the cloud.

2:27:32

>> We're going to have GPUs near you.

2:27:34

>> Google Stadia was big there.

2:27:34

And then >> Live was Microsoft game and is still Microsoft's actually still pulling it uh still pushing it very heavily.

2:27:41

awesome tech uh but not mass adoption. >> Yeah.

2:27:46

>> But if you look at you know a lot of these technologies are being really successful in letting people get more creative and test things out.

2:27:54

>> A lot of the use cases that we see for v 0ero and vibe coding are >> almost like a communication tool.

2:27:58

Like I want to prototype something.

2:27:59

I want to see what the what's possible.

2:28:01

I want to explore the latent space.

2:28:03

And I think those world models are going to be incredible just to inspire what the future of games could look like, right?

2:28:10

uh just getting ideas for actually then shipping them in a real uh 3D engine model.

2:28:15

I think short term I think long-term all bets are off.

2:28:16

Someone was saying in the chat, you know, junior devs are are roasted um or barbecued.

2:28:20

I think that's not quite true. >> Okay.

2:28:25

>> Uh same for like 3D uh engine developers.

2:28:28

>> Give us the bullcase for junior devs staying off the barbecue.

2:28:31

So the bullcase for I think people in general is that you move from I mean the progression in the industry has been assistant >> to agent >> Mhm.

2:28:43

>> to team of agents, agent orchestrator.

2:28:46

>> It's still really useful to have a human be the one that's sort of like managing the team. >> Yeah.

2:28:51

>> So you're moving from like junior dev to junior manager.

2:28:54

Uh especially as these tools become more agentic uh in in the new version of Ezero that's coming up really soon.

2:29:01

You're starting to notice that Vzero sort of splits the task between a little team.

2:29:07

>> You have the designer of the team.

2:29:07

You have the PM of the team that's sort of working on the spec. You have the architect. You have the engineer.

2:29:13

I don't know if you saw Cloud Code announced.

2:29:17

I think it's like SL security review. >> Yeah.

2:29:20

>> You think of think of that as having a security team or team of agents or security researcher at your disposal.

2:29:24

So junior dev as like a vertical skill might be a little barbecued but junior nench manager so I think it's just going to be the junior dev is so much more powered in this world if you allow yourself to be and you keep up with what these tools can do and and and I think you you stay you know at the cutting edge. >> Yeah.

2:29:43

I mean the obvious bull case is if you're like if someone's a college student today they can learn to code truly AI natively.

2:29:51

They don't have to say, "Oh, we're an AI native organization now.

2:29:55

We have to up upskill and kind of retrain people how to think."

2:29:59

They can just naturally start to think with these >> capab Alman post about how we'll look back on, you know, 93% of humanity was subsistence farming.

2:30:08

And if you ask those people how what they think about our email jobs, they'd be like, "You guys are crazy."

2:30:13

And it's almost like in the near future, midterm future, maybe even long-term future, it's like the number of individual contributors will be extremely low and almost everyone will be a manager and you'll become a manager much faster.

2:30:26

You'll just be managing agents and then you'll be managing people who manage agents.

2:30:30

But the job of almost everyone will become managerial.

2:30:34

May maybe that's what happens. I don't know. I'm not 100%.

2:30:35

But that that's what that made me think.

2:30:38

Someone asked me yesterday, you know, what do you think the future of uh the market of monitors looks like? Like does it stay flat?

2:30:45

Do people get more monitors because they're going to like Dogecoin trader analyst when like >> in the future everyone has the hedge fund six monitor set up or >> in the future everybody's just going to be at work on their phone.

2:30:59

I mean I've noticed that that what you know when I was an individual contributor I had three monitors.

2:31:05

I was programming on all the screens and now I r I mean I use my laptop during the show and then most of my work is done on my phone phone calls and then and then firing off tech messages.

2:31:16

Yeah, maybe maybe we actually shift away from monitors and go further into voice interfaces.

2:31:21

You oh I call the lead of my agents and then that agent relays it to some >> I'm very optimistic on voice by the way because I've now seen it.

2:31:29

Uh, I did what we're we're we're cooking on a on a a better mobile experience for Vzero. >> Sure.

2:31:36

>> And I was going back and forth with my head of mobile and he was talking to Vzero and I was writing down and a pretty f fast typer >> but he beat me with voice using the local model on the phone.

2:31:47

So there's still the question of like edge latency versus cloud latency kind of like what we talked about with 3D.

2:31:52

But I do think voice is going to play an increasingly uh exciting role in in programming which is kind of wild.

2:31:58

I would have never imagined.

2:31:59

I I've always been about like typing benchmarks in WPMs. Uh voice is coming. >> Yeah. Yeah. Yeah.

2:32:07

>> How do you think about competition broadly in developer tooling code gen?

2:32:09

I mean it it right now it seems like there's just so much demand.

2:32:16

>> It feels like massive TAM expansion moment.

2:32:17

Every company's ripping >> TAM expansion moment, but at the same time winners will emerge.

2:32:21

Uh obviously you're you're playing to win.

2:32:25

Um, and yeah, I'm curious, you know, >> yeah, on some level we're playing both sides of the bat.

2:32:33

We what we announced today that's really exciting is Vzero with GPD5 support.

2:32:40

>> So you can go to vzero.

2:32:40

dev/gpbd5 and we'll use GPD5 in combination with our model pipeline that makes it really good at by coding especially for nontechnical folks.

2:32:50

But we also on the Verscell AI cloud side of things, we open sourced basically you can create your own vibe coding platform powered by any model.

2:33:02

>> I was joking about this with Tyler.

2:33:02

Vibe code me a vibe coding platform please like one. Make no mistakes. >> Yeah.

2:33:08

V code me a billion dollar company. Yeah. No mistakes.

2:33:10

Uh but basically we are giving people that as a starter kit. >> Sure.

2:33:15

And B, by the way, the fundamental question that a CEO asked me the other day was, is VIP coding a product or a feature >> or is it both? You know, it's TBD.

2:33:26

>> The case for feature is okay, so there's going to be lots of systems of record.

2:33:32

>> Think Salesforce, Snowflake, >> data bricks, >> and increasingly they're going to incorporate codegen capabilities into their platforms.

2:33:39

their platforms. they can use a lot of these capabilities that we just open sourced and you'll go to their the existing place where you have the data kind of like what we've talked about for decades of like are you bringing computer to the data are you bringing vibes to the data right are you bringing

2:33:54

codegen to your own platform so >> you used to used to bring like a bu like a you know dashboard builder and it would have a couple widgets and now I could just potentially if I'm plugged into some sort of data source some system of record I could say vibe code this app on top of it there's some tool

2:34:10

retools played in this space zappier a little bit but yeah I mean this feels like you know we're getting we're not fully in the just the pixels are generated but we're you know generative UI generative application on top and that and that being bespoke and ad hoc >> I also think it's important to understand the line between consumer

2:34:29

vibe coding and just generating ephemeral software and websites and things like that versus enterprises which will have a lot of different use cases when I look at the when I look at the vibe coding market and I see businesses that are that that are almost entirely consumers just creating things for fun. I think that has to be a tough

2:34:46

I think that has to be a tough business because it's a hyper competitive market and consumers are flaky.

2:34:52

They'll create something you know for fun but they'll churn in month two because you know it's not they're not running a real business.

2:34:59

Whereas a a business knows, hey, we'll pay for this on on a long-term basis because we're we have a use for it all the time from this product manager to an engineer over here to somebody in marketing, etc. >> Yeah.

2:35:12

The the other side of the of the equation is how do you make these VIP coding tools work really well for enterprises?

2:35:19

Frankly, the most surprising emergent thing that I've learned is just how much demand there is in enterprises for VIP coding.

2:35:25

And this is because a lot of the the traditional thing has been the people that understand the business are sitting over here.

2:35:33

>> The people that understand the code are sitting over here and their communication is fraught with peril.

2:35:39

Like they don't speak the same language.

2:35:42

They kind of like resent one another.

2:35:42

I love to tell this story.

2:35:44

I was meeting with a CEO of a very successful company.

2:35:47

He was telling me that engineers like asking a feature to his own engineers felt like petitioning the government. >> Yeah.

2:35:55

even though he's the CEO, it's like he's struggling to like make the case >> and please like get me in your next sprint, get me this feature.

2:36:02

So that coding actually solves that problem.

2:36:06

>> All of the PMs, designers, marketers, business users that previously only had access to what like Jira and uh you know to-do lists and project management tools and writing PRDS and so those kinds of things.

2:36:21

They they weren't able to ship PRs.

2:36:22

they weren't able to, you know, ship software and now they can.

2:36:24

And so the the opportunity is how do you actually make this secure?

2:36:30

>> How do you make it high quality?

2:36:32

>> How do you create a guard rails?

2:36:32

And those are those are tricky problems.

2:36:34

And I'm I'm really happy that some of them are easy to overcome and at least for us >> and some of them are active areas of research, but I think the the enterprises really have a strong case for this. >> Yeah.

2:36:46

Can you walk me through like tool use?

2:36:47

I mean, we were talking to the OpenAI folks about GPT5 being like really like a summation of like standing on the shoulders of giants.

2:36:53

You get a Python ripple, you get a web browser, you get, you know, the ability to kind of run cron jobs.

2:36:57

Now, there's voice and, you know, all sorts of different tools kind of wrapped up into one multiple models.

2:37:02

You can trigger reasoning chains if it wants.

2:37:04

It can do all these different stuff.

2:37:06

And that's actually the benefit of like this isn't just a bigger model.

2:37:09

It's like a lot of it's a it's a next version of a thing.

2:37:13

It's more like switching from the iPhone 12 to 13 than going from the iPhone to the iPhone 3G.

2:37:18

It's not just a new technology that's in there.

2:37:20

Um, but in the in the in the world of vibe coding, what are the tools that you want to think about adding?

2:37:28

I know that basically every vibe coding platform uh you know recommends a database.

2:37:32

Um, but I was we were talking to Harley at Shopify yesterday and there's a world where if I go to a vibe coding platform and I say I'm building an e-commerce website, it should probably just be like, hey, I'm going to do Shopify under the hood and I'll vibe code the landing page on top.

2:37:47

But how are you thinking about the landscape of like tools that you could pull in o full open because there's open source repos that are like full projects that you could pull in and then just start customizing on top of.

2:37:58

It's kind of this big continuum.

2:38:00

>> Yeah, there's a couple layers on on the foundation model layer.

2:38:02

What you want is a model that is exceptional at tool calling.

2:38:07

>> Whether it has built-in tools or whether you register them yourself, this is a like sort of silent word has been going on.

2:38:13

Like if you talk to devs, what are you optimizing for? Tool calling quality. Why?

2:38:18

Because to demystify the word agent, what an agent is, it's a loop of tool calling that builds up context over time.

2:38:28

>> That's all an agent is.

2:38:28

So let's to give you an example concretely of B 0.

2:38:30

Vzero is becoming more and more agentic over time.

2:38:35

One of the things that it can do is it can take a screenshot of the thing that's building and reflect on it.

2:38:41

>> So today I live vibecoded to an audience of web 3 and crypto engineers >> and I told Vzero, hey make this dark mode and initially Vzero does me dirty.

2:38:53

He's like he changes some things with dark mode and then it kind of astonished me because I was like oh I have to now explain to this audience.

2:38:59

It then takes a screenshot, looks at it, and keeps fixing it.

2:39:04

And I was like, this is literally a developer that's alive on autopilot.

2:39:08

And the reason it's on autopilot is because he has access to these tools like looking at the web browser. Another one is research.

2:39:13

I vive I've coded an example of build me a Substack clone for cryptocurrency news.

2:39:22

And the agent didn't know what the cryptocurrency news were.

2:39:24

So I started doing research on the internet of okay Ethereum passed certain price and whatever.

2:39:30

So and then you're talking about the tools over the internet.

2:39:32

So to demystify another topic MCP is really exciting because it's a new protocol for registering tools that your agent doesn't locally have.

2:39:43

So those tools that I just talked about we gave them to VZ.

2:39:46

Here's a deep research tool.

2:39:46

Here's the uh screenshotting tool.

2:39:49

And those will likely become the new services.

2:39:54

When you think about like AWS of today, if if AWS was an AI cloud, uh which is kind of what we're trying to build at Versell, like you think a lot of those tools are going to become as a service, like bring me the research as a service, bring me browsing and screenshotting as a service and so on.

2:40:08

But then you have MCP, which allows you to okay, I need to sell something online.

2:40:12

All right, so now there's an MCP for Shopify.

2:40:15

Now there's an MCP for Stripe.

2:40:19

>> Uh there's even crypto MCP.

2:40:19

So it's really exciting like now it's like the ultimate choice for a builder and you don't have to go and learn all these things.

2:40:27

You don't have to this is almost like a discontinuity of the valley trend of like if we build amazing documentation they will come.

2:40:34

This is more so if the agent picks you, they will come, right?

2:40:39

And so there's a lot of uh figuring out right now like how do I make my infrastructure, how do I make my product to be loved by these agents?

2:40:48

And the MCP promises to be one of these first uh things that you are in control of.

2:40:53

>> That makes a ton of sense. >> Last question.

2:40:54

Uh someone on your team named Josh is in the chat.

2:40:56

He wants to know what what does he need to do to get a Twitter badge?

2:41:00

Oh >> well yeah 100k downloads of the AICLI.

2:41:04

I think we've been talking.

2:41:07

Uh okay been thrown down. >> Thank you.

2:41:12

>> It's on your work cut out for you.

2:41:14

>> It's burned into the immutable record of this live stream and the future training runs.

2:41:19

>> Best of luck you accountable now.

2:41:21

>> We're going to hold you accountable to that GMO. Great seeing you. Great to see you. We'll talk to you soon. Congratulations.

2:41:27

>> Let me tell you about Finn. ai.

2:41:27

The number one AI agent for customer service.

2:41:30

Number one in performance benchmarks.

2:41:32

Number one in competitive bake offs. Number one in IR in G2.

2:41:38

Number one in having an Irish founder. >> That's right.

2:41:41

>> And we will invite our next guest to the stream from factory. ai. Welcome to the stream. How are you doing? Good to see you. >> Hey, how's it going? Glad to be here. >> Great. Thanks so much.

2:41:52

Kick us off with an introduction on you and the company.

2:41:56

>> Yeah, my name is Eno, uh, co-founder, CTO at Factory.

2:41:58

uh we are building a platform for enterprise software developers to perform what we call agent-driven software development.

2:42:05

So basically more than just code bringing agents into every stage of the software development life cycle.

2:42:12

So think coding, code review, maintenance, incident response, documentation.

2:42:17

Uh we think agents should be a part of all of this and we think that they should be driving a lot of that menial component while you think at the high level about how to plan and structure the work.

2:42:30

>> There's so many different >> like enterprise is a narrow cate it's you know oh not consumer I guess but uh it's such a wide it's such a wide category. Is there a beach head?

2:42:39

Is there a certain type of project within within different industries or specific industry that's getting a special an especially large amount of value out of factory these days? >> Yeah, totally.

2:42:51

I think that one thing that we see a lot and typically when we say enterprise we're thinking greater than 1,000 engineers, right? Like 2,000 3,000.

2:42:59

And one reason why we focus on that larger scale, you tend to have uh these large organizations where people are the the bottleneck is not code, right?

2:43:11

The bottleneck is how do we plan a migration of 185 code bases to this new framework and there are 3,000 developers that are going to touch this over the next 6 months and an SI just told us the quote is $80 million to do it.

2:43:29

it. uh and we have to figure out how >> replatforming broadly is one of the major major uh tasks uh for many many enterprises right >> 100% modernization and migration is huge >> yeah yeah that makes a lot of sense how do you estimate the market that market size and is that is that what you guys

2:43:49

are leading with on the GTM side in terms of trying to find these legacy companies that are maybe not even using cursor yet I mean we who who we we talked to the CEO of GitHub yesterday and what >> 50% didn't he say or >> it was like at least half of their user base is not using any AI tools. >> Yeah. >> Yeah. >> Yeah. Totally.

2:44:09

I I think that the the the thing that we hear often we pretty much only deploy into companies that have already tried an AI native IDE or have an autocomplete tool deployed.

2:44:20

And I think that the thing that we hear often is you you sort of hear like these numbers thrown around like 5x 10x uh and then in practice when you adopt an AI IDE you see 10% 15%.

2:44:30

IDE you see 10% 15%. And so a lot of people are sort of saying like what is the delta there like what causes that transition and our our sort of argument here is that there is a workflow change that's actually required to really adopt agents in the life cycle right and so if you're just sort of like accelerating an

2:44:50

individual developer uh that you can go a little bit faster but if you are able to parallelize and automate at scale that is going to be that larger introduction of change and so if you imagine the market here There are companies where you know 5 or 10% of global payment transactions run on some cobalt system that was written 40 years ago. Every developer is gone and they

2:45:12

Every developer is gone and they are it's a ticking time bomb.

2:45:15

Like at some point it needs to go to Java but there's nobody who even knows how to do that.

2:45:21

And so those are the types of projects where the market is so enormous because you know half the business runs on this legacy system uh hundreds of billions of dollars.

2:45:30

put it all in lisp, skip Java, go straight to lisp. Um, >> yeah, exactly. Python, right?

2:45:37

>> Python would be the logical one.

2:45:37

Um, uh, I'm sorry, we're running behind, so we're going to have to cut this short, but, uh, I want to know more about how the enterprise coding agent market will develop.

2:45:49

we could see one world where we wind up with, you know, GCP, Azure, um, AWS, like, you know, pretty comparable competitive.

2:45:59

They've all had really great margins.

2:46:01

It's been this oligopoly.

2:46:03

There's another world where you could see more specialization.

2:46:05

One of these companies goes deep into high security environments or oil and gas or financial environments or specializing based on specific programming languages.

2:46:16

um as as the market develops like how do you think it'll play out? >> Yeah, great question.

2:46:22

I I think that what's very clear is that the bulk of very large enterprise has a lot of similar problems, refactors, migrations, modernization.

2:46:31

So, uh a platform like factory is able to deploy into that and solve problems quickly.

2:46:36

Uh I I think that there's likely to be like that sort of 8020 where there are going to be these very specialized providers that only focus on one sort of problem.

2:46:44

Um, and that will represent maybe like 20% of what's out there.

2:46:49

Uh, and so it won't be like necessarily black or white, but we do think that the bulk of enterprises have a lot of similar needs.

2:46:55

Uh, especially when you just get cross a certain threshold of number of engineers, scale of code base. >> Sure. Sure. Yeah.

2:47:03

I mean, we we we even see that with the the the clouds where, you know, obviously there's the hyperscalers, but then there are neoclouds and we talked to Armada where they'll they'll send you a shipping container with a bunch of racks inside and put it in stranded energy.

2:47:14

So, there will obviously be the a long tail here. Uh that's a great take.

2:47:18

Thank you so much for stopping by.

2:47:20

Have a great rest of your day and uh enjoy the GBT upgrade. We'll talk to you soon. >> Have fun out there.

2:47:26

>> Really quickly, let me tell you about Adio customer relationship magic.

2:47:28

Adio is the AI native CRM that builds, scales, and grows your company to the next level.

2:47:34

And we will be joined by our next guest from Augment. Welcome to the stream. How are you doing, Guy? >> Great.

2:47:42

Thanks so much for having me.

2:47:43

>> And that's his name, by the way, if you're listening. His name is Guy.

2:47:45

I'm not just calling him Guy.

2:47:46

Uh anyway, please introduce yourself and what do you do?

2:47:50

What does your company do?

2:47:52

>> Yeah, so I'm Guy Gerari from Augment Code.

2:47:54

I'm a co-founder and the chief scientist and we build AI coding assistants for large teams with large code bases and so you can use augment code to do question answering to do development to do refactoring to do migrations all the tasks that you do except that our product understands your large code base really well and so that means less prompting for you and uh faster and better results out of the agent. Today GPT5 launches.

2:48:17

It's kind of a rising tide.

2:48:21

Feels like it lifts all boats.

2:48:23

Every company gets access to it.

2:48:25

We've interviewed a number of companies that are building on top of GPT >> around GPT4, >> I guess.

2:48:30

But but in general, uh how do you think you can use GPT5?

2:48:34

Are there any pockets of value that you think you can uniquely take advantage of? >> Yeah, great question.

2:48:43

So we've been we've been triing the model for the past few weeks and what we found is that the GPT5 is a very thoughtful model.

2:48:49

It likes to make a lot of tool calls.

2:48:53

Uh it likes to ask clarifying questions of the user before starting to make code changes.

2:49:01

Uh and so the place where I reach out for GP5 is typically if I need to make large changes or if I'm trying to answer a very difficult question about the codebase, I will let GPT5 take a crack at it.

2:49:12

It will churn for a while making lots of tool calls just making sure it got it right and probably find all the different places in the code where it actually needs to make a change and so I will typically let it run in the background and come back to it and I will often get a high quality result out of it.

2:49:30

>> Are there any features or integrations that you're hoping GPT5 will roll out in the future?

2:49:37

We talked to a couple people who were like like we want models that have access to as many tools as possible.

2:49:44

Uh and and you can see with the MCP boom more people are trying to make their uh their services, their products accessible to these models.

2:49:51

Uh is there anything that you see as um potential lowhanging fruit to just add to the capabilities?

2:50:02

So I think for us we work hard on developing our own integrations and our own tools building them into the product rather than relying on um GPD5 or other model vendors uh to do so.

2:50:12

We have worked closely with open AAI to improve the prompting around our tools so that the agent kind of works uh flawlessly.

2:50:18

I think um the thing that would be very nice I think one of the previous guests mentioned a screenshot tool.

2:50:25

I think that's a very yeah that's a very nice way to close the loop on front-end software development >> just like we saw how on backend software development running the tests automatically really helps the the agent uh iterate until it gets to working code.

2:50:41

Uh so I think having more support for screenshotting uh and things like that that close the front end gap would be very nice to see.

2:50:49

>> I I wasn't aware that that that that screenshots weren't flowing through.

2:50:51

I feel like when I've when I've triggered uh operator, I'm getting a a view a web view into the website, but um I wasn't I I wasn't aware that that wasn't like being passed through easily in the API and you still kind of needed to build that yourself.

2:51:06

Um what where else um we were just talking about this like um where are the biggest p pockets of value right now for AI coding tools generally?

2:51:17

Obviously, everyone knows like the vibe coder who's just the designer who's learning how to use uh software for the first time.

2:51:23

Then there's the experienced uh developer going from a 10x to 100x with better code completion.

2:51:28

Then there's the enterprise that's uh you know maybe doing replatforming.

2:51:31

Where else are the interesting pockets of value um that are maybe on the horizon to be unlocked with new models? >> Yeah.

2:51:41

>> Yeah. So on top of everything you mentioned certainly the the inner loop of software development that's where we've spent most of our time at uh augment code developing product for um yes you can have a senior developer starting using agents starting to use multiple agents in parallel and unlock 10x or more productivity gains what

2:51:59

we're starting to see now with our tools is the beginning of automating software development life cycle tasks so with with augment code we have a CLI tool now where you can take the full power of our context engine and the agent, the thing that really understands your codebase and you can start automating tasks in the background. And so we're seeing more

2:52:18

And so we're seeing more and more developers saying, "Oh, this is great.

2:52:21

Like, I can break out of the IDE now.

2:52:24

I'm using the agent that's already familiar to me, but I'm starting to automate code reviews.

2:52:27

I'm starting to automate incident response.

2:52:29

I'm starting to automate um looking at production logs and automatically assigning tickets based on error logs that I'm seeing.

2:52:35

all kinds of new uh uh automation use cases that we're seeing just because agents have gotten so good and kind of really understands your codebase.

2:52:46

>> Are there are there high stakes pockets of software engineering work that most of the AI tooling has kind of stayed away from?

2:52:52

I'm imagining like the uh the the high stakes database migration.

2:52:59

Where where is the the the kind of sticky um part of the industry?

2:53:02

sticky um part of the industry? I was reading a blog post by someone who was doing like very advanced cyber security pen testing and they were saying like just the creativity of the models wasn't quite there yet to really come up with the to really act and embody like a

2:53:20

white hat hacker who was going for a bug bounty but uh where where are the pockets of still like intractability where I guess if you are you know in the in the individual contributor you love just just you know coding from scratch that's where you want to stay for at least the next couple of months. >> Yeah, I think still the attention of all

2:53:38

>> Yeah, I think still the attention of all the models we've seen and all the agents we've seen around making proper design and architecture decisions.

2:53:48

Um, that's still high stakes and still the ability is not there because if you do complete vibe coding and you just let the agent go and do whatever it wants, in the beginning it looks amazing.

2:54:01

The code works and it's all really good.

2:54:02

But once you get to low tens of thousands of lines, the bad decisions that were often made around the design and architecture start to show up and development slows down.

2:54:15

So that's where we still see a limitation of today's agents and where you still have to supervise the agent fairly closely uh in order to make sure that you don't get stuck later on.

2:54:24

Uh perhaps this will change in a year but today I would say all these decisions that you make around how the code is structured uh still requires close supervision and still high stakes because it can really slow your project down if you let it go autonomously for long enough.

2:54:40

Yeah, >> that makes sense.

2:54:41

Well, thank you so much for stopping by.

2:54:43

We will talk to you soon.

2:54:44

Have a good rest of your day. >> Thanks so much. >> Cheers.

2:54:47

>> Let's check in with Tyler on the timeline.

2:54:49

Tyler's manning the timeline. How are the vibes?

2:54:52

Are there any new posts that have hit the timeline?

2:54:53

Are we still in turmoil or has the narrative settled?

2:54:58

>> The vibes are are picking up a little bit.

2:55:00

You're starting to see people post like, "Oh, this is something I made."

2:55:03

Um, now you can see on Marina it's number one. >> Wait, wait, wait.

2:55:07

So, what's going on with the Poly Market then?

2:55:09

>> So, Poly Market is still uh still Google heavy.

2:55:13

Yeah, I think I guess they're just pricing in a Gemini 3. >> Okay.

2:55:17

>> Um, everyone's >> I'm not exactly sure, honestly.

2:55:18

I was actually very surprised to see that it was number one. >> Yeah.

2:55:22

>> Um, but yeah, maybe later we can show some of the posts. >> Yeah. Yeah. Yeah. Yeah. That' be great. >> Cool stuff.

2:55:26

>> Um, well, in the meantime, before our next guest, let's tell you about eight Sleep Get a Pod 5-year warranty, 30-day risk-free trial, free returns, and free shipping.

2:55:35

And we will uh have our next guest join us from Code Rabbit. How are you doing? Good to meet you. I'm good.

2:55:43

>> Good to meet you as well.

2:55:43

Thanks for having me here.

2:55:45

>> Uh what's your reaction to GPT5?

2:55:45

How long have you been playing with it?

2:55:48

What are the biggest improvements that you've noticed?

2:55:53

>> Yeah, I would say mind-blowing, right?

2:55:53

I mean, we have been playing, our team has been playing for like a few weeks now.

2:55:57

U tested a few snapshots and it's amazing.

2:56:02

It's a generational leap.

2:56:02

We would say like we have been using open AI models like I don't know how much you know about Kodak.

2:56:09

about Kodak. It's been like couple of years we have been on open AI anthropic um and our product is a very reasoning heavy product like one of the very few use cases where you are have a PhD style problems where we have to do code

2:56:23

reviews that's what code raid does like users open up pull request our agent uh uses reasoning models to find issues like race conditions or security issues and so on >> um so yeah so we've been testing GPD5 on some of the hardest pull requests we have in our golden data data set. So we

2:56:38

So we maintain a data set where we track progress of different models and progress of AI in general.

2:56:44

So we have like many problems that no model is able to use uh solve so far like I mean GPD5 but so far it has the highest score.

2:56:51

We would say it's like almost 2x better than the next 03 or solid or opus at this time.

2:56:59

>> What's the customer value there?

2:56:59

You think that uh like all the customers will just notice that the product gets better?

2:57:06

Are you going to upsell folks?

2:57:08

Are you going to uh like how do you play this given that this model is now in public availability?

2:57:12

Every company, every competitor can can access it as well.

2:57:17

>> Yeah, there's no upsell.

2:57:17

Like that's the thing with AI.

2:57:18

I mean, for the same price or even like better prices, you're getting much more AI, much better A B A B A B A B A B A B A B A B A B A BI like that's the whole idea how fast this space is evolving.

2:57:26

Um so yeah, from the pricing point of view, we don't see like this to be like a separate plan or something in our product.

2:57:33

I mean the pro for the same price per month customers will now just get better quality of results with code rabbit.

2:57:42

>> Um what's next for the business?

2:57:42

Um what kind of customers are you going after?

2:57:47

Who do you think has been on the fence and this release is going to be the thing that gets them to actually jump into the world of AI?

2:57:55

Yeah, we could track the topline metric like one of the things we track very closely in the company is like how many signups to the paid customers we get >> and that number has been constantly improving since GPD4 uh GPD4 turbo GPD4 actually dipped.

2:58:12

So there was a time when GP4 was almost like a Windows Vista of Elius like um it's funny like how we kind of trusted the evals and we thought it's the same model but in in a way it was imperial in many ways.

2:58:22

Um then we saw a huge improvement after 01 came out.

2:58:27

O1 preview was a game changanger for us.

2:58:29

Even at that time our conversion doubled actually right I mean so we went like more like close to 30% success in getting the paid users and now with JPD5 we are hoping we can see another big jump in the in the number of people who start becoming paid customers and and how many people churn.

2:58:47

how many people churn. So those are the real numbers like one is like wibes like how people like respond to the model and we get angry tweets or not though that's the other part but the other thing is like the actual revenues uh whether it moves the needle for us and that needs to be seen like one of the things we have seen even though you test these models in a lab they be it's not like a

2:59:07

huge data set but once you actually are in the wild you see hallucinations some of those scale issues at scale pop up so those are something we'll still be observing over the next few days to see whether it's like spar only um 80% of the cases, but then if the false positive rate, the hallucinations are too high, then also it's not a great model, but that remains to be seen. >> Yep, that makes a lot of sense. Uh well,

2:59:25

>> Yep, that makes a lot of sense.

2:59:25

Uh well, thank you so much for stopping by.

2:59:28

Congratulations on a new new tool in the tool chest, >> new toy.

2:59:34

>> Uh we will talk to you soon.

2:59:34

Have a great rest of your day. >> Cheers. >> Goodbye.

2:59:38

>> And let me tell you about public.

2:59:38

com investing for those who take it seriously.

2:59:41

They got multi-asset investing, industryleading yields.

2:59:45

They're trusted by millions. millions.

2:59:47

>> Um the the chat is going wild about public trading the S&P 6,900.

2:59:54

I think that comes from someone talking about like the non-MAGG7 stocks or something.

2:59:59

There's been uh people benchmarking the the Mag 7 versus the >> the big news while we were live or earlier today, Trump signed an executive order that is opening up 401ks to digital assets and private equity.

3:00:13

>> What what's what's crypto doing? Is it ripping?

3:00:17

>> Bitcoin is up a couple points.

3:00:17

Last time >> this point, you know, where's it going to go? It's already so high.

3:00:23

>> I mean, it's just like there there's been so many catal it could go up, it it could go down. >> Yep.

3:00:29

>> We have to wait and see.

3:00:30

>> Um Tyler, anything else notable from the timeline? What have people built?

3:00:33

I see this uh GPT5 just oneshotted a Minecraft clone.

3:00:39

>> Yeah, I think that's one of the cooler things I've seen.

3:00:41

Okay, so this is so it wrote it wrote code to generate this game that it's not generating the pixels.

3:00:47

You could do so many different things like you could generate you could generate a video, you could generate a world model, generate code that generates a game engine, you could generate code that runs on Unreal Engine.

3:00:57

I don't even know what they're using >> like um in on actual chatbt in there's like a native like it's like a music player.

3:01:04

It's almost like Garage Band.

3:01:06

You can say like if you prompt to like build I saw Samwin tweet about this.

3:01:08

you prompt like to to do some kind of like beat or something, >> it'll like make an interactive like GarageBand almost interface in there. That's cool.

3:01:16

I was playing with that earlier.

3:01:18

>> Yeah, I I do wonder how many of these features that we're seeing like where does OpenAI want to keep things in the B2B world and let other companies build versus just build it as as a consumer app?

3:01:28

Like yeah, um like will chat GPT eventually just let me push a website?

3:01:34

Like will it will it become a vibe coding platform?

3:01:36

um at least like a basic one like it's not it's not the most advanced >> coding environment but it can definitely write some code and execute it for you and do some stuff. >> Yeah.

3:01:47

Well, it's funny because like it used to be >> you would have a like so there was like GBT 3.

3:01:51

5 or something and people >> on top of that built a vibe coding thing.

3:01:56

So you could use that to build your own vibe coding thing but now you can just go straight from >> chat GBT to build your vibe coding platform. >> Yeah.

3:02:03

But now soon maybe it'll just be the vibe cutting platform.

3:02:07

>> The surface area of this stuff is is very interesting.

3:02:09

Clearly they're going after health care and therapy.

3:02:10

It's it's interesting that they've kind of stayed away from legal even maybe that's just the the the dynamic of the sales process and the dynamic of that particular market.

3:02:22

Um but I mean increasingly you can you can just ask more and more questions of Chachi PT.

3:02:27

So uh the the the consumer to business like bleed over.

3:02:31

Um there's certainly a world where just giving everyone in your organization chat GPT is a substitute for a bunch of different SAS products.

3:02:41

So be interesting to see where that developed.

3:02:42

What what are you thinking about that?

3:02:43

>> Near says there are concerns that the number used to represent our AI's intelligence does not in fact represent its intelligence.

3:02:49

Worry not to address these allegation.

3:02:51

We've added three new numbers. >> Near. Yeah. Near.

3:02:56

Near is building something that's like not particularly benchmarkable, right? Isn't it a companion? It's beyond benchmarks. >> Beyond benchmarks.

3:03:03

Well, in in completely other news, uh, Anderl opens a Taiwan office and begin selling AI powered attack drones to Taiwan.

3:03:10

Palmer Lucky has said he wants to turn Taiwan into a prickly porcupine.

3:03:14

We're in the age of spiky intelligence.

3:03:17

That spiky intelligence will be onboarded onto the AI powered attack drones and deployed in Taiwan to keep it safe.

3:03:25

What else is going on in the timeline while we wait for our next guest from OpenAI to join?

3:03:33

>> Spor says, "Raise your hand if you were not automated today." I'll raise my hand.

3:03:37

>> I was not automated today. We survived. We made it through.

3:03:39

Um, Sebastian Bubac says, "Here at OpenAI, we've crash we've cracked pre-training, then reasoning, and now we're experimenting with new set of techniques that maximally leverage their interaction.

3:03:50

GPT5 is just the first step in this direction.

3:03:52

We're excited to uh uh incredibly excited to see where scaling this up will lead us.

3:03:59

And it's the unicorn test, I believe.

3:04:02

And the latest unicorn is really really good.

3:04:03

That is that is a creative interpretation.

3:04:06

And I think it has a draw this with like SVGs.

3:04:07

Anyway, we can talk to our next guest about it.

3:04:11

>> Last post Jirro Ticket says, "I went to the permanent underclass party and everyone knew you."

3:04:18

>> Anyway, uh back to the serious interviews.

3:04:20

Welcome to the stream, Max. Good to see you. How are you doing? >> What's happening?

3:04:24

>> Nice to meet you guys. Yeah. Uh doing well.

3:04:26

It's a relief to have this launch out in the world.

3:04:28

I think it's uh you know, we've been working on this for the last few months now and it's exciting to let the whole world see what we've had.

3:04:36

>> Yeah, >> just a few months.

3:04:38

>> It's been uh I don't know.

3:04:38

It's been a little while.

3:04:41

>> What's the what what's the actual launch day like?

3:04:43

Because you're actually getting this out into the world.

3:04:44

uh the GPUs are on fire or about to be on fire warming up but uh is that is that out of your purview?

3:04:52

There is a different team for that fortunately.

3:04:55

>> So right so I run uh I ran a lot of the research for GBD5.

3:04:58

I don't necessarily handle the deployment but I do get dragged in when the GPUs are on fire.

3:05:02

Um I I think we're we're moderately burning right now. >> Okay. Okay.

3:05:08

>> Like a two alarm fire. >> Yeah. Yeah. Yeah.

3:05:10

Uh is it is it materially different?

3:05:11

I mean, this is a launch day, but uh we'll probably discover like the Studio Gibli capability once it gets out into the long tail of like, you know, hundreds of millions of people try it.

3:05:21

Someone comes out with some genius thing, then everyone's doing that and then the GPUs because I I feel like the Studio Gibli thing happened like a few days after the launch of images and chat GBT like >> it. It did.

3:05:33

It was um it was pretty fast, but it within about a week.

3:05:35

I think in this case, we're going to we're going to see that here. >> Okay.

3:05:40

Um, I think coding, you know, if I had to take my bets for what the Studio Gibli thing is going to be, it's it's coding.

3:05:44

Um, that's the place where I think GBD5 is like most tangibly a huge leap ahead of GPD4 and ahead of 03.

3:05:49

Do you think there's a chance that that that the coding will mean a Studio Giblly style uh meme or or kind of like uh and and what I mean by that is that is that like image generation is incredibly valuable in the context of like Hollywood will be using AI to chroma key and and rotoscope in a professional environment.

3:06:11

But yeah, what was special about Studio Gibli was that anyone was making these custom images and I could imagine a world where, you know, even going from like the levels IO example of like I vibe coded a flight simulator.

3:06:25

If we wind up in a studio Giblly moment for coding, I would imagine it's like everyone built their own game today.

3:06:32

>> I think that's pretty much it. Yeah.

3:06:32

So I I don't know if you guys watched the live.

3:06:35

That was that was one of the things we had on the live stream.

3:06:37

Like you can just go into chatbt.

3:06:38

Um, if you try it right now, it might or might not work because the GP5 rollout is still ongoing.

3:06:44

But if you have five, you can just tell it like basically make me a game. >> Yeah.

3:06:49

>> And it will make it and you can actually play it in Chad PT. >> That's amazing. >> So I Yeah. >> discover that.

3:06:56

And you don't The thing is like like with Studio Giblly, right?

3:06:58

Like for for Gibli, you don't have to know how to draw to make it work.

3:07:00

For this one, you don't have to know how to code. >> Yes.

3:07:02

But >> can you share Can you share that chat and someone else can play the same game?

3:07:07

How how does the kind of sharing mechanism work?

3:07:08

Yeah, there's a there you can do the share link.

3:07:10

Um we're we're I think going to try to make sharing for these a lot better over the next few days.

3:07:15

That was P2 after the P1 and P 0 of making the GPUs not completely melt. >> Yeah. Yeah. Yeah.

3:07:21

>> But yeah, I we will try to make it much more terrible. >> Yeah. Yeah.

3:07:25

I mean the the Studio Giblly thing is so interesting because it was it's not just that the model capability was there but it's also like the prompt was two words and and and it was so reliable that you always got a good result and you could personalize it.

3:07:39

So even if it wasn't like I've seen people build Doom.

3:07:44

I've seen people you know you can just buy Doom. It's a real game. You can build it.

3:07:48

Um, but if you build it and I'm like, "Oh, that's cool.

3:07:50

He did it in a vibe code environment or in chatbt, like that's awesome, but I don't necessarily want to go do that for myself."

3:07:56

But as soon as it becomes personal, which is what the studio, like I had to see what I looked like as Studio Gibli, I had to see what my favorite photo looked like, my favorite meme looked like in Studio Gibly.

3:08:04

And once that happens with games, people will will eventually uh, you know, there'll be this like mimemetic explosion and you'll see the GPUs will truly be on fire. >> Yeah.

3:08:14

I mean, I think even today you could probably with GBD5 do Doom, but all of the characters are like all the enemies are head shot of your friends.

3:08:21

Like, >> here we're going now. Now, we're real close. Yeah, we're real close.

3:08:24

It's going to be something that's personal, something that, you know, you can express your own creativity through because I think people, they still latch on to that.

3:08:31

They don't just want, you know, a copy of what already exists. They want something new.

3:08:35

And and and the Studio Gibbly moment was just new enough.

3:08:37

Anyway, we should talk about actual research.

3:08:39

We should talk about post training.

3:08:40

Um, what uh what's the thing you're most proud of?

3:08:43

thing you're most proud of? like like what can you give us on without you know immediately getting poached what what can you give us on on the actual innovation that went into GPT5 from a post post training perspective what are like the the kind of keywords and and paths in the tech tree that we should be

3:08:59

digging into over the next few years to understand how this works >> you know I would say the thing that is most impressive to me about GBT5 is how much getting all of the details right matters >> like when I look at GBD5 You know, we had an early version of this thing a while ago that was kind of okay, but clearly did not meet our bar for revolutionary. And we're trying to

3:09:20

And we're trying to figure out, you know, why why is that not as good as it should be?

3:09:24

And the team basically just went off and did a deep dive over a couple of months of just completely rebuilding the post-training stack for this model.

3:09:30

And turns out that when you do that, you get what would have taken, you know, another order of magnitude worth of pre-training improvements to to produce.

3:09:43

How much are you thinking in pro training in research about let's forget the benchmarks and just focus on user satisfaction like NPS score basically or like user user minutes or any of these other uh the real benchmarks. >> Yeah.

3:09:59

The intangibles revenue people using Yeah.

3:10:02

the feeling and the and the joy and the actual value that's delivered.

3:10:05

Um because Studio Gibli was a delightful moment. It wasn't a benchmark. >> Yeah.

3:10:12

I I think so that was something that we took very seriously for GBD5.

3:10:15

It's like look at what people are actually doing with tragedy and look at where the model is failing them. >> Mhm.

3:10:21

>> Either in the sense that the model is like sort of like you said it's not enjoyable to use. >> Yeah.

3:10:26

>> And so we we did I think make a lot of progress on that.

3:10:27

Like GBD5 is much more engaging than our previous really smart models.

3:10:33

Like 03 I I don't know if you guys talked to 03 in the past. It's a bit bland. >> Sure.

3:10:39

And GBD5 I think has a lot more character is a lot more more interesting.

3:10:42

But then also like I think for we really care about just actually being accurate like giving if a user is trying to do something economically valuable with our model.

3:10:54

We want to make sure it lands correctly. >> Yeah.

3:10:56

>> And so what we did there is just like look at act the actual distributions of what people are doing with our models in the real world.

3:11:01

Figure out where the models are going wrong.

3:11:02

Build interventions to target it.

3:11:04

And that was where, you know, we got, I think, the most impressive improvements in GBD5.

3:11:10

Like 03 would just get things wrong and not tell you it wasn't sure it was it was incorrect.

3:11:16

And GBD5 is much much better about like actually being honest when it thinks it might not know. >> Yeah.

3:11:21

How explicit is are all the different pieces of the post-training pipeline?

3:11:25

Like you have you have uh you know safety postraining, you have uh stop hallucinating, give me the real facts.

3:11:34

you have make sure the text the the the flavor the tone is pleasant.

3:11:36

Um there's so many different things to optimize for.

3:11:41

How much of that is like try and just blend it all up into one thing versus like explicit passes chunk it out like split it up.

3:11:48

H how how much can you decompose the problem?

3:11:52

So, you know, my background is in reinforcement learning.

3:11:56

And I think you when you look at something like this, the magic is in the reward function, right?

3:12:02

It it's in what you're actually telling the model to be good at.

3:12:06

>> And so fixing things like hallucinations to a huge extent is a essentially a function of just fixing the reward function, >> actually making it so that the model is reliably penalized for saying something that's false.

3:12:16

And if you do that, all of a sudden the model stops saying things that are false. Ditto for safety, right?

3:12:21

Um, you know, on the live stream, Sachi talked a bit about the the way we've changed safety for this model.

3:12:26

And to a huge extent, it's it's just a function of we we're actually putting out a paper today on the new safety stack for this model.

3:12:34

And the core insight in that paper is just figure out what you actually want to optimize for, which in our case is helpfulness conditional on not saying something that's actually dangerous or harmful.

3:12:45

You know, write that down, figure out what that means as a reward function, then optimize it for it. >> Mhm.

3:12:50

It's really not magic at all.

3:12:50

It's just again it's it's what I said earlier.

3:12:53

You got to get the details right.

3:12:54

You know, if at any part of that process you screw it up, the model will be unusable.

3:12:58

What's your current thinking on spiky intelligence?

3:13:03

And is there some is there some flywheel that you can get started where you're identifying low points that aren't spiky enough and then you're like almost automatically setting up the infrastructure the eval to then RL against to create a spike.

3:13:23

>> I think GPT5 was a preview of what's possible in that respect in the future. Yeah.

3:13:30

>> A step in that direction.

3:13:32

Do you think that there's a world where you get to a place where you're kind of It's weird because we're not hammering down the nails of the spikes, we're adding spikes, but um this weird metaphor that we're stretching a little bit too far, but uh is there a world

3:13:45

where um where where you can be doing post- training or just adding capabilities in a more iterative uh cadence so that as soon as you identify something, the the response can be yeah, we don't need to wait until GPT6 to fix this. we can just add this capability

3:14:00

we can just add this capability because hey, we just found a pocket of users who are who are trying to do a thing and they're not super happy with the results and let's add this capability. >> Yeah, I I think so.

3:14:11

Um I mean I think we you know we are going to launch other models between now and GBT6.

3:14:15

I think it's relatively common knowledge but we do update the model in chatbt reasonably often.

3:14:21

>> Yeah, people talk about it all the time. >> Yeah, exactly.

3:14:23

And you know I think we are now in a world where we can conceivably update that model and have it get materially better on capabilities too. Yeah.

3:14:30

>> Not just on, you know, the personality is a little bit better than it was before. >> Yeah.

3:14:35

>> Going back to your note on the new uh paper that I guess you guys are releasing today when you talk about optimizing for helpfulness.

3:14:42

Is there is part of that avoiding the model? you know, reinforcing.

3:14:50

There's times when you want to reinforce and and give kind of like confidence to the user that they're going down the right sort of like thought process and things like that, but then there's like a point where it can get too extreme in terms of maybe convincing a user of something that may be totally untrue.

3:15:04

Is that is that what the paper gets at or or am I reading?

3:15:10

>> We So, it's not specifically about this.

3:15:12

Although I will say we we do explicitly train the model to not lead users down bad paths.

3:15:17

That's something that I think we've started taking much more seriously over the last few months as we've realized Sam talked about this a little bit um I think back in May, but CHBT is just way more important for people's lives now than it was a year ago or especially two years ago.

3:15:33

And we do have to actually be very cognizant of what effects our models have on users.

3:15:36

So yeah, we we do we do very actively train models to not lead users down the right path.

3:15:42

Don't fact check me on the releasing today.

3:15:44

I know we're releasing it. I believe it is soon.

3:15:47

>> I think it's today, but I've also been in a hole dealing with launch all day.

3:15:51

>> Yeah, we're not big on fact checks here.

3:15:52

We're big on the truth zone, which is just the vibes.

3:15:55

And >> the vibes are we'll be publishing some information about the new safety setup >> at some point. That's great.

3:16:00

>> at some point. That's great. Um >> yeah, I think I think a large part of the conversation around safety should be how reliant and and how useful the product has become to users and then the new level of care that you have to

3:16:18

provide versus a while ago when it was just like people saying making a cute image or generating some text that they were going to use in an email or an internal document and and realizing this this vector of usage which is this like companion confidant that uh is is becoming so prevalent. >> Talk to me about post training for uh

3:16:38

>> Talk to me about post training for uh big big partners, enterprises, government organizations.

3:16:44

Um what is transferring from the research that you're doing to something that can be uh offered as a as an enterprise level product? >> Yeah.

3:16:58

So we do openAI does partner with external companies to do essentially custom post training.

3:17:02

Um that that is a thing that we do and from that perspective the stuff we do just directly transfers.

3:17:07

I'll also say that we we've >> put a lot of work into trying to make our models as general as possible.

3:17:14

>> But to as large an extent as possible if you want to get really good results from our model you can do it right on the API just by actually telling the model what you want it to do. >> Yeah. >> Right.

3:17:22

Like GBD5 I think is pretty comfortably our most durable model ever.

3:17:27

Y >> we've heard a lot of really positive feedback about this, especially from like folks like Curser. >> Yeah.

3:17:31

So if I came to you and I was like, I'm an enterprise and I need to generate a lot of Studio Giblies, you'd be like, what are you doing? Just prompt it.

3:17:39

Um but what are the examples of of of companies and organizations?

3:17:43

Is it just is it just private information, private data sets that aren't available on uh on the open web or is it specifically like there is enough data out there but there's just not the economic incentive for your team to go and RL on you know Gas Station bench or whatever we're talking about here hypothetically.

3:18:05

>> I think the answer is both.

3:18:07

>> Yeah, it's definitely both.

3:18:07

Um because yeah, we're not going to target, you know, as you said, gas station bench >> because what people are doing with not not on our own right now probably because it's not mostly what people are doing with tragic. >> Exactly.

3:18:19

>> You have some application that's super valuable to you. >> Yeah. Yeah.

3:18:23

>> We can be convinced that it's important. >> Yeah. Yeah. Yeah.

3:18:26

>> It's just not what our users are already trying to do.

3:18:28

>> What's the state of reward hacking and and and fighting that in in RL environments?

3:18:34

>> Um you know, I think we've actually made a lot of progress.

3:18:36

There was some discussion of this around 03 that 03 was like a little bit deceptive in ways that felt reward hacky and >> GBD5 is dramatically less deceptive than 03 was.

3:18:47

>> What's an example of how that would manifest?

3:18:48

Like do you have like a canonical uh case study? >> Yeah.

3:18:52

I mean the canonical thing is like you ask 03 to write you some code and instead of actually writing some code it changes some unit test >> changes the test case, right?

3:18:59

Which is which is kind of hilarious.

3:19:01

It's like one of the funniest things that an AI has ever done.

3:19:04

And I understand that it's very bad and it's not what we want, but it is just like it's kind of cheeky in my mind. >> It's kind of cheeky.

3:19:10

It's also like, you know, I've I feel like if you spend enough time around real software engineers, they do actually do stuff like this pretty often.

3:19:15

So >> I have 100% done that.

3:19:18

>> I I was going to say I also have done that.

3:19:19

Uh >> for for formal reasons, I won't say that I did it at OpenA, but back when I Well, I definitely did that. >> Yeah, of course. Of course. This is natural.

3:19:28

>> What do you think GPT6 looks like?

3:19:28

You guys, you mentioned that you're going to be shipping, you know, updates to five, but what what what are you most excited about?

3:19:36

Uh where where are you most excited about uh going forward?

3:19:39

>> And and just really quickly, give us the date that GBC 6 launches.

3:19:45

>> Oh man, hopefully we uh hopefully hopefully six launches as a complete surprise to everyone.

3:19:49

I think that would be ideal. >> Like a Beyonce album. >> Well, yeah.

3:19:52

Hopefully five just makes it and says, "Hey, it's ready now if you want if you want to hit."

3:19:55

>> Yeah, I think that would that would be a great thing for six, actually.

3:19:57

I would love for six to do all of the launchcoms and and to do the live stream.

3:20:00

That would be really great.

3:20:04

>> Live streaming is uh that's the real AGI test >> for sure. For sure.

3:20:08

>> I feel like we're not that far off actually. I don't know. >> We're getting there.

3:20:12

>> I mean, video synthesis maybe, but you know, >> talking through a script for 30 minutes. Come on.

3:20:17

Models got to be able to do that >> for sure.

3:20:18

Well, yeah, that'll be the the the the next Sora launch or something.

3:20:22

We'd love to have you back on.

3:20:22

But thank you so much for taking the time today. We'll talk to you soon.

3:20:26

>> Great to talk to you guys soon. >> Congratulations. Cheers. Bye.

3:20:28

>> Congrats on the launch.

3:20:29

>> Let me tell you about adquick. com.

3:20:29

Out ofome advertising made easy and measurable.

3:20:31

Say goodbye to the headaches of out ofome advertising.

3:20:33

Only adqu combines technology out of home expertise and data to enable efficient seamless ad buying across the globe.

3:20:39

And we have Scott Woo from Cognition coming in the studio for the fourth fifth time.

3:20:46

I can't keep track anymore.

3:20:46

Thank you for taking Scott. Thank you for coming. Miss you. >> How's it going? >> It is fantastic. >> Got to be honest.

3:20:54

Great week to be an application layer company.

3:20:55

I got to tell you guys, >> I was about to say the best thing ever.

3:21:01

Another win for Scott Woo. Wow. Wow. Wow. >> 4. 1. >> Yes. Uh so, yeah. How how big is this?

3:21:07

Is it are we in the are we in the Uber lift uh territory where you know, you're going to be uh you know, in price competition between anthropic and open AI going back and forth like what what is the real benefit to your business right now from today? >> Yeah. Yeah, for sure.

3:21:21

So, first of all, obviously massive capability gains across the board.

3:21:24

I think really really impressive work that OpenAI has put together.

3:21:28

together. um you know people have talked about uh what's going on in the AI coding model race and and I think by a lot of accounts you know anthropic has generally been ahead for for for a lot of the last year honestly and I think at this point open AI is is very clearly

3:21:42

you know has very clearly caught up and it's it's pretty neck and neck I'd say between the two right now and so um very exciting to to to see all of this unfold and to see what's next but but I think from our perspective um yeah I mean code is is just such a um such a core capabilities pill use case I'll call it. Um and so you know being able to work

3:21:59

Um and so you know being able to work with smarter and smarter models um and do a lot of the work that we do it just means that both Devon and Windsor can be a lot more capable um a lot more intelligent can predict what you want to write or what you want to do um uh with a lot of higher accuracy.

3:22:14

>> Yeah, it's almost it's almost like surprising that given the like cultural rigor at cognition that you're not doing fundamental frontier research.

3:22:24

So can you walk me through like what is the what is the focus of being an application layer company?

3:22:33

Is it is it UI go to market?

3:22:37

I'm sure it's all of these, but in terms of the the hardcore software engineering, like what is important to get right at some point there's fine-tuning and post- training, but is that moving back into the purview of the foundation labs or is there still work that you want to do on top of the models or on top of the APIs? >> Yeah.

3:22:57

Yeah, it's a great question.

3:22:57

Like I I mean I think the uh the core of being, you know, an applied lab is is really just focusing on a very particular use case on delivering real real just very direct results.

3:23:08

And I think um you know like I I think the foundation labs are obviously, you know, incredible at training base models and and all this pre-training and and all the work that they do there.

3:23:17

I think from our perspective, we um we we want to work on a lot of very particular capabilities that apply to software engineering in particular.

3:23:26

and then obviously you know run the whole stack from there to building a product figuring out the interface and and the UX and then obviously bringing that to market and selling that.

3:23:34

Um on on the capability side, there's a lot of particular stuff where um you know, one way to put it is I think the base IQ is very much already there in the models and you can see the the raw problem solving ability and I mean we've gotten some pretty insane results.

3:23:47

You know, getting a gold medal at the IMO or all these other things, right?

3:23:52

>> You called that, by the way, >> you called I think the first I mean we were one point away to be fair a year ago, right?

3:23:59

So it was it it was on the way I'd say.

3:24:02

Um but but but but yeah so so so you know you can really see the general intelligence improving it with every single model generation.

3:24:08

On the other hand for Devon obviously um you know it's a very clear like step up in the general intelligence but also you want to be able to you know if you ask Devon to to go debug your Kubernetes or to go and you know look look into your error logs and figure what figure out what went wrong or or or things like that.

3:24:24

there's often a lot of very specific capabilities and and that's where we find that you know the post- training of the RL is is is most effective there and a and a lot of the kind of various work around the models that that turns out to be useful. >> What about speed?

3:24:38

A lot of people that have gotten access to GPT5 are are uh at least in our chat are reporting that it just feels really really quick.

3:24:46

How how is that uh over time going to impact the I think a lot of people you know if they're using Devon today task Devon with something and then maybe they go work on something else for a little bit or they're running multiple agents concurrently but at some point the agent could get so fast that you're just sort of like watching it and work in real time and you actually want to be engaged. But uh are we there yet? Is it still a ways out? What do you think?

3:25:14

>> Yeah, it's a great question.

3:25:14

I I think in general I think async will continue on as a paradigm even as the models get faster and faster.

3:25:21

faster and faster. One of the reasons that it should by the way is because there are a lot of real world uh thresholds that start to matter like at some point you're actually spending less time on token generation um in the Devon

3:25:32

life cycle and you're spending more time on every time Devon runs the command to go install packages or Devon running the unit tests or like Devon pulling up the front end by itself or or or things like that that obviously take real world time, right? Um I think we are honestly

3:25:43

Um I think we are honestly getting closer and closer to that threshold.

3:25:47

But yeah, so so long story short, I I think like um in the asynchronous mode, uh yeah, these things will get faster.

3:25:53

You know, we'll see those gains or we'll be able to spend a lot more time, for example, thinking about a single um um problem relative to the amount of like real world clock time that gets spent.

3:26:03

Uh I think for the synchronous use cases is where we'll see things really really um um you know exploited with with speed which is you know windsurf and cascade for example um where where we see the speed gains really really matter.

3:26:16

>> Uh speaking of windsurf give us the update on the wind the chat wants to know about the windf uh and the 80hour demand.

3:26:22

Uh how have the the buyout offers gone?

3:26:25

What's the internal response been?

3:26:28

Where'd that idea even come from? >> Yeah. Yeah.

3:26:32

Look, people people are stoked honestly.

3:26:33

Um, and I think I think from our perspective, it's ob obviously really important to to kind of just like um unite and get to the point where we can just be one culture and and one kind of shared uh set of values and and this is how things are at Cognition.

3:26:48

It's you know it's it's it's a pretty busy time like we we are at the inflection point of code and and we work like that too.

3:26:57

Um and and so I think a lot of it for folks is is just kind of like >> um you know we want to make sure folk folks who who who who really want to do this with us, you know, make that conscious decision to opt in and for for anyone who doesn't.

3:27:08

Obviously, we totally understand that there a lot of talented folks that maybe that's just not the right thing for them right now or, you know, not at this time.

3:27:14

Um um and so wanted to make sure that they were well well taken care of, too.

3:27:19

>> And to be clear with the buyout offer, that's on top of the actual acquisition deal that already went through.

3:27:24

They already got >> they already got fully vested.

3:27:26

So yeah, I was thinking of the roller coaster.

3:27:29

It's like you have the OpenAI deal, then the Google deal, then the the Cognition deal, and then they're like, "Wait, these guys work really, really hard.

3:27:35

I don't know if I'm cut out for this."

3:27:36

And they come back up again where they're like, "Wait, I can just go, you know, take a sabbatical and and figure out my next thing." So it's a great outcome. >> Yeah. Yeah.

3:27:45

No, for obviously is, you know, overall is a killer team that's been through a lot and so um wanted to make sure that they're well taken care of. >> Yeah. >> That's fantastic.

3:27:54

Um uh any anything else you can tell us about the integration of Devon and Windsurf?

3:27:59

How are the teams getting along?

3:28:01

How do you see the products playing together in the long term?

3:28:04

Obviously Crossell seems really obvious.

3:28:06

They had the go to market team as well, but uh but how else are you thinking about the interaction maybe over the longer term there? >> Yeah. Yeah. Yeah. For sure. Yeah.

3:28:13

A lot of obvious integration on the team as you mentioned with Crossout and so on.

3:28:17

you mentioned with Crossout and so on. I I think the thing that's really exciting on products um which which I think actually comes along with the the these capabilities increases is you know as the capabilities keep getting better you start to take on harder and harder tasks with AI and with full agentic workflows right and I think there's an interesting

3:28:34

thing that happens where for a lot of the harder tasks you really actually do want to go back and forth between a synchronous and an asynchronous mode you know a and that's for a few reasons you know one of the reasons obviously is because there's a lot of review and and a lot of like looking at the pieces and and thinking about the the um uh you know all all the minutia and the details of what you're implementing. I think

3:28:53

I think another big reason for it is, you know, when you get started on on a larger project, you know, let's say you're you're sitting down as an engineer and you're saying, "All right, I'm going to go build this whole project today."

3:29:01

You yourself don't actually know all the trade-offs you want to make, all the decisions that you want to make and so on, right?

3:29:06

And so having a format where, you know, for the decisions that need you to be there and you're involved setting the kind of the strategy or or figuring out high level what should happen, you're able to do that in a nice synchronous environment, which is naturally the the Windsorf IDE, right?

3:29:20

And then for the parts of the task that you can actually hand off and have an agent work on, you're you're giving that to Devon.

3:29:25

Um, and figuring out how you go back and forth between those is is is super interesting.

3:29:29

So, wave 12 on the way soon.

3:29:31

Hope hopefully we'll we'll have a lot more more to share. >> Last question.

3:29:34

Uh, yeah, hit the soundboard, Jordy, for that. We're wave 12. Wave 12. Fantastic.

3:29:37

Um, uh, last question, we'll let you go.

3:29:40

Uh, what is your probability that AI will get a perfect score on the IMO next year? Interesting.

3:29:50

Um, so we, by the way, we just had the IOI, which is the the programming version, like the programming Olympiad.

3:29:55

Um, and I think there's a good chance that we'll have a golden medal at the IOI for for this year announced as well.

3:30:00

I think perfect score for next year.

3:30:03

>> Wait, uh, we as in humanity or we as in cognition >> as in humanity. Yes. Yes. Yes. And AI perfect score. Yeah. Uh, sorry.

3:30:08

And AI gold medal, right?

3:30:11

Uh, perfect score in the IMO next year.

3:30:15

>> I think it's got to be north of 50.

3:30:16

Honestly, I I would put it around like 75% or so. We'll see.

3:30:21

>> Well, thank you so much.

3:30:21

I we'll be following you closely.

3:30:22

Uh and uh good luck to you and uh congrats on all the progress. Very fantastic. We'll talk to you soon. >> Awesome, guys. Talk. >> Bye.

3:30:31

>> Let me tell you about bezel. Getbbezzle. com.

3:30:32

Your bezel concier is a is available now to source you any watch on the planet. Seriously, any watch.

3:30:37

And we are joined by our next guest, Claire Vo from Chat PRD.

3:30:41

Welcome to the stream, Cla. How you doing? What's going on?

3:30:47

>> I'm It's a fun day today, isn't it? It >> is a fun day.

3:30:49

Uh, what was your reaction to the stream?

3:30:50

Um, what was your reaction to GPT5?

3:30:54

>> You know, GPT5, the first thing I said and I got a little early access is I said it's a developer for developers by developer.

3:31:00

This thing is built to be a software engineer.

3:31:03

>> You've seen a long string of your guests come on and really speak about the coding abilities of it.

3:31:07

And what I think is interesting about this particular model, especially because we're seeing them deprecate the old models in the chat GPT experience, >> and we're seeing a lot of positive feedback, but I do think there are drawbacks to a model that's so clearly tuned to a developer use case.

3:31:21

And as somebody who's building an application um that isn't focused on agentic coding, I have noticed some personality quirks that are going to be really interesting to see how they shake out um >> as we roll out this model to our users.

3:31:39

>> Walk me through those.

3:31:39

What are the what's the timeline?

3:31:41

How how much like how much time do you have to kind of move users over to five before >> Yeah. Yeah.

3:31:49

So I I mean I think we have tons of time from the API side to move move users.

3:31:53

And in fact, you know, our strategy at ChatPD is not to just upgrade to the latest model.

3:31:58

I know Zach at Warp said like why wouldn't you want the latest intelligence?

3:32:02

And the reality is because we're doing a lot of business strategy and business writing.

3:32:06

I actually want to validate with our users that they're they're getting the quality of strategic thinking output writing that they really want.

3:32:13

So we actually AB test every single model roll out and really evaluate for user quality, token generation, all those things.

3:32:22

And you know, looking early on, it yaps.

3:32:24

Man, this thing just wants to go through tokens.

3:32:29

Right now, I'm seeing four to 10x the number of tokens generated between the, you know, four generation models and five.

3:32:35

And when you're in a business context, you do not always want longer >> words, you know, and so it'll be really interesting. there.

3:32:43

It is certainly focused on execution.

3:32:47

So I, you know, I've heard a lot from the OpenAI team, it's steerable. Yes.

3:32:50

And its natural inclination is to drive you towards like how, what, very tactical, very specific.

3:32:57

And and so if you're trying to zoom back out at a um strategic level or focus on a business initiative, it's actually a little harder to tune in that direction.

3:33:06

So, you know, I think there's a lot of positive things for me as somebody who uses Agentic coding platforms, who writes a lot of code.

3:33:12

It's my daily driver now. I love it.

3:33:15

Um, but for other use cases, I think it's going to take some time to figure out if it really is optimal in use cases where intelligence actually isn't the differentiating capability.

3:33:29

>> It's very interesting to think uh the the best the best product manager is not the one that writes the most the longest doc. >> No.

3:33:37

and you don't send your engineer into your executive meeting like I and I I really am looking forward to the time where we're not getting these numberbased models where actually I can get like GPT developer >> or GT GPT strategist where they're pre tuned and trained for the role they're going to play as opposed to general purpose but clearly oriented towards a set of set of tasks.

3:34:03

And I just think if you look at this model, it was oriented towards um engineer software engineering at least in my experience.

3:34:12

>> So have you been tempted to launch any type of agent like agentic coding products?

3:34:19

You are you guys are obviously responsible for creating documentation and if you look at the other guests that have joined today, many of them are competing with each other in different ways and trying to own different parts of the stack.

3:34:33

you guys have seemingly stayed really really laser focused and no one else uh is doing anything like you're doing at least on the show today.

3:34:40

But talk about like picking your your lane and kind of like >> optimizing.

3:34:49

>> Yeah, we're integrated with a lot of those platforms.

3:34:51

So a lot of the kind of like prototyping platforms, vzero.

3:34:53

dev, lovable, all those we integrate.

3:34:56

We just released our MCP.

3:34:58

So I use chat purity pretty consistently inside cursor through our MCP.

3:35:02

So I think of we we think of ourselves as the product pair to the AI engineer.

3:35:07

Now what's really interesting about my experience with GPT5 is the one place it actually does really well is technical specs and that's a place where u chat PRD has sort of bridged into engineering execution.

3:35:21

often our product managers are generating a PRD or some sort of business document and they're actually going the next layer and developing a technical spec.

3:35:28

The GPT5 technical specs fed into these agent coding frameworks or prototyping frameworks output output much higher quality assets on that end.

3:35:39

So I do almost think there's going to be this kind of like right model for right use case especially in our kind of business and so we think of ourselves as integrating.

3:35:47

The one thing I have thought about with GPT5, it's the first one where it feels really simple to just go ahead and roll your own agent coding framework or um prototyping framework inside of our application. So never say never.

3:36:00

It's something that we get asked for a lot, but we were we're friends we're good friends with almost all your guests on your show today.

3:36:04

And so we like we like the role we play in terms of being the product manager pair to all these AI engineers.

3:36:12

>> Yeah, that makes sense.

3:36:16

>> What are you looking for next?

3:36:16

What am I looking for next?

3:36:18

I mean, in terms of um model capabilities, what I think is really interesting about OpenAI and why I'm really committed to the OpenAI ecosystem, even though I test and use a variety of models, is I think developer support is a real differentiator here.

3:36:35

So we spend a lot of time talking about model capabilities and for application developers certainly ones that are doing more complex applications like agentic coding model capabilities really matter like core IQ of the model matters but the other thing that matters you know as

3:36:50

somebody who has built developer tooling products it's developer experience matters the primitives in these APIs matters and so what I'm really pushing the open AI team to think about which is in addition to the core intelligence the model. What are the developer tools you

3:37:04

What are the developer tools you need around these models to really make them a platform on which a variety of applications can build?

3:37:12

And I do think that OpenAI has disproportionately invested in developer experience, but I'm always looking for like give me better out of the box tooling, give me more control over these models, um, give me more hosted services, all those things that as an application developer are just going to make it easier to deploy these models in production beyond the core kind of intelligence of the models themselves. What was your read on 4. 5?

3:37:35

Is there a world where, you know, I'm I'm I'm thinking about the the the product manager versus the engineer.

3:37:40

You have your 03 go crunch some really hard reasoning and then you have four five turn it into uh you know stronger pros or or like more you know a human language. >> Yeah.

3:37:53

So I did a lot of experimentation around 405 and 41.

3:37:56

45 was my favorite pros writer >> by far.

3:38:01

It was was loved from a business writing perspective.

3:38:04

I thought the pros was the most natural. It was really slow. Like >> untenably slow.

3:38:10

And so the um the compromise we made in our testing is we ultimately ended up with 41 as the fan favorite for business writing when we were balancing off both quality of pros and intelligence as well as performance which for application developers is a real consideration. So I landed on 41.

3:38:34

41 is the model that's being tested right now against GPD5 in chat purity.

3:38:39

And one of the things that I have to go do now is figure out how to get chat GPT or GPT5 to stop writing it.

3:38:43

It writes a lot and it only wants to write in bullet points.

3:38:48

So I've got to go back into our prompts and figure out how to direct it to be a little bit more businessoriented.

3:38:56

>> Bullet point maximalist. >> It's the new mdash.

3:38:58

I'm telling you, you will not be able to stop seeing it.

3:39:00

It it just all it wants to do is write a bullet point and call a tool.

3:39:04

like it I was using it in cursor and it just kept maxing out >> my tool calls.

3:39:09

I'm like you do not need to read 50 files to to do this.

3:39:11

So I do think you know application developers are really going to have to think about how they slot this into their current workflows.

3:39:20

There's definitely tuning that needs to happen.

3:39:21

But I am telling you you are going to see a lot of bullet points when this thing rolls out. >> Yeah.

3:39:26

>> In 60 seconds where is product management going?

3:39:29

A lot of people talk about the, you know, examples of product managers that are starting to ship code themselves, ship whole features, products, but I'm sure those are edge cases to date, but but where do you feel like it's going based on on your user base?

3:39:47

>> Yeah, I mean, it's going to go one direction or the other.

3:39:49

Product managers are either going to develop the hard skills to do the design, the go to market and the engineering job to some extent because some of these other jobs are definitely going away for product managers or my favorite use case engineers and designers are going to get tools like chat pd or these prototyping tools or cursor and they're going to be able to actually do the product management job.

3:40:11

And so what I think is we're going to see a new type of role emerge which is a much more generalist role where people maybe have a specialist capability and they're augmenting that product thinking or they're augmenting that technical thinking with with AI.

3:40:23

But I don't think there's going to be product managers as they were, you know, five or 10 years ago for much longer. >> Makes sense.

3:40:31

Well, thank you so much for stopping by. Great. Great chatting you. >> Thanks for having me. >> We'll talk soon. Bye. >> Cheers.

3:40:36

>> Up next, we have Brad Lycap, the chief operating officer of OpenAI.

3:40:39

Welcome to the stream, Brad.

3:40:41

Also, Jordy, your post saying, "I'm updating my timelines.

3:40:45

You now have four years to escape the permanent underclass has over 4,000 likes." >> There we go. >> Absolutely banging.

3:40:53

>> Thousand likes for every year. Love to see. >> Love it.

3:40:54

Anyway, Brad, how you doing, >> Brad?

3:40:56

What's going on, >> guys? How are you? Good.

3:40:58

>> Congratulations on the launch.

3:40:58

Um, what uh what are the biggest takeaways for today from your side?

3:41:03

I'd love to know about what it actually means to be the COO of Open AI.

3:41:07

OpenAI does so many different things.

3:41:11

Consumer internet company, API business, enterprise, there's all sorts of stuff building data centers.

3:41:17

What What is your actual role?

3:41:19

>> My role is kind of whatever the company needs me to do.

3:41:21

Um I play everything from like you know PM when I need to to like uh you know saleserson when I need to.

3:41:28

Um uh that's kind of the fun part of the job for me.

3:41:31

Um on this launch in particular uh it was really fun.

3:41:33

I spent a lot of time last few weeks with customers with partners.

3:41:37

uh getting a feel for GPD5 relative to what they were previously using.

3:41:42

In some cases, those were OpenAI models.

3:41:44

In some cases, they were other models.

3:41:45

Um but you know, I've been OpenAI a long time.

3:41:48

I've been in OpenAI 7 years.

3:41:49

Uh so I've seen GBT3, I've seen GBD4, and then to be able to see GBD5.

3:41:54

Uh and you know, just I think the joy from of people being able to use it in production and seeing how much better it is. Uh that's the best part.

3:42:02

>> Greg told us earlier about the era having to pay people to use the early versions of the product.

3:42:06

You guys have come a long way since then.

3:42:11

>> Yeah, we had like three customers with GPT3 or something like that.

3:42:12

Um, and so it was easy to manage, easy to talk to all of them.

3:42:16

They actually were like tired of us calling them being like, "Is it is it good? Is it getting better?"

3:42:21

Um, and uh, so now it's, you know, we're fortunate that we've got more than that. But, um, it's cool.

3:42:25

I mean, the diversity of use cases, I think the number of things that people, uh, are able to use it for.

3:42:30

We've got everything from the team at Amgen, you know, big pharma life sciences using it for, uh, clinical workflows there.

3:42:37

We've got teams at Uber's, uh, you know, building it for customer support.

3:42:40

Um, teams at Notion and Cursor building it into products that people use every day.

3:42:44

So, uh, I think that's the power of it is is it just more and more covers the surface area of things people do uh, you know, with these tools.

3:42:51

I'm not sure how much you touch organizational design at OpenAI, but I'd be interested to hear your thoughts on how those companies that you mentioned should be thinking about AI changing their org structure.

3:43:05

Is it sort of like a horizontal crossf functional service layer like um you know a finance team that touches a lot of different elements of the business or should most companies be thinking about standing up a dedicated like AI implementation team?

3:43:20

How do we get a a chat box on every product that we already ship?

3:43:24

Like how do you think about those trade-offs if if you were talking to a you know a friend at a Fortune 500 company that was thinking about their AI strategy?

3:43:32

>> Yeah, you know, it's an interesting question.

3:43:34

Um I think it was maybe said earlier on the show.

3:43:37

The thing we see is just people can do more.

3:43:39

And so there's like this much wider latitude that you get if you're an individual person at an individual company where especially as you get bigger, you know, maybe more bureaucratic organizations that have a lot of different functions, a lot of different levels, you have to rely on a lot of other people in the org to get stuff done.

3:43:55

You've got to rely on your data science team to do data analysis.

3:43:58

You've got to rely on your design team to do mockups.

3:43:59

You've got to rely on your marketing team to do copy.

3:44:01

And I think what we see with AI is it just accelerates people to get to a great V1 of everything. Mhm.

3:44:08

>> So if you're a high agency individual uh and you want to get stuff done um you're no longer gated on people that uh you know you otherwise would be and I think that should enable organizations to move a lot faster and I think it should enable the the people at organizations that really drive them uh to do a lot more and we see that consistently chatbt

3:44:26

enterprise I think that is consistently what we hear and we we seek those people out when we deploy chatbt enterprise um we find those like you know two or three people at the organizations who are just the like AI superstars and champions and then try and actually use them as these kind of touch points for the rest of the or to learn from. >> How are you you how are you personally

3:44:43

>> How are you you how are you personally using AI these days?

3:44:47

>> Uh you know I my biggest challenge I think day-to-day is context switching.

3:44:50

If you look at my calendar from like top to bottom it's like I I joke like I you know with my wife I like have to like go show up to work like wearing like a lab coat and then I like take the lab coat off and like put some like sunglasses on and a film school jacket and you know then I'm talking to like a media company and then I like take that off and so so I go through the costume changes.

3:45:04

Um, and I think what what I actually mostly use it for is just to help with bridging me from kind of thing to thing.

3:45:10

Um, to kind of put me in the mindset of being able to work with customers, help customers.

3:45:16

Um, GBD5 is incredibly good at this kind of structured reasoning of how do we actually take what is this very diverse set of things that models like GBD5 can do and then apply them in domains that I don't think about every day.

3:45:28

And so it gives me this launching off point to be able to talk with with uh with leaders and with customers much more fluently about how we can help their organizations >> within let's say a set of companies like the Fortune 500.

3:45:39

What is AI adoption look like across the spectrum?

3:45:43

because I'm sure that there there's companies that you talk to that are truly, you know, adopting AI in the way that John was mentioning like trying to become AI native, changing their entire organizational approach.

3:45:57

And then there's companies that just want to buy software to say that they can that they're becoming AI native.

3:46:03

So what what is that spectrum look like in in practice?

3:46:09

>> Yeah, it is a wide spectrum.

3:46:09

So um at the top level we're seeing just like amazing uh appetite for wanting to adopt tools for people and I think that's like the easiest place to start.

3:46:19

Typically that's where we steer organizations uh if they're starting at zero is just give your people the best tools.

3:46:26

Um you may have seen we you know we've grown chatbt uh work which is our enterprise and team product from 3 million seats to 5 million seats now uh from from June till till now.

3:46:37

So um torid growth there and we don't see any abatement in in demand there.

3:46:41

Um if anything it's accelerated from from last year and so I think people and organizations are starting to realize that like at a minimum you need to make sure people have the best tools.

3:46:51

Um what's cool about GBT5 now is it also enables people to use the best tools at every point.

3:46:55

And so if you're in an organization you're not fumbling with the model picker.

3:46:59

You're not trying to figure out when to use a reasoning model.

3:47:01

You're not trying to figure out kind of the art of prompting to get the perfect thing.

3:47:04

all of that stuff is abstracted and it's kind of taken care of for you and you can have confidence that your people are actually using the best models at any at any given point.

3:47:13

Beneath that it gets a little more complicated.

3:47:14

So um more and more organizations I think are starting to grasp uh how the tools can actually help uh in the business process.

3:47:20

So whether that's in customer support, whether it's in research, whether it's in uh software engineering and data science, um you're seeing these tools more and more adopted in the enterprise.

3:47:30

I think there's still a quality gap though.

3:47:32

I think we've we now are just breaking into what I would call the kind of era of models that have capabilities that are good enough to make a dent in the types of problems businesses care about.

3:47:42

Um, businesses care a lot about things like reliability, right?

3:47:45

They think they care about accuracy.

3:47:47

They care about the resil the resiliency of the model to recover from tool use errors and to be able to string together these very long kind of multi-tool multi-step workflows.

3:47:57

Um so GPD5 is a step on all those things and I expect that that will enable us to be able to do more and more things in the business process.

3:48:03

the business process. Do you think those customers that you just mentioned will stick with uh um this idea of like GPT4 level workloads will stay on GPT4 and maybe there'll be cost savings but those workloads will stick around for a very

3:48:18

long time and then you'll develop almost new capabilities, new workflows, new uh new workloads that will be additive but um the enterprises will stick or will they want to is is everything so fresh that they'll want to just like rewrite everything with the latest and greatest. >> More often than not, I think it's the

3:48:36

>> More often than not, I think it's the latter.

3:48:37

I think you want you want to rewrite everything.

3:48:39

Um, one of the cool things we did here was we were able to keep the pricing on GPD5 at the level of 03 pricing.

3:48:45

And so, um, you know, if if you're cost sensitive, uh, you don't really have an excuse to to not upgrade.

3:48:52

GBD5 is faster than than 03 and and 41.

3:48:55

So, we've improved on latency for sensitive use cases that are speed sensitive, latency sensitive.

3:48:58

Um, and obviously the intelligence bar has gone up.

3:49:02

So, um, you know, unless you've got really a very kind of narrow and specific workflow where you've got a model like 41 that kind of is okay, there's really not a reason, I think, that people wouldn't upgrade. >> Yeah.

3:49:12

Do we need like a three-dimensional paro frontier right now that matches not just cost and and capability, but also cost, capability, and latency or something?

3:49:20

Is that is that something that you're seeing a lot of demand from in the enterprise? >> Yeah, 100%.

3:49:24

Um, we actually measure it that way.

3:49:26

So we we look at those three vectors and it's always kind of an optimization function along those three those three uh those three axes.

3:49:34

>> We think we found that here uh it was actually in terms of where my work was over the last few weeks.

3:49:38

It was a lot of I mean there's a qualitative you know kind of you know really like manual process of collecting feedback because um everyone's got a little bit of a different preference and we can only pick kind of one or two points on that curve and so just trying to kind of dial customer feedback namely developer feedback in for us on where that that balance of things are.

3:49:55

um is a big part of our our process for picking uh picking all those all those points.

3:49:59

Uh and so we hope that we hope that people like it and it unlocks uh you know the kind of maximal use. >> That's great.

3:50:07

Jord, >> how are you thinking about uh open source?

3:50:09

Who you know, who's been most excited to get access to it and uh yeah, where do you see it going?

3:50:17

>> Yeah, I mean it's important to us.

3:50:17

Uh you know, I'm glad we've uh we've gotten this out.

3:50:21

It's been a huge team effort.

3:50:22

Um, I think there was a kind of a thing that like, you know, Open AAI doesn't like open source anymore.

3:50:26

It's like, no, we're just like really busy with a gazillion other things.

3:50:28

Um, so I think hopefully going forward we've got more of a a a leaned in vantage point on on on open source.

3:50:35

But um, it it unlocks a huge number of use cases.

3:50:37

I mean, if you think about kind of like uh, you know, government use cases, uh, you think about on-prem, you know, use cases where you're you're handling uh, sensitive data in very sensitive environments, you think about where you want to run models on the edge.

3:50:50

Um all these things right now are kind of inaccessible to us as a service provider to customers um because we just don't quite have models that kind of fit at those points.

3:50:58

Uh so this for us we think is huge TAM expansion and uh we're excited to be able to work with enterprises on on implementing that model which is is I think you know competitive hopefully with with our O3 class of models.

3:51:09

So >> what is the landscape like for companies that are helping to implement OpenAI products at various enterprises?

3:51:15

you have the, you know, big consulting groups that will give you an AI strategy.

3:51:22

Maybe they'll try to take it a step further, but I imagine there's a cottage industry of of uh, you know, firms that have sprung up to try to help organizations unlock the value beyond, hey, let's just get everybody a seat with uh, chat work.

3:51:38

Yeah, I think there will be this new industry that emerges um that is kind of separate and apart from kind of the legacy set of SIS and uh um you know consultants uh that is really AI fluent. They're very AI native.

3:51:51

I think um it's very hard to borrow I think paradigms from the last 20 years of software building and uh you know implementation um that are going to kind of map to what we're dealing with here.

3:52:02

we're dealing with here. um you're dealing with fundamentally probabilistic systems uh that are moving and increasing and improving at a rate you know of of now kind of collapsing to every few months um and I think this the the the nature of use cases changes quickly um where enterprises are focused on kind of deploying them changes

3:52:21

quickly and so uh I think it's just hard for kind of the legacy industries to keep up frankly um we've had a lot of success working with some of this kind of new breed of SI so the distills of the world and others that um really have been born, I think, in forged in the fire, so to speak, of uh of this kind of new um this new platform. And so, we

3:52:37

And so, we hope there's more of them.

3:52:39

Um we we we'd be excited to work with anyone that uh that wants to work with us on it.

3:52:44

There's more business than we can handle.

3:52:45

And so, um we're always happy to to spread the love.

3:52:49

>> Talk about the uh $1 uh chat GPT product for uh the government.

3:52:55

Were you involved in that at all?

3:52:58

>> I was involved in that.

3:52:58

Um we we wanted to do something that was meaningful for for US government.

3:53:02

Um it's been a a real big focus of ours lately.

3:53:05

Um I think our our view is the government has got to start to modernize.

3:53:09

Um we've got to make sure that the tools that we use in the private sector are also in the hands of folks serving us in the public sector.

3:53:17

Um and we wanted to make that really simple.

3:53:18

So um we made chach you know basically equivalent to chach enterprise free.

3:53:22

Uh it's a dollar per year per agency.

3:53:24

Hopefully we can afford that.

3:53:26

Um, and uh, we wanted to make that uh, you know, available to anyone that wanted to use it uh, and standardize through GSA.

3:53:32

So, we're super appreciative of the partnership with them.

3:53:34

Uh, and more I think that we can do on that front.

3:53:37

>> How is that different than just like if I'm a government employee, I can just go to google.

3:53:41

com and I have access to that and Google right now.

3:53:43

I mean, Scott Kapor was saying that he can't use it. Yeah. So, yeah. Why? Yeah. Yeah.

3:53:49

just just talk to me about how how it's different to to offer uh chat as an actual service with a contract that you're that you're uh you know vending in you're actually they are a client versus just if you put up a website every government employee can access the web to some degree or would it be blocked like what why does it need to be like a deal at all as opposed to just like everyone just uses it? >> Yeah.

3:54:15

So um part of it is just making sure that government employees can access it.

3:54:18

So in some places obviously you know you can put blockers in place that wouldn't prevent access.

3:54:23

>> We hear a lot of stories by the way of people like going out on their lunch break to their car in the parking lot and like you know pulling up catchy on their phone and like throwing a bunch of stuff in there just to like because they know it'll get them through the day faster.

3:54:33

And >> um we've done work by the way with governments with the state of Pennsylvania and other places where we've seen dramatic increases you know things like two to three hours a day saved per employee um given the nature of the work that they do and and how helpful chatd can be.

3:54:47

And so this lets us have an interface into them as a customer.

3:54:51

It lets our team engage with them in a a direct way.

3:54:52

We can see how they're using the product and can help them use it better.

3:54:56

Um, and so that's that's important for us is like we got to build on that foundation with them.

3:55:00

>> And then presumably it also allows the the government to define like security and privacy in their world as opposed to if you're just like some website out there.

3:55:08

They their choice is only block or don't block as opposed to actually uh you know communicate with you.

3:55:12

This is okay to train on this is not etc.

3:55:14

like keep everything private, etc. , etc.

3:55:18

>> Yeah, I mean, we don't we don't train on on on enterprise data um at all.

3:55:20

So, yeah, you know, you're safe there, but um the Yeah, I mean, for us, like just being able to uh to treat them as a customer, right?

3:55:28

To treat them as a user and um you go, you know, you mentioned earlier like we were talking about kind of like there being these points of success at every organization that um you know, you've got people who are like way more sophisticated and using these tools than others.

3:55:40

Um we want to be able to see those people and amplify them and the government's no different.

3:55:43

Um, there are people that we've worked with in government who are incredibly sophisticated in how they use AI tools and our goal is to get everyone there.

3:55:51

>> How do you think about the group of users that are active students?

3:55:53

They've been on summer break.

3:55:56

You guys have been busy over summer.

3:55:58

Are you thinking about uh and and you recently launched uh I forget the exact name for the product.

3:56:03

I think it was like chatbt learning.

3:56:05

How are you thinking about that cohort and unlocking new capabilities for them uh this coming year? Yeah.

3:56:12

So we launched something called study mode um which was in our core chatbt product and um it was a little bit of an experiment.

3:56:19

We wanted to see if you change the way the model behaves when it can kind of uh when it knows you you want to be in a learning mode um if that can actually enhance outcomes for students.

3:56:29

Um where we we have all these kind of studies that have been done very like anecdotally about ChachiPT's ability to um to drive student outcomes and learning outcomes.

3:56:40

Um, so here we kind of took a little bit more of an intentional approach of if you actually model the take the model and actually use it in a more Socratic style where it can actually kind of quiz you, it can withhold certain information that it wants you to be able to to empirically deduce.

3:56:52

Um, it wants you to reason about problems and it kind of reasons with you as a partner. Um, so far so good.

3:56:58

Uh, it's it's really cool.

3:57:00

Um, and learning is kind of the killer use case of TACPt.

3:57:02

Uh, and so I think you know to be able to actually launch something that is in some sense extends that kind of killer use case.

3:57:08

um is has been really cool and and the student feedback so far, even on summer break, has been positive.

3:57:14

>> Well, we'll let you get back to your day.

3:57:15

What's what's next on your agenda?

3:57:17

Are you putting on the lab coat or the suit and tie and going to Washington? >> Uh good question.

3:57:21

Uh you know, today I'm I'm mostly with the team uh and talking to customers and um uh maybe tomorrow I'll get back to the lab code, but uh in the meantime, >> I appreciate you taking the time to talk today. So, >> yeah.

3:57:34

Well, thank you so much for taking the time to talk to us.

3:57:36

We will talk to you soon.

3:57:37

have a great rest of your day.

3:57:39

>> And the timeline has been in turmoil because President Trump says he will be imposing a 100% tariff on all semiconductors coming into the United States.

3:57:50

Uh it started with widespread tariffs on chips and then turned into export controls.

3:57:53

This is from the Kobe letter.

3:57:56

Is this a red flag moment?

3:57:56

I don't know why we have the red flag. >> I felt like it.

3:57:59

Ben was getting the flag.

3:58:01

>> Getting the flag and video potentially affected.

3:58:03

Taiwan says TSMC exempt from Trump's 100% chip tariff. Uh very unclear.

3:58:10

The story is obviously still developing.

3:58:13

Um and Dylan Midik says, "You're telling me that this level of monitoring the situation is free."

3:58:21

And it it's a picture of you in front of the whiteboard uh monitoring the chat GBT versus the timeline.

3:58:28

Today >> we're monitoring.

3:58:29

Um, Illinois has banned AI therapy, making it the first state to regulate the use of AI in mental health services.

3:58:37

>> Interesting headline that's coming out.

3:58:39

Um, >> which is interesting because the product can just be used for therapy like the user can choose to do that.

3:58:48

It's not necessarily It's kind of hard to ban outright.

3:58:49

Like maybe you can ban it in a clinical setting. >> Yep.

3:58:54

I wonder how they define this.

3:58:55

There's probably a a loophole if I know anything about how these bans are implemented.

3:58:59

Um, but yeah, m maybe it's like if you're if you're in the clinical setting, you can't be you can't use it, but then people will just use it independently like >> therapist just on their phone.

3:59:08

>> They're going to be going to the car.

3:59:09

>> They're going to be having they're going to No, they're just going to have it listening to the conversation.

3:59:13

>> They're going to be like, "Oh, what should I do right now? What should I say? >> What should I say?"

3:59:16

How does that make you feel?

3:59:19

>> That's what it's going to tell you.

3:59:19

Uh, Celsius nearly doubles revenue year-over-year.

3:59:23

This is the energy drink revenue of $739 million versus 632 million consensus.

3:59:32

North America grew 87% international grew 27%.

3:59:34

But here's the real kicker.

3:59:37

Alani new acquisition is the primary driver of growth.

3:59:39

Alani new added $300 million in revenue and retail sales are up.

3:59:44

So um wow, what a what performance.

3:59:47

But yeah, I mean that was the expectation when they bought Alani new is that they would I I I guess this is like the first moment they they um they rolled them in probably.

3:59:53

Um but uh huge growth for Celsius as they become multi-product multi- multi-consumer uh company.

4:00:02

Um what else is going on in the timeline? We have one last guest.

4:00:05

I think you might have to hop on with Taipei.

4:00:08

Um so feel free to jump when you need to.

4:00:11

Um Tyler, anything going on on the timeline we should be monitoring?

4:00:16

We are of course monitoring the situation.

4:00:19

>> Um I've been So So when Max was on, he was talking about like how you can like make a little game, right?

4:00:23

So I've been working on uh like a Bloons Tower Defense game. >> Okay. How's it going?

4:00:28

>> So it's going pretty well.

4:00:28

Um >> I'm I'm making another change, but then maybe I can screen record and and do a little share. >> Yeah.

4:00:34

>> Yeah, that'd be great.

4:00:34

You could share with the with the folks, too. Yeah.

4:00:35

Um I like this post from Ray Sullivan.

4:00:37

These GPT5 numbers are insane.

4:00:40

And it's a chart of GPT version versus number.

4:00:43

And then uh once it gets to four, it goes 4. 1, 4. 2, 4. 3, 4. 4, 4. 5.

4:00:48

So the the fifth one is a massive massive bar.

4:00:54

>> We need an analysis of the charts from today.

4:00:57

It seems like there was multiple that were >> kind of odd or hallucinated or or off.

4:01:05

>> It's interesting that multiple of them snuck up >> just in sheets.

4:01:07

The popular convenience store chain with 750 locations is now offering 50% off purchases paid with Bitcoin and crypto daily from 3 to 7 p. m.

4:01:18

What a wild move by Sheets.

4:01:18

Well, >> well Ben is in the waiting waiting room. Let's bring him in.

4:01:27

>> Let's bring in Ben Hilac. How you doing? Good to see you. >> Doing well. How are you guys doing? >> We're doing well.

4:01:31

I'm just going to say hello.

4:01:32

I got to take off and talk with Taipei.

4:01:34

I'm going to let John take it from here. Absolutely. I'll close out the show.

4:01:37

Let's have a fantastic conversation. >> Give me the update.

4:01:39

Uh, how's the day been for you?

4:01:41

What were your expectations? Did this meet succeed?

4:01:42

Uh, did it underwhelm you? How you doing?

4:01:47

>> Well, um, so I've actually had access for a couple weeks.

4:01:50

So, we actually did a video.

4:01:51

I'm not sure if you've seen it, but, uh, OpenAI brought a couple of, uh, folks from the Twitter sphere to their office a couple weeks ago to try and some other >> Yep. Yep. Yep.

4:02:01

Um, I think that uh it pretty much exactly meets my expectation as far as like how uh how it's been received.

4:02:10

Um, and I've tweeted about this as well, but I think that >> it's really really good at like oneshotting things.

4:02:17

Um, you know, I think like it's better than I think other models we've seen, but I think it's actually sort of a distraction in a lot of ways.

4:02:23

Um, I think that the things it's a lot better at are a a lot harder to describe and b I don't think the the harnesses for it really exist yet.

4:02:33

Um, I think >> harnesses.

4:02:37

So, >> uh, what the way I've been describing it is that I think I've seen, you know, web search existed in chatbt for a really long time, right?

4:02:45

Like it was able to like call a tool, search the web.

4:02:49

>> Y >> obviously like deep research was very different than that, right?

4:02:52

Like what we saw was it was like actually like calling you know searching the web it was like reasoning about those results changing its kind of course like course correcting the middle.

4:03:02

So like intermediate reasoning is like is the is the term for it.

4:03:06

>> Um and they really trained it how to search the web well.

4:03:08

Um I think GPT5 does that for like a whole plethora of tools.

4:03:15

>> The interesting thing is that a lot of products like I think a lot of the agent products that exist today were kind of built wrong.

4:03:21

uh like they weren't built that they didn't build their tools the right way.

4:03:25

right way. Um and we've seen this before like if you look at like you know the first um you know kind of infrastructure for agents was lang chain like way back when like two or three yeah it was it was you know it was uh it was early but it was wrong right and so like anybody that you know they've iterated since right they have like lang graph a better

4:03:44

implementation but the first imp implementation of lang chain was like again early but wrong and so if you built your product on lang chain like you had to you know significantly change it >> I think we will see a similar thing happen for GP5 five, you know, it's not just like, you know, uh, change the string in git, you know, from, you know, 40 to 5 or something and push and now you, you know. Yeah. Yeah. >> Yeah.

4:04:04

You know, that meme about like, oh, like Sam Hullman stood on stage and like just like, you know, killed 75 startups, Google just killed 100 startups, Apple just killed Partle with their new thing or whatever.

4:04:14

Uh, did any of that happen today?

4:04:17

It feels like it feels like this is like the Lang chain needing to change their strategy.

4:04:22

uh that happened a while ago.

4:04:24

I haven't identified anything.

4:04:27

It feels like, you know, Scott Wuh hopped on and said like, you know, great day to be an application layer company.

4:04:32

The foundation models got better.

4:04:34

>> It's uh it's more tools in my tool chest.

4:04:37

I'm extremely happy and uh and and I'm I'm more confident than ever. And I believe him.

4:04:42

I believe that he was he doesn't see today as like fundamentally needing to change his business model.

4:04:48

>> I I think that's true actually.

4:04:48

I think that um people have been you know there's a lot of people building agents right now.

4:04:54

I think a lot of them have not been feasible for some of the reasons that GPT5 starts to address.

4:04:59

So I think it is I think that what it means is that the entire architecture behind agents will get a lot simpler like it feels like a a good day for people building applications.

4:05:12

Um >> yeah it's not immediate that there's like some you know uh like company or something that got killed today. Yeah. Yeah. Yeah.

4:05:20

I mean, in general, it feels like, you know, Dwar Cash up updated his timelines.

4:05:23

There's just been a general idea that like we've we've maxed out pre-training, we've kind of maxed out post-training.

4:05:31

We're now in the let's reap the reward of this.

4:05:33

And we've seen it in uh like the incredible financial performance, the incredible usage numbers.

4:05:40

Uh you know, millions and millions, hundreds of millions of people are using CHP 30 30 minutes a day.

4:05:43

I'm I love the product and yet it feels it feel it feels like the what have you done for me lately meme. It's totally like okay. Yeah.

4:05:54

We went from the iPhone 4 to the iPhone 5 today. >> Yes.

4:05:58

>> Still really an important technology, great company, but like I want another iPhone 1. >> Yes. Yeah. Yeah. Yeah.

4:06:03

I No, I totally get what you're saying.

4:06:05

Um, I think that like I I I wrote a piece about this with Swix, but >> yeah, >> I it really actually changed the way I see that path to AGI.

4:06:14

Like I think before using it a lot, I kind of was like, "Okay, we need like bigger bigger models.

4:06:19

They're going to like get smarter or something."

4:06:22

>> Um, >> I think like I I had this realization.

4:06:27

So, I was watching it like solve um I had this like really weird um like dependency conflict with yarn.

4:06:31

Like we have like a mono repo.

4:06:34

It's like part of the problem also with this discourse is like um the sort of problems it gets good at solving are just like not sexy things to talk about.

4:06:41

They're not things that even you'll understand.

4:06:43

And I'm like, we have this issue with our like the way we structured things and like >> but like um a a couple weeks ago I was watching it like I had this problem.

4:06:50

No other model would solve it.

4:06:53

>> And um I watched it sort of like poke around like it started running this like yarn y command in a bunch of different directories in between.

4:06:59

It's like reasoning and like correctly reasoning about like why what and why and what it was learning and it you know taking little actions in between seeing what happened.

4:07:08

happened. M >> um I think what I realized is that like um you know if you imagine like humans without tools like if we never had any tools we're never even able to write things down like would you be able to tell that we're intelligent uh would we have like you know learn to speak etc like I I just like don't >> you know even if we could not have ever invented fire right it's like it's like

4:07:29

where would we be right now there there feels like there's a similar like I actually think a lot of the next year is just going to be how do you get these models to do things better is like you know I think it's next year >> uh in your yarn uh example um you said like you were you were having it I assume gpt 5 like work on the problem was that wrapped in a coding tool did you just go to chat.com and give it your

4:07:53

com and give it your GitHub repo like like talk to me like what was the actual user experience from your side >> yeah so this was in cursor >> um I think the codeex CLI the new version of the codeex CLI which they just released today is also really really really good.

4:08:09

>> Um I think that you will really only see a significant difference >> in places where it can sort of like explore its environment is the way I would put it.

4:08:19

Like when I was watching it like go bounce around my repo and like like I felt almost like I was watching something navigate like a little like video game like Pokemon or something like that.

4:08:29

That's kind of what it felt like.

4:08:30

Like it's kind of like I'm going to go over here. I'm going to see this. Okay, wait a minute.

4:08:32

That conflicts with what I just saw over here.

4:08:35

Like where should I go next? Do you know what I mean?

4:08:37

Like it felt very um novel uh is like what I would say. Yeah. >> Yeah. Yeah. Yeah.

4:08:42

Um what uh so yeah, I mean how are you using it?

4:08:45

What what um uh where do you see it going?

4:08:48

Do you see it like uh just like a little bump of a tailwind today or or what's your read on like uh like how you'll be using GPT5 going forward?

4:08:59

>> I mean yeah there's two huge things.

4:08:59

So like one thing that like really got missed today is that uh they also released GD5 nano which is like an incredibly good model actually.

4:09:08

Um so like we're not talking about it but it's half the cost for input tokens than flash light or sorry yeah I think it's actually half the cost of input tokens than flash light and it's a really good model like it's like 40 level for a lot of like writing and stuff like that.

4:09:24

Um and uh so yeah, we'll we'll be using that uh probably in the short term.

4:09:29

I think it'll be interesting to see how other providers react.

4:09:32

Like I'm sure Google will cut their prices as a result, but it is the cheapest like hosted uh model.

4:09:38

I think that I I don't think anyone's serving at any other model for those prices for that matter.

4:09:44

>> Yeah, >> that makes sense.

4:09:45

Um what else are you looking for for the rest of the year?

4:09:47

Uh probably no GPT6 on the horizon, but what are you looking out for?

4:09:51

I mean, it seems like Google is expected to respond with Gemini 3 soon.

4:09:55

What else are you tracking in the in the world of AI these days?

4:10:00

>> It's a great question.

4:10:00

Um, I think that, yeah, that's going to be wildly interesting.

4:10:04

I think what Google does will tell us a lot.

4:10:07

>> I think that they you've probably seen it, but you know, they released this like world model uh yesterday.

4:10:10

We're kind of not talking about it anymore.

4:10:13

I mean like if those videos I haven't tried it myself.

4:10:16

If those videos are real like that's that's one of the most mind-blowing things I've seen in the last like you know decade or something.

4:10:24

So like if that's real like that's extremely interesting and I think has all the stuff that's going on with role models right now has like huge implications for like everything like from robotics just like so many different fields.

4:10:34

So >> super super interested in that.

4:10:36

And the other thing is I actually just think that like again I'm I'm actually really bullish on Chip T5.

4:10:42

I think that the way it was received today is like just about how I expected it.

4:10:47

Like and the reason is like when I say harness again I'm like I think that like canvas in chatbt is pretty bad is like my would be my take like you know it's a tough product to make but like uh yeah like does really poorly with like long files crashes sometimes like that sort of like I think that we don't have >> the the product layer around GG5 doesn't exist yet.

4:11:06

So I think we're going to see some really really interesting products um that are built around it.

4:11:10

Yeah, it's always hard when you go from like a a binary qualitative in yourrface improvement GPT like CH GPT was like we passed the touring test and now the next test is like >> super intelligence that self-replicates smarter than every single person knows everything.

4:11:26

It's like the bar is like we really moved the the goalposts you know >> 100%.

4:11:31

I think that there was like a lot of you know discourse around the model as well like leading up to it which I think didn't help you know but like the way that I would think about it is like I think that you know depending there's some percentage of the way through automating software engineering that we've made it like let's say it's like 70% or something 75%.

4:11:50

>> Um the tough part is like that last like 25% is um a the hardest it's like the least um sort of decipherable to like explain to people.

4:11:58

It's the least like um universal like like if I'm just like oh make a you know one of the examples I did I I made a personal website it's like all Mac OS 9 themed in like 20 minutes with GT5.

4:12:09

Um, and so it's really fun, right? You get it. Like my mom gets it. Like I can show it. I I can share it. You get it.

4:12:16

You know, my mom I can't explain any of the like the very specific ways that 5 like helps in our specific codebase, our specific problem, whatever.

4:12:25

Um, so I think that like it'll be less and these launches will probably get less and less uh sort of interesting from a so like from a what it does for software engineering as that gap gets closed.

4:12:38

like I you know what's the last 5% of software engineering like I you know like I it's probably not going to be that interesting to me.

4:12:46

>> Um >> do you think they'll be on an annual release cadence now?

4:12:47

Like Apple >> updated all of their iOS all their operating system nomenclature to be like we are now on 26 because it's the year it's like a car model like like >> I don't think you can plan it.

4:12:59

I don't think you can plan ahead.

4:13:00

Like that's the interesting thing is like I think that you know there there's people that say that GT4.

4:13:05

5 was supposed to be GT5. Yep.

4:13:07

Um, and like I think that it sort of came out and they're like, "Eh, it's like, you know, it I I actually love 4. 5.

4:13:14

I think it's a really fun model, but um, >> well, it's clear that like improvements come in many places just like with the with the iPhone, like the latest iPhone, you buy that because it doesn't it's not just like the one with the new screen, it has a slightly better camera, slightly lighter, longer battery life.

4:13:29

It's like an ensemble of improvements that then they add up.

4:13:32

And I think that that feels like what we're getting here today and what we will get in the future is like this little like we did a little extra RL over here.

4:13:39

This tool is now sharper. It has new capabilities.

4:13:41

We added multimodal like you know the video generation got better and this feature got better etc etc.

4:13:48

got better etc etc. And I think that like what a model is is still going to change a lot and like how we value like so just give an example like 40 was sort of this big thing you know where they talked about it being like natively

4:14:00

multimodal you know taking in even like video at some point video in video out like audio in audio out and like you know you haven't heard that from GT5 yet like you can't talk to it on advanced voice mode like it doesn't it doesn't

4:14:14

generate images like you know what I mean there's no at least yet native image generation how it works under these model capabilities like seems quite possible like the best model for writing natural language might not or like writing um creative you know

4:14:37

creatively might not be the same model that writes you know really good rust code like these might be different models um so I don't know we'll see >> yeah create image here is now tucked next to deep research agent etc. But I

4:14:46

But I would hope that you can call that from the actual chat interface.

4:14:52

>> You can call it from the GT5 chat.

4:14:52

It's just using it's using um GPT image one I think is actually the name of the model.

4:14:58

So it's it's a dedicated image generation model which I think is maybe 40. I I don't totally know. >> Yeah.

4:15:02

I I just I I I don't particularly care.

4:15:05

I'm not looking for one model to rule them all.

4:15:07

I'm fine if with models calling different tools. It seems fine. Um >> yes. >> Anyway, uh fun day. Thanks for hopping on. We'll talk to you soon. >> Of course. Anytime. Talk. >> Have a good one. Bye.

4:15:17

And that's our show today, folks.

4:15:19

Leave us five stars on Apple Podcast and Spotify.

4:15:21

And thank you for tuning in to the GPT5 Giga stream.

4:15:27

We're on hour four and a half.

4:15:30

Uh we've enjoyed hanging out with you, Tyler.

4:15:32

Anything else from the timeline? Close it out for me.

4:15:33

Timeline's still in turmoil.

4:15:37

>> Show the little game I made. >> Okay.

4:15:38

Yeah, let's show Tyler's game. Can we do that? Is that >> You got it. The Tyler Sour Defense. Okay.

4:15:44

>> This was This was one shot. Okay.

4:15:44

I didn't >> Wait, what do you mean one shot? One prompt.

4:15:47

You said you were working on it.

4:15:50

>> I was, but then it's like, wasn't >> the definition?

4:15:52

Oh, so you went back to a single prompt. Got you.

4:15:54

This >> I made a change, but then I've realized like, okay, this is not as good.

4:15:57

So, I just went back to the first one. >> Okay.

4:16:00

So, yeah, my my my question is I mean, this this seems actually like it's it's like the game engine.

4:16:05

I don't know what it's using under the hood, what it's do you know?

4:16:09

Did it write like WebGL code or did it write like >> I think it's just it's just like JS. >> Okay.

4:16:14

And it's just like HTML canvas. >> That's pretty crazy. >> Yeah.

4:16:18

>> Um you'd think it would use some like 2D engine off the shelf or something, but um my my question is like what that won't go viral cuz that is less impressive than just the Tower Defense app that I can get in the app store. >> For sure.

4:16:33

Yeah, >> but it's like maybe if I take my, you know how those like control net images went viral where people would take their corporate logo and then they'd throw that through control net and it would be like the TBPN logo overlaid over like a forest and like the trees would look like the logo. >> Yeah.

4:16:48

So maybe like it's tower defense but it's my logo or something like that and like the the the enemies are like w like moving through something like that. I don't know.

4:16:57

There's just got to be a way to personalize it and make it so every single game is a unique snowflake that you want to go and experience that one.

4:17:04

You want to look at it, you want to spend some time in it. I don't know.

4:17:07

>> Yeah, >> it's hard cuz it's like >> it's still, you know, predicting the next token.

4:17:11

It's not like image the 40 image generation was like kind of a it wasn't novel, I guess, cuz there was image generation, but it was like such a massive improvement. This is >> Yeah.

4:17:21

>> Like >> there's not any clear massive step change here.

4:17:23

It's a little bit better in a lot of ways. >> Yeah. So, >> oh well.

4:17:27

Well, we'll have to play with it more.

4:17:28

Let us know what you think about GPT5 and we will see you tomorrow. Have a great day. Thank you so much. Bye.