Leopold Aschenbrenner — 2027 AGI, China/US super-intelligence race, & the return of history

0:00

What will be at stake will  not just be cool products But whether liberal democracy survives, Whether the CCP survives, what the world order for the next century is going to be The CCP is going to have an all out effort to infiltrate American AI labs Billions of dollars, thousands of people CCP is going to try to out-build us.

0:13

People don’t realize like how intense state level espionage can be When we have like literal superintelligence They can like Stuxnet the chinese data centers You really think that will be like a private company And the government wouldn’t be like “oh my god what is going on?

0:24

” I do think it is incredibly important that these clusters are in the united states I mean would you do the manhattan project in the UAE?

0:30

2023 was the moment for me when it went from AGI as this sort of theoretical,  abstract thing, and you’d make the models to like, I see it, I feel it.

0:38

I can see the cluster where  it’s trained on, like the rough combination of algorithms, the people, like how it’s happening,  and I think most of the world is not; most of the people who feel it are like right here Today I’m chatting with my friend Leopold Aschenbrenner.

0:51

He grew up in Germany and graduated  as valedictorian of Columbia when he was 19.

0:51

After that, he had a very interesting gap year which  we’ll talk about.

0:58

Then, he was on the OpenAI superalignment team, may it rest in peace.

1:04

Now, with some anchor investments — from Patrick and John Collison, Daniel Gross, and Nat  Friedman — he is launching an investment firm.

1:17

Leopold, you’re off to a slow start but  life is long.

1:17

I wouldn’t worry about it too much.

1:21

You’ll make up for it in due  time.

1:21

Thanks for coming on the podcast. Thank you.

1:26

I first discovered your  podcast when your best episode had a couple of hundred views.

1:31

It’s been amazing to  follow your trajectory. It’s a delight to be on.

1:38

In the Sholto and Trenton episode, I mentioned  that a lot of the things I’ve learned about AI I’ve learned from talking with them.

1:44

The  third, and probably most significant, part of this triumvirate has been you.

1:47

We’ll  get all the stuff on the record now.

1:53

Here’s the first thing I want to get on the  record.

1:53

Tell me about the trillion-dollar cluster.

1:59

I should mention this for the context of  the podcast.

1:59

Today you’re releasing a series called Situational Awareness.

2:04

We’re going to  get into it.

2:04

First question about that is, tell me about the trillion-dollar cluster.

2:07

Unlike most things that have recently come out of Silicon Valley, AI is an industrial  process.

2:13

The next model doesn’t just require some code.

2:19

It’s building a giant new cluster.

2:19

It’s building giant new power plants.

2:19

Pretty soon, it’s going to involve building giant new fabs.

2:25

Since ChatGPT, this extraordinary techno-capital acceleration has been set into  motion.

2:32

Exactly a year ago today, Nvidia had their first blockbuster earnings call.

2:37

It went up 25% after hours and everyone was like, "oh my God, AI is a thing."

2:42

Within a year, Nvidia  data center revenue has gone from a few billion a quarter to $25 billion a quarter and continues  to go up.

2:50

Big Tech capex is skyrocketing. It’s funny.

2:59

There’s this crazy scramble going on,  but in some sense it’s just the continuation of straight lines on a graph.

3:04

There’s this long-run  trend of almost a decade of training compute for the largest AI systems growing by about  half an order of magnitude, 0. 5 OOMs a year. Just play that forward.

3:16

GPT-4 was reported  to have finished pre-training in 2022.

3:23

On SemiAnalysis, it was rumored to have a cluster  size of about 25,000 A100s.

3:23

That’s roughly a $500 million cluster.

3:33

Very roughly, it’s 10 megawatts.

3:33

Just play that forward half a year.

3:33

By 2024, that’s a cluster that’s 100 MW and 100,000  H100 equivalents with costs in the billions.

3:52

Play it forward two more years.

3:52

By 2026, that’s  a gigawatt, the size of a large nuclear reactor.

3:59

That’s like the power of the Hoover Dam.

3:59

That costs tens of billions of dollars and requires a million H100 equivalents.

4:02

By 2028, that’s a cluster that’s ten GW.

4:06

That’s more power than most US states.

4:06

That’s 10 million H100 equivalents, costing hundreds of billions of dollars.

4:12

By 2030, you get the trillion-dollar cluster using 100 gigawatts, over 20% of US electricity  production.

4:18

That’s 100 million H100 equivalents.

4:25

That’s just the training cluster.

4:25

There are more  inference GPUs as well.

4:25

Once there are products, most of them will be inference GPUs.

4:30

US power  production has barely grown for decades.

4:39

Now we’re really in for a ride.

4:39

When I had Zuck on the podcast, he was claiming not a plateau per se, but that AI  progress would be bottlenecked by this constraint on energy.

4:51

Specifically, he was like, "oh,  gigawatt data centers, are we going to build another Three Gorges Dam or something?"

4:55

According to public reports, there are companies planning things on the scale of  a 1 GW data center.

5:01

With a 10 GW data center, who’s going to be able to build that?

5:06

A 100 GW  center is like a state project.

5:06

Are you going to pump that into one physical data center?

5:11

How  is it going to be possible? What is Zuck missing?

5:18

Six months ago, 10 GW was the talk of the town.

5:18

Now, people have moved on. 10 GW is happening.

5:25

There’s The Information report on OpenAI and  Microsoft planning a $100 billion cluster. Is that 1 GW? Or is that 10 GW?

5:31

I don’t know but if you try to map out how expensive the 10 GW cluster would be,  that’s a couple of hundred billion.

5:35

It’s sort of on that scale and they’re planning it.

5:39

It’s not just my crazy take.

5:39

AMD forecasted a $400 billion AI accelerator market by 2027.

5:50

AI  accelerators are only part of the expenditures.

5:56

We’re very much on track for a $1 trillion of  total AI investment by 2027.

5:56

The $1 trillion cluster will take a bit more acceleration.

6:04

We  saw how much ChatGPT unleashed.

6:04

Every generation, the models are going to be crazy  and shift the Overton window.

6:15

Then the revenue comes in.

6:15

These are  forward-looking investments.

6:15

The question is, do they pay off?

6:19

Let’s estimate the GPT-4 cluster  at around $500 million.

6:19

There’s a common mistake people make, saying it was $100 million for GPT-4.

6:27

That’s just the rental price.

6:27

If you’re building the biggest cluster, you have to build and pay  for the whole cluster.

6:34

You can’t just rent it for three months. Can’t you?

6:40

Once you’re trying to get into the  hundreds of billions, you have to get to like $100 billion a year in revenue.

6:42

This  is where it gets really interesting for the big tech companies because their revenues  are on the order of hundreds of billions. $10 billion is fine.

6:50

It’ll pay off  the 2024 size training cluster.

6:56

It’ll really be gangbusters with Big Tech when  it costs $100 billion a year.

6:56

The question is how feasible is $100 billion a year from  AI revenue?

6:59

It’s a lot more than right now.

7:05

If you believe in the trajectory of  AI systems as I do, it’s not that crazy.

7:12

There are like 300 million Microsoft Office  subscribers. They have Copilot now.

7:12

I don’t know what they’re selling it for.

7:19

Suppose  you sold some AI add-on for $100/month to a third of Microsoft Office subscribers.

7:25

That’d  be $100 billion right there. $100/month is a lot.

7:30

That’s a lot for a third of Office subscribers.

7:30

For the average knowledge worker, it’s a few hours of productivity a month.

7:35

You have  to be expecting pretty lame AI progress to not hit a few hours of productivity a month.

7:40

Sure, let’s assume all this.

7:40

What happens in the next few years?

7:47

What can the AI trained  on the 1 GW data center do?

7:47

What about the one on the 10 GW data center?

7:56

Just map out  the next few years of AI progress for me.

8:01

The 10 GW range is my best guess  for when you get true AGI.

8:01

Compute is actually overrated. We’ll talk about that.

8:08

By 2025-2026, we’re going to get models that are basically smarter than most college  graduates.

8:14

A lot of the economic usefulness depends on unhobbling.

8:21

The models are smart  but limited.

8:21

There are chatbots and then there are things like being able to use a computer  and doing agentic long-horizon tasks.

8:35

By 2027-2028, it’ll get as smart as the smartest  experts.

8:35

The unhobbling trajectory points to it becoming much more like an agent than a chatbot.

8:46

It’ll almost be like a drop-in remote worker.

8:53

This is the question around the economic returns.

8:53

Intermediate AI systems could be really useful, but it takes a lot of schlep to  integrate them.

8:59

There’s a lot you could do with GPT-4 or GPT-4.

9:04

5 in a business  use case, but you really have to change your workflows to make them useful.

9:07

It’s a very Tyler  Cowen-esque take.

9:07

It just takes a long time to diffuse.

9:12

We’re in SF and so we miss that.

9:12

But in some sense, the way these systems want to be integrated is where you get this kind  of sonic boom.

9:21

Intermediate systems could have done it, but it would have taken schlep.

9:28

Before  you do the schlep to integrate them, you’ll get much more powerful systems that are unhobbled.

9:32

They’re agents, drop-in remote workers.

9:32

You’re interacting with them like coworkers.

9:39

You  can do Zoom calls and Slack with them.

9:39

You can ask them to do a project and they go off and  write a first draft, get feedback, run tests on their code, and come back.

9:51

Then you can tell them  more things.

9:51

That’ll be much easier to integrate.

10:00

You might need a bit of overkill to make  the transition easy and harvest the gains.

10:05

What do you mean by overkill?

10:05

Overkill on model capabilities?

10:08

Yeah, the intermediate models could do it but  it would take a lot of schlep.

10:08

The drop-in remote worker AGI can automate cognitive  tasks.

10:13

The intermediate models would have made the software engineer more productive.

10:20

But will the software engineer adopt it?

10:24

With the 2027 model, you just don’t need the  software engineer.

10:24

You can interact with it like a software engineer, and it’ll  do the work of a software engineer.

10:32

The last episode I did was with John Schulman. I was asking about this.

10:32

We have these models that have come out in the last year and none seem  to have significantly surpassed GPT-4, certainly not in an agentic way where they interact  with you as a coworker.

10:46

They’ll brag about a few extra points on MMLU.

10:51

Even with GPT-4o,  it’s cool they can talk like Scarlett Johansson (I guess not anymore) but  it’s not like a coworker.

11:13

It makes sense why they’d be good at answering  questions.

11:13

They have data on how to complete Wikipedia text.

11:18

Where is the equivalent training  data to understand a Zoom call?

11:18

Referring back to your point about a Slack conversation,  how can it use context to figure out the cohesive project you’re working on?

11:30

Where is that training data coming from?

11:38

A key question for AI progress in the next few  years is how hard it is to unlock the test time compute overhang.

11:44

Right now, GPT-4 can do a few  hundred tokens with chain-of-thought.

11:44

That’s already a huge improvement.

11:52

Before, answering a  math question was just shotgun.

11:52

If you tried to answer a math question by saying the first thing  that comes to mind, you wouldn’t be very good.

12:03

GPT-4 thinks for a few hundred tokens.

12:03

If  I think at 100 tokens a minute, that’s like what GPT-4 does.

12:17

It’s equivalent to me thinking  for three minutes.

12:17

Suppose GPT-4 could think for millions of tokens.

12:24

That’s +4 OOMs on test  time compute on one problem. It can’t do it now. It gets stuck. It writes some code.

12:31

It can do a  little bit of iterative debugging, but eventually gets stuck and can’t correct its errors. There’s a big overhang.

12:37

In other areas of ML, there’s a great paper on AlphaGo, where you can  trade off train time and test time compute.

12:45

If you can use 4 OOMs more test time compute,  that’s almost like a 3. 5x OOM bigger model.

12:54

Again, if it’s 100 tokens a minute, a few million  tokens is a few months of working time.

12:54

There’s a lot more you can do in a few months of working  time than just getting an answer right now.

13:00

The question is how hard is it to unlock that?

13:03

In the short timelines AI world, it’s not that hard.

13:11

The reason it might not be  that hard is that there are only a few extra tokens to learn.

13:16

You need to learn things like  error correction tokens where you’re like “ah, I made a mistake, let me think about that again.

13:22

”  You need to learn planning tokens where it’s like “I’m going to start by making a plan.

13:25

Here’s my  plan of attack.

13:25

I’m going to write a draft and now I’m going to critique my draft and think about  it.

13:30

” These aren’t things that models can do now, but the question is how hard it is.

13:35

There are two paths to agents.

13:35

When Sholto was on your podcast, he talked about  scaling leading to more nines of reliability. That’s one path.

13:48

The other path is the  unhobbling path.

13:48

It needs to learn this System 2 process.

13:53

If it can learn that, it can  use millions of tokens and think coherently. Here’s an analogy.

14:04

When you drive, you’re  on autopilot most of the time.

14:04

Sometimes you hit a weird construction zone or intersection.

14:13

Sometimes my girlfriend is in the passenger seat and I’m like “ah, be quiet for a moment,  I need to figure out what’s going on.

14:18

” You go from autopilot to System 2 and  you’re thinking about how to do it.

14:23

Scaling improves that System 1 autopilot.

14:29

The brute force  way to get to agents is improving that system.

14:29

If you can get System 2 working, you can  quickly jump to something more agentified and test time compute overhang is unlocked.

14:43

What’s the reason to think this is an easy win?

14:49

Is there some loss function that easily  enables System 2 thinking?

14:49

There aren’t many animals with System 2 thinking.

15:00

It took a long  time for evolution to give us System 2 thinking.

15:05

Pre-training has trillions of tokens of Internet  text, I get that.

15:05

You match that and get all of these free training capabilities.

15:11

What’s  the reason to think this is an easy unhobbling?

15:22

First of all, pre-training is magical.

15:22

It gave us a huge advantage for models of general intelligence because you can predict the  next token.

15:29

But there’s a common misconception.

15:35

Predicting the next token lets the model learn  incredibly rich representations.

15:35

Representation learning properties are the magic of  deep learning.

15:39

Rather than just learning statistical artifacts, the models learn models  of the world.

15:44

That’s why they can generalize, because it learned the right representations.

15:49

When you train a model, you have this raw bundle of capabilities that’s useful.

15:55

The unhobbling  from GPT-2 to GPT-4 took this raw mass and RLHF’d it into a good chatbot. That was a huge win.

16:06

In the original InstructGPT paper, comparing RLHF vs.

16:13

non-RLHF models it’s like a 100x model  size win on human preference rating.

16:13

It started to be able to do simple chain-of-thought and  so on.

16:18

But you still have this advantage of all these raw capabilities, and there’s still  a huge amount you’re not doing with them.

16:28

This pre-training advantage is also the difference  to robotics.

16:28

People used to say it was a hardware problem.

16:35

The hardware is getting solved,  but you don’t have this huge advantage of bootstrapping with pre-training.

16:41

You don’t  have all this unsupervised learning you can do.

16:45

You have to start right away with RL self-play.

16:45

The question is why RL and unhobbling might work.

17:00

Bootstrapping is an advantage.

17:00

Your Twitter  bio is being pre-trained.

17:00

You’re not being pre-trained anymore.

17:05

You were pre-trained in  grade school and high school.

17:05

At some point, you transition to being able to learn by yourself.

17:10

You weren’t able to do it in elementary school.

17:18

High school is probably where it started and by  college, if you’re smart, you can teach yourself.

17:26

Models are just starting to enter that regime.

17:26

It’s a little bit more scaling and then you figure out what goes on top. It won’t be trivial.

17:32

A lot  of deep learning seems obvious in retrospect.

17:41

There’s some obvious cluster of ideas.

17:41

There  are some ideas that seem a little dumb but work.

17:46

There are a lot of details you have to  get right.

17:46

We’re not going to get this next month.

17:50

It’ll take a while to figure out.

17:50

A while for you is like half a year.

17:55

I don’t know, between six months and three  years. But it's possible.

17:55

It’s also very related to the issue of the data wall.

18:06

Here’s one  intuition on learning by yourself.

18:06

Pre-training is kind of like the teacher lecturing to  you and the words are flying by.

18:15

You’re just getting a little bit from it.

18:24

That's not what you do when you learn by yourself.

18:28

When you learn by yourself,  say you're reading a dense math textbook, you're not just skimming through it once.

18:32

Some wordcels just skim through and reread and reread the math textbook and they memorize.

18:37

What you do is you read a page, think about it, have some internal monologue going on, and have  a conversation with a study buddy.

18:45

You try a practice problem and fail a bunch of times.

18:49

At some point it clicks, and you're like, "this made sense."

18:53

Then you read a few more pages.

18:53

We've kind of bootstrapped our way to just starting to be able to do that  now with models.

18:59

The question is, can you use all this sort of self-play, synthetic  data, RL to make that thing work.

19:08

Right now, there's in-context learning, which is super  sample efficient.

19:18

In the Gemini paper, it just learns a language in-context.

19:23

Pre-training, on  the other hand, is not at all sample efficient.

19:30

What humans do is a kind of in-context  learning.

19:30

You read a book, think about it, until eventually it clicks.

19:34

Then you somehow  distill that back into the weights.

19:34

In some sense, that's what RL is trying to do.

19:39

RL is super  finicky, but when it works it's kind of magical.

19:46

It's the best possible data for the model.

19:46

It’s when you try a practice problem, fail, and at some point figure it out in a way  that makes sense to you.

19:51

That's the best possible data for you because it's the  way you would have solved the problem, rather than just reading how somebody else solved  the problem, which doesn't initially click.

20:07

By the way, if that take sounds familiar it's  because it was part of the question I asked John Schulman.

20:10

It goes to illustrate the thing  I said in the intro.

20:10

A bunch of the things I've learned about AI comes from these dinners we  do before the interviews with me, you, Sholto, and a couple of others.

20:20

We’re like, “what should  I ask John Schulman, what I should ask Dario.

20:20

” Suppose this is the way things  go and we get these unhobblings— And the scaling.

20:31

You have this baseline of this  enormous force of scaling. GPT-2 was amazing.

20:31

It could string together plausible sentences, but it  could barely do anything.

20:38

It was kind of like a preschooler.

20:43

GPT-4, on the other hand, could write  code and do hard math, like a smart high schooler.

20:49

This big jump in capability is explored in the  essay series.

20:49

I count the orders of magnitude of compute and scale-up of algorithmic progress.

20:53

Scaling alone by 2027-2028 is going to do another preschool to high school jump on top of GPT-4.

21:00

At  a per token level, the models will be incredibly smart.

21:07

They'll gain more reliability, and with the  addition of unhobblings, they'll look less like chatbots and more like agents or drop-in remote  workers.

21:11

That's when things really get going.

21:20

I want to ask more questions about this but let's  zoom out.

21:20

Suppose you're right about this.

21:20

This is because of the 2027 cluster which is at 10 GW? 2028 is 10 GW.

21:28

Maybe it'll be pulled forward. Something like a 5.

21:37

5 level by 2027, whatever  that's called.

21:37

What does the world look like at that point?

21:45

You have these remote workers who can  replace people.

21:45

What is the reaction to that in terms of the economy, politics, and geopolitics?

21:51

2023 was a really interesting year to experience as somebody who was really following the AI stuff.

22:00

What were you doing in 2023? OpenAI.

22:06

When you were at OpenAI in 2023, it was a  weird thing.

22:06

You almost didn't want to talk about AI or AGI.

22:20

It was kind of a dirty word.

22:20

Then  in 2023, people saw ChatGPT for the first time, they saw GPT-4, and it just exploded.

22:25

It triggered huge capital expenditures from all these firms and an explosion in revenue  from Nvidia and so on.

22:31

Things have been quiet since then, but the next thing has been in the  oven.

22:39

I expect every generation these g-forces to intensify.

22:44

People will see the models.

22:44

They won’t have counted the OOMs so they're going to be surprised. It'll be kind of crazy.

22:50

Revenue is going to accelerate.

22:50

Suppose you do hit $10 billion by the end of this year.

22:54

Suppose it just continues on the trajectory of revenue doubling every six months.

22:58

It's not  actually that far from $100 billion, maybe by 2026.

23:04

At some point, what happened to Nvidia  is going to happen to Big Tech. It's going to explode.

23:10

A lot more people are going to feel it.

23:10

2023 was the moment for me where AGI went from being this theoretical, abstract thing.

23:21

I see  it, I feel it, and I see the path. I see where it's going.

23:28

I can see the cluster it's trained on,  the rough combination of algorithms, the people, how it's happening.

23:34

Most of the world is not there  yet.

23:34

Most of the people who feel it are right here.

23:38

A lot more of the world is going to start  feeling it.

23:38

That's going to start being intense. Right now, who feels it?

23:48

You can go on Twitter  and there are these GPT wrapper companies, like, "whoa, GPT-4 is going to change our business."

23:53

I'm so bearish on the wrapper companies because they're betting on stagnation.

23:57

They're  betting that you have these intermediate models and it takes so much schlep to integrate  them.

24:02

I'm really bearish because we're just going to sonic boom you.

24:06

We're going to get the  unhobblings.

24:06

We're going to get the drop-in remote worker.

24:09

Your stuff is not going to matter. So that's done.

24:09

SF, this crowd, is paying attention now.

24:17

Who is going to be paying  attention in 2026 and 2027?

24:17

Presumably, these are years in which hundreds of  billions of capex is being spent on AI.

24:30

The national security state is  going to start paying a lot of attention.

24:32

I hope we get to talk about that. Let’s talk about it now. What happens?

24:32

What is the immediate political reaction?

24:39

Looking  internationally, I don't know if Xi Jinping sees the GPT-4 news and goes, "oh, my God,  look at the MMLU score on that.

24:45

What are we doing about this, comrade?"

24:50

So what happens when he sees a remote worker replacement and it has $100  billion in revenue?

24:57

There’s a lot of businesses that have $100 billion in revenue, and people  aren't staying up all night talking about it.

25:06

The question is, when does the CCP and when does  the American national security establishment realize that superintelligence is going to be  absolutely decisive for national power?

25:11

This is where the intelligence explosion stuff  comes in, which we should talk about later. You have AGI.

25:21

You have this drop-in  remote worker that can replace you or me, at least for remote jobs.

25:25

Fairly quickly, you  turn the crank one or two more times and you get a thing that's smarter than humans.

25:37

Even more than just turning the crank a few more times, one of the first jobs to  be automated is going to be that of an AI researcher or engineer.

25:46

If you can automate  AI research, things can start going very fast.

25:54

Right now, there's already at this trend  of 0.

25:54

5 OOMs a year of algorithmic progress.

25:59

At some point, you're going to have GPU fleets  in the tens of millions for inference or more.

26:04

You’re going to be able to run 100 million human  equivalents of these automated AI researchers.

26:10

If you can do that, you can maybe do a decade's  worth of ML research progress in a year.

26:10

You get some sort of 10x speed up.

26:15

You can make  the jump to AI that is vastly smarter than humans within a year, a couple of years.

26:23

That broadens from there.

26:23

You have this initial acceleration of AI research.

26:29

You apply  R&D to a bunch of other fields of technology.

26:29

At this point, you have a billion super intelligent  researchers, engineers, technicians, everything.

26:42

They’re superbly competent at all things.

26:42

They're going to figure out robotics.

26:42

We talked about that being a software problem.

26:46

Well,  you have a billion super smart — smarter than the smartest human researchers — AI researchers  in your cluster.

26:51

At some point during the intelligence explosion, they're going to be able  to figure out robotics. Again, that’ll expand.

27:01

If you play this picture forward, it is fairly  unlike any other technology.

27:01

A couple years of lead could be utterly decisive in say, military  competition.

27:11

If you look at the first Gulf War, Western coalition forces had a 100:1 kill ratio.

27:19

They had better sensors on their tanks.

27:19

They had better precision missiles, GPS, and stealth.

27:26

They had maybe 20-30 years of technological lead.

27:33

They just completely crushed them.

27:33

Superintelligence applied to broad fields of R&D — and the industrial explosion that comes  from it, robots making a lot of material — could compress a century’s worth of technological  progress into less than a decade.

27:47

That means that a couple years could mean a Gulf War 1-style  advantage in military affairs.

27:52

That’s including a decisive advantage that even preempts nukes.

28:03

How do you find nuclear stealth submarines?

28:08

Right now, you have sensors and software to  detect where they are. You can do that. You can find them.

28:13

You have millions or billions  of mosquito-sized drones, and they take out the nuclear submarines.

28:18

They take out the mobile  launchers.

28:18

They take out the other nukes.

28:24

It’s potentially enormously destabilizing  and enormously important for national power.

28:29

At some point people are going to realize  that. Not yet, but they will.

28:29

When they do, it won’t just be the AI researchers in charge.

28:38

The CCP is going to have an all-out effort to infiltrate American AI labs.

28:45

It’ll involve  billions of dollars, thousands of people, and the full force of the Ministry of State  Security.

28:49

The CCP is going to try to outbuild us.

29:01

They added as much power in the last decade as an  entire US electric grid.

29:01

So the 100 GW cluster, at least the 100 GW part of it, is going to be  a lot easier for them to get.

29:05

By this point, it's going to be an extremely  intense international competition.

29:18

One thing I'm uncertain about in this picture  is if it’s like what you say, where it's more of an explosion. You’ve developed an AGI.

29:23

You  make it into an AI researcher.

29:23

For a while, you're only using this ability to make hundreds  of millions of other AI researchers.

29:32

The thing that comes out of this really frenetic process  is a superintelligence.

29:39

Then that goes out in the world and is developing robotics and helping  you take over other countries and whatever.

29:52

It's a little bit more gradual.

29:52

It's an  explosion that starts narrowly.

29:52

It can do cognitive jobs.

29:56

The highest ROI use for  cognitive jobs is to make the AI better and solve robotics.

29:59

As you solve robotics, now  you can do R&D in biology and other technology.

30:06

Initially, you start with the factory workers.

30:06

They're wearing the glasses and AirPods, and the AI is instructing them because you can make any  worker into a skilled technician.

30:11

Then you have the robots come in. So this process expands.

30:15

Meta's Ray-Bans are a complement to Llama.

30:21

With the fabs in the US, their constraint is  skilled workers.

30:21

Even if you don't have robots, you have the cognitive superintelligence  and can kind of make them all into skilled workers immediately.

30:30

That's a very  brief period. Robots will come soon.

30:34

Suppose this is actually how the tech  progresses in the United States, maybe because these companies are already generating  hundreds of billions of dollars of AI revenue At this point, companies are borrowing hundreds  of billions or more in the corporate debt markets.

30:47

Why is a CCP bureaucrat, some 60-year-old  guy, looking at this and going, "oh, Copilot has gotten better now" and now— This is much more than Copilot has gotten better now.

30:52

It’d require shifting the production of an entire country,  dislocating energy that is otherwise being used for consumer goods or something, and feeding all  that into the data centers.

31:08

Part of this whole story is that you realize superintelligence is  coming soon.

31:17

You realize it and maybe I realize it.

31:23

I'm not sure how much I realize it.

31:23

Will the national security apparatus in the United States and the CCP realize it?

31:27

This is a really key question.

31:27

We have a few more years of mid-game.

31:35

We have a few more  2023s.

31:35

That just starts updating more and more people.

31:39

The trend lines will become clear.

31:39

You will see some amount of the COVID dynamic.

31:50

COVID in February of 2020 honestly feels a lot  like today.

31:50

It feels like this utterly crazy thing is coming.

31:58

You see the exponential and yet most  of the world just doesn't realize it.

31:58

The mayor of New York is like, "go out to the shows," and "this  is just Asian racism."

32:05

At some point, people saw it and then crazy, radical reactions came.

32:16

By the way, what were you doing during COVID?

32:23

Was it your freshman or sophomore year? Junior.

32:29

Still, you were like a 17-year-old junior or  something right?

32:29

Did you short the market or something?

32:36

Did you sell at the right time? Yeah.

32:42

So there will be a March 2020 moment.

32:42

You can make the analogy you make in the series that this will cause a reaction like, “we  have to do the Manhattan Project again for America here.

32:58

” I wonder what the politics of this will  be like.

32:58

The difference here is that it’s not just like, “we need the bomb to beat the Nazis.

33:03

” We'll be building this thing that makes all our energy prices go up a bunch and it's automating a  lot of our jobs.

33:09

The climate change stuff people are going to be like, "oh, my God, it's making  climate change worse and it's helping Big Tech."

33:19

Politically, this doesn't seem like a dynamic  where the national security apparatus or the president is like, "we have to step on  the gas here and make sure America wins."

33:30

Again, a lot of this really depends on how  much people are feeling it and how much people are seeing it.

33:33

Our generation is so used to  peace, American hegemony and nothing matters.

33:51

The historical norm is very much one of extremely  intense and extraordinary things happening in the world with intense international competition.

33:54

There's a 20-year very unique period.

33:54

In World War II, something like 50% of GDP went  to war production.

34:09

The US borrowed over 60% of GDP.

34:14

With Germany and Japan I think  it was over 100%.

34:14

In World War I, the UK, France, and Germany all borrowed over 100% of GDP.

34:21

Much more was on the line.

34:21

People talk about World War I being so destructive with 20  million Soviet soldiers dying and 20% of Poland.

34:38

That happened all the time.

34:38

During  the Seven Years' War something like 20-30% of Prussia died.

34:43

In the Thirty Years' War,  up to 50% of a large swath of Germany died.

34:58

Will people see that the stakes here are  really high and that history is actually back?

35:04

The American national security state  thinks very seriously about stuff like this.

35:11

They think very seriously about competition with  China.

35:11

China very much thinks of itself on this historical mission of the rejuvenation of the  Chinese nation.

35:15

They think a lot about national power.

35:18

They think a lot about the world order.

35:18

There's a real question on timing.

35:18

Do they start taking this seriously when the intelligence  explosion is already happening quite late.

35:25

Do they start taking this seriously two years earlier?

35:30

That matters a lot for how things play out.

35:34

At some point they will and they will realize  that this will be utterly decisive for not just some proxy war but for major questions.

35:42

Can liberal democracy continue to thrive?

35:42

Can the CCP continue existing?

35:48

That will activate  forces that we haven't seen in a long time.

35:56

The great power conflict definitely seems  compelling.

35:56

All kinds of different things seem much more likely when you think  from a historical perspective.

36:02

You zoom out beyond the liberal democracy that  we’ve had the pleasure to live in America for say the last 80 years.

36:10

That includes  things like dictatorships, war, famine, etc.

36:16

I was reading The Gulag Archipelago and one of the  chapters begins with Solzhenitsyn saying how if you had told a Russian citizen under the tsars  that because of all these new technologies — we wouldn’t see some Great Russian revival with  Russia becoming a great power and the citizens made wealthy — you would see tens of millions of  Soviet citizens tortured by millions of beasts in the worst possible ways.

36:39

If you’d told them that  that would be the result of the 20th century, they wouldn’t have believed you.

36:44

They’d have called you a slanderer.

36:50

The possibilities for dictatorship with  superintelligence are even crazier as well.

36:55

Imagine you have a perfectly loyal military  and security force. No more rebellions.

36:55

No more popular uprisings.

37:00

You have perfect lie  detection.

37:00

You have surveillance of everybody.

37:07

You can perfectly figure out who's the dissenter  and weed them out.

37:07

No Gorbachev who had some doubts about the system would have ever risen to  power.

37:12

No military coup would have ever happened.

37:18

There's a real way in which part of why things  have worked out is that ideas can evolve.

37:18

There's some sense in which time heals a lot of wounds and  solves a lot of debates.

37:28

Throughout time, a lot of people had really strong convictions, but a lot  of those have been overturned over time because there's been continued pluralism and evolution.

37:38

Imagine applying a CCP-like approach to truth where truth is what the party says.

37:44

When  you supercharge that with superintelligence, that could just be locked in and enshrined for a long  time.

37:48

The possibilities are pretty terrifying.

37:56

To your point about history and living in America  for the past eight years, this is one of the things I took away from growing up in Germany.

38:03

A  lot of this stuff feels more visceral.

38:03

My mother grew up in the former East, my father in the  former West.

38:08

They met shortly after the Wall fell.

38:12

The end of the Cold War was this extremely pivotal  moment for me because it's the reason I exist.

38:16

I grew up in Berlin with the former Wall.

38:16

My great-grandmother, who is still alive, is very important in my life.

38:24

She was born in 1934  and grew up during the Nazi era.

38:24

In World War II, she saw the firebombing of Dresden from  this country cottage where they were as kids.

38:37

Then she spent most of her life in  the East German communist dictatorship.

38:42

She'd tell me about how Soviet tanks came  when there was the popular uprising in 1954.

38:48

Her husband was telling her to get home  really quickly and get off the streets.

38:52

She had a son who tried to ride a motorcycle  across the Iron Curtain and then was put in a Stasi prison for a while.

38:58

Finally, when  she's almost 60, it was the first time she lived in a free country, and a wealthy country.

39:06

When I was a kid, the thing she always really didn't want me to do was get involved in  politics.

39:17

Joining a political party had very bad connotations for her.

39:22

She raised  me when I was young.

39:22

So it doesn't feel that long ago. It feels very close.

39:30

There’s one thing I wonder about when we're talking today about the CCP.

39:34

The people  in China who will be doing their version of this project will be AI researchers who are somewhat  Westernized.

39:40

They’ll either have gotten educated in the West or have colleagues in the West.

39:48

Are they going to sign up for the CCP project that's going to hand over control to Xi  Jinping?

39:57

What's your sense of that?

39:57

Fundamentally, they're just people, right?

40:03

Can't you convince  them about the dangers of superintelligence?

40:07

Will they be in charge though?

40:07

In some  sense, this is also the case in the US.

40:13

This is like the rapidly depreciating  influence of the lab employees.

40:13

Right now, the AI lab employees have so much power.

40:17

You  saw this November event. It’s so much power.

40:25

Both are going to get automated and they're  going to lose all their power.

40:25

It'll just be a few people in charge with their armies of  automated AIs.

40:29

It’s also the politicians and the generals and the national security  state.

40:35

There are some of these classic scenes from the Oppenheimer movie.

40:40

The  scientists built it and then the bomb was shipped away and it was out of their hands.

40:44

It's good for lab employees to be aware of this.

40:50

You have a lot of power now, but  maybe not for that long. Use it wisely.

40:58

I do think they would benefit from some  more organs of representative democracy.

41:01

What do you mean by that?

41:01

In the OpenAI board events, employee power is exercised in a very direct  democracy way.

41:05

How some of that went about really highlighted the benefits of representative  democracy and having some deliberative organs. Interesting.

41:14

Let's go back to the $100 billion  revenue question.

41:14

The companies are trying to build clusters that are this big.

41:24

Where  are they building it?

41:24

Say it's the amount of energy that would be required for a small or  medium-sized US state.

41:28

Does Colorado then get no power because it's happening in the United  States?

41:33

Is it happening somewhere else?

41:37

This is the thing that I always find funny,  when you talk about Colorado getting no power.

41:40

The easy way to get the power would be to  displace less economically useful stuff.

41:45

Buy up the aluminum smelting plant that has a  gigawatt.

41:45

We're going to replace it with the data center because that's important.

41:49

That's not  actually happening because a lot of these power contracts are really locked in long-term.

41:53

Also, people don't like things like this.

41:59

In practice what it requires, at least  right now, is building new power. That might change.

42:03

That's when things get really  interesting, when it's like, “no, we're just dedicating all of the power to the AGI.

42:06

” So right now it's building new power. 10 GW is quite doable.

42:10

It's like a few percent of US  natural gas production.

42:10

When you have the 10 GW training cluster, you have a lot more inference.

42:18

100 gigawatts is where it starts getting pretty wild.

42:23

That's over 20% of US electricity  production.

42:23

It's pretty doable, especially if you're willing to go for natural gas.

42:30

It is incredibly important that these clusters are in the United States.

42:37

Why does it matter that it's in the US?

42:43

There are some people who are trying to  build clusters elsewhere.

42:43

There's a lot of free-flowing Middle Eastern money that's trying  to build clusters elsewhere.

42:48

This comes back to the national security question we talked about.

42:54

Would you do the Manhattan Project in the UAE?

42:59

You can put the clusters in the US and you can  put them in allied democracies.

42:59

Once you put them in authoritarian dictatorships, you create this  irreversible security risk.

43:07

Once the cluster is there, it's much easier for them to exfiltrate  the weights.

43:14

They can literally steal the AGI, the superintelligence.

43:18

It’s like they got a direct  copy of the atomic bomb.

43:18

It makes it much easier for them.

43:25

They have weird ties to China.

43:25

They  can ship that to China. That's a huge risk.

43:29

Another thing is they can just seize the compute.

43:29

The issue here is people right now are thinking of this as ChatGPT, Big Tech product  clusters.

43:35

The clusters being planned now, three to five years out, may well be the AGI,  superintelligence clusters.

43:39

When things get hot, they might just seize the compute.

43:45

Suppose we put 25% of the compute capacity in these Middle Eastern dictatorships. Say they seize that.

43:50

Now it's a ratio of compute of 3:1.

43:54

We still have more, but even with  only 25% of compute there it starts getting pretty hairy.

44:01

3:1 is not that great of a ratio.

44:01

You can do a lot with that amount of compute.

44:06

Say they don't actually do this.

44:06

Even if  they don't actually seize the compute, even if they actually don't steal the weights,  there's just a lot of implicit leverage you get.

44:13

They get seats at the AGI table.

44:13

I  don't know why we're giving authoritarian dictatorships the seat at the AGI table.

44:21

There's going to be a lot of compute in the Middle East if these deals go through. First of all, who is it?

44:26

Is it just every single Big Tech company trying  to figure it out over there?

44:35

It’s not everybody, some.

44:35

There are reports, I think Microsoft. We'll get into it.

44:36

So say the UAE gets a bunch of compute because we're building the clusters there.

44:43

Let's say they have 25% of the compute.

44:43

Why does a compute ratio matter?

44:49

If it's about them being  able to kick off the intelligence explosion, isn't it just some threshold where you have  100 million AI researchers or you don't?

45:01

You can do a lot with 33 million extremely  smart scientists.

45:01

That might be enough to build the crazy bio weapons.

45:09

Then  you're in a situation where they stole the weights and they seized the compute.

45:13

Now they can make these crazy new WMDs that will be possible with superintelligence.

45:19

Now you've just proliferated the stuff that’ll be really powerful.

45:22

Also, 3x  on compute isn't actually that much.

45:29

The riskiest situation is if we're in some sort  of really neck and neck, feverish international struggle.

45:45

Say we're really close with the CCP  and we're months apart.

45:45

The situation we want to be in — and could be in if we play our cards  right — is a little bit more like the US building the atomic bomb versus the German project years  behind.

45:55

If we have that, we just have so much more wiggle room to get safety right.

46:02

We're going to be building these crazy new WMDs that completely undermine nuclear  deterrence.

46:05

That's so much easier to deal with if you don't have somebody right on your  tails and you have to go at maximum speed. You have no wiggle room.

46:18

You're worried  that at any time they can overtake you.

46:22

They can also just try to outbuild you.

46:22

They  might literally win.

46:22

China might literally win if they can steal the weights, because they  can outbuild you.

46:27

They may have less caution, both good and bad caution in terms of  whatever unreasonable regulations we have.

46:40

If you're in this really tight race, this  sort of feverish struggle, that's when there's the greatest peril of self-destruction.

46:43

Presumably the companies that are trying to build clusters in the Middle East realize this.

46:49

Is it just that it’s impossible to do this in America?

46:53

If you want American companies  to do this at all, do you have to do it in the Middle East or not at all?

46:56

Then you just  have China build a Three Gorges Dam cluster. There’s a few reasons.

47:00

People  aren’t thinking about this as the AGI superintelligence cluster.

47:02

They’re just  like, “ah, cool clusters for my ChatGPT.

47:02

” If you’re doing ones for inference, presumably  you could spread them out across the country or something.

47:15

The ones they’re building,  they’re going to do one training run in a single thing they’re building.

47:19

It’s just hard to distinguish between inference and training compute.

47:23

People  can claim it’s inference compute, but they might realize that actually this is  going to be useful for training compute too.

47:33

Because of synthetic data and things like that?

47:33

RL looks a lot like inference, for example.

47:33

Or you just end up connecting them in time.

47:38

It's a  lot like raw materials.

47:38

It's like placing your uranium refinement facilities there.

47:44

So there are a few reasons.

47:44

One, they don't think about this as the AGI  cluster.

47:47

Another is just that there’s easy money coming from the Middle East.

47:50

Another one is that some people think that you can't do it in the US.

47:57

We actually  face a real system competition here.

47:57

Some people think that only autocracies that can do  this with top-down mobilization of industrial capacity and the power to get stuff done fast.

48:08

Again, this is the sort of thing we haven't faced in a while.

48:13

But during the Cold War, there was  this intense system competition. East vs. West Germany was this.

48:19

It was West Germany as liberal  democratic capitalism vs. state-planned communism.

48:27

Now it's obvious that the free world  would win.

48:27

But even as late as 1961, Paul Samuelson was predicting that the Soviet  Union would outgrow the United States because they were able to mobilize industry better.

48:36

So there are some people who shitpost about loving America, but then in private they're  betting against America.

48:43

They're betting against the liberal order.

48:46

Basically, it's just a bad  bet.

48:46

This stuff is really possible in the US.

48:53

To make it possible in the US, to some degree we  have to get our act together.

48:53

There are basically two paths to doing it in the US.

48:58

One is you just  have to be willing to do natural gas.

48:58

There's ample natural gas.

49:02

You put your cluster in West  Texas.

49:02

You put it in southwest Pennsylvania by the Marcellus Shale.

49:06

The 10 GW cluster is super  easy.

49:06

The 100 GW cluster is also pretty doable.

49:13

I think natural gas production in the United  States has almost doubled in a decade.

49:13

You do that one more time over the next seven years, you  could power multiple trillion-dollar data centers.

49:25

The issue there is that a lot of people made these  climate commitments, not just the government.

49:25

It's actually the private companies themselves,  Microsoft, Amazon, etc.

49:29

, that have made these climate commitments.

49:33

So they won't do  natural gas.

49:33

I admire the climate commitments, but at some point the national interest  and national security is more important.

49:43

The other path is doing green energy  megaprojects.

49:43

You do solar and batteries and SMRs and geothermal.

49:48

If we want to do that,  there needs to be a broad deregulatory push.

49:57

You can't have permitting take a decade.

49:57

You  have to reform FERC.

49:57

You have to have blanket NEPA exemptions for this stuff.

50:02

There are inane state-level regulations.

50:02

You can build the solar panels and batteries next to your  data center, but it'll still take years because you actually have to hook it up to the state  electrical grid.

50:11

You have to use governmental powers to create rights of way to have multiple  clusters and connect them and have the cables. Ideally we do both.

50:25

Ideally we do natural gas  and the broader deregulatory green agenda.

50:28

We have to do at least one.

50:28

Then this  stuff is possible in the United States.

50:36

Before the conversation I was reading a good book  about World War II industrial mobilization in the United States called Freedom's Forge.

50:41

I’m thinking  back on that period, especially in the context of reading Patrick Collison’s Fast and the progress  study stuff.

50:49

There’s this narrative out there that we had state capacity back then and people just  got shit done but that now it's a clusterfuck.

50:59

It wasn’t at all the case!

50:59

It was really interesting.

50:59

You had people from the Detroit auto industry side, like  William Knudsen, who were running mobilization for the United States.

51:08

They were extremely competent.

51:08

At the same time you had labor organization and agitation, which is very analogous to the climate  change pledges and concerns we have today.

51:21

They would literally have these strikes,  into 1941, costing millions of man-hours worth of time when we're trying to make tens  of thousands of planes a month.

51:28

They would just debilitate factories for trivial concessions  from capital that were pennies on the dollar.

51:43

There were concerns that the auto companies  were trying to use the pretext of a potential war to prevent paying labor the money it  deserves.

51:50

So with what climate change is today, you might think, "ah, America's fucked.

51:58

We're  not going to be able to build this shit if you look at NEPA or something,” I didn't realize  how debilitating labor was in World War II. It wasn’ just that.

52:06

Before 1939, the American  military was in total shambles.

52:06

You read about it and it reads a little bit like the German  military today.

52:12

Military expenditures were I think less than 2% of GDP.

52:17

All the European countries  had gone, even in peacetime, above 10% of GDP.

52:22

It was rapid mobilization starting  from nothing.

52:22

We were making no planes.

52:26

There were no military contracts.

52:26

Everything  had been starved during the Great Depression.

52:30

But there was this latent capacity.

52:30

At some  point the United States got its act together.

52:35

This applies the other way around too with China.

52:35

Sometimes people count them out a little bit with the export controls and so on.

52:43

They're able to  make 7-nanometer chips now.

52:43

There's a question of how many they could make.

52:48

There's at least  a possibility that they're going to mature that ability and make a lot of 7-nanometer chips.

52:52

There's a lot of latent industrial capacity in China.

52:57

They are able to build a lot of power  fast.

52:57

Maybe that isn't activated for AI yet.

52:57

At some point, the same way the United States and  a lot of people in the US government are going to wake up, the CCP is going to wake up.

53:08

Companies realize that scaling is a thing.

53:23

Obviously their whole plans are contingent on  scaling.

53:23

So they understand that in 2028 we're going to be building 10 GW data centers.

53:27

At that point, the people who can keep up are Big Tech, potentially at the edge of their  capabilities, sovereign wealth fund-funded things, and also major countries like America and China. What's their plan?

53:40

With the AI labs, what's their plan given this landscape?

53:50

Do they not want  the leverage of being in the United States?

53:58

The Middle East does offer capital, but America  has plenty of capital.

53:58

We have trillion-dollar companies.

54:05

What are these Middle Eastern  states?

54:05

They're kind of like trillion-dollar oil companies.

54:07

We have trillion-dollar companies  and very deep financial markets.

54:07

Microsoft could issue hundreds of billions of dollars of  bonds and they can pay for these clusters.

54:16

Another argument being made, which is worth  taking seriously, is that if we don't work with the UAE or with these Middle Eastern  countries, they're just going to go to China.

54:28

They're going to build data centers and  pour money into AI regardless.

54:28

If we don't work with them, they'll just support China.

54:31

There's some merit to the argument in the sense that we should be doing benefit-sharing  with them.

54:40

On the road to AGI, there should be two tiers of coalitions.

54:46

There should be a narrow  coalition of democracies that's developing AGI.

54:52

Then there should be a broader coalition of other  countries, including dictatorships, and we should offer them some of the benefits of AI.

55:00

If the UAE wants to use AI products, run Meta recommendation engines,  or run the last-generation models, that's fine.

55:10

By default, they just wouldn't  have had this seat at the AGI table.

55:10

So they have some money, but a lot of people have money.

55:15

The only reason they're getting this seat at the AGI table and giving these dictators this leverage  over this extremely important national security technology, is because we're getting  them excited and offering it to them.

55:39

Who specifically is doing this?

55:39

Who are the  companies who are going there to fundraise?

55:44

It’s been reported that Sam Altman is trying  to raise $7 trillion or whatever for a chip project.

55:49

It's unclear how many of the clusters  will be there, but definitely stuff is happening.

55:55

There’s another reason I'm a little suspicious  of this argument that if the US doesn't work with them, they'll go to China.

56:00

I've heard from  multiple people — not from my time at OpenAI, and I haven't seen the memo — that at some  point several years ago, OpenAI leadership had laid out a plan to fund and sell AGI by  starting a bidding war between the governments of the United States, China, and Russia.

56:18

It's surprising to me that they're willing to sell AGI to the Chinese and Russian governments.

56:24

There's also something that feels eerily familiar about starting this bidding war and then  playing them off each other, saying, "well, if you don't do this, China will do it." Interesting. That's pretty fucked up. Suppose you're right.

56:41

We ended up in this place  because, as one of our friends put it, the Middle East has billions or trillions of dollars up  for persuasion like no other place in the world.

56:53

With little accountability.

56:53

There’s no  Microsoft board. It's only the dictator.

57:04

Let's say you're right, that you shouldn't  have gotten them excited about AGI in the first place.

57:06

Now we're in a place where they  are excited about AGI and they're like, "fuck, we want to have GPT-5 while you're going to be off  building superintelligence.

57:12

This Atoms for Peace thing doesn't work for us."

57:16

If you're in this  place, don't they already have the leverage?

57:25

The UAE on its own is not competitive.

57:25

They're  already export-controlled.

57:25

You're not supposed to ship Nvidia chips over there.

57:31

It's  not like they have any of the leading AI labs.

57:34

They have money, but it's hard  to just translate money into progress.

57:38

But I want to go back to other things you've  been saying in laying out your vision.

57:38

There's this almost industrial process of putting in  the compute and algorithms, adding that up, and getting AGI on the other end.

57:48

If it's  something more like that, then the case for somebody being able to catch up rapidly seems  more compelling than if it's some bespoke...

57:58

Well, if they can steal the algorithms and if they  can steal the weights, that’s really important.

58:09

How easy would it be for an actor to steal  the things that are not the trivial released things, like Scarlett Johansson's voice, but the  RL things we're talking about, the unhobblings? It’s all extremely easy.

58:20

They don’t make the claim  that it’s hard.

58:20

DeepMind put out their Frontier Safety Framework and they lay out security levels,  zero to four.

58:28

Four is resistant to state activity.

58:34

They say, we're at level zero.

58:34

Just recently,  there was an indictment of a guy who stole a bunch of really important AI code and went to China  with it.

58:41

All he had to do to steal the code was copy it, put it into Apple Notes, and export  it as a PDF.

58:46

That got past their monitoring.

58:51

Google has the best security of any of the AI  labs probably, because they have the Google infrastructure.

58:55

I would think of the security of  a startup.

58:55

What does security of a startup look like? It's not that good. It's easy to steal.

59:02

Even if that's the case, a lot of your post is making the argument for why we are going  to get the intelligence explosion.

59:09

If we have somebody with the intuition of an Alec Radford to  come up with all these ideas, that intuition is extremely valuable and you can scale that up.

59:19

If it's just intuition, then that's not going to be just in the code, right?

59:32

Also because  of export controls, these countries are going to have slightly different hardware.

59:36

You're going  to have to make different trade-offs and probably rewrite things to be compatible with that.

59:41

Is it just a matter of getting the right pen drive and plugging it into the gigawatt  data center next to the Three Gorges Dam and then you're off to the races?

59:50

There are a few different things, right?

59:54

One threat model is just them stealing  the weights themselves.

59:54

The weights one is particularly insane because they can just steal  the literal end product — just make a replica of the atomic bomb — and then they're ready to go.

1:00:03

That one is extremely important around the time we have AGI and superintelligence because China  can build a big cluster by default.

1:00:09

We'd have a big lead because we have the better scientists,  but if we make the superintelligence and they just steal it, they're off to the races.

1:00:19

Weights are a little bit less important right now because who cares if they steal the GPT-4 weights.

1:00:23

We still have to get started on weight security now because if we think there’s AGI by 2027, this  stuff is going to take a while.

1:00:30

It's not just going to be like, "oh, we do some access control."

1:00:35

If you actually want to be resistant to Chinese espionage, it needs to be much more intense.

1:00:39

The thing that people aren't paying enough attention to is the secrets.

1:00:45

The compute stuff is  sexy, but people underrate the secrets.

1:00:45

The half an order of magnitude a year is just by default,  sort of algorithmic progress. That's huge.

1:00:58

If we have a few years of lead, by default, that's a  10-30x, 100x bigger cluster, if we protect them.

1:01:09

There's this additional layer of the data wall.

1:01:09

We have to get through the data wall.

1:01:09

That means we actually have to figure out some sort of  basic new paradigm.

1:01:12

So it’s the “AlphaGo step two.

1:01:15

” “AlphaGo step one” learns from human  imitation.

1:01:15

“AlphaGo step two” is the kind of self-play RL thing that everyone's working  on right now.

1:01:20

Maybe we're going to crack it.

1:01:27

If China can't steal that, then they're stuck.

1:01:27

If they can steal it, they're off to the races.

1:01:33

Whatever that thing is, can I literally write it  down on the back of a napkin?

1:01:33

If it's that easy, then why is it so hard for them to figure  it out?

1:01:39

If it's more about the intuitions, then don't you just have to hire Alec  Radford?

1:01:42

What are you copying down?

1:01:45

There are a few layers to this.

1:01:45

At the top is the  fundamental approach.

1:01:45

On pre-training it might be unsupervised learning, next token prediction,  training on the entire Internet.

1:01:56

You actually get a lot of juice out of that already.

1:02:00

That one's very quick to communicate.

1:02:04

Then there's a lot of details that matter,  and you were talking about this earlier.

1:02:07

It's probably going to be somewhat obvious in  retrospect, or there's going to be some not too complicated thing that'll work, but there's  going to be a lot of details to get that.

1:02:17

If that's true, then again, why do we  think that getting state-level security in these startups will prevent China  from catching up?

1:02:23

It’s just like, "oh, we know some sort of self-play RL will  be required to get past the data wall."

1:02:33

It's going to be solved by  2027, right? It's not that hard.

1:02:36

The US, and the leading labs in the United States,  have this huge lead.

1:02:36

By default, China actually has some good LLMs because they're just using open  source code, like Llama.

1:02:42

People really underrate both the divergence on algorithmic progress and  the lead the US would have by default because all this stuff was published until recently.

1:02:55

Look at Chinchilla Scaling laws, MoE papers, transformers.

1:03:00

All that stuff was published.

1:03:00

That's why open source is good and why China can make some good models.

1:03:05

Now, they're  not publishing it anymore.

1:03:05

If we actually kept it secret, it would be a huge edge.

1:03:10

To your point about tacit knowledge and Alec Radford, there's another layer at the  bottom that is something about large-scale engineering work to make these big training  runs work.

1:03:19

That is a little bit more like tacit knowledge, but China will be able to figure  that out.

1:03:24

It's engineering schlep, and they're going to figure out how to do it.

1:03:27

Why can't they figure that out, but not how to get the RL thing working? I don't know.

1:03:29

Germany during World War II went down the wrong path with heavy  water.

1:03:36

There's an amazing anecdote in The Making of the Atomic Bomb about this.

1:03:41

Secrecy was one of the most contentious issues early on.

1:03:46

Leo Szilard really thought a nuclear  chain reaction and an atomic bomb were possible.

1:03:56

He went around saying, "this is going to be of  enormous strategic and military importance."

1:04:01

A lot of people didn't believe it or thought,  "maybe this is possible, but I'm going to act as though it's not, and science should be open."

1:04:05

In the early days, there had been some incorrect measurements made on graphite as a moderator.

1:04:13

Germany thought graphite wasn't going to work, so they had to do heavy water.

1:04:20

But then Enrico  Fermi made new measurements indicating that graphite would work.

1:04:27

This was really important.

1:04:27

Szilard assaulted Fermi with another secrecy appeal and Fermi was pissed off, throwing a  temper tantrum.

1:04:33

He thought it was absurd, saying, "come on, this is crazy."

1:04:39

But Szilard persisted,  and they roped in another guy, George Pegram.

1:04:45

In the end, Fermi didn't publish it. That was just in time.

1:04:45

Fermi not publishing meant that the Nazis didn't figure out graphite  would work.

1:04:51

They went down the path of heavy water, which was the wrong path.

1:04:55

This is a key reason why the German project didn't work out. They were way behind.

1:04:59

We face a similar situation now.

1:04:59

Are we just going to instantly leak how to get past the data  wall and what the next paradigm is? Or are we not?

1:05:13

The reason this would matter is if being  one year ahead would be a huge advantage.

1:05:18

In the world where you deploy AI over time  they're just going to catch up anyway.

1:05:23

I interviewed Richard  Rhodes, the guy who wrote The Making of the Atomic Bomb.

1:05:29

One of the anecdotes  he had was when the Soviets realized America had the bomb.

1:05:33

Obviously, we dropped it in Japan.

1:05:33

Lavrentiy Beria — the guy who ran the NKVD, a famously ruthless and evil guy — goes to the  Soviet scientist who was running their version of the Manhattan Project.

1:05:46

He says, "comrade, you  will get us the American bomb."

1:05:46

The guy says, "well, listen, their implosion device actually is  not optimal.

1:05:53

We should make it a different way."

1:05:57

Beria says, "no, you will get us the American  bomb, or your family will be camp dust."

1:06:03

The thing that's relevant about that anecdote is  that the Soviets would have had a better bomb if they hadn't copied the American design, at least  initially.

1:06:09

That suggests something about history, not just for the Manhattan Project.

1:06:15

There's  often this pattern of parallel invention because the tech tree implies that a certain  thing is next — in this case, a self-play RL — and people work on that and are going to  figure it out around the same time.

1:06:26

There's not going to be that much gap in who gets it first.

1:06:32

Famously, a bunch of people invented the light bulb around the same time.

1:06:37

Is it the case  that it might be true but the one year or six months makes the difference?

1:06:43

Two years makes all the difference.

1:06:45

I don't know if it'll be two years though.

1:06:45

If we lock down the labs, we have much better scientists. We're way ahead. It would  be two years.

1:06:50

Even six months, a year, would make a huge difference.

1:06:56

This gets back  to the intelligence explosion dynamics.

1:06:56

A year might be the difference between a system that's  sort of human-level and a system that is vastly superhuman.

1:07:05

It might be like five OOMs.

1:07:05

Look at the current pace.

1:07:05

Three years ago, on the math benchmark — these are really  difficult high school competition math problems — we were at a few percent, we couldn't  solve anything. Now it's solved.

1:07:21

That was at the normal pace of AI progress.

1:07:26

You didn't  have a billion superintelligent researchers.

1:07:30

A year is a huge difference, particularly after  superintelligence.

1:07:30

Once this is applied to many elements of R&D, you get an industrial  explosion with robots and other advanced technologies.

1:07:40

A couple of years might  yield decades worth of progress.

1:07:40

Again, it’s like the technological lead  the U. S.

1:07:44

had in the first Gulf War, when the 20-30 years of technological lead  proved totally decisive. It really matters.

1:07:51

Here’s another reason it really matters.

1:07:51

Suppose  they steal the weights, suppose they steal the algorithms, and they're close on our tails.

1:07:56

Suppose we still pull out ahead.

1:07:56

We're a little bit faster and we're three months ahead.

1:08:01

The world in which we're really neck and neck, we only have a three-month lead, is incredibly  dangerous.

1:08:07

We're in this feverish struggle where if they get ahead, they get to dominate, maybe  they get a decisive advantage.

1:08:12

They're building clusters like crazy.

1:08:19

They're willing to throw  all caution to the wind. We have to keep up.

1:08:24

There are crazy new WMDs popping up.

1:08:24

Then we're  going to be in the situation where it's crazy new military technology, crazy new WMDs, deterrence,  mutually assured destruction keeps changing every few weeks.

1:08:34

It's a completely unstable,  volatile situation that is incredibly dangerous.

1:08:39

So you have to look at it from the point of  view that these technologies are dangerous, from the alignment point of view.

1:08:42

It might be  really important during the intelligence explosion to have a six-month wiggle room to be like, “look,  we're going to dedicate more compute to alignment during this period because we have to get it  right.

1:08:51

We're feeling uneasy about how it's going.

1:08:51

” One of the most important inputs to  whether we will destroy ourselves or whether we will get through this incredibly  crazy period is whether we have that buffer.

1:09:09

Before we go further, it's very much worth noting  that almost nobody I talk to thinks about the geopolitical implications of AI.

1:09:19

I have some  object-level disagreements that we'll get into, things I want to iron out.

1:09:26

I  may not disagree in the end.

1:09:30

The basic premise is that if you keep scaling, if  people realize that this is where intelligence is headed, it's not just going to be the same old  world.

1:09:37

It won't just be about what model we're deploying tomorrow or what the latest thing  is.

1:09:42

People on Twitter are like, "oh, GPT-4 is going to shake your expectations" or whatever.

1:09:48

COVID is really interesting because when March 2020 hit, it became clear to the world  — presidents, CEOs, media, the average person — that there are other things happening  in the world right now but the main thing we as a world are dealing with right now is COVID. Soon it will be AGI.

1:10:07

This is the quiet period.

1:10:14

Maybe you want to go on vacation.

1:10:14

Maybe now is the  last time you can have some kids.

1:10:14

My girlfriend sometimes complains when I’m off doing work that  I don’t spend enough time with her.

1:10:24

She threatens to replace me with GPT-6 or whatever.

1:10:32

I'm like,  “GPT-6 will also be too busy doing AI research.

1:10:32

” Why aren't other people talking  about national security?

1:10:45

I made this mistake with COVID.

1:10:45

In February  of 2020, I thought it was going to sweep the world and all the hospitals would  collapse.

1:10:51

It would be crazy, and then it'd be over.

1:10:55

A lot of people thought this kind  of thing at the beginning of COVID.

1:10:55

They shut down their office for a month or whatever.

1:10:58

The thing I just really didn't price in was societal reaction.

1:11:02

Within weeks, Congress  spent over 10% of GDP on COVID measures.

1:11:10

The entire country was shut down. It was crazy.

1:11:10

I didn't sufficiently price it in with COVID.

1:11:18

Why do people underrate it?

1:11:18

Being in the  trenches actually gives you a less clear picture of the trend lines.

1:11:27

You don’t have  to zoom out that much, only a few years.

1:11:32

When you're in the trenches, you're trying  to get the next model to work.

1:11:32

There's always something that's hard.

1:11:35

You might underrate  algorithmic progress because you're like, "ah, things are hard right now," or "data wall"