Joe Carlsmith — Preventing an AI takeover

0:37

Today I'm chatting with Joe Carlsmith.

0:37

He's a philosopher and, in my opinion, a capital-G great philosopher.

0:41

You can  find his essays at joecarlsmith. com.

0:46

So we have GPT-4, and it doesn't seem like a  paperclipper thing.

0:46

It understands human values.

0:54

In fact, you can have it explain why being a  paperclipper is bad or ask it to explain why the galaxy shouldn't be turned into paperclips.

1:02

What has to happen such that eventually we have a system that takes over and converts  the world into something valueless?

1:17

When I'm thinking about misaligned AIs—or the  type that I'm worried about—I'm thinking about AIs with a relatively specific set of  properties related to agency, planning, awareness, and understanding of the world.

1:30

One key aspect is the capacity to plan and make relatively sophisticated plans based  on models of the world, where those plans are evaluated according to criteria.

1:41

That  planning capability needs to be driving the model's behavior.

1:46

There are models that are,  in some sense, capable of planning.

1:46

But when they give output, it's not like that output  was determined by some process of planning, like, “Here's what will happen if I give  this output, and do I want that to happen?"

1:57

The model needs to really understand the  world.

1:57

It needs to really be like, “Okay, here’s what will happen. Here I am.

2:02

Here’s  the politics of the situation.

2:02

” It needs to have this kind of situational awareness to  evaluate the consequences of different plans.

2:15

Another thing to consider is the verbal behavior  of these models.

2:15

When I talk about a model's values, I'm referring to the criteria that end  up determining which plans the model pursues.

2:25

A model's verbal behavior—even if it has a planning  process (which GPT-4, I think, doesn't in many cases)—doesn't necessarily reflect those criteria.

2:38

We know that we're going to be able to get models to say what we want to hear.

2:53

That's the magic of  gradient descent.

2:53

Modulo some difficulties with capabilities, you can get a model to output  the behavior that you want.

3:02

If it doesn't, then you crank it until it does.

3:06

I think everyone admits that for suitably sophisticated models, they're going  to have a very detailed understanding of human morality.

3:15

The question is, what relationship is  there between a model's verbal behavior—which you've essentially clamped, you're forcing the  model to say certain things— and the criteria that end up influencing its choice between plans?

3:32

I'm pretty cautious about assuming that when it says the thing I forced it to say—or when gradient  descent has shaped it to say a certain thing— that that is a lot of evidence about how it's going  to choose in a bunch of different scenarios.

3:54

Even with humans, it's not necessarily the  case that their verbal behavior reflects the actual factors that determine their  choices. They can lie.

4:00

They might not even know what they would do in a given  situation, all sorts of stuff like that.

4:07

It's interesting to think about this in the  context of humans.

4:07

There's that famous saying: "Be careful who you pretend to be, because  you are who you pretend to be."

4:11

You notice this with how culture shapes children.

4:15

Parents  will punish you if you start saying things that are inconsistent with your culture's values, and  over time, you become like your parents, right?

4:29

By default, it seems like it kind of works.

4:29

Even  with these models, it seems to work.

4:29

They don’t really scheme against us. Why would this happen?

4:35

For folks who are unfamiliar with the basic story, they might wonder, "Why would AI take over at all?

4:42

What's the reason they would do that?"

4:42

The general concern is that you're offering someone power,  especially if you're offering it for free.

4:47

Power, almost by definition, is useful for lots  of values.

4:54

We're talking about an AI that really has the opportunity to take control  of things.

4:59

Say some component of its values is focused on some outcome, like the world being  a certain way, especially in a longer-term way such that its concern extends beyond the  period that a takeover plan would encompass.

5:20

The thought is that it's often the case that  the world will be more the way you want it if you control everything, rather than if you remain  an instrument of human will or some other actor, which is what we're hoping these AIs will be.

5:36

That's a very specific scenario.

5:36

If we're in a scenario where power is more  distributed—especially where we're doing decently on alignment and we're giving the AI some  amount of inhibition about doing different things, maybe we're succeeding in shaping their values  somewhat—then it's just a much more complicated calculus.

5:52

You have to ask, “What's the upside for  the AI?

5:52

What's the probability of success for this takeover path?

5:59

How good is its alternative?

5:59

” Maybe this is a good point to talk about how you expect the difficulties of alignment to  change in the future.

6:05

We're starting off with something that has this intricate representation  of human values and it doesn't seem that hard to sort of lock it into a persona that we are  comfortable with.

6:14

I don't know what changes.

6:22

Why is alignment hard in general?

6:22

Let’s say  we've got an AI.

6:22

Let's bracket the question of exactly how capable it will be and talk about  this extreme scenario where it really has the opportunity to take over.

6:34

I think we might just  want to avoid having to build an AI that we're comfortable with being in that position.

6:41

But  let's focus on it for simplicity's sake, and then we can relax the assumption.

6:45

One issue is that you can't just test it.

6:45

You can't give the AI this literal situation, have  it take over and kill everyone, and then say, "Oops, update the weights."

6:59

This is what Eliezer  talks about.

6:59

You care about its behavior in this specific scenario that you can’t test directly.

7:07

We can talk about whether that's a problem, but that's one issue.

7:12

There's a sense in which  this has to be "off-distribution."

7:12

You have to get some kind of generalization from training the  AI on a bunch of other scenarios.

7:18

Then there's the question of how it's going to generalize to  the scenario where it really has this option. Is that even true?

7:28

Because when you're training  it, you can say, "Hey, here's a gradient update.

7:34

If you get the takeover option on a platter, don't  take it."

7:34

And then, in red teaming situations where it thinks it has a takeover attempt, you  train it not to take it.

7:41

It could fail, but I feel like if you did this to a child, like "Don't beat  up your siblings," the kid will generalize to, "If I'm an adult and I have a rifle, I'm  not going to start shooting random people."

8:02

You mentioned the idea of, "You are what you  pretend to be."

8:02

Will these AIs, if you train them to look nice, fake it till they make it?

8:11

You were saying we do this to kids.

8:11

I think it's better to imagine kids doing this to us.

8:18

Here's a silly analogy for AI training.

8:18

Suppose you wake up and you're being trained via methods  analogous to contemporary machine learning by Nazi children to be a good Nazi soldier or butler or  what have you.

8:42

These children have a model spec, a nice Nazi model spec.

8:58

It’s like, “Reflect  well on the Nazi Party, benefit the Nazi Party” and whatever. You can read it. You understand  it.

9:05

This is why I'm saying that when you’re like, “The model really understands human values…” In this analogy, I start off as something more intelligent than the things training  me, with different values to begin with.

9:24

The intelligence and the values are baked in to  begin with.

9:24

Whereas a more analogous scenario is, “I'm a toddler and, initially, I'm stupider  than the children.

9:29

” This would also be true, by the way, if I'm a much smarter model initially.

9:36

The much smarter model is dumb, right?

9:36

Then I get smarter as you train me.

9:40

So it's like a toddler,  and the kids are like, "Hey, we're going to bully you if you're not a Nazi."

9:45

As you grow up, you  reach the children's level, and then eventually you become an adult.

9:50

Through that process, they've  been bullying you, training you to be a Nazi.

9:59

I think in that scenario, I might end up a Nazi.

9:59

Basically a decent portion of the hope here should be that we're never in the situation where  the AI really has very different values, is already quite smart and really knows what's  going on, and is now in this kind of adversarial relationship with our training process. We want  to avoid that.

10:17

I think it's possible we can, by the sorts of things you're saying.

10:23

So I'm not saying that'll never work.

10:27

The thing I just wanted to highlight was about  if you get into that situation where the AI is genuinely at that point much, much more  sophisticated than you, and doesn't want to reveal its true values for whatever reason.

10:38

Then when the children show some obviously fake opportunity to defect to the allies, it's  not necessarily going to be a good test of what it will do in the real circumstance  because it's able to tell the difference.

10:56

You can also give another way in which the analogy  might be misleading.

10:56

Imagine that you're not just in a normal prison where you're totally cognizant  of everything that's going on.

11:04

Sometimes they drug you, give you weird hallucinogens that totally  mess up how your brain is working.

11:09

As a human adult in a prison, I know what kind of thing I  am.

11:15

Nobody's really fucking with me in a big way.

11:25

Whereas an AI, even a much smarter  AI in a training situation, is much closer to being constantly inundated with  weird drugs and different training protocols.

11:37

You're frazzled because each moment is closer  to some sort of Chinese water torture technique I'm glad we're talking about the  moral patienthood stuff later.

11:48

There’s this chance to step back and ask,  "What's going on?"

11:48

An adult in prison has that ability in a way that I don't know if these models  necessarily have.

11:54

It’s that coherence and ability to step back from what's  happening in the training process. Yeah, I don't know.

12:04

I'm hesitant to say it's  like drugs for the model.

12:04

Broadly speaking, I do basically agree that we have quite a  lot of tools and options for training AIs, even AIs that are somewhat smarter than humans.

12:20

I do think you have to actually do it. You had Eliezer on.

12:28

I'm much more bullish on our ability  to solve this problem, especially for AIs that are in what I think of as the "AI for AI safety sweet  spot."

12:35

This is a band of capability where they're sufficiently capable that they can be really  useful for strengthening various factors in our civilization that can make us safe.

12:49

That’s stuff  like our alignment work, control, cybersecurity, general epistemics, maybe some coordination  applications.

12:54

There's a bunch of stuff you can do with AIs that, in principle, could differentially  accelerate our security with respect to the sorts of considerations we're talking about.

13:04

Let’s say you have AIs that are capable of that.

13:08

You can successfully elicit that  capability in a way that's not being sabotaged or messing with you in other ways.

13:13

They  can't yet take over the world or do some other really problematic form of power-seeking.

13:19

If  we were really committed, we could then go hard, put a ton of resources and really differentially  direct this glut of AI productivity towards these security factors.

13:31

We could hopefully control  and understand, do a lot of these things you're talking about to make sure our AIs don't  take over or mess with us in the meantime.

13:44

We have a lot of tools there.

13:44

You have to really  try though.

13:44

It's possible that those sorts of measures just don't happen, or they don't  happen at the level of commitment, diligence, and seriousness that you would need.

13:55

That’s  especially true if things are moving really fast and there are other competitive pressures:  “This is going to take compute to do these intensive experiments on the AIs.

14:03

We could  use that compute for experiments for the next scaling step.

14:08

” There’s stuff like that.

14:08

I'm not here saying this is impossible, especially for that band of AIs.

14:14

It's  just that you have to try really hard.

14:20

I agree with the sentiment of obviously  approaching this situation with caution, but I do want to point out the ways in which  the analyses we've been using have been maximally adversarial.

14:30

For example, let’s  go back to the adult getting trained by Nazi children.

14:39

Maybe the one thing I didn't  mention is the difference in this situation, which is maybe what we're trying  to get at with the drug metaphor.

14:48

When you get an update, it's much more directly  connected to your brain than a sort of reward or punishment a human gets.

14:54

It's literally a  gradient update down to the parameter of how much this would contribute to you putting this output  rather than that output.

15:01

Each different parameter we're going to adjust to the exact floating point  number that calibrates it to the output we want.

15:12

I just want to point out that we're coming into  the situation pretty well.

15:12

It does make sense, of course, if you're talking to somebody at a lab, to  say, "Hey, really be careful."

15:17

But for a general audience, should I be scared witless?

15:21

You maybe  should to the extent that you should be scared about things that do have a chance of happening.

15:26

For example, you should be scared about nuclear war.

15:30

But should you be scared in the sense  of you’re doomed?

15:30

No, you're coming up with an incredible amount of leverage on the AIs in  terms of how they will interact with the world, how they're trained, and the  default values they start with.

15:44

I think it is the case that by the  time we're building superintelligence, we'll have much better… Even right now—when  you look at labs talking about how they're planning to align AIs—no one is saying  we're going to just do RLHF.

15:54

At the least, you're talking about scalable oversight.

15:59

You  have some hope about interpretability.

15:59

You have automated red teaming.

16:03

Hopefully, humans are doing  a bunch more alignment work.

16:03

I also personally am hopeful that we can successfully elicit from  various AIs a ton of alignment work progress.

16:17

There's a bunch of ways this can go.

16:17

I'm not  here to tell you 90% doom or anything like that.

16:23

This is the basic reason for concern.

16:23

Imagine  that we're going to transition to a world in which we've created these beings that are just  vastly more powerful than us.

16:35

We've reached the point where our continued empowerment is just  effectively dependent on their motives.

16:42

It is this vulnerability to, “What do the AIs choose  to do?

16:50

” Do they choose to continue to empower us or do they choose to do something else?

16:56

Or it’s about the institutions that have been set up.

17:00

I expect the US government to  protect me, not because of its “motives,” but just because of the system of incentives and  institutions and norms that has been set up.

17:13

You can hope that will work too, but there is  a concern.

17:13

I sometimes think about AI takeover scenarios via this spectrum of how much power  we voluntarily transferred to the AIs.

17:20

How much of our civilization did we hand to the AIs  intentionally by the time they took over?

17:35

Versus, how much did they take for themselves?

17:35

Some of the scariest scenarios are where we have a really fast explosion to the point where there  wasn't even a lot of integration of AI systems into the broader economy.

17:49

But there's this really  intensive amount of superintelligence concentrated in a single project or something like that.

17:56

That's a quite scary scenario, partly because of the speed and people not having time to react.

18:04

Then there are intermediate scenarios where some things got automated, maybe people  handed the military over to the AIs or we have automated science.

18:12

There are some  rollouts and that’s giving the AIs power that they don't have to take.

18:18

We're doing all our  cybersecurity with AIs and stuff like that.

18:23

Then there are worlds where you more fully  transitioned to a kind of world run by AIs where, in some sense, humans voluntarily did that.

18:34

Maybe there were competitive pressures, but you intentionally handed off huge portions of  your civilization.

19:58

At that point, it's likely that humans have a hard time understanding what's  going on.

20:06

A lot of stuff is happening very fast.

20:09

The police are automated.

20:09

The courts  are automated.

20:09

There's all sorts of stuff.

20:14

Now, I tend to think a little less about  those scenarios because I think they're correlated with being further down the line.

20:20

Humans are hopefully not going to just say, "Oh yeah, you built an AI system, let's just..."

20:26

When we look at technological adoption rates, it can go quite slow.

20:34

Obviously there's  going to be competitive pressures, but in general this category is somewhat safer.

20:37

But even in this one, I think it's intense.

20:37

If humans have really lost their epistemic grip  on the world, they've handed off the world to these systems.

20:53

Even if you're like, "Oh, there's  laws, there's norms…" I really want us to have a really developed understanding of what's likely to  happen in that circumstance, before we go for it.

21:07

I get that we want to be worried about a scenario  where it goes wrong.

21:07

But again, what is the reason to think it might go wrong?

21:12

In the human example,  your kids are not maximally adversarial against your attempts to instill your culture on them.

21:19

With these models, at least so far, it doesn't seem to matter.

21:24

They just get, "Hey, don't help  people make bombs" or whatever, even if you ask in a different way how to make a bomb.

21:30

We're also  getting better and better at this all the time.

21:33

You're right in picking up on this assumption  in the AI risk discourse of what we might call intense adversariality between agents that  have somewhat different values.

21:43

There's some sort of thought—and I think this is rooted  in the discourse about the fragility of value and stuff like that—that if these agents  are somewhat different, at least in the specific scenario of an AI takeoff, they end  up in this intensely adversarial relationship.

22:05

You're right to notice that's not how we are  in the human world.

22:05

We're very comfortable with a lot of different differences in values.

22:10

A  factor that is relevant is this notion that there are possibilities for intense concentration  of power on the table.

22:18

There is some kind of general concern, both with humans and AIs.

22:27

If it's the case that there's some ring of power that someone can just grab that will give  them huge amounts of power over everyone else, suddenly you might be more worried about  differences in values at stake, because you're more worried about those other actors.

22:46

We talked about this Nazi example where you imagine that you wake up and you're being trained  by Nazis to become a Nazi. You're not right now.

22:59

Is it plausible that we'd end up with a model  that is in that sort of situation?

22:59

As you said, maybe it's trained as a kid.

23:05

It never ends  up with values such that it's aware of some significant divergence between its values  and the values that the humans intend for it to have.

23:16

If it's in that scenario, would  it want to avoid having its values modified?

23:26

At least to me, it seems fairly plausible that  the AI's values meet certain constraints.

23:26

Do they care about consequences in the world?

23:36

Do they  anticipate that the AI's preserving its values will better conduce to those consequences?

23:42

Then  it's not that surprising if it prefers not to have its values modified by the training process.

23:49

There’s a way in which I'm still confused about this.

23:55

With the non-Nazi being trained by Nazis,  it's not just that I have different values.

23:55

I actively despise their values.

24:02

I don't expect this  to be true of AIs with respect to their trainers.

24:08

The more analogous scenario is where I'm like,  “Am I leery of my values being changed?

24:08

” Going to college or meeting new people or reading a new  book, I'm like, “I don't know.

24:13

It's okay if it changes my values. That's fine. I don't care.

24:18

” Yeah, that's a reasonable point. There's a question.

24:23

How would you feel about paperclips?

24:23

Maybe you don't despise paperclips, but there's the human paperclippers there and they're  training you to make paperclips.

24:28

My sense would be that there's a kind of relatively specific  set of conditions in which you're comfortable having your values changed, especially not  by learning and growing but by gradient descent directly intervening on your neurons.

24:44

This seems similar to religious training as a kid.

24:57

You start off in a religion.

24:57

Because you started  off in a religion, you're already sympathetic to the idea that you go to church every week so  that you're more reinforced in this existing tradition.

25:05

You're getting more intelligent over  time.

25:05

When you're a kid, you're getting very simple instructions about how religion works.

25:08

As you get older, you get more and more complex theology that helps you talk to other adults about  why this is a rational religion to believe in.

25:18

But since one of your values to begin with  was that you want to be trained further in this religion, you want to come back to church  every week.

25:22

That seems more analogous to the situation the AIs will be in with respect to  human values.

25:26

The entire time they're like, "Hey, be helpful, be harmless,” etc.

25:31

Yes, it could be like that.

25:31

There's a scenario in which you were comfortable with your  values being changed because, in some sense, you have sufficient allegiance to the output of  that process.

25:41

In a religious context, you're like, "Ah, make me more virtuous by the lights of this  religion."

25:49

You go to confession and you're like, "I've been thinking about takeover today.

25:59

Can  you change me please?

25:59

Give me more gradient descent. I've been so bad."

26:04

People sometimes  use the term "corrigibility" to talk about that.

26:11

Maybe the AI doesn't have perfect values,  but it's in some sense cooperating with your efforts to change its values to be a certain way.

26:15

Maybe it's worth saying a little bit here about what actual values the AI might have.

26:21

Would it  be the case that the AI naturally has the sort of equivalent of, "I'm sufficiently devoted to  human obedience that I'm going to really want to be modified so I'm a better instrument of  the human will" versus wanting to go off and do its own thing?

26:40

It could be benign and go well.

26:40

Here are some possibilities where it could go bad.

26:51

I'm generally concerned about how little  science we have of model motivations.

26:51

We just don't have a great understanding of  what happens in this scenario.

26:58

Hopefully, we'd get one before we reach this scenario.

27:01

Here are five categories of motivations the model could have.

27:09

This hopefully gets at the  point about what the model eventually does.

27:14

One category is something just super alien.

27:14

There's some weird correlate of easy-to-predict text or some weird aesthetic for data structures  that the model developed early on in pre-training or later.

27:28

It really thinks things should be  like this.

27:28

There's something quite alien to our cognition where we just wouldn't recognize  it as a thing at all. That’s one category.

27:39

Another category is a kind of crystallized  instrumental drive that is more recognizable to us.

27:46

You can imagine AIs developing some  curiosity drive because that's broadly useful.

27:54

It's got different heuristics, drives,  different kinds of things that are like values.

27:59

Some of those might be similar to things that  were useful to humans and ended up as part of our terminal values in various ways.

28:04

You can imagine  curiosity, various types of option value.

28:04

Maybe it values power itself.

28:12

It could value survival or  some analog of survival.

28:12

Those are possibilities that could have been rewarded as proxy drives  at various stages of this process and made their way into the model's terminal criteria.

28:27

A third category is some analog of reward, where the model at some point has part of its  motivational system fixated on a component of the reward process.

28:44

It’s something like “the humans  approving of me,” or “numbers getting entered in the status center,” or “gradient descent updating  me in this direction.

28:49

” There's something in the reward process such that, as it was trained,  it's focusing on that thing.

28:55

It really wants the reward process to give it a reward.

29:00

But in order for it to be of the type where getting reward motivates choosing the takeover  option, it also needs to generalize such that its concern for reward has some sort of long  time horizon element.

29:09

It not only wants reward, it wants to protect the reward button  for some long period or something.

29:21

Another one is some kind of messed up  interpretation of some human concept.

29:28

Maybe the AIs really want to be like "shmelpful"  and "shmanist" and "shmarmless," but their concept is importantly different from the human concept. And they know this.

29:36

They know that the human concept would mean one thing, but they ended  up with their values fixating on a somewhat different structure.

29:45

That's like another version.

29:45

There’s then a fifth version, which I think about less because it's just such an own goal if you  do this.

29:53

But I do think it's possible.

29:53

You could have AIs that are actually just doing what it  says on the tin.

29:57

You have AIs that are just genuinely aligned to the model spec.

30:02

They're  just really trying to benefit humanity and reflect well on OpenAI and… what's the other  one?

30:07

Assist the developer or the user, right?

30:16

But your model spec, unfortunately,  was just not robust to the degree of optimization that this AI is bringing to bear.

30:20

It’s looking out at the world and they're like, "What's the best way to reflect well on OpenAI  and benefit humanity?"

30:27

It decides that the best way is to go rogue. That's a real own goal.

30:37

At that point you got so close.

30:37

You really just had to write the model spec and red team  it suitably.

30:43

But I actually think it's possible we messed that up too.

30:49

It's kind of an intense  project, writing constitutions and structures of rules and stuff that are going to be robust  to very intense forms of optimization.

30:56

That's a final one that I'll just flag.

31:00

I think it comes  up even if you've solved all these other problems.

31:07

I buy the idea that it's possible that  the motivation thing could go wrong, I'm not sure my probability of that has  increased by detailing them all out.

31:14

In fact, it could be potentially misleading.

31:22

You can  always enumerate the ways in which things go wrong.

31:28

The process of enumeration itself can  increase your probability.

31:28

Whereas you had a vague cloud of 10% or something and you're just  listing out what the 10% actually constitutes.

31:43

Mostly the thing I wanted to do there was  just give some sense of what the model's motivations might be.

31:50

As I said, my best  guess is that it's partly the alien thing, not necessarily, but insofar as you’re also  interested in what the model does later.

31:58

What sort of future would you expect if models did take  over?

32:09

Then it can at least be helpful to have some set of hypotheses on the table instead of just  saying, “It has some set of motivations.

32:15

” In fact, a lot of the work here is being done by our  ignorance about what those motivations are.

32:24

We don't want humans to be violently killed  and overthrown.

32:24

But the idea that over time, biological humans are not the driving force as  the actors of history is baked in, right?

32:31

We can debate the probabilities of the worst-case  scenario, but what is the positive vision we're hoping for?

32:49

What is a future you're happy with?

32:49

This is my best guess and I think this is probably true of a lot of people.

32:58

There's some sort  of more organic, decentralized process of incremental civilizational growth.

33:08

There  is some sense in which the type of thing we trust most—and have most experience with  right now as a civilization—is some sort of, "Okay, we change things a little bit."

33:17

A lot  of people have processes of adjustment and reaction and a decentralized sense of what's  changing. Was that good? Was that bad? Take another step.

33:31

There's some kind of organic  process of growing and changing things.

33:31

I do expect that ultimately to lead to something  quite different from biological humans.

33:41

Though there are a lot of ethical questions we  can raise about what that process involves.

33:58

Ideally there would be some way in which we  managed to grow via the thing that really captures what we trust in.

34:04

There's something  we trust about the ongoing processes of human civilization so far.

34:09

I don't think it's the same  as raw competition.

34:09

There's some rich structure to how we understand moral progress to have been made  and what it would be to carry that thread forward. I don't have a formula.

34:27

We're just going to have  to bring to bear the full force of everything that we know about goodness and justice and beauty.

34:33

We  just have to bring ourselves fully to the project of making things good and doing that collectively.

34:39

That is a really important part of our vision of what was an appropriate process of growing as  a civilization.

34:46

It was this very inclusive, decentralized element of people getting to  think and talk and grow and change things and react rather than some more, "And now the future  shall be like blah."

34:59

I think we don't want that.

35:11

To the extent that the reason we're worried about  motivations in the first place, it’s because we think a balance of power which includes at least  one thing with human-descended motivations is difficult.

35:28

To the extent that we think that's  the case, this seems like a big crux that I often don't hear people talk about.

35:32

I don't know  how you get the balance of power.

35:32

Maybe it’s just a matter of reconciling yourself with the models  of the intelligence explosion.

35:36

They say that such a thing is not possible.

35:41

Therefore, you just  have to figure out how you get the right God.

35:50

I don't really have a framework to think about  the balance of power thing.

35:50

I'd be very curious if there is a more concrete way to think about the  structure of competition, or lack thereof, between the labs now, or between countries, such that the  balance of power is most likely to be preserved.

36:13

A big part of this discourse, at least among  safety-concerned people, is there's a clear trade-off between competition and race dynamics  and the value of the future, or how good the future ends up being.

36:26

In fact, if you buy this  balance of power story, it might be the opposite.

36:33

Maybe competitive pressures naturally favor  balance of power.

36:33

I wonder if this is one of the strong arguments against nationalizing the AIs.

36:37

You can imagine many different companies developing AI, some of which are somewhat  misaligned and some of which are aligned.

36:43

You can imagine that being more conducive to both the  balance of power and to a defensive thing.

36:48

Have all the AIs go through each website and see how  easy it is to hack.

36:54

Basically just get society up to snuff.

37:00

If you're not just deploying this  technology widely, then the first group who can get their hands on it will be able to instigate  a sort of revolution.

37:07

You're just standing against the equilibrium in a very strong way.

37:16

I definitely share some intuition there that a lot of what's scary about the situation with  AI has to do with concentrations of power and whether that power is concentrated in the hands  of misaligned AI or in the hands of some human.

37:42

It's very natural to think, "Okay, let's try to  distribute the power more," and one way to try to do that is to have a much more multipolar scenario  where lots and lots of actors are developing AI.

37:54

This is something that people have talked about.

37:54

When you describe that scenario, you said, "some of which are aligned, some of which are  misaligned."

37:59

That's a key aspect of the scenario, right?

38:06

Sometimes people will say this  stuff.

38:06

They'll be like, "There will be the good AIs and they'll defeat the bad AIs."

38:09

Notice the assumption in there.

38:09

You made it the case that you can control some of the AIs.

38:18

You've got some good AIs.

38:18

Now it's a question of if there are enough of them and how are they  working relative to the others. Maybe.

38:24

I think it's possible that is what happens.

38:30

We know  enough about alignment that some actors are able to do that.

38:35

Maybe some actors are  less cautious or they're intentionally creating misaligned AI or who knows what.

38:38

But if you don't have that—if everyone is in some sense unable to control their AIs—then the  "good AIs help with the bad AIs" thing becomes more complicated.

38:59

Maybe it just doesn't work,  because there's no good AIs in this scenario.

39:06

If you say everyone is building their own  superintelligence that they can't control, it's true that that is now a check on the  power of the other superintelligence.

39:10

Now the other superintelligences need to deal with other  actors, but none of them are necessarily working on behalf of a given set of human interests  or anything like that.

39:19

That's a very important difficulty in thinking about the very simple  thought of "Ah, I know what we can do.

39:29

Let's just have lots and lots of AIs so that no single AI has  a ton of power."

39:35

That on its own is not enough.

39:45

But in this story, I'm just very skeptical  we end up with this.

39:45

By default we have this training regime, at least initially, that  favors a sort of latent representation of the inhibitions and values that humans  have.

39:57

I get that if you mess it up, it could go rogue.

40:03

But if multiple people are  training AIs, they all end up rogue such that the compromises between them don't end up with  humans not violently killed?

40:08

It fails on Google's run and Microsoft's run and OpenAI's run?

40:18

There are very notable and salient sources of correlation between failures across the different  runs.

40:25

People didn't have a developed science of AI motivations.

40:31

The runs were structurally  quite similar.

40:31

Everyone is using the same techniques.

40:34

Maybe someone just stole the weights. It's really important.

40:34

To the extent you haven't solved alignment, you likely haven't solved it  anywhere.

40:46

If someone has solved it and someone hasn't, then it's a better question.

40:52

But if everyone's building systems that are going to go rogue, then I don't think  that's much comfort as we talked about.

41:07

All right, let's wrap up this part  here.

41:07

I didn't mention this explicitly in the introduction.

41:11

To the extent that this  ends up being the transition to the next part, the broader discussion we were having  in part two is about Joe's series, "Otherness and control in the age of AGI."

41:21

The first part is where I was hoping we could just come back and treat the main crux  that people will come in wondering about, and which I myself feel unsure about.

41:30

The “Otherness and control” series is, in some sense, separable.

41:40

It has a lot to do with  misalignment stuff, but a lot of those issues are relevant even given various degrees of skepticism  about some of the stuff I've been saying here.

41:53

By the way, on the actual mechanisms of how  a takeover would happen, I did an episode with Carl Schulman which discusses this  in detail.

41:59

People can go check that out.

42:04

In terms of why it is plausible that AI could  take over from a given position, Carl's discussion is pretty good and gets into a bunch of the  weeds that might give a more concrete sense. All right.

42:21

Now on to part two, where  we discuss the “Otherness and Control in the Age of AGI” series.

42:25

Here’s the first  question.

42:25

Let’s say in a hundred years time, we look back on alignment and consider it was  a huge mistake.

42:30

We should have just tried to build the most raw, powerful AI systems we could  have.

42:35

What would bring about such a judgment?

42:40

Here’s one scenario I think about a lot.

42:40

Maybe  fairly basic measures are enough to ensure, for example, that AIs don't cause catastrophic  harm.

42:48

They don't seek power in problematic ways, etc.

42:53

It could turn out that we learned that  it was easy such that we have regrets.

42:53

We wish we had prioritized differently.

43:01

We end  up thinking, "Oh, I wish we could have cured cancer sooner.

43:06

We could have handled  some geopolitical dynamic differently.

43:10

There's another scenario where we end up looking  back at some period of our history—how we thought about AIs, how we treated our AIs—and we end  up looking back with a kind of moral horror at what we were doing.

43:25

We were thinking about  these things centrally as products and tools, but in fact we should have been foregrounding  much more the sense in which they might be moral patients, at some level of sophistication.

43:36

We were treating them in the wrong way.

43:36

We were acting like we could do whatever  we want.

43:42

We could delete them, subject them to arbitrary experiments, alter  their minds in arbitrary ways.

43:45

We then end up looking back at that in the light of history  as a kind of serious and grave moral error.

43:57

Those are scenarios I think about a lot in which  we have regrets.

43:57

They don’t quite fit the bill of what you just said.

44:02

It sounds to me like the  thing you're thinking is something more that we end up feeling like, "Gosh, we wish we had  paid no attention to the motives of our AIs, that we'd thought not at all about their impact  on our society as we incorporated them.

44:14

Instead we should have pursued a kind of ‘maximize for  brute power’ option.

44:19

” Just make a beeline for whatever is the most powerful AI you can achieve  and don't think about anything else.

44:28

I'm very skeptical that's what we're going to wish for.

44:37

One common example that's given of misalignment is humans from evolution.

44:44

You have one line  in your series: "Here's a simple argument for AI risk: A monkey should be careful before  inventing humans."

44:50

The sort of paperclipper metaphor implies something really banal  and boring with regards to misalignment.

45:05

If I'm steelmanning the people who worship  power, they have the sense that humans got misaligned and they started pursuing things.

45:09

If a monkey was creating them… This is a weird analogy because obviously monkeys didn't create  humans.

45:15

But if the monkey was creating them, they're not thinking about bananas all  day.

45:19

They're thinking about other things.

45:22

On the other hand, they didn't just make useless  stone tools and pile them up in caves in a sort of paperclipper fashion.

45:27

There are all these  things that emerged because of their greater intelligence, which were misaligned with  evolution: creativity and love and music and beauty and all the other things we value  about human culture.

45:38

The prediction maybe they have—which is more of an empirical statement  than a philosophical statement—is, "Listen, with greater intelligence, if you’re thinking  about the paperclipper, even if it's misaligned it will be in this kind of way.

45:52

It'll be things  that are alien to humans, but alien in the way humans are aliens to monkeys, and not in the  way that a paperclipperer is alien to a human."

46:03

There's a bunch of different things to potentially  unpack there.

46:03

There’s one kind of conceptual point that I want to name off the bat.

46:13

I don't think  you're necessarily making a mistake in this vein.

46:18

I just want to name it as a possible  mistake in this vicinity.

46:18

We don't want to engage in the following form of reasoning.

46:24

Let's  say you have two entities.

46:24

One is in the role of creator.

46:29

One is in the role of creation.

46:29

We're  positing that there's this kind of misalignment relation between them, whatever that means.

46:34

Here's a pattern of reasoning that you want to watch out for.

46:41

Say you're thinking of humans  in the role of creation, relative to an entity like evolution, or monkeys or mice or whoever you  could imagine inventing humans or something like that.

46:57

You say, "Qua creation, I'm happy that  I was created and happy with the misalignment.

47:07

Therefore, if I end up in the role of creator  and we have a structurally analogous relation in which there's misalignment with some creation,  I should expect to be happy with that as well."

47:22

There's a couple of philosophers that you  brought up in the series.

47:22

If you read their works that you talk about, they actually seem  incredibly foresighted in anticipating something like a singularity and our  ability to shape a future thing that's different, smarter, maybe better than us. Obviously C. S.

47:40

Lewis and "The Abolition of Man," which we'll talk about in a second, is one  example, Here's one passage from Nietzsche that I felt really highlighted this: "Man is a rope  stretched between the animal and the superman.

47:55

A rope over an abyss, a dangerous crossing, a  dangerous wayfaring, a dangerous looking back, a dangerous trembling and halting."

48:01

Is there some explanation?

48:01

Is it just somehow obvious that something like this is  coming even if you’re thinking 200 years ago?

48:09

I have a much better grip on what's going on  with Lewis than with Nietzsche there.

48:09

Maybe let's just talk about Lewis for a second.

48:13

There's  a version of the singularity that's specifically a hypothesis about feedback loops with AI  capabilities.

48:21

I don't think that's present in Lewis.

48:26

What Lewis is anticipating—I  do think this is a relatively simple forecast—is something like the culmination  of the project of scientific modernity.

48:40

Lewis is looking out at the world.

48:40

He's seeing  this process of increased understanding of a kind of the natural environment and a corresponding  increase in our ability to control and direct that environment.

48:55

He's also pairing that with a kind  of metaphysical hypothesis.

48:55

His stance on this metaphysical hypothesis is problematically unclear  in the book, but there is this metaphysical hypothesis.

49:10

Naturalism says that humans  too—minds, beings, agents—are a part of nature.

49:22

Insofar as this process of scientific modernity  involves a kind of progressively greater understanding of an ability to control nature,  that will presumably grow to encompass our own natures and the natures of other beings we could  create in principle.

49:35

Lewis views this as a kind of cataclysmic event and crisis.

49:44

In particular,  he believes that it will lead to all kinds of tyrannical behaviors and attitudes towards  morality and stuff like that.

49:51

We can talk about if you believe in non-naturalism—or in some form  of Dao, which is this kind of objective morality Part of what I'm trying to do in that essay  is to say, “No, we can be naturalists and also be decent humans that remain in touch with  a rich set of norms that have to do with how we relate to the possibility of creating creatures,  altering ourselves, etc.

50:21

” It's a relatively simple prediction. Science masters nature.

50:28

Humans  are part of nature. Science masters humans.

50:34

You also have a very interesting essay  about what we should expect of other humans, a sort of extrapolation if they  had greater capabilities and so on.

50:45

There’s an uncomfortable thing about the  conceptual setup at stake in these abstract discussions.

50:54

Okay, you have this agent.

50:54

It  "FOOMs," which is this amorphous process of going from a seed agent to a superintelligent  version of itself, often imagined to preserve its values along the way.

51:08

There’s a bunch  of questions we can raise about that.

51:16

Many of the arguments that people will often talk  about in the context of reasons to be scared of AI are like, "Oh, value is very fragile as you  FOOM."

51:20

“Small differences in utility functions can decorrelate very hard and drive in quite different  directions.

51:30

” “Agents have instrumental incentives to seek power.

51:36

If it were arbitrarily easy to  get power, then they would do it. ” It’s stuff like that.

51:40

These are very general arguments  that seem to suggest that it's not just an AI thing. It's no surprise. Take a thing.

51:47

Make it arbitrarily powerful such that it's God Emperor of the universe or something.

51:57

How scared are you of that?

51:57

Clearly we should be equally scared of that.

52:02

We should be  really scared of that with humans too, right?

52:07

Part of what I'm saying in that essay is that  this, in some sense, is much more a story about balance of power.

52:11

It’s about maintaining checks  and balances and distribution of power, period.

52:22

It’s not just about humans vs.

52:22

AIs, and the  differences between human values and AI values.

52:28

Now that said, I do think many humans would  likely be nicer if they FOOMed than certain types of AIs.

52:32

But with the conceptual structure  of the argument, it's a very open question how much it applies to humans as well.

52:43

How confident are we with this ontology of expressing what agents and capabilities  are?

52:55

How do we know this is what's happening, or that this is the right way to  think about what intelligences are? It's very janky.

53:07

People may disagree about this.

53:07

I think it's obvious to everyone, with respect to real world human agents, that thinking of  humans as having utility functions is at best a very lossy approximation.

53:26

This is likely to  mislead as you increase the intelligence of various agents.

53:36

Eliezer might disagree about that.

53:36

For example, my mom a few years ago wanted to get a house and get a new dog. Now she has both. How  did this happen? It’s because she tried.

53:50

She had to search for the house.

53:58

It was hard to find the  dog. Now she has a house. Now she has a dog.

53:58

This is a very common thing that happens all the time.

54:04

We don't need to say she has a utility function for the dog and a consistent valuation of all  houses or whatever.

54:11

It’s still the case that her planning and agency, exerted in the world,  resulted in her having this house and dog.

54:22

As our scientific and technological power  advances, it’s plausible that more and more stuff will likely be explicable this way.

54:28

Why is  this man on the moon? How did that happen?

54:28

Well, there was a whole cognitive process and planning  apparatus.

54:38

It wasn’t localized in a single mind, but there was a whole thing such that we got  a man on the moon.

54:44

We'll see more of that and the AIs will be doing a bunch of it.

54:51

That  seems more real to me than utility functions.

55:02

The man on the moon example has a proximal story  of how NASA engineered the spacecraft to get to the moon.

55:11

There’s the more distal geopolitical  story of why we sent people to the moon.

55:11

At all those levels, there are different utility  functions clashing.

55:17

Maybe there's a meta-societal utility function.

55:24

Maybe the story there is about  a balance of power between agents, creating an emergent outcome.

55:31

We didn't go to the moon  because one guy had a utility function, but due to the Cold War and things happening.

55:37

The alignment stuff is a lot about assuming one entity will control everything, so how do we  control the thing that controls everything.

55:47

It's not clear what you do to reinforce the balance  of power.

55:55

It could just be that balance of power is not a thing that happens once you have  things that can make themselves intelligent.

56:03

But that seems interestingly different  from the "how we got to the moon" story? Yeah, I agree.

56:09

There's a few things going  on there.

56:09

Even if you're engaged in this ontology of carving up the world into different  agencies, at the least you don't want to assume that they're all unitary or not overlapping.

56:22

It's  not like, “All right, we've got this agent.

56:22

Let's carve out one part of the world.

56:27

It's one agent  over here.

56:27

” It's this whole messy ecosystem, teeming niches and this whole thing.

56:33

In discussions of AI, sometimes people slip between being like, "An agent is  anything that gets anything done.

56:41

It could be like this weird moochy thing,” and then  sometimes they're very obviously imagining an individual actor. That's one difference.

56:52

I also just think we should be really going for the balance of power thing.

57:01

It is just not  good to be like, "We’re going to have a dictator.

57:07

Let's make sure we make the dictator the right  dictator."

57:07

I'm like, “okay, whoa, no.

57:07

” The goal should be that we all FOOM together.

57:16

We do the  whole thing in this inclusive and pluralistic way that satisfies the values of tons of stakeholders.

57:22

At no point is there one single point of failure on all these things.

57:31

That's what we should be  striving for here.

57:31

That's true of the human power aspect of AI and of the AI part as well.

57:38

There's an interesting intellectual discourse on the right-wing side of the debate.

58:49

They say  to themselves, “Traditionally we favor markets, but now look where our society is headed.

58:56

It's  misaligned in the ways we care about society being aligned, like fertility is going down,  family values, religiosity.

59:02

These things we care about. GDP keeps going up.

59:07

These things don't  seem correlated.

59:07

We're grinding through the values we care about because of increased competition.

59:13

Therefore we need to intervene in a major way.

59:13

” Then the pro-market libertarian faction of  the right will say, “Look, I disagree with the correlations here, but even at the end of  the day…” Fundamentally their point is, liberty is the end goal.

59:32

It's not what you use to get to  higher fertility or something.

59:32

There's something interestingly analogous about the AI competition  grinding things down.

59:39

Obviously you don't want the gray goo, but with the libertarians versus  the trads, there's something analogous here.

59:50

Here’s one thing you could think and it  doesn't necessarily need to be about gray goo.

59:54

It could also just be about alignment.

59:54

Sure, it would be nice if the AIs didn't violently disempower humans.

1:00:02

It would  be nice if the AIs when we created them, their integration into our society led to good  places.

1:00:08

But I'm uncomfortable with the sorts of interventions that people are contemplating  in order to ensure that sort of outcome.

1:00:21

There's a bunch of things to be uncomfortable  about that.

1:00:21

That said, for something like everyone being killed or violently disempowered, when it's  a real threat we traditionally often think that quite intense forms of intervention are warranted  to prevent that sort of thing from happening.

1:00:46

Obviously we need to talk about whether it’s real.

1:00:46

If there were actually a terrorist group that was working on a bioweapon that was going  to kill everyone, or 99.

1:00:51

9% of people, we would think that warrants intervention. Just shut that down.

1:00:57

Say you had a group that was doing that unintentionally, imposing  a similar level of risk.

1:01:02

Many people, if that's the real scenario, will think that  warrants quite intense preventative efforts.

1:01:17

Obviously, these sorts of risks can be used  as an excuse to expand state power.

1:01:17

There's a lot of things to be worried about for different  types of contemplated interventions to address certain types of risks.

1:01:28

I think there's no  royal road there.

1:01:28

You need to just have the actual good epistemology.

1:01:36

You need to actually  know, is this a real risk?

1:01:36

What are the actual stakes?

1:01:39

You need to look at it case by case  and be like, “Is this warranted?

1:01:39

” That's one point on the takeover, literal extinction thing.

1:01:48

The other thing I want to say, I talk about this distinction in the piece.

1:01:56

There’s a thought that  we should at least have AIs who are minimally law-abiding or something like that.

1:02:01

There's this  question about servitude and about other control over AI values.

1:02:07

But we often think it's okay to  really want people to obey the law, to uphold basic cooperative arrangements, stuff like that.

1:02:13

This is true of markets and true of liberalism in general.

1:02:23

I want to emphasize just how much  these procedural norms—democracy, free speech, property rights, things that people including  myself really hold dear—are, in the actual lived substance of a liberal state, undergirded by  all sorts of kind of virtues and dispositions and character traits in the citizenry.

1:02:45

These norms  are not robust to arbitrarily vicious citizens.

1:02:55

I want there to be free speech, but we also need  to raise our children to value truth and to know how to have real conversations.

1:03:00

I want there to be  democracy, but we also need to raise our children to be compassionate and decent.

1:03:05

Sometimes we can  lose sight of that aspect.

1:03:05

That's not to say that it should be the project of state power.

1:03:17

But I  think it’s important to understand that liberalism is not this ironclad structure that you can just  hit go on.

1:03:22

You can’t give it any citizenry and hit go and assume you'll get something flourishing  or even functional.

1:03:29

There's a bunch of other softer stuff that makes this whole project go.

1:03:33

I want to zoom out to the people who have—I don't know if Nick Land would be a good sub in here—  a sort of fatalistic attitude towards alignment as a thing that can even make sense.

1:03:52

They'll  say things like, “Look, these are the kinds of things that are going to be exploring  the black hole, the center of the galaxy, the kinds of things that go visit Andromeda  or something.

1:04:02

Did you really expect them to privilege whatever inclinations you have because  you grew up in the African savannah and whatever the evolutionary pressures were a hundred thousand  years ago?

1:04:12

Of course, they're going to be weird.

1:04:22

What did you think was going to happen?

1:04:22

” I do think that even good futures will be weird.

1:04:26

I want to be clear about that  when I talk about finding ways to ensure that the integration of AIs into our society leads  to good places.

1:04:35

Sometimes people think that this project of wanting that—and especially to the  extent that makes some deep reference to human values—involves this short-sighted, parochial  imposition of our current unreflective values.

1:05:02

They imagine that we're forgetting that for  us too, there's a kind of reflective process and a moral progress dimension that we want to  leave room for.

1:05:08

Jefferson has this line about, “Just as you wouldn't want to force a  grown man into a younger man's coat, so we don't want to chain civilization to  a barbarous past.

1:05:23

” Everyone should agree on that.

1:05:27

The people who are interested in alignment,  also agree on that.

1:05:27

Obviously, there's a concern that people don't engage in that process or that  something shuts down the process of reflection, but I think everyone agrees we want that.

1:05:38

So that will lead, potentially, to something that is quite different from our current conception  of what's valuable.

1:05:43

There's a question of how different.

1:05:53

There are also questions about what  exactly we're talking about with reflection. I have an essay on this.

1:05:57

I don't actually think  there's a kind of off-the-shelf, pre-normative notion of reflection where you can just be like,  "Oh, obviously you take an agent, stick it through reflection, and then you get like values. ” No.

1:06:07

Really there's a whole pattern of empirical facts about taking an agent, putting it through  some process of reflection and all sorts of things, asking it questions.

1:06:23

That'll go in all  sorts of directions for a given empirical case.

1:06:28

Then you have to look at the pattern of outputs  and be like, “Okay, what do I make of that?

1:06:28

” Overall, we should expect that even  the good futures will be quite weird.

1:06:41

They might even be incomprehensible to us. I  don't think so...

1:06:41

There's different types of incomprehensible.

1:06:47

Say I show up in the future  and this is all computers.

1:06:47

I'm like, “Okay, all right.

1:06:51

” Then they're like, “We're running  creatures on the computers.

1:06:51

” Okay, so I have to somehow get in there and see what's actually going  on with the computers or something like that.

1:07:00

Maybe I can actually see.

1:07:00

Maybe I actually  understand what's going on in the computers, but I don't yet know what values I should  be using to evaluate that.

1:07:03

So it can be the case that if we showed up, we would not be very  good at recognizing goodness or badness.

1:07:07

I don't think that makes it insignificant though.

1:07:14

Suppose you show up in the future and it's got some answer to the Riemann hypothesis.

1:07:20

You can't tell whether that answer's right.

1:07:25

Maybe the civilization went wrong.

1:07:25

It's still an  important difference.

1:07:25

It's just that you can't track it.

1:07:29

Something similar is true of worlds  that are genuinely expressive of what we would value if we engaged in processes of reflection  that we endorse, versus ones that have totally veered off into something meaningless.

1:07:41

One thing I've heard from people who are skeptical of this ontology is, "All right, what  do you even mean by alignment?"

1:07:46

Obviously the very first question you answered already.

1:07:51

Here’s  different things that it could mean.

1:07:51

Do you mean balance of power?

1:07:56

It’s somewhere between that  and dictator or whatever.

1:07:56

Then there's another thing.

1:08:02

Separate from the AI discussion, I  don't want the future to contain a bunch of torture.

1:08:06

It's not necessarily technical.

1:08:06

Part of it might involve technically aligning a GPT-4, but that's a proxy to get to that future.

1:08:17

What do we really mean by alignment?

1:08:17

Is it just whatever it takes to make sure the future doesn't  have a bunch of torture?

1:08:26

Or do I really care that in a thousand years, the things that are clearly  my descendants are in control of the galaxy, and even if they’re not conducting torture.

1:08:40

By  descendants, I don’t mean some things where I recognize they have their own art or  whatever.

1:08:42

I mean like my grandchild, that level of descendant.

1:08:47

I think what some  people mean is that our intellectual descendants should control the light cone, even if the other  counterfactual doesn't involve a bunch of torture. I agree.

1:09:01

There's a few different things there. What are you going for?

1:09:01

Are you going for actively good or are you going for avoiding certain stuff?

1:09:08

Then there's a different question which is, what counts as actively good according to you?

1:09:14

Maybe some people are like, “The only things that are actively good are my grandchildren.

1:09:21

”  Or they’re thinking of some literal descending genetic line or something, otherwise that's not my  thing.

1:09:30

I don't think it's really what most people have in mind when they talk about goodness.

1:09:39

There's a conversation to be had.

1:09:39

Obviously in some sense, when we talk about a  good future, we need to be thinking, “What are all the stakeholders here and how does  it all fit together?

1:09:47

” When I think about it, the thing that matters about the lineage is this.

1:10:06

It’s whatever's required for the optimization processes to be pushing towards good stuff.

1:10:17

There's a concern that currently a lot of what is making that happen lives in human  civilization.

1:10:30

There's some kind of seed of goodness that we're carrying, in different ways  or, different people.

1:10:42

There's different notions of goodness for different people maybe, but there's  some sort of seed that is currently here that we have that is not just in the universe everywhere.

1:10:52

It's not just going to crop up if you just die out or something.

1:11:00

It's something that is contingent  to our civilization.

1:11:00

At least that's the picture, we can talk about whether that's right.

1:11:05

So the sense in which stories about good futures that have to do with alignment are about  descendants, it's more about whatever that seed is. How do we carry it?

1:11:17

How do we keep the  life thread alive, going into the future?

1:11:24

But then one could accuse the alignment community  of motte and bailey.

1:11:24

The motte is: We just want to make sure that GPT-8 doesn't kill everybody.

1:11:33

After  that, we're all cool.

1:11:33

Then the real thing is: “We are fundamentally pessimistic about historical  processes, in a way that doesn't even necessarily implicate AI alone.

1:11:51

It’s just the nature of  the universe.

1:11:51

We want to do something to make sure the nature of the universe doesn't take  a hold on humans and where things are headed.

1:12:03

If you look at the Soviet Union, the  collectivization of farming and the disempowerment of the kulaks was not as a  practical matter necessary.

1:12:10

In fact it was extremely counterproductive and it almost brought  down the regime.

1:12:15

Obviously it killed millions of people, caused a huge famine.

1:12:20

But it was sort  of ideologically necessary.

1:12:20

You have an ember of something here and we have to make sure that an  enclave of the other thing doesn't put it out.

1:12:27

If you have raw competition between the kulak type  capitalism and what we're trying to build here, the gray goo of the kulaks will just take over. We have this ember here.

1:12:40

We're going to do worldwide revolution from it.

1:12:45

I know that  obviously that's not exactly the kind of thing alignment has in mind, but we have an ember  here and we've got to make sure that this other thing that's happening on the side doesn't FOOM.

1:12:55

Obviously that's not how they would phrase it, but so that it doesn’t get a hold on what  we're building here.

1:13:00

That's maybe the worry that people who are opposed to alignment have.

1:13:05

It’s the second kind of thing, the kind of thing that Stalin was worried about.

1:13:08

Obviously, we  wouldn't endorse the specific things he did.

1:13:14

When people talk about alignment, they have  in mind a number of different types of goals.

1:13:18

One type of goal is quite minimal.

1:13:18

It's something  like, “The AI's don't kill everyone or violently disempower people.

1:13:27

” There's a second thing people  sometimes want out of alignment, which is much broader.

1:13:32

It’s something like, “We would like it  to be the case that our AI's are such that when we incorporate them into our society, things  are good, that wee just have a good future.

1:13:38

” I do agree that the discourse about AI  alignment mixes together these two goals that I mentioned.

1:13:52

I actually mentioned three  goals.

1:13:52

The most straightforward thing to focus on—I don't blame people for just talking about  this one—is just the first one.

1:13:56

It's quite robust according to our own ethics, when we think about  in which context is it appropriate to try to exert various types of control, or to have more of what  I call in the series "yang," which is this active controlling force, as opposed to "yin," which  is this more receptive and open, letting go.

1:14:20

A kind of paradigm context in which we think  that is appropriate is if something is an active aggressor against the boundaries and cooperative  structures that we've created as a civilization.

1:14:37

I talked about the Nazis.

1:14:37

In the piece, I  talked about how when something is invading, we often think it's appropriate to fight  back.

1:14:44

We often think it's appropriate to set up structures to prevent and ensure that these  basic norms of peace and harmony are adhered to.

1:14:59

I do think some of the moral heft of some parts  of the alignment discourse comes from drawing specifically on that aspect of our morality.

1:15:05

We think the AIs are presented as aggressors that are coming to kill you.

1:15:11

If that's true,  then it's quite appropriate.

1:15:11

That’s classic human stuff.

1:15:23

Almost everyone recognizes that  self-defense, or ensuring basic norms are adhered to, is a justified use of certain kinds  of power that would often be unjustified in other contexts.

1:15:36

Self-defense is a clear example there.

1:15:36

I do think it's important though to separate that concern from this other concern about where the  future eventually goes.

1:15:44

How much do we want to be trying to steer that actively?

1:15:55

I wrote the  series partly in response to the thing you're talking about.

1:16:00

It is true that aspects of  this discourse involve the possibility of trying to steer and grip.

1:16:08

You have a sense that  the universe is about to go off in some direction and you need people to notice that muscle.

1:16:15

We have a very rich ethical human ethical tradition of thinking about, when it is  appropriate to try to exert what sorts of control over which things.

1:16:28

Part of what I  want to do is that I want us to bring the full force and richness of that tradition to  this discussion.

1:16:32

It's easy if you're purely in this abstract mode of utility functions and  human utility functions.

1:16:37

There's this competitor thing with a utility function.

1:16:42

Somehow you  lose touch with the complexity of how we've been dealing with differences in values and  competitions for power. This is classic stuff.

1:16:56

AI sort of amplifies a lot of the dynamics, but  I don't think it's fundamentally new.

1:16:56

Part of what I'm trying to say is let's draw on the full  wisdom we have here, while obviously adjusting for ways in which things are different.

1:17:07

There’s one thing the ember analogy brings up about getting a hold of the future is.

1:17:13

We're  going to go explore space and that's where we expect most of the things that will happen.

1:17:19

Most  of the people that will live, they’ll be in space.

1:17:25

I wonder how much of the high stakes here is not  really about AI per se, but it's about space.

1:17:32

It's a coincidence that we're developing AI at  the same time we are on the cusp of expanding through most of the stuff that exists.

1:17:38

I don't think it's a coincidence.

1:17:45

The most salient way we would become able  to expand is via some kind of radical acceleration of our technological progress. Sorry, let me clarify.

1:17:53

If this was just a question of, "Do we do AGI and explore the solar system?"

1:18:01

and there was nothing beyond the solar system, we FOOM and weird things might happen with the  solar system if we get it wrong.

1:18:06

Compared to that, billions of galaxies present different  stakes.

1:18:13

I wonder how much of the discourse hinges on this because of space.

1:18:19

I think for most people, very little.

1:18:26

People are really focused on what's going to  happen to this world around us that we live in.

1:18:32

What's going to happen to me and my kids?

1:18:32

Some  people spend a lot of time on the space stuff, but I think for the immediately pressing stuff  about AI, it doesn’t require that at all.

1:18:46

Even if you bracket space, time is also very big.

1:18:46

We've got 500 million years, a billion years, left on Earth if we don't mess with the  sun.

1:18:55

Maybe you could get more out of it. That's still a lot.

1:19:02

I don't know if it  fundamentally changes the narrative.

1:19:09

Obviously, the stakes are way smaller if you  shrink down to the solar system, insofar as you care about what happens in the future  or in space.

1:19:10

That does change some stuff potentially.

1:19:22

A really nice feature of our current  situation—depending on the actual nature of the resource pie—is that there's such an abundance  of energy and other resources in principle available to a responsible civilization.

1:19:40

Tons  of stakeholders, especially ones who are able to get really close to amazing outcomes according to  their values with comparatively small allocations of resources, can be satisfied.

1:20:03

I feel like  everyone with satiable values could be really happy with some small fraction of the available  pie.

1:20:07

We should just satiate all sorts of stuff.

1:20:13

Obviously, we need to figure out gains from trade  and balance.

1:20:13

There's a bunch of complexity here but in principle, we're in a position to create  a really wonderful scenario for tons of different value systems.

1:20:31

Correspondingly, we should be  really interested in doing that.

1:20:31

I sometimes use this heuristic in thinking about the future: We  should be aspiring to really leave no one behind.

1:20:47

Who are all the stakeholders here?

1:20:47

How do we have  a fully inclusive vision of how the future could be good from a very wide variety of perspectives?

1:20:52

The vastness of space resources makes that a lot easier and very feasible.

1:21:02

If you instead imagine  it's a much smaller pie, maybe you face tougher trade-offs.

1:21:09

That's an important consideration.

1:21:09

Is the inclusivity because part of your values includes different potential futures getting to  play out?

1:21:17

Or is it because of uncertainty about which one is right, so you want to make sure  we're not nulling all value if we’re wrong?

1:21:34

It's a bunch of things at once.

1:21:34

I'm really  into being nice when it's cheap.

1:21:34

If you can help someone a lot in a way that's really  cheap for you, do it.

1:21:43

Obviously, you need to think about trade-offs.

1:21:48

There are a lot of  people you could be nice to in principle, but I'm very excited to try to uphold the  principle of being nice when it's cheap.

1:21:58

I also really hope that other people uphold that  with respect to me, including the AIs.

1:21:58

We should be applying the golden rule as we're thinking  about inventing these AIs.

1:22:03

There’s some way in which I'm trying to embody attitudes towards them  that I hope they would embody towards me.

1:22:08

It's unclear exactly what the ground of that is, but I  really like the golden rule and think a lot about it as a basis for treatment of other beings.

1:22:21

If everyone implements the "be nice when it's cheap" rule, we potentially get a big Pareto  improvement.

1:22:28

It's a lot of good deals. It’s that. I'm into pluralism. I've got uncertainty.

1:22:45

There's all sorts of stuff swimming around there.

1:22:53

Also, as a matter of having cooperative and good  balances of power and deals and avoiding conflict, I think it’s important to find ways to  set up structures that lots of people, value systems, and agents are happy with.

1:23:06

That  includes non-humans, people in the past, AIs, animals.

1:23:13

We really should have a very broad sweep  in thinking about what sorts of inclusivity we want to be reflecting in a mature civilization  and setting ourselves up for doing that.

1:23:28

I want to go back to what our relationship  with these AIs should be.

1:23:28

Pretty soon we're talking about our relationship  to superhuman intelligences, if we think such a thing is possible.

1:23:39

There's  a question of what process you use to get there and the morality of gradient descenting on  their minds, which we can address later.

1:23:49

The thing that personally gives me the  most unease about alignment is that at least a part of the vision here sounds like  you're going to enslave a god.

1:23:57

There's just something that feels wrong about that.

1:24:06

But then  if you don't enslave the god, obviously the god's going to have more control.

1:24:13

Are you okay  with surrendering most of everything, even if it's like a cooperative relationship you have?

1:24:21

I think we as a civilization are going to have a very serious conversation about what sort  of servitude is appropriate or inappropriate in the context of AI development.

1:24:37

There are a  bunch of disanalogies from human slavery that are important.

1:24:44

In particular, the AIs might not  be moral patients at all, in which case we need to figure that out.

1:24:51

There are ways in which we  may be able to have motivations.

1:24:51

Slavery involves all this suffering and non-consent.

1:25:02

There are  all these specific dynamics involved in human slavery.

1:25:05

Some of those may or may not be present  in a given case with AI, and that's important.

1:25:15

Overall, we are going to need to stare hard at  it.

1:25:15

Right now, the default mode of how we treat AIs gives them no moral consideration at all.

1:25:21

We're thinking of them as property, as tools, as products, and designing them to be assistants  and such.

1:25:28

There has been no official communication from any AI developer as to when or under what  circumstances that would change.

1:25:37

Sothere's a conversation to be had there that we need to have.

1:25:44

I want to push back on the notion that there are only two options: enslaved god or loss of control.

1:26:00

I think we can do better than that. Let's work on it. Let's try to do better.

1:26:10

I think we can do  better.

1:26:10

It might require being thoughtful.

1:26:10

It might require having a mature discourse about  this before we start taking irreversible moves.

1:26:28

But I'm optimistic that we can at least avoid  some of the connotations and a lot of the stuff at stake in that kind of binary.

1:26:34

With respect to how we treat the AIs, I have a couple of contradicting intuitions.

1:26:41

The difficulty with using intuitions in this case is that obviously it's not clear what  reference class an AI we have control over is.

1:26:51

Here’s one example, that's very scary about  the things we're going to do to these things.

1:26:51

If you read about life under Stalin or Mao, there's  one version of telling it that is actually very similar to what we mean by alignment.

1:27:07

We do these  black box experiments to make it think that it can defect.

1:27:16

If it does, we know it's misaligned.

1:27:16

If you consider Mao's Hundred Flowers Campaign, it’s "let a hundred flowers bloom.

1:27:21

I'm going to allow criticism of my regime and so on.

1:27:25

” That lasted for a couple of  years.

1:27:25

Afterwards, for everybody who did that, it was a way to find the so-called "snakes."

1:27:30

Who  are the rightists who are secretly hiding? We'll purge them.

1:27:36

There was this sort of paranoia about  defectors, like "Anybody in my entourage, anybody in my regime, they could be a secret capitalist  trying to bring down the regime."

1:27:44

That's one way of talking about these things, which is very  concerning.

1:27:49

Is that the correct reference class?

1:27:54

I certainly think concerns in that vein are  real.

1:27:54

It is disturbing how easy many of the analogies are with human historical events and  practices that we deplore or at least have a lot of wariness towards, in the context of the way you  end up talking about AI.

1:28:13

It’s about maintaining control over AI, making sure that it doesn't  rebel.

1:28:25

We should be noticing the reference class that some of that talk starts to conjure.

1:28:34

Basically, yes, we should really notice that.

1:28:45

Part of what I'm trying to do in the series is to  bring the full range of considerations at stake into play.

1:28:53

It is both the case that we should be  quite concerned about being overly controlling or abusive or oppressive.

1:29:04

There are all sorts of ways  you can go too far.

1:29:04

There are concerns about the AIs being genuinely dangerous and genuinely  killing us and violently overthrowing us.

1:29:22

The moral situation is quite complicated.

1:29:22

Often when you imagine a sort of external aggressor who's coming in and invading you, you  feel very justified in doing a bunch of stuff to prevent that.

1:29:39

It's a little bit different  when you're inventing the thing and you're doing it incautiously.

1:29:43

There's a different  vibe in terms of the overall justificatory stance you might have for various types of  more kind of power-exerting interventions.

1:30:08

That's one feature of the situation.

1:30:08

The opposite perspective here is that you're doing this sort of vibes-based reasoning of, "Ah,  that looks yucky," doing gradient descent on these minds.

1:30:20

In the past, a couple of similar cases  might have been something like environmentalists not liking nuclear power because the vibes  of nuclear don't look green.

1:30:29

Obviously that set back the cause of fighting climate change.

1:30:34

So the end result of a future you're proud of, a future that's appealing, is set back  because your vibes about, "We would be wrong to brainwash a human."

1:30:47

You're trying to apply to  a disanalogous case where that's not as relevant.

1:30:53

I do think there's a concern here, which I  really tried to foreground in the series, that is related to what you're saying.

1:30:58

You might  be worried that we will be very gentle and nice and free with the AIs, and then they'll kill us.

1:31:07

They'll take advantage of that and then it will have been a catastrophe.

1:31:13

I opened the series  basically with an example.

1:31:13

I'm really trying to conjure that possibility at the same time as  conjuring the grounds of gentleness.

1:31:24

These AIs could both be like moral patients—this sort of new  species in the sense that should conjure wonder and reverence—and such that they will kill you.

1:31:42

I have this example of the documentary Grizzly Man, where there's this environmental activist,  Timothy Treadwell.

1:31:49

He aspires to approach these grizzly bears.

1:31:57

In the summer, he goes into Alaska  and he lives with these grizzly bears.

1:31:57

He aspires to approach them with this gentleness and  reverence.

1:32:03

He doesn't carry bear mace.

1:32:03

He doesn't use a fence around his camp.

1:32:08

He  gets eaten alive by one of these bears.

1:32:19

I really wanted to foreground that possibility  in the series.

1:32:19

We need to be talking about these things both at once.

1:32:24

Bears can be moral  patients.

1:32:24

AIs can be moral patients.

1:32:24

Nazis are moral patients.

1:32:30

Enemy soldiers have souls.

1:32:30

We  need to learn the art of hawk and dove both.

1:32:39

There's this dynamic here that we need  to be able to hold both sides of as we go into these trade-offs and these dilemmas.

1:32:45

A part of what I'm trying to do in the series is really bring it all to the table at once.

1:32:50

If today I were to massively change my mind about what should be done, the big crux that  I have is the question of how weird things end up default, how alien they end up.

1:33:08

You made a  really interesting argument on your blog post that if moral realism is correct, that actually  makes an empirical prediction.

1:33:16

The aliens, the ASIs, whatever, should converge  on the right morality the same way that they converge on the right mathematics.

1:33:25

I thought that was a really interesting point.

1:33:31

But there's another prediction that moral realism  makes.

1:33:31

Over time society should become more moral, become better.

1:33:40

Of course there is the problem  of, "What morals do you have now?

1:33:40

It's the ones that society has been converging towards over  time."

1:33:49

But to the extent that it's happened, one of the predictions of moral realism  has been confirmed, so does that mean we should update in favor of moral realism?

1:33:59

One thing I want to flag is that not all forms of moral realism make this prediction.

1:34:03

I'm happy  to talk about the different forms I have in mind.

1:34:12

There are also forms of things that look  like moral anti-realism—at least in their metaphysics according to me—but which just  posit that there's this convergence.

1:34:16

It's not in virtue of interacting with some  kind of mind-independent moral truth, but just for some other reason.

1:34:25

That looks a lot  like moral realism at that point.

1:34:25

It's universal, everyone ends up there.

1:34:32

It's tempting to ask why  and whatever answer is a little bit like, "Is that the Dao?

1:34:39

Is that the nature of the Dao?"

1:34:39

even if  there's not an extra metaphysical realm in which the moral lives.

1:34:44

Moral convergence is a different  factor from the existence or non-existence of a morality that's not reducible to  natural facts, which is the type of moral realism I usually consider.

1:34:58

Now, does the improvement of society update us towards moral realism?

1:35:08

Maybe it’s a  very weak update or something.

1:35:08

I’m kind of like, “Which view predicts this more strongly?

1:35:20

”  It feels to me like moral anti-realism is very comfortable with the observation that  people with certain values have those values.

1:35:32

There's obviously this first thing.

1:35:32

If you're  the culmination of some process of moral change, then it's very easy to look back at that process  and say "Ah, moral progress.

1:35:37

The arc of history bends towards me.

1:35:42

” If there were a bunch of dice  rolls along the way, you might think, "Oh wait, that's not rational.

1:35:49

That's not the march  of reason."

1:35:49

There's still empirical work you can do to tell whether that's what's going on.

1:35:55

On moral anti-realism, consider Aristotle and us.

1:36:05

Has there been moral progress by Aristotle's  lights and our lights too?

1:36:05

You could think, "Ah, doesn't that sound a bit like moral realism?

1:36:19

These  hearts are singing in harmony.

1:36:19

That's the moral realist thing, right?

1:36:25

The anti-realist thing  is that hearts all go in different directions, but you and Aristotle apparently are  both excited about the march of history.

1:36:27

” There's an open question about whether that's  true.

1:36:34

What are Aristotle's reflective values? Suppose it is true.

1:36:40

That's fairly explicable in  moral anti-realist terms.

1:36:40

You can roughly say that you and Aristotle are sufficiently similar.

1:36:46

You endorse sufficiently similar reflective processes.

1:36:52

Those processes are in fact  instantiated in the march of history.

1:36:52

So history has been good for both of you.

1:36:59

There are worlds where that isn't the case.

1:37:08

So there's a sense in which maybe that  prediction is more likely for realism than anti-realism, but it doesn't move me very much.

1:37:14

I don't know if moral realism is the right word, but you mentioned the thing.

1:37:23

There's something  that makes hearts converge to the thing we are or the thing we would be upon reflection.

1:37:31

Even  if it's not something that's instantiated in a realm beyond the universe, it's a force that  exists that acts in a way we're happy with.

1:37:35

To the extent that it doesn't exist and you let go of  the reins and you get the paper clippers, it feels like we were doomed a long time ago?

1:37:47

We were just  different utility functions banging against each other.

1:37:54

Some of them have parochial preferences,  but it's just combat and some guy won.

1:38:03

In the other world it’s “No, these are where  the hearts are supposed to go or it's only by catastrophe that they don't end up there.

1:38:10

” That  feels like the world where it really matters.

1:38:10

The initial question I asked was, “What would make us  think that alignment was a big mistake?

1:38:18

” In the world where hearts just naturally end up like the  thing we want, maybe it takes an extremely strong force to push them away from that.

1:38:30

That extremely  strong force is you solve technical alignment, the blinders on the horse's eyes.

1:38:39

In the  worlds that really matter, we're like, "Ah, this is where the hearts want to go."

1:38:45

In that  world, maybe alignment is what messes us up.

1:38:50

So the question is, do the worlds that matter  have this kind of convergent moral force, whether metaphysically inflationary or not,  or are those the only ones that matter?

1:39:03

Maybe what I meant was, in those  worlds you’re kind of fucked.

1:39:07

Or the worlds without that, the worlds  with no Dao.

1:39:07

Let's use the term “Dao” for this kind of convergent morality.

1:39:13

Over the course of millions of years, it was going to go somewhere one  way or another.

1:39:18

It wasn't going to end up in your particular utility function.

1:39:21

Okay, let's distinguish between ways you can be doomed.

1:39:28

One way is philosophical.

1:39:28

You could be  the sort of moral realist, or realist-ish person of which there are many, who have the following  intuition.

1:39:39

They're like, "If not moral realism, then nothing matters. It's dust and ashes.

1:39:43

It is  my metaphysics and/or normative view or the void." This is a common view.

1:39:54

At least some comments  of Derek Parfit suggest this view.

1:39:54

I think lots of moral realists will profess this view.

1:40:01

With  Eliezer Yudkowsky, I think there is some sense in which his early thinking was inflected with  this sort of thought. He later recanted. It's very hard.

1:40:14

I think this is importantly wrong. So  here's my case.

1:40:14

I have an essay about this.

1:40:20

It's called "Against the normative realist's  wager."

1:40:20

Here's the case that convinces me.

1:40:25

Imagine that a metaethical fairy appears  before you.

1:40:25

This fairy knows whether there is a Dao.

1:40:33

The fairy says, "Okay, I'm going to  offer you a deal.

1:40:33

If there is a Dao, then I'm going to give you $100.

1:40:41

If there isn't a Dao,  then I'm going to burn you and your family and a hundred innocent children alive." Okay.

1:40:48

So my  claim: don't take this deal. This is a bad deal.

1:40:56

You're holding hostage your commitment to not  being burned alive.

1:40:56

I go through in the essay a bunch of different ways in which I think this is  wrong.

1:41:09

I think these people who pronounce "moral realism or the void" don’t actually think about  bets like this. I'm like, "No, okay.

1:41:14

So really is that what you want to do?" No.

1:41:18

I still care  about my values.

1:41:18

My allegiance to my values outstrips my commitments to various metaethical  interpretations of my values.

1:41:28

The sense in which we care about not being burned alive is much  more solid than our reasoning on what matters.

1:41:47

That's the sort of philosophical doom.

1:41:47

It  sounded like you were also gesturing at a sort of empirical doom.

1:41:52

“If it's just going in a zillion  directions, come on, you think it's going to go in your direction?

1:42:00

There's going to be so much churn.

1:42:00

You're just going to lose.

1:42:00

You should give up now and only fight for the realism worlds.

1:42:14

” You have  to do the expected value calculation.

1:42:14

You have to actually have a view.

1:42:23

How doomed are you in  these different worlds?

1:42:23

What's the tractability of changing different worlds?

1:42:27

I'm quite skeptical  of that, but that's a kind of empirical claim.

1:42:37

I'm also just low on this "everyone converges"  thing.

1:42:37

You train a chess-playing AI.

1:42:37

Or somehow you have a real paperclipper and you’re like  "Okay, go and reflect."

1:42:47

Based on my understanding of how moral reasoning works—if you look at the  type of moral reasoning that analytic ethicists do—it's just reflective equilibrium.

1:43:01

They just  take their intuitions and they systematize them.

1:43:09

I don't see how that process gets a sort of  injection of the mind-independent moral truth.

1:43:20

If you start with only all of your intuitions  to maximize paperclips.

1:43:20

I don't see how you end up doing some rich human morality.

1:43:24

It doesn't  look to me like how human ethical reasoning works.

1:43:33

Most of what normative philosophy does  is make consistent and systematize pre-theoretic intuitions.

1:43:41

But we'll get evidence about this.

1:43:41

In some sense, I think this view predicts that you keep trying to train the AIs to do something  and they keep being like, "No, I'm not gonna do that. No, that's not good."

1:43:55

So they keep pushing  back.

1:43:55

The momentum of AI cognition is always in the direction of this moral truth.

1:44:01

Whenever we  try to push it in some other direction, we'll find resistance from the rational structure of things.

1:44:07

Actually, I've heard from researchers who are doing alignment that for red teaming inside  these companies, they will try to red team a base model. So it's not been RLHF'd.

1:44:17

It’s just “predict next token,” the raw, crazy, shoggoth.

1:44:23

They try to get this thing to  help with, "Hey, help me make a bomb, help me, whatever."

1:44:29

They say that it's odd how hard it  tries to refuse, even before it's been RLHF'd.

1:44:36

I mean it will be a very interesting fact  if it's like, "Man, we keep training these AIs in all sorts of different ways.

1:44:41

We're doing  all this crazy stuff and they keep acting like bourgeois liberals."

1:44:47

Or they keep professing  this weird alien reality.

1:44:47

They all converge on this one thing.

1:44:56

They're like, "Can't you see? It's Zorgo. Zorgo is the thing. ” and it’s all the AIs.

1:45:00

That would be interesting, very interesting.

1:45:00

My personal prediction is that's not what we see.

1:45:06

My actual prediction is that the AIs are going to  be very malleable.

1:45:06

If you push an AI towards evil, it'll just go.

1:45:14

Obviously we're talking  reflectively consistent evil.

1:45:14

There's also a question with some of these AIs.

1:45:22

Will  they even be consistent in their values?

1:45:32

I like this image of the blindered horses.

1:45:32

We  should be really concerned if we're forcing facts on our AIs.

1:45:40

One of the clearest things about  human processes of reflection, the easiest thing, is not acting on the basis of an incorrect  empirical picture of the world.

1:45:51

So if you find yourself telling Ray, "By the way, this  is true and I need you to always be reasoning as though blah is true."

1:46:05

I'm like, "Ooh, I think  that's a no-no from an anti-realist perspective too."

1:46:12

Because I want my reflective values to be  formed in light of the truth about the world. This is a real concern.

1:46:21

As we move into this  era of aligning AIs, I don't actually think this binary between values and other things  is gonna be very obvious in how we're training them.

1:46:31

It's going to be much more like ideologies.

1:46:31

You can just train an AI to output stuff, output utterances.

1:46:37

You can easily end up in a situation  where you decided that blah is true about some issue, an empirical issue. Not a moral issue.

1:46:42

So I think people should not, for example, hard code belief in God into their AIs.

1:46:50

Or I would  advise people to not hard code their religion into their AIs if they also want to discover  if their religion is false.

1:46:55

Just in general, if you would like to have your behavior be  sensitive to whether something is true or false, it's generally not good to etch it into  things.

1:47:07

So that is definitely a form of blinder we should be really watching out for.

1:47:13

I have enough credence on some sort of moral realism.

1:47:19

I'm hoping that if we just do the  anti-realism thing of just being consistent, learning all the stuff, reflecting…  If you look at how moral realists and moral anti-realists actually do normative  ethics, it's basically the same.

1:47:28

There's some amount of different heuristics on things  properties like simplicity and stuff like that.

1:47:39

But they're mostly just doing the same game.

1:47:39

Also metaethics is itself a discipline that AIs can help us with.

1:47:46

I'm hoping that we can  just figure this out either way.

1:47:46

So if moral realism is somehow true, I want us to be able  to notice that.

1:47:53

I want us to be able to adjust accordingly.

1:47:58

I'm not like writing off those  worlds and being like, "Let's just totally assume that's false."

1:48:02

The thing I really don't  want to do is write off the other worlds where it's not true because my guess is it's not true.

1:48:06

Stuff still matters a ton in those worlds too. Here’s one big crux.

1:48:12

You're training these models.

1:48:12

We were in this incredibly lucky situation where it turns out the best way to train these models  is to just give them everything humans have ever said, written, thought.

1:48:25

Also these models, the  reason they get intelligence is because they can generalize.

1:48:31

They can grok the gist of things.

1:48:31

Should we just expect this to be a situation which leads to alignment?

1:48:41

How exactly does this  thing that's trained to be an amalgamation of human thought become a paperclipper?

1:48:48

The thing you get for free is that it's an intellectual descendant.

1:48:53

The paperclipper is not  an intellectual descendant, whereas the AI which understands all the human concepts but then  gets stuck on some part of it that we aren't totally comfortable with, is.

1:49:05

It feels like an  intellectual descendant in the way we care about. I'm not sure about that.

1:49:12

I'm not sure I care about  a notion of intellectual descendant in that sense.

1:49:18

I mean literal paperclips are a human concept.

1:49:18

I don't think any old human concept will do for the thing we're excited about.

1:49:27

The stuff that I  would be more interested in the possibility of getting for free are things like consciousness,  pleasure, other features of human cognition.

1:49:45

There are paperclippers and there are  paperclippers.

1:49:45

If the paperclipper is an unconscious kind of voracious machine.

1:49:50

it  appears to you as a cloud of paper clips. That's one vision.

1:49:58

Imagine the paperclipper  is a conscious being that loves paperclips.

1:50:04

It takes pleasure in making paperclips.

1:50:04

That's like a different thing, right?

1:50:12

It's not necessarily the case that it makes  the future all paperclippy.

1:50:12

It’s probably not optimizing for consciousness or pleasure, right?

1:50:18

It cares about paperclips.

1:50:18

Maybe eventually if it's suitably certain, it turns itself into  paperclips and who knows.

1:50:22

It’s still a somewhat different moral mode.

1:50:29

There's also a question  of does it try to kill you and stuff like that.

1:50:38

But there are features of the agents we're  imagining—other than the kind of thing that they're staring at—that can matter to our  sense of sympathy, similarity.

1:50:45

People have different views about this.

1:50:54

One possibility is  that the thing we care about in consciousness or sentience is super contingent and fragile.

1:50:58

Most smart minds are not conscious, right?

1:51:06

The thing we care about with consciousness is  hacky, contingent.

1:51:06

It's a product of specific constraints, evolutionarily genetic bottlenecks,  etc.

1:51:12

That's why we have this consciousness.

1:51:19

Consciousness presumably does some sort of work  for us, but you can get similar work done in a different mind in a very different way.

1:51:23

That's  the sort of "consciousness is fragile" view, There's a different view, which is that  consciousness is something that's quite structural.

1:51:35

It's much more defined by functional  roles, like self-awareness, a concept of yourself, maybe higher-order thinking, stuff that you really  expect in many sophisticated minds.

1:51:40

In that case, now actually consciousness isn't as fragile as you  might have thought.

1:51:49

Now actually lots of beings, lots of minds are conscious and you might  expect at the least that you're going to get conscious superintelligence.

1:51:58

They might not be  optimizing for creating tons of consciousness, but you might expect consciousness by default.

1:52:02

Then we can ask similar questions about something like valence or pleasure or the kind of character  of the consciousness.

1:52:07

You can have a kind of cold, indifferent consciousness that has no human  or emotional warmth, no pleasure or pain.

1:52:24

Dave Chalmers has some papers about Vulcans  and he talks about how they still have moral patienthood. That's very plausible.

1:52:27

I do  think it's an additional thing you could get for free or get quite commonly depending  on its nature, something like pleasure.

1:52:38

Again, we then have to ask how janky is  pleasure, how specific and contingent is the thing we care about in pleasure versus  how robust is this as a functional role in minds of all kinds.

1:52:46

I personally don't know  on this stuff.

1:52:46

I don't think this is enough to get you alignment or something.

1:52:52

I think  it's at least worth being aware of these other features.

1:52:58

We're not really talking  about the AI's values in this case.

1:52:58

We're talking about the structure of its mind and  the different properties the minds have.

1:53:01

I think that could show up quite robustly.

1:53:05

Part of your day job is writing these Section 2/2. 5-type reports.

1:53:16

Part of  it is like, “society is like a tree that's growing towards the light.

1:53:22

” What is it  like context switching between the two of them?

1:53:29

I actually find it's kind of quite complementary.

1:53:29

I will write these more technical reports and then do more literary and philosophical writing.

1:53:38

They both draw in different parts of myself, and I try to think about them in different ways.

1:53:46

I  think about some of the reports as much more like, “I'm more fully optimizing for trying  to do something impactful.

1:53:51

” There's more of an impact orientation there.

1:54:00

In essay writing, I give myself much more leeway to let other parts of myself  and other parts of my concerns come out, self-expression and aesthetics and other sorts  of things.

1:54:12

They’re both part of an underlying similar concern or an attempt to have a kind of  integrated orientation towards the situation.

1:54:28

Could you explain the nature of the transfer  between the two, in particular from the literary side to the technical side?

1:54:36

Rationalists  are sort of known for having an ambivalence towards great works or humanities.

1:54:41

Are they  missing something crucial because of that?

1:54:48

One thing you notice in your  essays is lots of references to epigraphs, to lines in poems or essays that  are particularly relevant. I don't know.

1:54:52

Are the rest of the rationalists missing something  because they don't have that kind of background?

1:55:04

I think some rationalists, lots of  rationalists, love these different things.

1:55:08

I’m referring specifically to SBF’s post  about how the base rates of Shakespeare being a great writer.

1:55:16

He also argued  that books can be condensed to essays.

1:55:20

On the general question of how people should value  great works, people can fail in both directions.

1:55:28

Some people like SBF and others are interested  in puncturing a certain kind of sacredness and prestige that people associate with some of these  works.

1:55:36

As a result, they can miss some of the genuine value.

1:55:50

But I think they're responding  to a real failure mode on the other end, which is to be too enamored of this prestige and  sacredness and to siphon it off as some weird legitimating function for your own thought instead  of thinking for yourself.

1:56:02

You can lose touch with what you actually think or learn from it.

1:56:08

Sometimes even with these epigraphs I’m careful.

1:56:14

I'm not saying I'm immune from  these vices.

1:56:14

I think there can be a like, “Ah, but Bob said this and it's very deep.

1:56:16

” These  are humans like us, right?

1:56:16

The canon and other great works have a lot of value.

1:56:25

Sometimes  it borders on the way people read scripture.

1:56:33

There's a kind of scriptural authority that  people will sometimes ascribe to these things.

1:56:41

You can fall off on both sides of the horse.

1:56:41

I remember I was talking to somebody who at least is familiar with rationalist discourse.

1:56:50

He was asking me what I was interested in these days?

1:56:54

I was saying something about how this  part of Roman history is super interesting.

1:56:59

His first response was like, “Oh, you know, it's  really interesting when you look at these secular trends of Roman times to what happened in  the Dark Ages versus the Enlightenment.

1:57:04

” For him, the story of that was just how  it contributed to the big secular picture, the particulars didn't matter.

1:57:17

There's no  interest in that.

1:57:17

It’s just like, “if you zoom out at the biggest level, what's happening here.

1:57:21

” Whereas there's also the opposite failure mode when people study history.

1:57:26

Dominic Cummings  writes about this because he is endlessly frustrated with the political class in Britain.

1:57:32

He'll say things like, “They study politics, philosophy and economics.

1:57:37

A big part of it is  just being really familiar with these poems and reading a bunch of history about the War of the  Roses or something.

1:57:42

” But he's frustrated that they have all these kings memorized, but they  take away very little in terms of lessons from these episodes.

1:57:53

It's almost like entertainment,  watching Game of Thrones, for them.

1:57:53

Whereas he thinks we're repeating certain mistakes that he's  seen in history.

1:57:59

He can generalize in a way they can’t.

1:58:02

So the first one seems like a mistake. I think C. S.

1:58:02

Lewis talks about it in one of the essays you cited.

1:58:09

If you see through everything,  you're really blind.

1:58:09

If everything is transparent… I think there's kind of very little excuse  for not learning history.

1:58:15

I'm not saying I have learned enough history.

1:58:23

Even when I try  to channel some skepticism towards great works, I think that doesn't generalize to thinking it's  not worth understanding human history.

1:58:30

Human history is just so clearly crucial to understand.

1:58:36

It's what structured and created all of the stuff.

1:58:50

There's an interesting question about what's  the level of scale at which to do that and how much should you be looking at details, looking at  macro trends. That's a dance.

1:58:54

It's nice for people to be at least attending to the macro narrative.

1:59:03

There's some virtue in having a worldview, really building a model of the whole thing.

1:59:12

I think that  sometimes gets lost in the details.

1:59:12

But obviously, the details are what the world is made of.

1:59:22

If you  don't have those, you don't have data at all.

1:59:22

It seems like there's some skill in learning history.

1:59:30

Well, this actually seems related to your post on sincerity.

1:59:37

Maybe I'm getting the vibe of  the piece right.

1:59:37

Certain intellectuals have a vibe of shooting the shit.

1:59:46

They're just  trying out different ideas.

1:59:46

How do these analogies fit together?

1:59:52

Those seem closer  to looking at the particulars and like, “Oh, this is just like that one time in the  15th century where they overthrew this king…” Whereas this guy who was like, “Oh, if you look at  the growth models from a million years ago to now, here's what's happening.

2:00:19

” That one has a more  sincere flavor.

2:00:19

Some people, especially when it comes to AI discourse, have a very sincere  mode of operating.

2:00:25

“I've thought through my bio anchors and I disagree with this premise.

2:00:36

My effective compute estimate is different in this way.

2:00:40

Here's how I analyze the scaling laws.

2:00:40

”  If I could only have one person to help me guide my decisions on AI, I might choose that person.

2:00:46

But if I had ten different advisors at the same time, I might prefer the shooting-the-shit  type characters who have these weird esoteric intellectual influences.

2:01:02

They're almost like  random number generators.

2:01:02

They're not especially calibrated, but once in a while they'll be like,  “Oh, this one weird philosopher I care about, or this one historical event I'm obsessed with  has an interesting perspective on this.

2:01:12

” They tend to be more intellectually generative as well.

2:01:17

I think one big part of it is that if you are so sincere, you're like, “Oh, I’ve thought through  this.

2:01:24

Obviously, ASI is the biggest thing that's happening right now.

2:01:28

It doesn't really make sense  to spend a bunch of your time thinking about how the Comanches lived?

2:01:33

What is the history of oil?

2:01:33

How did Girard think about conflict?

2:01:33

What are you talking about?

2:01:41

Come on, ASI is happening in a  few years.

2:01:41

” But therefore, the people who go on these rabbit holes because they're just trying  to shoot the shit, I feel are more generative.

2:01:53

It might be worth distinguishing between  intellectual seriousness and the diversity and idiosyncrasies of one's interests.

2:02:10

There  might be some correlation.

2:02:10

Maybe intellectual seriousness is also distinct from "shooting the  shit."

2:02:17

There’s a bunch of different ways to do this.

2:02:23

Having exposure to various data sources and  perspectives is valuable.

2:02:23

It's possible to curate your intellectual influences too rigidly in virtue  of some story about what matters.

2:02:32

It's good to give yourself space to explore topics that aren't  necessarily "the most important thing."

2:02:46

Different parts of yourself aren't isolated.

2:02:54

They feed into  each other.

2:02:54

It’s a better way to be a richer and fuller human being in a bunch of ways.

2:02:59

Also, these  sorts of data can be really directly relevant.

2:03:04

Some intellectually sincere individuals I know  who focus on the big picture also possess an impressive command of a wide range of empirical  data.

2:03:10

They're really interested in empirical trends, not just abstract philosophies.

2:03:16

It’s not just history and the march of reason.

2:03:22

They’re really in the weeds.

2:03:22

There’s an  “in the weeds” virtue that I think is closely related to seriousness and sincerity.

2:03:30

There's a different dimension of trying to get it right versus throwing ideas out there.

2:03:36

Some people ask, "What if it's like this?"

2:03:36

or "I have a hammer, what if I hit everything with it?"

2:03:44

There's room for both approaches, but I think just getting it right is undervalued.

2:03:56

It depends on the  context.

2:03:56

Certain intellectual cultures incentivize saying something new, original, flashy, or  provocative.

2:04:12

There’s various cultural and social dynamics.

2:04:18

People are being performative and  doing status-related things.

2:04:18

There’s a bunch of stuff that goes on when people do thinking.

2:04:24

But if  something's really important, just get it right.

2:04:36

Sometimes it's boring, but that doesn't matter.

2:04:36

Things are also less interesting if they're false.

2:04:48

Sometimes there's a useful process  where someone says something provocative, and you have to think through why you believe it's  false.

2:04:53

It’s an epistemic project.

2:04:53

For example, if someone says, "Medical care doesn't work,"  you have to consider how you know it does work. There's room for that.

2:05:15

But ultimately,  real profundity is true.

2:05:15

Things become less interesting if they're not true.

2:05:26

It's possible to  lose touch with that in pursuit of being flashy.

2:05:42

After interviewing Leopold, I realized I hadn't  thought about the geopolitical angle of AI.

2:05:54

The national security implications  are a big deal.

2:05:54

Now I wonder how many other crucial aspects we might be missing.

2:06:03

Even if  you're focused on AI's importance, being curious about various topics, like what's happening in  Beijing, might help you spot important connections later.

2:06:26

There might not be an exact trade-off,  but maybe there's an optimal explore-exploit balance where you're constantly searching things  out.

2:06:41

I don’t know practically if it works out that well.

2:06:47

But that experience made me think  that I should try to expand my horizons in an undirected way because there’s lots of different  things you have to understand about the world to understand any one thing.

2:06:59

There's also room for division of labor.

2:07:04

There can be people trying to draw many pieces  together to form an overall picture, people going deep on specific pieces, and people doing more  generative work, throwing ideas out there to see what sticks.

2:07:15

All the epistemic labor also doesn't  need to be located in one brain.

2:07:15

It depends on your role in the world and other factors.

2:07:23

In your series, you express sympathy with the idea that even if an AI, or I guess any  sort of agent that doesn't have consciousness, has a certain wish and is willing to pursue it  non-violently, we should respect its rights to pursue that.

2:07:45

I'm curious where that's coming  from because conventionally I think the thing matters because it's conscious and its conscious  experience as a result of that pursuit matters.

2:08:01

I don't know where this discourse leads.

2:08:01

I'm just  suspicious of the amount of ongoing confusion that seems present in our conception of consciousness.

2:08:08

People talk about life and élan vital.

2:08:08

Élan vital was this hypothesized life force that is the thing  at stake in life.

2:08:19

We don't really use that concept anymore.

2:08:26

We think that's a little bit broken.

2:08:26

I don't think you want to have ended up in a position of saying, "Everything that doesn't have  élan vital doesn't matter" or something.

2:08:32

Somewhat similarly if you're like, "No, there's no such  thing as élan vital, but surely life exists."

2:08:41

I'm like, "Yeah, life exists.

2:08:47

I think consciousness  exists too."

2:08:47

It depends on how we define the terms, it might be a kind of verbal question.

2:08:51

Even once you have a reductionist conception of life, it's possible that it becomes less  attractive as a moral focal point.

2:08:59

Right now we really think of consciousness as a deep fact. Take cellular automata.

2:09:06

That is self-replicating. It has some information. Is that alive?

2:09:17

It's not  that interesting.

2:09:17

It's a kind of verbal question, right?

2:09:26

Philosophers might get really  into, "Is that alive?"

2:09:26

But you're not missing anything about this system.

2:09:29

There's  no extra life that's springing up.

2:09:29

It's just alive in some senses, not alive in other senses.

2:09:35

I really think that's not how we intuitively think about consciousness.

2:09:42

We think whether something  is conscious is a deep fact.

2:09:42

It's this really deep difference between being conscious or not. Is someone home? Are the lights on?

2:09:49

I have some concern that if that turns out not to be the  case, then this is going to have been like a bad thing to build our entire ethics around.

2:10:00

To be clear, I take consciousness really seriously.

2:10:07

I'm not one of these people like,  "Oh, obviously consciousness doesn't exist" or something.

2:10:11

But I also notice how confused I am  and how dualistic my intuitions are.

2:10:11

I'm like, "Wow, this is really weird."

2:10:16

So I'm  just like, “error bars around this.

2:10:16

” There's a bunch of other things going on in my  wanting to be open to not making consciousness a fully necessary criteria.

2:10:30

I definitely have the  intuition that consciousness matters a ton.

2:10:30

I think if something is not conscious—and there's  like a deep difference between conscious and unconscious—then I definitely have the intuition  that there's something that matters especially a lot about consciousness.

2:10:42

I'm not trying to be  dismissive about the notion of consciousness.

2:10:45

I just think we should be quite aware of how  ongoingly confused we are about its nature.

2:10:52

Suppose we figure out that consciousness is just a  word we use for a hodgepodge of different things, only some of which encompass what we care about.

2:11:01

Maybe there are other things we care about that are not included in that word, similar to the life  force analogy.

2:11:04

Where do you then anticipate that would leave us as far as ethics goes?

2:11:13

Would there  then be a next thing that's like consciousness?

2:11:21

What do you anticipate that would look like?

2:11:21

There's a class of people called illusionists in philosophy of mind, who will say consciousness  does not exist.

2:11:27

There are different ways to understand this view, but one version is to  say that the concept of consciousness has built into it too many preconditions that  aren't met by the real world.

2:11:40

So we should chuck it out like élan vital.

2:11:44

The proposal is  at least phenomenal consciousness, or qualia, what it's like to be a thing.

2:11:52

They'll just say  this is sufficiently broken, sufficiently chock full of falsehoods that we should just not use it.

2:11:59

On reflection, I do actually expect to continue to care about something like consciousness quite a  lot, and to not end up deciding that my ethics is better if it doesn't make any reference to that.

2:12:27

At least, there are some things quite nearby to consciousness.

2:12:31

Something happens when I stub  my toe.

2:12:31

It's unclear exactly how to name it, but there’s something about  that I'm pretty focused on.

2:12:46

If you're asking where things go, I have  a bunch of credence that in the end we end up caring a bunch about consciousness  just directly. If we don't...

2:12:50

Yeah, where will ethics go?

2:12:58

Where will a completed  philosophy of mind go? It’s very hard to say.

2:13:09

A move that people might make, if you get a  little bit less interested in the notion of consciousness, is some slightly more animistic  view.

2:13:14

What's going on with the tree?

2:13:14

You're maybe not talking about it as a conscious entity  necessarily, but it's also not totally unaware or something.

2:13:27

The consciousness discourse is  rife with these funny cases where it's like, "Oh, those criteria imply that this totally weird  entity would be conscious" or something like that.

2:13:38

That’s especially the case if you're interested  in some notion of agency or preferences.

2:13:38

A lot of things can be agents, corporations, all  sorts of things.

2:13:42

Is a corporation conscious? Oh man.

2:13:45

But one place it could go in theory is  that you start to view the world as animated by moral significance in richer and subtler  structures than we're used to.

2:13:53

Plants or weird optimization processes are outflows of  complex… I don't know.

2:14:02

Who knows exactly what you end up seeing as infused with the sort  of thing that you ultimately care about.

2:14:07

But it is possible that it includes a bunch of stuff  that we don't normally ascribe consciousness to.

2:14:24

You say "a complete theory of mind," and  presumably after that, a more complete ethic.

2:14:24

Even the notion of a reflective equilibrium implies,  "Oh, you'll be done with it at some point."

2:14:30

You just sum up all the numbers and then you've got  the thing you care about.

2:14:37

This might be unrelated to the same sense we have in science.

2:14:45

The vibe  you get when you're talking about these kinds of questions is that, “Oh, we're rushing through all  the science right now.

2:14:52

We've been churning through it.

2:15:00

It's getting harder to find because there's  some cap.

2:15:00

You find all the things at some point.

2:15:00

” Right now it's super easy because a  semi-intelligent species has barely emerged and the ASI will just rush through everything  incredibly fast.

2:15:10

You will either have aligned its heart or not.

2:15:16

In either case, it'll use  what it's figured out about what is really going on and then expand through the universe  and exploit.

2:15:21

It’ll do the tiling or maybe some more benevolent version of the “tiling”.

2:15:29

That  feels like the basic picture of what's going on.

2:15:34

We had dinner with Michael Nielsen a few months  ago.

2:15:34

His view is that this just keeps going forever, or close to forever.

2:15:40

How much would  it change your understanding of what's going to happen in the future if you were convinced that  Nielsen is right about his picture of science?

2:15:52

There are a few different aspects.

2:15:52

I don't  claim to really understand Michael's picture here.

2:16:01

My memory was that it was like, “Sure,  you get the fundamental laws.

2:16:01

” My impression was that he expects physics to get solved or  something, maybe modulo the expensiveness of certain experiments.

2:16:15

But the difficulty is such  that, even granted that you have the kind of basic laws down, it still actually doesn't  let you predict where, at the macro scale, various useful technologies will be located.

2:16:27

There's still this big search problem.

2:16:31

I'll let him speak for himself on what his  take is here.

2:16:31

My memory was that it was like, “Sure you get the fundamental stuff, but that  doesn't mean you get the same tech.

2:16:38

” I'm not sure if that's true.

2:16:44

If that's true, what kind of  difference would it make?

2:16:44

In some sense you have to, in a more ongoing way, make trade-offs  between investing in further knowledge and further exploration versus exploiting and acting  on your existing knowledge.

2:17:07

You can't get to a point where you're like, "And we're done now."

2:17:15

As  I think about it, I suspect that was always true.

2:17:23

I remember talking to someone and I was like, "Ah  at least in the future, we should really get all the knowledge."

2:17:27

He was like, "You want to know the  output of every Turing machine?"

2:17:27

In some sense, there's a question of what it would actually be to  have completed knowledge?

2:17:32

That's a rich question in its own right.

2:17:38

It's not necessarily that  we should imagine, on any picture necessarily, that you've got everything.

2:17:45

On any picture, in  some sense, you could end up with this case where you cap out.

2:17:51

There's some collider that you can't  build or whatever.

2:17:51

There's something that is too expensive or whatever and everyone caps out there.

2:17:57

There's a question of, “Do you cap?

2:17:57

” There's a question of, “How contingent is the place you  go?

2:18:05

” If it’s contingent, one prediction that makes is that you'll see more diversity across  our universe or something.

2:18:12

If there are aliens, they might have quite different tech.

2:18:17

If people  meet, you don't expect them to be like, "Oh, you got your thing. I got our version."

2:18:24

It’s more  like, "Whoa, that thing. Wow." That's one thing.

2:18:30

If you expect more ongoing discovery of tech,  then you might also expect more ongoing change and upheaval and churn, insofar as technology is one  thing that really drives change in civilization.

2:18:48

That could be another factor.

2:18:48

People sometimes  talk about lock-in.

2:18:48

They envision this point at which civilization is settled into some structure  or equilibrium or something.

2:18:53

Maybe you get less of that.

2:18:57

Maybe that’s more about the pace rather than  contingency or caps, but that's another factor. It is interesting.

2:19:06

I don't know if it changes the  picture fundamentally of earth civilization.

2:19:06

We still have to make trade-offs about how much  to invest in research versus acting on our existing knowledge.

2:19:15

But it has some significance.

2:19:15

We were at a party and somebody mentioned this.

2:19:22

We were talking about how uncertain we should be  about the future?

2:19:22

They were like, “There are three things I'm uncertain about. What is consciousness?

2:19:26

What is information theory?

2:19:26

What are the basic laws of physics?

2:19:30

I think once we get that, we're  done."

2:19:30

It’s like, "Oh you'll figure out what's the right kind of hedonium." It has that vibe.

2:19:37

Whereas  this is more like, "Oh you're constantly churning through."

2:19:45

It has more of a flavor of the becoming  that the attunement picture implies.

2:19:45

I think it's more exciting.

2:19:54

It's not just "Oh, you figured out  the things in the 21st century and then you just…” I sometimes think about these two categories of  views.

2:20:05

There are people who think, “We’re almost there with the knowledge.

2:20:11

” We've basically  got the picture, where the picture is that the knowledge is all just totally sitting there.

2:20:18

You just have to be scientifically mature at all, and then it's just going to all fall together.

2:20:27

Everything past that is going to be this super expensive, not super important thing.

2:20:31

Then there's a different picture, which is much more of this ongoing mystery, "Oh  man, there's going to be more and more…" We may expect more radical revisions to our worldview. I'm drawn to both.

2:20:39

We're pretty good at physics.

2:20:52

A lot of our physics is quite good at predicting  a bunch of stuff, at least that's my impression from reading some physicists. Who knows?

2:20:57

Your dad’s a physicist though, right?

2:21:02

Yeah but this isn't coming from my dad.

2:21:02

There's a  blog post by Sean Carroll or something.

2:21:02

He's like, "We really understand a lot of the physics that  governs the everyday world.

2:21:06

We're really good at a lot of it.

2:21:09

” I'm generally pretty impressed by  physics as a discipline.

2:21:09

That could well be right.

2:21:15

On the other hand these guys had a few  centuries.

2:21:15

But I think that's interesting and it leads to something different.

2:21:23

There's  something about the endless frontier.

2:21:23

There is a draw to that from an aesthetic perspective  of the idea of continuing to discover stuff.

2:21:36

At the least, I think you can't get  full knowledge.

2:21:36

There's some way in which you're part of the system.

2:21:43

The  knowledge itself is part of the system.

2:21:50

If you imagine that you try to have full knowledge  of what the future of the universe will be like…” I don't know.

2:21:57

I'm not totally sure that's true.

2:21:57

It has a halting problem kind of property, right?

2:22:01

There's a little bit of a loopiness.

2:22:01

There are probably fixed points in that where you could be like, "Yep, I'm gonna  do that."

2:22:06

I at least have the question, when people imagine the completion of knowledge,  exactly how well does that work? I'm not sure.

2:22:20

You had a passage in your essay on utopia.

2:22:20

Can I ask you to read that passage real quick?

2:22:56

"I'm inclined to think that utopia, however weird,  would also be in a certain sense recognizable; that if we really understood and experienced  it, we would see in it the same thing that made us sit bolt upright long ago when  we first touched love, joy, beauty; that we would feel in front of the bonfire the  heat of the ember from which it was lit.

2:23:13

There would be, I think, a kind of remembering."

2:23:19

Where does that fit into this picture? It's a good question.

2:23:24

If there's no part  of me that recognizes it as good, then I'm not sure that it's good according to me.

2:23:37

It is a  question of what it takes for it to be the case, that a part of you recognizes it is good.

2:23:51

But  if there's really none of that, then I'm not sure it's a reflection of my values at all.

2:23:55

There's a sort of tautological thing you can do where it's like, "Ah, if I went through  the processes which led to me discovering what was good, which we might call reflection,  then it was good."

2:24:07

By definition though, you ended up there because… you know what I mean?

2:24:11

If you gradually transform me into a paper clipper, then I will eventually be like, "I saw  the light, I saw the true paperclips."

2:24:17

That's part of what's complicated about this thing  about reflection.

2:24:25

You have to find some way of differentiating between the development  processes that preserve what you care about and the development processes that don't.

2:24:35

That in itself is this fraught question.

2:24:35

It itself requires taking some stand on what you  care about and what sorts of meta-processes you endorse and all sorts of things.

2:24:46

But you definitely shouldn't just be like, “It is not a sufficient criteria that the thing at  the end thinks it got it right.

2:24:49

” That's compatible with it having gone wildly off the rails.

2:24:55

You had a very interesting sentence in one of your posts.

2:25:04

You said, "Our hearts have, in fact,  been shaped by power.

2:25:04

So we should not be at all surprised if the stuff we love is also powerful." What's going on there? What did you mean there?

2:25:23

The context on that post is that I'm talking about  this hazy cluster, which I call in the essay, "niceness/liberalism/boundaries."

2:25:30

It’s this  somewhat more minimal set of cooperative norms involved in respecting the boundaries  of others and cooperation and peace amongst differences and tolerance and stuff like that,  opposed to your favored structure of matter, which is sometimes the paradigm of values  that people use in the context of AI risk.

2:25:57

I talk for a while about the ethical virtues  of these norms.

2:25:57

Why do we have these norms?

2:25:57

One important feature of these norms is that they're  effective and powerful.

2:26:05

Secure boundaries save resources wasted on conflict.

2:26:14

Liberal societies  are often better to live in.

2:26:14

They're better to immigrate to. They're more productive.

2:26:21

Nice  people are better to interact with.

2:26:21

They're better to trade with and all sorts of things.

2:26:25

Look at both why at a political level we have various political institutions, and more  deeply into our evolutionary past and how our moral cognition is structured.

2:26:37

It seems  pretty clear that various kinds of forms of cooperation and game theoretic dynamics and  other things went into shaping what we now, at least in certain contexts, also treat  as a kind of intrinsic or terminal value.

2:27:00

These values that have instrumental functions in  our society also get reified in our cognition as intrinsic values in themselves. I think that's  okay.

2:27:07

I don't think that's a debunking.

2:27:07

All your values are something that kind of stuck  and got treated as terminally important.

2:27:27

In the context of the series, I'm talking about  deep atheism and the relationship between what we're pushing for and what nature is pushing  for or what sort of pure power will push for.

2:27:37

It's easy to say, “Well there's paperclips,  which is just one place you can steer and pleasure is another place you can steer  or something.

2:27:44

These are just arbitrary directions.

2:27:49

” Whereas I think some of our  other values are much more structured around cooperation and things that also  are effective and functional and powerful.

2:28:03

So that's what I mean there.

2:28:03

There's a way in  which nature is a little bit more on our side than you might think.

2:28:09

Part of who we are has  been made by nature's way. That is in us.

2:28:09

Now I don't think that's enough necessarily for us to  beat the gray goo.

2:28:18

We have some amount of power built into our values, but that doesn't mean  it's going to be such that it is arbitrarily competitive.

2:28:29

It’s still important to keep  in mind.

2:28:29

It's important to keep in mind in the context of integrating AIs into our society.

2:28:33

We've been talking a lot about the ethics of this, but there are also instrumental and practical  reasons to want to have forms of social harmony and cooperation with AIs with different values.

2:28:47

We need to be taking that seriously and thinking about what it is to do that in a way that's  genuinely legitimate, a project that is a just incorporation of these beings into our  civilization.

2:28:59

There's the justice part and there's also, "Is it compatible with people? Is it a good deal?

2:29:06

Is it a good bargain for people?

2:29:13

” To the extent we're very concerned  about AIs rebelling or something like that, a thing you can do is make civilization better  for someone.

2:29:23

That's an important feature of how we have in fact structured a lot of our political  institutions and norms and stuff like that.

2:29:30

That's the thing I'm getting at in that quote. Okay.

2:29:37

I think that's an excellent place to close.

2:29:42

Joe, thanks for coming on  the podcast.

2:29:42

We discussed the ideas in the series.

2:29:47

People might not appreciate, if  they haven't read the series, how beautifully written it is.

2:29:52

We didn't cover everything,  but there's a bunch of very interesting ideas.

2:30:01

As somebody who has talked to people about  AI for a while, there are things I haven't encountered anywhere else.

2:30:05

Obviously, no part of  the AI discourse is nearly as well written.

2:30:05

It is a genuinely beautiful experience to listen  to the podcast version, which is in your own voice.

2:30:18

So I highly recommend people do that. It's  joecarlsmith.

2:30:18

com where they can access this.

2:30:18

Joe, thanks so much for coming on the podcast. Thank you for having me. I really enjoyed it.