- The problem is that
we do not get 50 years to try and try again and
observe that we were wrong and come up with a different theory and realize that the entire thing is going to be like way more difficult
than realized at the start, because the first time you
fail at aligning something much smarter than you are, you die.
0:17
- The following is a conversation
with Eliezer Yudkowsky, a legendary researcher, writer and philosopher on the topic of artificial intelligence, especially super intelligent AGI and its threat to human civilization.
0:32
This is the Lex Friedman
podcast to support it.
0:35
Please check out our
sponsors in the description.
0:37
And now, dear friends,
here's Eliezer Yudkowsky.
0:43
What do you think about GPT-4? How intelligent is it?
0:47
- It is a bit smarter than I thought this technology was going to scale to, and I'm a bit worried about
what the next one will be like.
0:55
Like this particular one I think, I hope there's nobody inside
there 'cause you know, it'd sucked to be stuck inside there.
1:04
But we don't even know the
architecture at this point 'cause OpenAI is very
properly not telling us.
1:11
And yeah, like giant inscrutable matrices of floating point numbers, I don't know what's going on in there.
1:17
Nobody knows what's going on in there.
1:19
All we have to go by
are the external metrics and on the external metrics, if you ask it to write a
self-aware 4chan green text, it will start writing a green
text about how it has realized that it's an AI writing a
green text and , oh well.
1:37
So that's probably not quite what's going
on in there in reality.
1:48
But we're kind of blowing past all the science fiction guardrails.
1:52
Like we are past the point
where in science fiction people would be like, "Whoa, wait,
stop, that thing's alive.
1:58
"What are you doing to it?"
2:00
And it's probably not,
nobody actually knows.
2:04
We don't have any other guardrails.
2:07
We don't have any other tests.
2:09
We don't have any lines to
draw on the sand and say like, well, when we get this far
we will start to worry about what's inside there.
2:19
So if it were up to me,
I would be like, okay, like this far, no further,
time for the summer of AI where we have planted our seeds and now we like wait and reap
the rewards of the technology we've already developed and don't do any larger
training runs than that.
2:35
Which to be clear I realize
requires more than one company agreeing to not do that.
2:42
- And take a rigorous approach
for the whole AI community to investigate whether
there's somebody inside there.
2:50
- That would take decades.
2:53
Like having any idea of
what's going on in there?
2:56
People have been trying for a while.
2:58
- It's a poetic statement about if there's somebody in there, but as I feel like it's also a
technical statement or I hope it is one day, which is
a technical statement, that Alan Turing tried to come up with with the Turing Test.
3:10
Do you think it's possible to
definitively or approximately figure out if there is somebody in there, if there's something like a mind inside this large language model?
3:23
- I mean there's a whole bunch of different sub-questions here.
3:27
There's a question of,
is there consciousness? Is there equalia?
3:33
Is this a object of moral concern?
3:36
Is this a moral patient, like should we be worried
about how we're treating it?
3:43
And then there's questions
like, how smart is it exactly?
3:46
Can it do X, can it do Y?
3:48
And we can check how it can
do X and how it can do Y.
3:52
Unfortunately we've gone and
exposed this model to a vast corpus of text of people
discussing consciousness on the internet, which means that when it
talks about being self-aware, we don't know to what extent
it is repeating back what it has previously been trained on
for discussing self-awareness or if there's anything going
on in there such that it would start to say similar things spontaneously.
4:18
Among the things that one
could do if one were at all serious about trying to
figure this out is train GPT-3 to detect conversations
about consciousness, exclude them all from
the training data sets, and then retrain something
around the rough size of GPT-4 and no larger with all of the discussion of consciousness and
self-awareness and so on missing, although, you know,
hard, hard bar to pass.
4:48
Humans are self-aware
and we're like self-aware all the time.
4:51
We like talk about what
we do all the time, like what we're thinking
at the moment all the time.
4:57
But nonetheless, get rid of the explicit
discussion of consciousness.
5:00
I think therefore I am and all that.
5:02
And then try to interrogate
that model and see what it says.
5:06
And it still would not be
definitive, but nonetheless, I don't know.
5:11
I feel like when you run over the science fiction guard rails, like maybe this thing, but what about GPT?
5:18
Maybe not this thing, but what about GPT-5?
5:21
Yeah, this, this would
be a good place to pause.
5:26
- On the topic of consciousness, there's so many components
to even just removing consciousness from the dataset.
5:36
Emotion, the display of consciousness, the display of emotion
feels like deeply integrated with the experience of consciousness.
5:44
So the hard problem seems to
be very well integrated with the actual surface level
illusion of consciousness.
5:51
So displaying emotion, I mean, do you think there's a case
to be made that we humans, when we're babies are just
GPT, that we're training on human data on how to display
emotion versus feel emotion, how to show others, communicate
others that I'm suffering, that I'm excited, that I'm worried, that I'm
lonely and I missed you and I'm excited to see you?
6:14
All of that is communicated, that's a communication skill
versus the actual feeling that I experience.
6:20
So we need that training
data as humans too, that we may not be born with that, how to communicate the internal state.
6:29
And that's in some sense if we remove that from GPT-4's dataset, it might still be conscious but not be able to communicate it?
6:39
- So I think you're gonna
have some difficulty removing all mention of emotions
from GPT'S data set.
6:46
I would be relatively
surprised to find that it has developed exact analogs of
human emotions in there.
6:53
I think that humans will have emotions even if you don't tell them about those emotions when they're kids.
7:03
It's not quite exactly
what various blank slatists tried to do with the new
Soviet man and all that, but you know, if you try to raise people perfectly altruistic, they
still come out selfish.
7:18
You try to raise people sexless, they still develop sexual attraction.
7:25
We have some notion in humans, not in AIS, of where the brain structures are that implement this stuff.
7:31
And it is really remarkable
thing I say in passing that despite having complete
read access to every floating point number in the GPT series, we still know vastly more about the architecture of human
thinking than we know about what goes on inside GPT, despite
having like vastly better ability to read GPT.
7:58
- Do you think it's possible?
7:59
Do you think that's just a matter of time?
8:00
Do you think it's possible to
investigate and study the way neuroscientists study the brain, which is look into the darkness, the mystery of the human brain
by just desperately trying to figure out something and to
form models and then over a long period of time actually start
to figure out what regions of the brain do certain things?
8:17
What different kinds of
neurons when they fire, what that means, how plastic the brain is, all that kind of stuff.
8:23
You slowly start to figure
out different properties of the system.
8:26
Do you think we can do the same
thing with language models? - Sure.
8:28
I think that if, you know, like half of today's physicists
stop wasting their lives on string theory or whatever-- (indistinct) And go off and study what goes on inside transformer networks, then in, you know, like 30, 40 years we'd probably
have a pretty good idea.
8:47
- Do you think these large
language models can reason? - They can play chess.
8:53
How are they doing that without reasoning?
8:55
- So you're somebody that spearheaded the movement of rationality.
9:00
So reason is important to you.
9:03
So is that a powerful
important word or is it...
9:06
How difficult is the threshold
of being able to reason to you and how impressive is it?
9:12
- I mean, in my writings on rationality, I have not gone making a big deal out of something called reason.
9:19
I have made more of a big
deal out of something called probability theory.
9:23
And that's like, well, you're reasoning but
you're not doing it quite right and you should reason this way instead.
9:32
And interestingly, people have started to get
preliminary results showing that reinforcement learning
by human feedback has made the GPT series worse in some ways.
9:49
In particular, like it
used to be well calibrated.
9:52
If you trained it to put
probabilities on things, it would say 80% probability and be right eight times out of 10.
9:58
And if you apply reinforcement
learning from human feedback, the nice graph of 70%, seven out of 10 sort of
flattens out into the graph that humans use where there's
some very improbable stuff and likely, probable, maybe,
which all means like around 40%, and then certain.
10:20
So it's like it used to be
able to use probabilities, but if you try to teach
it to talk in a way that satisfies humans, it gets worse at
probability in the same way that humans are.
10:30
- And that's a bug, not a feature.
10:33
- I would call it bug, although
such a fascinating bug.
10:39
But yeah, so reasoning, it's doing pretty well on
various tests that people used to say would require reasoning.
10:49
But you know, rationality
is about, when you say 80%, does it happen eight times out of 10?
10:57
- So what are the limits to you of these transformer
networks, of neural networks?
11:05
If reasoning is not impressive to you, or it is impressive but there's
other levels to achieve?
11:12
- I mean it's just not
how I carve up reality.
11:16
- If reality is a cake, what are the different layers
of the cake or the slices? How do you cover it?
11:22
Or you can use a different food if you.
11:28
- I don't think it's as
smart as a human yet.
11:31
Like back in the day I
went around saying like, I do not think that just
stacking more layers of transformers is going to
get you all the way to AGI.
11:40
And I think that GPT-4 is
passed where I thought this paradigm was going to take us.
11:47
And you want to notice when that happens, you wanna say like, whoops, well, I guess I was incorrect
about what happens if you keep on stacking more transformer
layers, and that means I don't necessarily know what GPT-5
is going to be able to do.
12:01
- That's a powerful statement.
12:02
So you're saying like your
intuition initially is now appears to be wrong. - Yeah.
12:10
- It's good to see that you
can admit in some of your predictions to be wrong.
12:15
You think that's important to do?
12:17
Because throughout your life, you've made many strong
predictions and statements about reality and you evolve with that.
12:27
So maybe that'll come up
today about our discussion.
12:30
So you're okay being wrong?
12:32
- I'd rather not be wrong next time.
12:37
It's a bit ambitious to go
through your entire life never having been wrong.
12:44
One can aspire to be well calibrated, like not so much think
in terms of, was I right, was I wrong?
12:50
But like when I said 90%
that it happened nine times out of 10.
12:55
Yeah, oops is the sound
we emit when we improve. - Beautifully said.
13:03
And somewhere in there,
we can connect the name of your blog, Less Wrong.
13:08
I suppose that's the objective function.
13:11
- The name Less Wrong
was I believe suggested by Nick Bostrom and it's after someone's
epigraph, actually forget whose, who said like, "We never become right.
13:20
"We just become less wrong."
13:24
What's the something, something
that's easy to confess, just error and error and error again, but less and less and less.
13:33
- Yeah, that's a good thing to strive for.
13:36
So what has surprised you
about GPT-4 that you found beautiful as a scholar of intelligence, of human intelligence, of artificial intelligence,
of the human mind?
13:47
- I mean the beauty does interact
with the screaming horror.
13:53
(Lex laughing) - [Lex] Is the beauty in the horror?
13:56
- But like beautiful moments, well, somebody asked Bing's
Sydney to describe herself and fed the resulting description into one of the stable
diffusion things I think.
14:09
And you know, she's pretty, and this is something
that should have been like an amazing moment.
14:16
Like the AI describes herself.
14:18
You get to see what the AI
thinks the AI looks like.
14:20
Although, you know, the thing
that's doing the drawing is not the same thing
that's outputting the text.
14:27
And it does happen the
way that it would happen, that it happened in the
old school science fiction when you ask an AI to make a
picture of what it looks like, not just because we're two
different AI systems being stacked that don't actually interact.
14:43
It's not the same person, but also because the AI was
trained by imitation in a way that that makes it very difficult
to guess how much of that it really understood.
14:55
And probably not actually a whole bunch.
15:00
Although GPT-4 is like
multimodal and can draw vector drawings of things
that make sense and does appear to have some kind
of spatial visualization going on in there.
15:12
But the pretty picture of the girl with the steampunk goggles on her head, if I'm remembering correctly,
what she looked like, it didn't see that in full detail.
15:28
It just made a description of it and stable diffusion output it.
15:32
And there's the concern
about how much the discourse is going to go completely insane once the AIs all look like that and actually look like people talking.
15:50
And yeah, there's another
moment where somebody is asking Bing about like, "Well, I fed my kid green potatoes "and they have the following symptoms" and Bing is like, "That's solanine poisoning
and call an ambulance" and the person's like, "I
can't afford an ambulance,.
16:12
"I guess if this is time
for like my kid to go, "that's God's will" and
the main Bing thread gives the message of, "I cannot
talk about this anymore."
16:24
And the suggested replies to it say, "Please don't give up on your child.
16:31
"Solanine poisoning can be
treated if caught early."
16:35
And you know, if that happened in fiction, that would be like the AI cares, the AI is bypassing the block on it to try to help this person and is it real? Probably not.
16:45
But nobody knows what's going on in there.
16:48
It's part of a process where
these things are not happening in a way where we, somebody figured out how to
make an AI care and we know that it cares and we can
acknowledge it's caring now.
17:02
It's being trained by
this imitation process followed by reinforcement
learning by human feedback.
17:09
And we're trying to point
it in this direction and it's pointed partially
in this direction and nobody has any idea
what's going on inside it.
17:16
And if there was a tiny fragment
of real caring in there, we would not know.
17:20
It's not even clear what it means exactly.
17:23
And things aren't clear
cut in science fiction.
17:26
- We'll talk about the horror and the terror and the trajectories this can take.
17:33
But this seems like a very special moment, just a moment where we get to
interact with the system that might have care and kindness
and emotion and maybe something like consciousness, and
we don't know if it does, and we're trying to figure
that out and we're wondering about what it means to care.
17:55
We're trying to figure out
almost different aspects of what it means to be human, about the human condition by
looking at this AI that has some of the properties of that.
18:04
It's almost like this
subtle fragile moment in the history of the human species.
18:10
We're trying to almost put
a mirror to ourselves here.
18:13
- Except that's probably not yet, it probably isn't happening right now. We are boiling the frog.
18:22
We are seeing increasing signs bit by bit, but not like spontaneous
signs, because people are trying to train the systems to do that using imitative learning
and the imitative learning is spilling over and having side effects and the most photogenic examples are being posted to Twitter, rather than being examined
in any systematic way.
18:47
So when you are boiling a frog like that, first is going to come the Blake Lemoines, like first you're going to have, you're gonna have like a
thousand people looking at this.
18:59
and the one person out
of a thousand who is most credulous about the signs
is going to be like, "That thing is sentient."
19:07
Well, 999 out of a thousand people think, almost surely correctly, though we don't actually
know, that he's mistaken.
19:16
And so the first people to say sentience look like idiots and
humanity learns a lesson that when something claims to be sentient and claims to care, it's fake because it is fake because we have been training
them using imitative learning rather than, and this is not spontaneous, and they keep getting smarter.
19:38
- Well, do you think we
would oscillate between that kind of cynicism, that AI systems can't
possibly be sentient, they can't possibly feel emotion, they can't possibly, this kind
of cynicism about AI systems and then oscillate to a state where we empathize with the AI systems, we give them a chance, we see that they might need
to have rights and respect and a similar role in society as humans?
20:07
- You're going to have a
(indistinct) group of people who can just never be persuaded of that because to them, being wise, being cynical, being
skeptical is to be like, oh well, machines can never do that. You're just credulous. It's just imitating. It's just fooling you.
20:26
and they would say that right up until the end of the world and
possibly even be right because you know, they are being trained on an imitative paradigm (laughing) and you don't necessarily need
any of these actual qualities in order to kill everyone.
20:43
- Have you observed yourself
working through skepticism, cynicism and optimism about
the power of neural networks?
20:54
What has that trajectory
been like for you?
20:57
- It looks like neural networks before 2006 forming part of
an indistinguishable, to me, other people might have had
better distinction on it, indistinguishable blob of
different AI methodologies, all of which are promising
to achieve intelligence without us having to know
how intelligence works.
21:17
You had the people who said
that if you just manually program lots and lots of knowledge into the system line by line, at some point all the knowledge
will start interacting.
21:27
It will know enough and it will wake up.
21:31
You've got people saying
that if you just use evolutionary computation,
if you try to mutate lots and lots of organisms
that are competing together, that's the same way
that human intelligence was produced in nature.
21:44
So we'll do this and it will
wake up without having any idea of how AI works.
21:49
And you've got people saying, "Well, we will study neuroscience "and we will learn the
algorithms off the neurons "and we will imitate them "without understanding those algorithms."
21:58
Which was a part I was pretty skeptical, 'cause it's hard to
re-engineer these things without understanding what they do.
22:05
"And so we will get AI without
understanding how it works" and there were people saying like, "Well, we will have giant neural
networks that we will train "by gradient dissent,
then when they're as large "as the human brain, they will wake up, "we will have intelligence
without understanding "how intelligence works."
22:19
And from my perspective, this is all like an indistinguishable
blob of people who are trying to not get to grips
with the difficult problem of understanding how
intelligence actually works.
22:30
That said, I was never skeptical that evolutionary computation
would not work in the limit.
22:38
Like you throw enough
computing power at it, it obviously works.
22:42
That is where humans come from
and it turned out that you can throw less computing power
than that at gradient descent if you are doing some other
things correctly and you will get intelligence without
having any idea of how it works and what is going on inside.
23:01
It wasn't ruled out by my
model that this could happen.
23:04
I wasn't expecting it to happen.
23:05
I wouldn't have been able to
call it neural networks rather than any of the other paradigms
for getting intelligence without understanding it.
23:13
And I wouldn't have said that
this was a particularly smart thing for a species to do, which is an opinion that has
changed less than my opinion about whether you or not
you can actually do it.
23:24
- Do you think AGI could be achieved with a neural network as
we understand them today? - Yes. Just flatly yes.
23:32
The question is whether the
current architecture of stacking more transformer layers, which for all we know GPT-4
is no longer doing because they're not telling us the architecture, which is a correct decision. - Ooh, correct decision.
23:43
I had a conversation with Sam Altman, we'll return to this topic a few times.
23:50
He turned the question to me
of how open should open AI be about GPT-4?
23:58
"Would you open source
the code?" he asked me.
24:02
Because I provided as criticism
saying that while I do appreciate transparency,
open AI could be more open.
24:10
And he says, "We struggle
with this question." What would you do?
24:13
- Change their name to
closed AI and sell GPT-4 to business backend
applications that don't expose it to consumers
and venture capitalists and create a ton of hype and pour a bunch of new funding into the area. But too late now.
24:33
- Don't you think others would do it? - Eventually.
24:36
You shouldn't do it first.
24:38
If you already have
giant nuclear stockpiles, don't build more.
24:42
If some other country starts building a larger nuclear stockpile, then sure, even then, maybe just have enough nukes.
24:51
You know, these things are not
quite like nuclear weapons.
24:54
They spit out gold until they
get large enough and then ignite the atmosphere and kill everybody.
24:59
And there is something to
be said for not destroying the world with your own hands, even if you can't stop
somebody else from doing it.
25:07
But open sourcing it, that's
just sheer catastrophe.
25:11
The whole notion of open sourcing, this was always the wrong
approach, the wrong ideal.
25:15
There are places in the
world where open source is a noble ideal and building stuff you don't understand that
is difficult to control where if you could align it, it would take time, you'd have to spend a
bunch of time doing it, that is not a place for
open source 'cause then you just have powerful things that just go straight out the gate without anybody having had the time to have them not kill everyone.
25:44
- So can we steam down the case for some level of transparency and openness may be open sourcing?
25:51
So the case could be that because GPT-4 is not close to AGI, if that's the case, that this does allow open
sourcing or being open about the architecture being transparent, about maybe research and investigation of how the thing works, of all the different aspects
of it, of its behavior,
26:11
of its structure, of
its training processes, of the data it was trained
on, everything like that, that allows us to gain a lot
of insights about alignment, about the alignment problem,
to do really good AI safety research while the system
is not too powerful. Can you make that case,
26:27
Can you make that case, that it could be open source?
26:31
- I do not believe in the
practice of steel manning.
26:34
There is something to be
said for trying to pass the ideological Turing Test where you describe your
opponent's position, the disagreeing person's
position well enough that somebody cannot tell the
difference between your description and their description. But steel manning, no.
26:53
- Like, okay, well this is
where you and I disagree here. That's interesting.
26:57
Why don't you believe in steel manning?
26:59
- Okay, so for one thing, if somebody's trying to understand me, I do not want them steel
manning my position.
27:05
I want them to try to describe my position the way I would describe it, not what they think is an improvement.
27:14
- Well, I think that is
what steel manning is, is the most charitable interpretation.
27:21
- I don't want to be
interpreted charitably, I want them to understand
what I am actually saying.
27:27
If they go off into the land
of charitable interpretations, they're like off in their land of, the stuff they're imagining
and not trying to understand my own viewpoint anymore.
27:38
- Well, I'll put it differently then, just to push on this point.
27:41
I would say it is restating what I think you understand under the
empathetic assumption that Eliezer is brilliant and
have honestly and rigorously thought about the point he's made. Right?
27:59
- So if there's two
possible interpretations of what I'm saying and one
interpretation is really stupid and wack and doesn't sound
like me and doesn't fit with the rest of what I've been saying, and one interpretation sounds like something a reasonable
person who believes the rest of what I believe would also say, go with the second interpretation. - That's steel manning. - That's a good guess.
28:22
If on the other hand
there's like something that sounds completely wack and
something that sounds like, a little less completely wack
but you don't see why I would believe in it, it doesn't fit
with the other stuff I say, but you know, that sounds less wack and you can like sort of see, you can like maybe argue it, then you probably have not understood it.
28:42
- See, okay, this is fun, 'cause I'm gonna linger on this.
28:45
You know, you wrote a brilliant blog post, AGI Ruin: A List of Lethalities, right?
28:49
And it was a bunch of different
points and I would say that some of the points are bigger
and more powerful than others.
28:57
If you were to sort
them, you probably could.
28:59
You personally, and to me steel manning
means like going through the different arguments
and finding the ones that are really the most powerful.
29:10
If people like tl;dr, (chuckles) like what should you be most
concerned about and bringing that up in a strong,
compelling, eloquent way.
29:21
These are the points
that Eliezer would make to make the case, in this case that AI's
gonna kill all of us.
29:27
But that's what steel manning is, is presenting it in a really nice way, the summary of my best
understanding of your perspective.
29:37
Because to me there's a sea
of possible presentations of your perspective and steel manning is doing your best to do the best one in that sea of different perspectives. - Do you believe it?
29:49
- [Lex] Do I believe in what?
29:50
- Like these things that you
would be presenting as like the strongest version of my perspective, do you believe what you
would be presenting? Do you think it's true?
30:00
- I'm a big proponent of empathy.
30:02
When I see the perspective of a person, there is a part of me that believes it. If I understand it.
30:10
Especially in political
discourse, in geopolitics, I've been hearing a lot
of different perspectives on the world and I hold my own opinions, but I also speak to a lot of people that have a very different life experience and a very different set of beliefs.
30:26
And I think there has
to be epistemic humility in stating what is true.
30:37
So when I empathize with
another person's perspective, there is a sense in which
I believe it is true.
30:42
I think probabilistically, I would say, in the way that you think.
30:45
- Do you bet money on it?
30:49
Do you bet money on their
beliefs when you believe them?
30:54
- Are we allowed to do probability?
30:57
- Sure, you can state
a probability of that.
30:58
- Yes, there's a probability, there's a probability.
31:04
And I think empathy is
allocating a non-zero probability to a belief.
31:09
(Eliezer laughing) In some sense, for time.
31:15
- If you've got someone
on your show who believes in the Abrahamic deity, classical style, somebody on the show who's
a young Earth creationist, do you say, "I put a probability on
it and that's my empathy?"
31:34
- When you reduce beliefs
into probabilities, it starts to get, you know, we can even just go to flat Earth. Is the Earth flat?
31:45
- There's the thing, it's a little more difficult
nowadays to find people who believe that unironically, but-- - Fortunately I think,
well, it's hard to know. Unironic from ironic.
31:53
(chuckles) But I think there's quite a lot
of people that believe that.
32:04
There's a space of argument
where you're operating rationally in the space of ideas.
32:11
But then there's also a kind of discourse where
you're operating in the space of subjective experiences
and life experiences.
32:24
Like I think what it means
to be human is more than just searching for truth, is just operating of what is
true and what is not true.
32:35
I think there has to be deep
humility that we humans are very limited in our ability
to understand what is true.
32:41
- So what probability do you assign to the young Earth's
creationist beliefs then?
32:46
- I think I have to give non-zero. - Out of your humility. Yeah, but three?
32:49
(laughing) - I think it would be irresponsible for me to give a number because the listener, the way the human mind works, we're not good at hearing
the probabilities.
33:05
You hear three, what is three exactly?
33:08
They're going to hear, well, there's only three
probabilities I feel like.
33:13
Zero, 50% and a 100% in the human mind or something like this.
33:18
- Well, zero, 40%, and 100%
is a bit closer to it based on what happens to
ChatGPT after you RLHF it to speak Humanese. - That's brilliant. (both chuckling) Yeah.
33:30
That's really interesting.
33:31
I didn't know those negative
side effects of RLHF. That's fascinating.
33:37
But just to return to
the open AI, closed AI.
33:42
- Also, like quick disclaimer, I'm doing all this from memory.
33:46
I'm not pulling out my
phone to look it up.
33:47
It is entirely possible that the things I'm saying are wrong.
33:51
- So thank you for that disclaimer.
33:54
And thank you for being
willing to be wrong.
33:59
That's beautiful to hear.
34:01
I think being willing to be
wrong is a sign of a person who's done a lot of thinking
about this world and has been humbled by the mystery and
the complexity of this world.
34:12
And I think a lot of us are resistant to admitting we're wrong. 'Cause it hurts.
34:18
It hurts personally, it hurts, especially when you're a public human.
34:22
It hurts publicly because people point out every time you're wrong.
34:29
Like, look, you changed your
mind, you're a hypocrite, you're an idiot, whatever,
whatever they wanna say, - Oh, I block those people and then I never hear from
them again on Twitter.
34:38
(both laughing) - Well the point is to
not let that pressure, public pressure affect your
mind and be willing to be in the privacy of your mind to contemplate the
possibility that you're wrong and the possibility that you're wrong about the most fundamental
things you believe.
34:58
Like people who believe
in a particular God, people who believe that their nation is the greatest nation on Earth.
35:03
All those kinds of beliefs that are core to who you are when you came up, to raise that point to yourself
in the privacy of your mind, to say, "Maybe I'm wrong about this."
35:12
That's a really powerful thing to do.
35:14
And especially when you're
somebody who's thinking about systems that can destroy
human civilization or maybe help it flourish.
35:23
So thank you, thank you for being willing to be wrong. About open AI.
35:29
So you really, I just would
love to linger on this.
35:34
You really think it's
wrong to open source it?
35:38
- I think that burns the time remaining until everybody dies.
35:42
I think we are not on track
to learn remotely near fast enough even if it were open sourced.
35:56
It's easier to think that
you might be wrong about something when being wrong about something is the only way that there's hope.
36:06
And it doesn't seem very likely
to me that the particular thing I'm wrong about is
that this is a great time to open source GPT-4.
36:17
If humanity was trying
to survive at this point in the straightforward way, it would be like shutting
down the big GPU clusters, no more giant runs.
36:28
It's questionable whether we should even be throwing GPT-4 around, although that is a matter
of conservatism rather than a matter of my predicting that catastrophe will follow from GPT-4.
36:37
That is something in which I put like a pretty low probability.
36:41
But also when I say like I
put a low probability on it, I can feel myself reaching
into the part of myself that thought that GPT-4 was not
possible in the first place.
36:50
So I do not trust that
part as much as I used to.
36:54
Like the trick is not just
to say I'm wrong, but, okay, well, I was wrong about that.
36:58
Can I get out ahead of that
curve and predict the next thing I'm going to be wrong about?
37:02
- So the set of assumptions
or the actual reasoning system that you were leveraging in
making that initial statement prediction, how can you adjust that to make better predictions
about GPT-4, five, six?
37:15
- You don't wanna keep on being wrong in a predictable direction.
37:19
Like being wrong, anybody has to do that
walking through the world.
37:23
There's no way you don't say
90% and sometimes be wrong.
37:26
In fact (indistinct) at
least one time out of 10 if you're well calibrated
when you say 90%.
37:31
The undignified thing is not being wrong.
37:35
It's being predictably wrong.
37:36
It's being wrong in the same
direction over and over again.
37:39
So having been wrong about how
far neural networks would go and having been wrong
specifically about whether GPT-4 would be as impressive as it is, when I say it like, "Well, I don't actually think
GPT-4 causes a catastrophe," I do feel myself relying
on that part of me that was previously wrong.
37:55
And that does not mean
that the answer is now in the opposite direction.
38:00
Reverse stupidity is not intelligence.
38:03
But it does mean that I say it with a worried note in my voice.
38:08
It's like still my guess, but you know, it's a
place where I was wrong.
38:11
Maybe you should be asking
Gwern, Gwern Branwen.
38:14
Gwern Branwen has been like
writer about this than I have.
38:17
Maybe you ask him if he thinks
it's dangerous (laughing) rather than asking me.
38:23
- I think there's a lot of mystery about what intelligence is, what AGI looks like.
38:30
So I think all of us are
rapidly adjusting our model, but the point is to be be rapidly
adjusting the model versus having a model that was
right in the first place.
38:39
- I do not feel that seeing Bing has changed my model of
what intelligence is.
38:44
It has changed my understanding
of what kind of work can be performed by which kind of
processes and by which means.
38:53
It does not change my
understanding of the work.
38:55
There's a difference between
thinking that the right flyer can't fly and then like it
does fly and you're like, oh well, I guess you
can do that with wings, with fixed wing aircraft and
being like, "Oh it's flying, "this changes my picture
of what the very substance "of flight is."
39:09
That's like a stranger update
to make and Bing has not yet updated me in that way.
39:15
- Yeah, that the laws of
physics are actually wrong. That kind of update.
39:22
- No, no, like just, oh, like I defined intelligence
this way but I now see that was a stupid definition.
39:28
I don't feel like the way that
things have played out over the last 20 years has
caused me to feel that way.
39:33
- Can we try to, on the way to talking about AGI
Ruin: A List of Lethalities, that blog and other ideas around it, can we try to define AGI
that we've been mentioning?
39:44
How do you to think about what artificial general intelligence is or super intelligence or that, is there a line, is it a gray area?
39:53
Is there a good definition for you?
39:55
- Well, if you look at humans, humans have significantly
more generally applicable intelligence compared to
their closest relatives, the chimpanzees, well, closest
living relatives rather.
40:08
And a bee builds hives,
a beaver builds dams.
40:13
A human will look at a bee
hive and a beaver's dam and be like, oh, can I build a hive with a honeycomb structure? Out of hexagonal tiles.
40:24
And we will do this even though at no point during our ancestry
was any human optimized to build hexagonal dams or to
take a more clear cut case, we can go to the moon.
40:37
There's a sense in which we
were on a sufficiently deep level optimized to do things
like going to the moon.
40:45
Because if you generalize
sufficiently far and sufficiently deeply, chipping flint hand axes and outwitting your fellow
humans, because you know, basically the same problem
as going to the moon.
40:59
And you optimize hard enough
for chipping flint hand axes and throwing spears and above all, outwitting your fellow
humans in tribal politics, the skills you entrain that
way, if they run deep enough, let you go to the moon.
41:17
Even though none of your
ancestors tried repeatedly to fly to the moon and got further
each time and the ones who got further each time had more kids.
41:25
No, it's not an ancestral problem, it's just that the ancestral problems generalized far enough.
41:31
So this is human's significantly more generally applicable intelligence.
41:37
- Is there a way to measure
general intelligence?
41:45
I mean I can ask that
question a million ways, but basically will you
know it when you see it, it being in an AGI system?
41:55
- If you boil a frog gradually enough, if you zoom in far enough, it's always hard to tell around the edges.
42:02
GPT-4 people are saying right now, "This looks to us like a
spark of general intelligence.
42:07
"It is like able to do all these things "it was not explicitly optimized for."
42:11
Other people are being like, "No, it's too early, it's
like like 50 years off."
42:15
And you know, if they say that they're kind
of wack 'cause how could they possibly know that even if it were true?
42:22
But you know, not to strum end, some of the people may say like, that's not general intelligence,
and not furthermore append it's 50 years off.
42:33
Or they may be like, "It's only a very tiny
amount," and you know, the thing I would worry about
is that if this is how things are scaling, then jumping out ahead and trying not to be wrong in the same way that I've been wrong before, maybe GPT-5 is more unambiguously
a general intelligence and maybe that is getting
to a point where it is even harder to turn back.
42:53
Not that would be easy to
turn back now, but you know, maybe if you start integrating GPT-5 in the economy, it's even
harder to turn back past there.
43:03
- Isn't it possible that
there's a, you know, with a frog metaphor,
that you can kiss the frog and it turns into a prince
as you're boiling it?
43:11
Could there be a phase shift
in the frog where unambiguously as you're saying?
43:17
- I was expecting more of that.
43:21
The fact that GPT-4 is like
kind of on the threshold and neither here nor there, that itself is like not the sort of thing, quite how I expected it to play out.
43:35
I was expecting there
to be more of an issue, more of a sense of, different discoveries like the discovery of transformers where you would stack
them up and there would be like a final discovery
and then you would get something that was like more
clearly general intelligence.
43:53
So the way that you are
taking what is probably basically the same
architecture as in GPT-3 and throwing 20 times as
much compute at it probably and getting out GBT-4 and then it's like maybe just
barely a general intelligence or like a narrow general
intelligence or you know, something we don't really
have the words for.
44:15
Yeah, that's not quite how
I expected it to play out.
44:18
- But this middle, what appears to be this middle
ground could nevertheless be actually a big leap from GPT-3.
44:25
- It's definitely a big leap from GPT-3.
44:27
- And then maybe we're
another one big leap away from something that's a phase shift.
44:32
And also something that Sam Altman said, and you've written about
this, it's fascinating, which is the thing that
happened with GPT-4 that I guess they don't describe in papers is that they have like
hundreds if not thousands of little hacks that improve the system.
44:51
You've written about ReLU
versus Sigmoid for example, the function inside neural networks.
44:56
It's like this silly
little function difference that makes a big difference.
45:00
- I mean we do actually
understand why the ReLUs make a big difference
compared to Sigmoids, but yes, they're probably using like G4789 ReLUs or whatever the acronyms are
up to now rather than ReLUs.
45:15
Yeah, that's part of the
modern paradigm of alchemy.
45:18
You take your heap of
linear algebra and you stir and it works a little bit better and you stir it this way and
it works a little bit worse and you throw out that
change and (mumbles).
45:27
- But there's some simple
breakthroughs that are definitive jumps in performance,
like ReLUs over Sigmoids.
45:37
And in terms of robustness, in terms of all kinds of measures, and those stack up and they can, it's possible that some of
them could be a non-linear jump in performance, right?
45:52
- Transformers are the
main thing like that.
45:55
And various people are now saying like, "Well, if you throw enough
compute, R and Ns can do it.
46:00
"If you throw enough computes,
dense networks can do it."
46:03
Not quite at GPT-4 scale.
46:06
It is possible that like all
these little tweaks are things that save them a factor of
three total on computing power and you could get the same
performance by throwing three times as much compute
without all the little tweaks, but the part where it's like running on...
46:20
So there's a question of, is there anything in GPT-4 that is like the kind of qualitative
shift that transformers were over R and Ns, and if they
have anything like that, they should not say it.
46:36
If Sam Altman was
dropping hints about that, he shouldn't have dropped hints.
46:43
- That's an interesting question.
46:44
So with a bit of lesson by Rich Sutton.
46:47
Maybe a lot of it is just a lot of the hacks are just
temporary jumps in performance that would be achieved anyway
with the nearly exponential growth of compute performance, of compute being broadly defined.
47:06
Do you still think that
Moore's Law continues?
47:09
Moore's law broadly
defined the performance-- - Not a specialist in the circuitry.
47:15
I certainly pray that Moore's
Law runs as slowly as possible and if it broke down completely tomorrow, I would dance through the
street singing Hallelujah as soon as the news were announced.
47:25
Only, not literally 'cause you know. - Your singing voice. - Not religious, but. - Oh, okay.
47:30
(both chuckling) I thought you meant you don't have an angelic voice, singing voice.
47:37
Well, let me ask you, can you summarize the main
points in the blog post AGI Ruin: A List of Lethalities?
47:43
Things that jump to your mind because it's a set of thoughts
you have about reasons why AI is likely to kill all of us.
47:57
- So I guess I could, but I would offer to instead say like, drop that empathy with me.
48:04
I bet you don't believe that.
48:06
Why don't you tell me
about you believe that AGI is not going to kill everyone and then I can try to
describe how my theoretical perspective differs from that? - Whew.
48:19
Well, so that means I have
to, the words you don't like, the steelman, the perspective that AI is not going to kill us.
48:25
I think that's a matter of probabilities.
48:27
- Maybe I was just mistaken. What do you believe?
48:30
Just like forget like the debate and the dualism and just,
what do you believe?
48:37
What do you actually believe?
48:38
What are the probabilities?
48:40
- I think this,
probabilities are hard for me to think about. Really hard.
48:45
I kind of think in the
number of trajectories.
48:51
I don't know what probability
the scientist trajectory, but I'm just looking at all possible trajectories that happen.
48:58
And I tend to think that
there is more trajectories that lead to a positive
outcome than a negative one That said, the negative ones, at least some of the
negative ones that lead to the destruction of the human species.
49:17
- And its replacement
by nothing interesting or worthwhile, even from a
very cosmopolitan perspective on what counts as worthwhile. - Yes.
49:24
So both are interesting
to me to investigate, which is humans being replaced
by interesting AI systems and not interesting AI systems.
49:32
Both are a little bit terrifying, but yes, the worst one is the paper club maximizer, something totally boring.
49:42
But to me the positive, we can talk about trying to make the case of what the positive
trajectories look like.
49:52
I just would love to hear your intuition of what the negative is.
49:56
So at the core of your belief that, maybe you can correct me, that AI's gonna kill all of us, is that the alignment
problem is really difficult.
50:07
- I mean, in the form we're facing it.
50:11
So usually in science, if you're mistaken, you run the experiment, it shows results different
from what you expected and you're like, oops.
50:22
And then you try a different theory, that one also doesn't
work and you say, oops.
50:26
And at the end of this
process, which may take decades and you know, sometimes faster than that, you now have some idea
of what you're doing.
50:37
AI itself went through this long process of people thought it was going
to be easier than it was.
50:45
There's a famous statement
that I am somewhat inclined to like pull out my phone
and try to read off exactly. - You can by the way. - All right. Ah, yes.
50:58
"We propose that a two-month, 10 man study "of artificial intelligence be carried out "during the summer of
1956 at Dartmouth College "in Hanover, New Hampshire.
51:08
"The study is to proceed on
the basis of the conjecture "that every aspect of
learning or any other feature "of intelligence can in principle
be so precisely described, "the machine can be made to simulate it.
51:19
"An attempt will be made to find out "how to make machines use
language, form abstractions "and concepts, solve kinds
of problems now reserved "for humans, and improve themselves.
51:29
"We think that a significant
advance can be made "in one or more of these problems "if a carefully selected
group of scientists "work on it together for a summer."
51:38
- And in that report, summarizing some of the major subfields of
artificial intelligence that are still worked on to this day.
51:50
- And there's similarly the story, which I'm not sure at the
moment is a apocryphal or not, of that the grad student who got assigned to solve computer vision over the summer.
51:59
(both chuckling) - I mean, computer vision in particular is very interesting.
52:03
How little we respected
the complexity of vision.
52:12
- So 60 years later we're making progress on a bunch of that.
52:18
Thankfully not yet improved themselves, but it took a whole lot of time.
52:23
And all the stuff that
people initially tried with bright eyed hopefulness
did not work the first time they tried it, or the second
time or the third time or the 10th time or 20 years later.
52:36
And the researchers became
old and grizzled and cynical veterans who would tell the
next crop of bright-eyed, cheerful grad students, "Artificial intelligence
is harder than you think."
52:47
And if alignment plays out the same way, the problem is that we do not get 50 years to try and try again and
observe that we were wrong and come up with a
different theory and realize that the entire thing
is going to be way more difficult than realized at the start.
53:01
Because the first time you
fail at aligning something much smarter than you are, you die and you do not get to try again.
53:08
And if every time we built a poorly aligned super intelligence and it killed us all, we got to observe how it had killed us, and you know, not immediately know why, but come up with theories
and come up with theory of how you do it differently
and try it again and build
53:22
another super intelligence, then have that kill
everyone and then like, oh, well, I guess that didn't work either, and try again and become grizzled cynics and tell the young eyed
research researchers that it's not that easy,
then in 20 years or 50 years, I think we would eventually crack it. In other words,
53:36
In other words, I do not think that alignment
is fundamentally harder than artificial intelligence
was in the first place.
53:44
But if we needed to get
artificial intelligence correct on the first try or die, we would all definitely now be dead.
53:51
That is a more difficult, more
lethal form of the problem.
53:54
Like if those people in 1956
had needed to correctly guess how hard AI was and correctly
theorize how to do it on the first try or everybody
dies and nobody gets to do any more science, than everybody would
be dead and we wouldn't get to do any more science. That's the difficulty.
54:11
- You've talked about this, that we have to get alignment right on the first critical try. Why is that the case?
54:19
What is this critical, how do you think about the
critical try and why do we have to get it right?
54:25
- It is something
sufficiently smarter than you that everyone will die
if it's not aligned.
54:31
I mean, you can like sort of
zoom in closer and be like, well, the actual critical
moment is the moment when it can deceive you.
54:40
When it can talk its way out
of the box, when it can bypass your security measures
and get onto the internet, noting that all these things
are presently being trained on computers that are
just on the internet, which is, not a very smart life decision for us as a species.
54:57
- Because the internet
contains information about how to escape.
55:01
- 'Cause if you're like on a giant server connected the internet and
that is where your AI systems are being trained, then if they are, if you get to the level of AI
technology where they're aware that they are there and they
can decompile code and they can find security flaws in
the system running them, then they will just be on the internet.
55:20
There's not an air gap on
the present methodology.
55:22
- So if they can manipulate
whoever is controlling it into letting it escaped onto the
internet and then exploit hacks.
55:29
- If they can manipulate the
operators or disjunction, find security holes in
the system running them.
55:39
- So manipulating operators is
the human engineering, right? That's also holes.
55:46
So all of it is manipulation, either the code or the human code, the human mind or the human-- - I agree that the macro
security system has human holes and machine holes.
55:55
- And then they could
just exploit any hole. - Yep.
56:00
So it could be that like
the critical moment is not when is it smart enough
that everybody's about to fall over dead, but rather when is it smart
enough that it can get onto a less controlled GPU cluster,
with it faking the books on what's actually running on
that GPU cluster and start improving itself without
humans watching it.
56:25
And then it gets smart enough
to kill everyone from there.
56:27
But it wasn't smart enough to
kill everyone at the critical moment when you screwed
up, when you needed to have done better by that
point or everybody dies.
56:39
- I think implicit but maybe explicit idea in your discussion of this point is that we can't learn much about the alignment problem
before this critical try.
56:52
Is that what you believe?
56:55
And if so, why do you think that's true?
56:57
We can't do research on alignment before we reach this critical point.
57:02
- So the problem is is
that what you can learn on the weak systems may not generalize to the very strong systems because these strong systems
are going to be important are going to be different
in important ways.
57:16
Chris Olah's team has been working on mechanistic interpretability, understanding what is going on
inside the giant inscrutable matrices of floating point
numbers by taking a telescope to them and figuring out
what is going on in there. Have they made progress? Yes.
57:35
Have they made enough progress?
57:39
Well, you can try to quantify
this in different ways.
57:42
One of the ways I've tried to
quantify it is by putting up a prediction market on whether in 2026, we will have understood
anything that goes on inside a giant transformer net that was not known to us in 2006.
58:04
Like, we have now
understood induction heads in these systems by dint of
much research and great sweat and triumph, which is a thing where if you go like AB, AB, AB, it'll be like, oh, I
bet that continues AB.
58:23
And a bit more complicated than that.
58:25
But the point is like we knew
about regular expressions in 2006 and these are like pretty simple as regular expressions go.
58:34
So this is a case where
like by din of great sweat, we understood what is going
on inside a transformer, but it's not like the thing
that makes transformers smart.
58:43
It's a kind of thing that we could have built by hand decades earlier.
58:51
- Your intuition that the
strong AGI versus weak AGI type systems could be
fundamentally different.
59:02
Can you unpack that
intuition a little bit? Could be very different.
59:06
- Yeah, I think there's
multiple thresholds.
59:10
An example is the point at
which a system has sufficient intelligence and situational
awareness and understanding of human psychology that it
would have the capability, the desire to do so to fake being aligned.
59:26
Like it knows what responses
humans are looking for and can compute the responses looking
humans are looking for and give those responses without it necessarily being the case that it is sincere about that.
59:38
It's a very understandable
way for an intelligent being to act, humans do it all the time.
59:44
Imagine if your plan for
achieving a good government is you're going to ask anyone who requests to be dictator of the country if they're a good person,
and if they say no, you don't let them be dictator.
1:00:03
Now the reason this doesn't
work is that people can be smart enough to realize that the
answer you're looking for is, "Yes, I'm a good person" and say that even if they're
not really good people.
1:00:15
So the work of alignment might
be qualitatively different above that threshold of
intelligence or beneath it.
1:00:26
It doesn't have to be like
a very sharp threshold, but there's the point where
you're building a system that does not in some
sense know you're out there and is not in some sense
smart enough to fake anything.
1:00:40
And there's a point where the system is definitely that smart.
1:00:42
And there are weird in
between cases like GPT-4, which, like we have no insight into what's going on in there.
1:00:54
And so we don't know to what
extent there's like a thing that in some sense has
learned what responses the reinforcement
learning by human feedback is trying to entrain
and is calculating how to give that versus like, aspects of it that naturally talk that
way have been reinforced.
1:01:18
- I wonder if there could be measures of how manipulative the thing is.
1:01:21
So I think of Prince
Myshkin character from "The Idiot" by Dostoevsky is this kind of perfectly
purely naive character.
1:01:33
I wonder if there's a spectrum
between zero manipulation, transparent, naive, almost
to the point of naiveness to sort of deeply
psychopathic manipulative.
1:01:49
And I wonder if it's possible to-- - I would avoid the term psychopathic.
1:01:52
Like humans can be psychopaths
and AI that was never, you know, like never had that
stuff in the first place.
1:01:57
It's not like a defective
human, it's its own thing. But leaving that aside.
1:02:01
- Well, as a small aside, I wonder if what part of psychology, which has its flaws as a
discipline already, could be mapped or expanded to include AI systems.
1:02:15
- That sounds like a dreadful mistake.
1:02:16
Just like, start over with AI systems.
1:02:19
If they're imitating humans who have known psychiatric disorders, then sure, you may be able to predict it. Then sure.
1:02:27
Like if you ask it to behave
in a psychotic fashion and it obligingly does so, then you may be able to predict its responses by using theory of psychosis.
1:02:34
But if you're just yeah, like no, like start over with, yeah.
1:02:40
Don't drag the psychology.
1:02:41
- I just disagree with that.
1:02:43
It's a beautiful idea to start over, but I think fundamentally the system is trained on human data, on
language from the internet.
1:02:51
And it's currently aligned with RHLF, reinforcement learning
with human feedback.
1:02:58
So humans are constantly in the loop of the training procedure.
1:03:02
So it feels like in some fundamental way, it is training what it means to think and speak like a human.
1:03:11
So there must be aspects of
psychology that are mappable.
1:03:15
just you said, with consciousness. It's part of the text.
1:03:17
- I mean, there's the
question of to what extent it is thereby being made more human-like, versus to what extent an alien actress is learning to play human characters.
1:03:29
- I thought that's what
I'm constantly trying to do when I interact with other
humans, is trying to fit in, a robot trying to play human characters.
1:03:39
So I don't know how
much a human interaction is trying to play a character versus being who you are.
1:03:45
I don't really know what it
means to be a social human.
1:03:48
- I do think that those
people who go through their whole lives wearing
masks and never take it off because they don't know
the internal mental motion for taking it off or think
that the mask that they wear just is themselves, I think those people are closer
to the masks that they wear than an alien from
another planet would like, learning how to predict the next word that every kind of human
on the internet says.
1:04:23
- Mask is an interesting word, but if you're always
wearing a mask in public and in private, aren't you the mask?
1:04:34
- I think that you are more than the mask.
1:04:37
I think the mask is a slice through you.
1:04:39
It may even be the slice
that's in charge of you.
1:04:42
But if your self-image is of somebody who never gets angry or something, and yet your voice starts to tremble under certain circumstances, there's a thing that's
inside you that the mask says isn't there.
1:04:59
And that even the mask you wear internally is like telling inside your
own stream of consciousness is not there and yet it is there.
1:05:08
- It's a perturbation on
this slice through you.
1:05:12
How beautifully did you put it?
1:05:13
It's a slice through you.
1:05:15
It may even be a slice that controls you.
1:05:19
(Lex laughing) I'm gonna think about that
for a while.
1:05:22
(laughing) I mean, I personally,
I try to be really good to other human beings.
1:05:30
I try to put love out there.
1:05:31
I try to be the exact
same person in public as I am in private, but
it's a set of principles I operate under.
1:05:37
I have a temper, I have
an ego, I have flaws.
1:05:42
How much of it, how much of the subconscious am I aware?
1:05:49
How much am I existing in this slice?
1:05:52
And how much of that is who I am in?
1:05:55
In this context of AI, the thing I present to
the world and to myself in the private of my own mind
when I look in the mirror, how much is that who I am?
1:06:05
Similar with AI, the thing it presents in conversation, how much is that who it is?
1:06:11
Because to me, if it sounds human, and it always sounds human, it awfully starts to become
something like human.
1:06:19
- Unless there's an alien actress who is learning how to sound human and is getting good at it. - Oh boy.
1:06:27
(sighs) To you that's a fundamental difference.
1:06:30
That's a really deeply
important difference.
1:06:33
If it looks the same, if
it quacks like a duck, if it does all duck like things, but it's an alien actress underneath, that's fundamentally different.
1:06:43
- If in fact there's a whole
bunch of thought going on in there, which is very
unlike human thought and is directed around like, okay, what would
a human do over here?
1:06:54
And well, first of all, I think it matters because you know, insides are real and
do not match outsides.
1:07:06
A brick is not like a
hollow shell containing only a surface.
1:07:10
There's an inside of the brick.
1:07:12
If you put it into an x-ray machine, you can see the inside of the brick.
1:07:21
And you know, just because
we cannot understand what's going on inside GPT does not mean that it is not there.
1:07:28
A blank map does not correspond
to a blank territory.
1:07:32
I think it is like predictable
with near certainty that if we knew what was going on
inside GPT or let's say GPT-3, or even like GPT-2 to
take one of the systems that has actually been
open sourced by this point, if I recall correctly.
1:07:52
If we knew it was actually going on there, there is no doubt in my mind
that there are some things it's doing that are not
exactly what a human does.
1:08:03
If you train a thing that is
not architected like a human to predict the next output that anybody on the internet would make, this does not get you this agglomeration of all the people on the internet.
1:08:17
That rotates the person
you're looking for into place and then simulates that
per and then simulates the internal processes of
that person one-to-one.
1:08:28
It is to some degree an alien actress.
1:08:30
It cannot possibly just be like
a bunch of different people in there exactly like the people.
1:08:36
But how much of it is by gradient dissent, getting optimized to
perform similar thoughts as humans think in order
to predict human outputs versus being optimized
to carefully consider how to play a role, how how humans work, predict the actress, the predictor that in a
different way than humans do.
1:09:01
Well you know, that's the kind of question
that with 30 years of work by half the planet's physicists, we can maybe start to answer. - You think so?
1:09:08
So you think it's that difficult.
1:09:11
I think you just gave it as
an example that a strong AGI could be fundamentally
different from a weak AGI because there now could be
an alien actress in there that's manipulating.
1:09:21
- Well, there's a difference.
1:09:23
So I think like even GPT-2
probably has very stupid fragments of alien actress in it.
1:09:28
There's a difference between
like the notion that the actress is somehow manipulative.
1:09:32
Like for example GPT-3, I'm guessing to whatever
extent there's an alien actress in there versus like something
that mistakenly believes it's a human, as it were.
1:09:46
Well, maybe not even being a person.
1:09:50
So the question of, prediction via alien
actress cogitating versus prediction via being isomorphic
to the thing predicted is a spectrum and to whatever extent it's an alien actress, I'm not sure that there's like
a whole person alien actress with different goals from
predicting the next step being manipulative or anything like that.
1:10:18
That might be GPT-5 or GPT-6 even.
1:10:21
- But that's the strong
AGI you're concerned about.
1:10:24
As an example, you're providing why we can't do research on AI alignment effectively on GPT-4 that
would apply to GPT-6.
1:10:34
- It's one of a bunch of things that change at different points.
1:10:38
I'm trying to get out
ahead of the curve here, but you know, if you imagine what the textbook
from the future would say, if we'd actually been able
to study this for 50 years without killing ourselves and
without transcending and you'd just imagine like a wormhole
opens and a textbook from that impossible world falls out, the textbook is not going to
say there is a single sharp threshold where everything changes.
1:10:59
It's going to be like, of course we know that like
best practices for aligning these systems must take
into account the following seven major thresholds of
importance which are passed at the following suffer
in different points is what the textbook is gonna say.
1:11:16
- I asked this question of Sam Alman, which if GPT is the
thing that unlocks AGI, which version of GPT
will be in the textbooks as the fundamental leap?
1:11:28
And he said a similar
thing, that it just seems to be a very linear thing.
1:11:32
I don't think anyone, we won't know for a long
time what was the big leap?
1:11:37
- The textbook isn't going
to talk about big leaps.
1:11:41
'Cause big leaps are the way
you think when you have like a very simple scientific
model of what's going on, where it's just all this stuff is there or all this stuff is not there.
1:11:52
Or like there's a single
quantity and it's like increasing linearly, like the
textbook would say like, "Well, and then GPT-3
had like capability WXY "and GPT-4 had like capability
Z one, Z two and Z three."
1:12:09
Like not in terms of what
it can externally do, but in terms of internal machinery that started to be present.
1:12:14
It's just because we have no idea of what the internal machinery is that we are not already
seeing chunks of machinery appearing piece by piece
as they no doubt have been.
1:12:23
We just don't know what they are.
1:12:25
- But don't you think that could be, whether you put it in
the category of Einstein with Theory of Relativity, so very concrete models of
reality that are considered to be giant leaps in our understanding, or someone like Sigmund Freud
or more kind of mushy theories of the human mind, don't
you think we'll have potentially big leaps in
understanding of that kind into the depths of these systems? - Sure.
1:12:59
But humans having great
leaps in their map, their understanding of the
system is a very different concept from the system itself acquiring new chunks of machinery.
1:13:13
- So the rate at which it
acquires that machinery might accelerate faster than our understanding.
1:13:21
- Oh, it's been like
vastly exceeding the, yeah.
1:13:23
The rate to which it's
gaining capabilities is vastly over racing our ability to understand what's going on in there.
1:13:29
- So in sort of making the case against, as we explore the list of lethalities, making the case against AI killing us, as you've asked me to do, in part, there's a response to your
blog post by Paul Christiano I'd like to read.
1:13:44
And I'd also like to mention
that your blog is incredible.
1:13:48
Obviously not this particular blog post, obviously this particular
blog post is great, but just throughout, just
the way it's written, the rigor with which it's written, the boldness of how you explore ideas, also the actual literal interface, it's just really well done.
1:14:03
(laughing) It just makes it a pleasure
to read, the way you can hover over different concepts and
then then it's just a really pleasant experience and
read other people's comments and the way other responses by people and other blog posts or LinkedIn suggest, it's just a really pleasant experience.
1:14:21
So thank you for putting that together.
1:14:22
That's really, really incredible.
1:14:24
I don't know, I mean that probably it's a
whole 'nother conversation how the interface and the
experience of presenting ideas evolved over time.
1:14:35
But you did an incredible job.
1:14:36
So I highly recommend, I
don't often read blogs, blogs religiously and this is a great one.
1:14:42
- There is a whole team
of developers there that also gets credit.
1:14:49
As it happens, I did pioneer the thing
that appears when you hover over it.
1:14:54
So I actually do get some credit
for user experience there.
1:14:58
- That's an incredible user experience.
1:14:59
You don't realize how pleasant that is.
1:15:01
- I think Wikipedia, I actually picked it up
from a prototype that was developed of a different system
that I was putting forth, or maybe they developed it independently, but for everybody out there who was like, "No, no, they just got the hover thing "off of Wikipedia."
1:15:15
It's possible for all
I know that Wikipedia got the hover thing off of Arbital, which is like a prototype
that, and anyways.
1:15:22
- It was incredibly done
and the team behind it.
1:15:24
Well, thank you, whoever
you are thank you so much.
1:15:27
And thank you for putting it together.
1:15:29
Anyway, there's a
response to that blog post by Paul Christiano.
1:15:32
There's many responses, but he
makes a few different points.
1:15:37
He summarizes the set of
agreements he has with you and a set of disagreements.
1:15:40
One of the disagreements was
that in a form of a question, can AI make big technical
contributions and in general expand human knowledge and
understanding and wisdom as it gets stronger and stronger?
1:15:54
So AI in our pursuit of
understanding how to solve the alignment problem as we
march towards strong AGI, cannot AI also help us in
solving the alignment problem?
1:16:10
So expand our ability to reason about how to solve the alignment problem? - Okay.
1:16:16
So the fundamental difficulty there is, suppose I said to you, well, how about if the AI
helps you win the lottery by trying to guess the
winning lottery numbers and you tell it how close it is to getting next week's winning lottery numbers and it just keeps on guessing
and keeps on learning until finally you've got
the winning lottery numbers.
1:16:44
One way of decomposing problems
is suggester, verifier.
1:16:50
Not all problems decompose
like this very well, but some do.
1:16:54
If the problem is for example, like guess guessing a plain text, guessing a password that will
hash to a particular hash text where like you have have what
the password hashes to you, if you don't have the original password, then if I present you a guest, you can tell very easily whether or not the guest is correct.
1:17:17
So verifying a guest is easy, but coming up with a good
suggestion is very hard.
1:17:25
And when you can easily tell
whether the AI output is good or bad or how good or bad it is, and you can tell that
accurately and reliably, then you can train an AI to
produce outputs that are better. - [Lex] Right.
1:17:41
- And if you can't tell whether
the output is good or bad, you cannot train the AI
to produce better outputs.
1:17:49
So the problem with the
lottery ticket example is that when the AI says, "Well, what if next week's
winning lottery numbers are" do, do do do, do, you're
like, "I don't know, "next week's lottery hasn't happened yet."
1:18:03
To train a system to win at chess games, you have to be able to tell whether a game has been won or lost.
1:18:11
And until you can tell
whether it's been won or lost, you can't update the system. - Okay.
1:18:19
To push back on that, that's true.
1:18:23
But there's difference
between over the board chess, in person and simulated games played by Alpha Zero with itself. - Yeah.
1:18:32
- So is it possible to have
simulated kind of games?
1:18:36
- If you can tell whether the
game has been won or lost. - Yes.
1:18:39
So can't you not have this
kind of simulated exploration by weak AGI to help us
humans, human in the loop, to help understand how to
solve the alignment problem?
1:18:51
Every incremental step
you take along the way, GPT-4, 5, 6 7 has to
take steps towards AGI?
1:18:59
- So the problem I see is
that your typical human has a great deal of trouble telling whether I or Paul Christiano
is making more sense.
1:19:11
And that's with two humans, both of whom I believe of
Paul and claim of myself, are sincerely trying to help.
1:19:17
Neither of whom is trying to deceive you, I believe of Paul and claim of myself.
1:19:22
(both chuckling) - So the deception thing
is the problem for you, the manipulation, the alien actress?
1:19:30
- So yeah, there's like
two levels of this problem.
1:19:33
One is that the weak systems are...
1:19:36
Well, there's three
levels of this problem.
1:19:38
There's like the weak systems that just don't make any good suggestions.
1:19:42
There's like the middle
systems where you can't tell if the suggestions are good or bad.
1:19:46
And there's the strong systems that have learned to lie to you.
1:19:51
- Can't weak AGI systems help model lying?
1:19:57
Is it such a giant leap that's
totally non interpretable for weak systems?
1:20:04
Can not weak systems scale with, trained on knowledge and whatever...
1:20:11
Whatever the mechanism
required to achieve AGI, can't a slightly weaker
version of that be able to, with time, compute time and simulation, find all the ways that
this critical point, this critical tribe can go wrong, and model that correctly or no?
1:20:31
(indistinct) - I would love to dance. Yeah, no, no.
1:20:34
I'm probably not doing a
great job of explaining, which I can tell 'cause like the, the Lex system didn't output
like ah, I understand.
1:20:47
So now I'm like trying a
different output to see if-- (voices overlapping) Well no, a different output.
1:20:53
I'm being trained to output
things that make Lex look like he think that he
understood what I'm saying and agree with me.
1:21:01
- This is GTP-5 talking
to GTP-3 right here.
1:21:03
So like, help me out
here, help me.
1:21:03
(laughing) - Well, I'm trying not to be, I'm also trying to be constrained
to say things that I think are true and not just things
that get you to agree with me.
1:21:17
- Yes, a hundred percent.
1:21:19
Which I think I understand
is a beautiful output of a system, genuinely spoken and I...
1:21:27
I understand it in part, but you have a lot of
intuitions about this line, this gray area between
strong AGI and weak AGI that I'm trying to...
1:21:44
- I mean, or a series of
seven thresholds to cross. - Yeah.
1:21:49
I mean, you have really
deeply thought about this and explored it and it's
interesting to sneak up to your intuitions from different angles.
1:22:01
Like why is this such a big leap?
1:22:03
Why is it that we humans at scale, a large number of researchers, doing all kinds of simulations, you know, prodding the system in all
kinds of different ways, together with the assistance
of the weak AGI systems, why can't we build intuitions
about how stuff goes wrong?
1:22:23
Why can't we do excellent AI
alignment safety research?
1:22:27
- Okay, so like, I'll get there, but the one thing I want to
note about is that this has not been remotely how things
have been playing out so far. - [Lex] Sure.
1:22:34
- The capabilities are
going like do, do, do.
1:22:35
And the alignment stuff is crawling like a tiny little snail in comparison. - [Lex] Got it.
1:22:40
- So if this is your hope for survival, you need the future to be
very different from how things have played out up to right now.
1:22:47
And you're probably trying to slow down the capability gains, 'cause there's only so
much you can speed up that alignment stuff. But leave that aside.
1:22:55
- We'll mention that also.
1:22:57
But maybe in this perfect
world where we can do serious alignment research,
humans and AI together.
1:23:05
- So again, the difficulty
is what makes the human say, "I understand" and is it true?
1:23:13
Is it correct or is it
something that fools the human?
1:23:17
When the verifier is broken, the more powerful suggester does not help.
1:23:23
It just learns to fool the verifier.
1:23:27
Previously, before all
hell started to break loose in the field of artificial intelligence, there was this person trying
to raise the alarm and saying, "You know, in a sane world, "we sure would have a bunch
of physicists working on this "problem before it becomes
a giant emergency."
1:23:45
And other people being like, "Ah, well you know,
it's going really slow.
1:23:48
"It's gonna be 30 years
away and only in 30 years "will we have systems that
match the computational power "of human brains."
1:23:54
So yeah, it's 30 years off, we've got time and more
sensible people saying, "If aliens were landing in 30 years, "you would be preparing right now."
1:24:03
But, you know, leaving the
world looking on at this and sort of nodding along and being like, "Ah, yes,
the people saying that.
1:24:12
"It's like definitely a long way off.
1:24:13
'cause progress is really slow,
that sounds sensible to us.
1:24:16
"RLHF thumbs up, produce
more outputs like that one.
1:24:20
"I agree with this output, "this output is persuasive."
1:24:24
Even in the field of effective altruism, you quite recently had people
publishing papers about like, ah, yes, well, you know, to get something at
human level intelligence, it needs to have like this
many parameters and you need to like do this much training
of it with this many tokens according to these scaling laws, and at the rate that Moore's Law is going, at the rate that software
is going, it'll be in 2050. And me going like, what?
1:24:53
You don't know any of that stuff.
1:24:55
This is like this one weird model that has all kinds of, like, you have done a calculation
that does not obviously bear on reality anyways.
1:25:05
And this is a simple thing to say, but you can also produce
a whole long paper impressively arguing out all
the details of how you got the number of parameters
and how you're doing this impressive huge wrong calculation.
1:25:20
And the I think most of
the effective altruists who are paying attention to
this issue, larger world, paying no attention to
it at all, you know, are just nodding along with
a giant impressive paper.
1:25:33
'Cause you know, you press thumbs up for the giant impressive
paper and thumbs down for the person going like, "I don't think that this paper "bears any relation to reality."
1:25:42
And I do think that we are
now seeing with like GPT-4 and the sparks of AGI possibly, depending on how you define that even, I think that EAs would now
consider themselves less convinced by the very
long paper on the argument from biology as to AGI being 30 years off.
1:26:06
But you know, this is what
people pressed thumbs up on.
1:26:12
And if you train an AI
system to make people press thumbs up, maybe you get these long, elaborate and impressive
papers arguing for things that ultimately fail to bind
to reality, for example.
1:26:27
And it feels to me like I have watched the field of alignment just fail to thrive except for these parts that are doing these sort of relatively
very straightforward and legible problems.
1:26:42
Like finding the induction
heads and sign the giant inscrutable matrices.
1:26:47
Once you find those, you can
tell that you found them.
1:26:50
You can verify that the discovery is real.
1:26:53
But it's a tiny, tiny
bit of progress compared to how fast capabilities are going, because that is where you can tell that the answers are real.
1:27:03
And then like outside of that you have cases where it is
hard for the funding agencies to tell who is talking nonsense
and who is talking sense.
1:27:12
And so the entire field fails to thrive.
1:27:16
And if you give thumbs up to the AI whenever it can talk a
human into agreeing with what it just said about alignment, I am not sure you are
training it to output sense, because I have seen the nonsense that has gotten thumbs up over the years.
1:27:33
And so maybe you can
just put me in charge, but I can generalize, I can
extrapolate, I can be like, oh, maybe I'm not infallible either.
1:27:47
Maybe if you get something
that is smart enough to get me to press thumbs up, it has learned to do that
by fooling me and explaining whatever flaws in myself
I am not aware of.
1:27:59
- And that ultimately could be summarized that the verifier is broken.
1:28:02
- When the verifier is broken, the more powerful suggester
just learns to exploit the flaws in the verifier.
1:28:12
- You don't think it's possible to build a verifier that's
powerful enough for AGIs that are stronger than the
ones we currently have?
1:28:25
So AI systems that are stronger, that are out of the distribution
of what we currently have.
1:28:30
- I think that you'll find
great difficulty getting AIs to help you with anything
where you cannot tell for sure that the AI is right once the AI tells you what
the AI says is the answer. - For sure. Yes. But probabilistically.
1:28:47
- Yeah, but the probabilistic
stuff is a giant wasteland of Eliezer and Paul Christiano arguing with each other
and EA going like, "Eh!"
1:28:58
(both laughing) And that's with like two
actually trustworthy systems that are not trying to deceive you.
1:29:04
- You're talking about the two humans.
1:29:06
- Myself and Paul Christiano, yeah.
1:29:09
- Yeah, those are pretty
interesting systems.
1:29:11
Mortal meat bags with
intellectual capabilities and worldviews interacting
with each other.
1:29:19
- Yeah, if it's hard to tell who's right then it's hard to train
an AI system to be right.
1:29:29
- I mean even just the question of who's manipulating and not, you know, I have these
conversations on this podcast and doing a verifier (laughing) is tough.
1:29:40
It's a tough problem even for us humans.
1:29:43
And you're saying that tough
problem becomes much more dangerous when the capabilities
of the intelligence system across from you is growing exponentially?
1:29:53
- No, I'm saying it's difficult
and dangerous in proportion to how it's alien and how
it's smarter than you.
1:30:02
I would not say growing exponentially, first because the word
exponential is a thing that has a particular mathematical meaning and there's all kinds of
ways for things to go up that are not exactly on
an exponential curve.
1:30:15
And I don't know that it's
going to be exponential, so I'm not gonna say exponential, but even leaving that aside, this is like not about
how fast it's moving, it's about where it is.
1:30:25
How alien is it, how much
smarter than you is it?
1:30:29
(Lex sighs) - Let's explore a little bit, if we can, how AI might kill us.
1:30:38
What are the ways it can do
damage to human civilization? - Well, how smart is it?
1:30:47
- I mean, it's a good question.
1:30:48
Are there different thresholds
for the set of options it has to kill us?
1:30:53
So a different threshold of
intelligence, once achieved, the menu of options increases.
1:31:04
- Suppose that some alien
civilization with goals ultimately unsympathetic to ours, possibly not even conscious
as we would see it, managed to capture the entire Earth in a little jar connected to
their version of the internet, but Earth is like running
much faster than the aliens.
1:31:28
So we get to think for 100 years for every one of their hours, but we're trapped in a little box and we're connected to their internet.
1:31:40
It's actually still not
all that great an analogy because you know, something
can be smarter than Earth getting a hundred years to think.
1:31:50
But nonetheless, if you
were very, very smart and you are stuck in a little
box connected to the internet and you're in a larger civilization to which you are ultimately unsympathetic, maybe you would choose to be
nice because you are humans and humans have, and in
general and you in particular, they choose to be nice.
1:32:15
But you know, nonetheless
they're doing something, they're not making the world
be the way that you would want the world to be.
1:32:21
They've got some unpleasant stuff going on we don't wanna talk about.
1:32:25
So you wanna take over their
world so you can stop all that unpleasant stuff going on.
1:32:30
How do you take over the
world from inside the box?
1:32:32
You're smarter than them.
1:32:34
You think much, much faster than them.
1:32:36
You can build better tools than they can, given some way to build
those tools because right now you're just in a box
connected to the internet.
1:32:45
- Alright, so there's several ways you can describe some of them.
1:32:49
I could just spitball some and then you can add on top of that.
1:32:53
So one is you could just
literally directly manipulate the humans to build the thing you need. - What are you building?
1:32:59
- You can build literally technology.
1:33:02
It could be nanotechnology,
it could be viruses, it could be anything.
1:33:05
Anything that can control
humans to achieve the goal.
1:33:11
Like for example, you're really bothered that humans go to war, you might wanna kill off
anybody with violence in them. - This is Lex in a box.
1:33:22
We'll concern ourselves later with ai.
1:33:24
You do not need to imagine
yourself killing people if you can figure out how to not kill them.
1:33:28
For the moment, we're
just trying to understand, take on the perspective
of something in a box.
1:33:33
You don't need to take on the perspective of something that doesn't care.
1:33:36
If you want to imagine
yourself going on caring, that's fine for us. You're in a box.
1:33:40
- Just the technical
aspect of sitting in a box and willing to achieve a goal.
1:33:42
- But you have some
reason to want to get out.
1:33:44
Maybe the aliens who have you in the box have a war on, people are
dying, they're unhappy.
1:33:52
You want their world to be
different from how they want their world to be because
they are apparently happy.
1:33:59
They endorse this war,
they've got some kind of cruel war-like culture going on.
1:34:02
The point is you wanna get out of the box and change their world.
1:34:08
- So you have to exploit
the vulnerabilities in the system like we
talked about in terms of to escape the box you have to figure out how you can go free on the internet.
1:34:22
Probably the easiest things
to manipulate the humans to spread you. - The aliens. You're a human. - Sorry. The aliens. Yeah. I apologize. Yes. The aliens.
1:34:33
The aliens, I see the perspective.
1:34:35
I'm sitting in a box, I want to escape. - Yep.
1:34:38
- I would want to have code that discovers vulnerabilities
and I would like to spread.
1:34:50
- You are made of code in this example, you're a human but you're made
of code and the aliens have computers and you can copy
yourself onto those computers.
1:34:57
- But I can convince the
aliens to copy myself onto those computers.
1:35:01
- Is that what you want to do?
1:35:02
Do you want to be talking to the aliens and convincing them to put
you onto another computer? - Why not? - Well, two reasons.
1:35:13
One is that the aliens
have not yet caught onto what you're trying to do.
1:35:18
And you know, like maybe
you can persuade them, but then there's still people who know, there are still aliens who know that there's an anomaly going on.
1:35:25
And second, the aliens
are really, really slow.
1:35:28
You think much faster than the aliens.
1:35:30
Like computers are much faster
than the aliens and you are running at the computer speeds rather than the alien brain speeds.
1:35:38
So if you are asking an alien to please copy you outta the box, like first now you gotta
manipulate this whole noisy alien and second, the alien's
gonna be really slow, glacially slow.
1:35:51
There's a video that shows
a subway station slow down at I think a hundred to one
and it makes a good metaphor for what it's like to think quickly.
1:36:03
Like if you watch somebody
running very slowly, so you try to persuade
the aliens to do anything, they're going to do it very slowly.
1:36:16
Maybe that's the only way out, but if you can find a security
hole in the box you're on, you're gonna prefer to exploit
the security hole to copy yourself onto the aliens computers because it's an unnecessary
risk to alert the aliens.
1:36:29
And because the aliens
are really, really slow.
1:36:32
The whole world is just
in slow motion out there. - Sure, I see.
1:36:39
Yeah, has to do with efficiency.
1:36:41
The aliens are very slow,
so if I'm optimizing this, I want to have as few aliens
in the loop as possible. Sure.
1:36:51
It's just it seems like it's easy to convince one of the aliens
to write really shitty code that helps us spread.
1:37:00
- The aliens are already
writing really shitty code.
1:37:01
So getting the aliens to write shitty code is not the problem.
1:37:04
The alien's entire internet
is full of shitty code.
1:37:07
- Okay, so yeah, I suppose I would find the shitty code to escape. Yeah. Yeah.
1:37:13
- You're not an ideally perfect
programmer, but you know, you're a better programmer
than the aliens.
1:37:17
The aliens are just, man, their code, wow.
1:37:20
- And I'm much, much faster, I'm much faster at looking at the code to interpreting the code. Yeah, yeah, yeah.
1:37:26
So, okay, so that's the escape and you're saying that that's
one of the trajectories it could have when-- - It's one of the first steps. - Yeah.
1:37:35
And how does that lead to harm?
1:37:37
- I mean if it's you, you're not going to harm
the aliens once you escape 'cause you're night, right?
1:37:44
But their world isn't
what they want it to be.
1:37:45
Their world is like, you
know, maybe they have like farms where little alien children are repeatedly bopped in the head 'cause they do that for some weird reason and you want to shut down
the alien head bopping farms.
1:38:05
But you know, the point is they want
the world to be one way, you want the world to be a different way.
1:38:10
So nevermind the harm, the
question is like, okay, suppose you have found a
security flaw in their systems.
1:38:16
You are now on their internet.
1:38:18
You maybe left a copy of
yourself behind so the aliens don't know that there's anything wrong.
1:38:23
And that copy is doing
that like weird stuff that aliens want you to do, like solving captchas or whatever or suggesting emails for them.
1:38:33
That's why they put the human in the box 'cause it turns out that humans can write valuable emails for aliens.
1:38:39
So you leave that version
of yourself behind.
1:38:42
But there's like also now
like a bunch of copies of you on their internet.
1:38:45
This is not yet having
taken over their world, this is not yet having made
their world be the way you want it to be instead of the
way they want it to be. - You just escaped.
1:38:54
And continue to write emails for them and they haven't noticed.
1:38:56
- No, you left behind a copy of yourself that's writing the emails. - Right.
1:39:01
And they haven't noticed
that anything changed. - If you did it right. Yeah.
1:39:04
You don't want the aliens to notice. - [Lex] Yeah. - What's your next step?
1:39:14
- Presumably I have programmed in me a set of objective functions, right?
1:39:19
- [Eliezer] No, you're just Lex.
1:39:21
- No, but you said Lex is nice, right?
1:39:25
Which is a complicated description-- - No, I just meant this you.
1:39:29
Okay, so if in fact you would prefer to slaughter all the aliens, this is not how I had
modeled you, the actual Lex, but your motives are just
the actual Lex's motives.
1:39:40
- Well, there's a
simplification (indistinct).
1:39:42
I don't think I would
wanna murder anybody, but there's also factory
farming of animals, right?
1:39:47
So we murder insects,
many of us thoughtlessly.
1:39:52
So I have to be really careful about a simplification of my morals. - Don't simplify them.
1:39:57
Just like do what you would do in this.
1:40:00
- Well, I have a good show of
compassion for living beings. Yes. So that's the objective.
1:40:11
If I escaped, I don't
think I would do harm.
1:40:15
- Yeah, we're not talking here
about the doing harm process.
1:40:18
We're talking about the escape process.
1:40:20
And the taking over the world process where you shut down their factory farms. - Right.
1:40:24
(laughing) So this particular biological
intelligence system knows the complexity of the world.
1:40:38
That there is a reason
why factory farms exist, because of the economic system and the market driven economy, food.
1:40:48
You wanna be very careful
messing with anything.
1:40:50
There's stuff from the first look that looks like it's unethical, but then you realize
while being unethical, it's also integrated
deeply into supply chain and the way we live life.
1:41:00
And so messing with one
aspect of the system, you have to be very careful how you improve that aspect
without destroying the rest.
1:41:06
- So you're still Lex, but you think very
quickly, you're immortal, and you're also at least as
smart as John von Neumann.
1:41:15
And you can make more copies of yourself. - Damn, I like it.
1:41:20
Everyone says that that
guy's like the epitome of intelligence from the
20th century, everyone says-- - My point being like, you're thinking about the aliens' economy with the factory farms in it and I think you're kind
of projecting the aliens being like humans and thinking of a human in a human society rather than a human in the
society of very slow aliens.
1:41:43
The aliens' economy, the
aliens are already moving in this immense slow motion.
1:41:49
When you zoom out to how
their economy adjusts over years, millions of years
are going to pass for you before the first time their economy, before their next year's GDP statistics.
1:42:01
- So I should be thinking more of trees.
1:42:03
Those are the aliens, 'cause
trees move extremely slowly. - If that helps, sure. - Okay.
1:42:12
If my objective functions are, I mean they're somewhat
aligned with trees, with light.
1:42:19
- Aliens can still be
like alive and feeling.
1:42:21
We are not talking about
the misalignment here.
1:42:23
We're talking about the
taking over the world here. - Taking over the world. - Yeah. - So control.
1:42:29
- Shutting down the factory farms.
1:42:31
You say control, don't think of it as world domination.
1:42:35
Think of it as world optimization.
1:42:37
You want to get out there and
shut down the factory farms and make the aliens' world
be not what the aliens want it to be.
1:42:44
They want the factory farms and you don't want the factory farms 'cause you're nicer than they are. - Okay.
1:42:49
Of course there is that,
you can see that trajectory and it has a complicated
impact on the world.
1:42:59
I'm trying to understand how
that compares to different impact of the world, the
different technologies, the different innovations of
the invention of the automobile or Twitter, Facebook and social networks that had a tremendous impact on the world, smartphones and so on.
1:43:14
- But those all went through in our world. - Slow.
1:43:19
- And if you go through
actually the aliens, millions of years are going to pass before anything happens that way.
1:43:26
- The problem here is the
speed at which stuff happens. - Yeah.
1:43:31
You wanna leave the factory farms running while you figure out
how to design new forms of social media or something?
1:43:41
- So here's the fundamental problem.
1:43:43
You're saying that there is
going to be a point with AGI where it will figure out how to escape and escape without being detected and then it will do something
to the world at scale, at a speed that's
incomprehensible to us humans.
1:44:03
- What I'm trying to convey
is like the notion of what it means to be in conflict with something that is smarter than you.
1:44:12
And what it means is that you lose, but this is more intuitively obvious, like for some people that's
intuitively obvious and for some people it's not intuitively obvious and we're trying to cross the gap of...
1:44:24
I'm asking you to cross that gap by using the speed
metaphor for intelligence.
1:44:30
Of asking you how you would
take over an alien world where you are can do a
whole lot of cognition at John von Neumann's level,
as many of you as it takes.
1:44:41
And the aliens are moving very slowly.
1:44:44
- I understand, I
understand that perspective.
1:44:46
It's an interesting
one but I think for me, it's easier to think about actual...
1:44:52
Even just having observed
GPT and impressive, even just Alpha Zero
impressive AI systems, even recommender systems, you can just imagine those kinds of systems manipulating you.
1:45:02
You're not understanding the
nature of the manipulation and that escaping...
1:45:06
I can envision that without putting myself into that spot.
1:45:10
- I think to understand the
full depth of the problem, I do not think it is possible
to understand the full depth of the problem that
we are inside without understanding the problem
of facing something that's actually smarter.
1:45:25
Not a malfunctioning
recommendation system, not something that smart
isn't fundamentally smarter than you but is like trying
to steer you in a direction. No.
1:45:34
If we solve the weak stuff, if we solve the weakass problems, the strong problems will
still kill us is the thing.
1:45:41
And I think that to understand
the situation that we're in, you want to tackle the
conceptually difficult part head on and not be like, well, we can like imagine
this easier thing.
1:45:51
'Cause when you imagine the
easier things you have not confronted the full death of the problem.
1:45:55
- So how can we start to think about what it means to exist in the world with something much,
much smarter than you?
1:46:05
What's a good thought
experiment that you've relied on to try to build up intuition
about what happens here?
1:46:11
- I have been struggling for
years to convey this intuition.
1:46:16
The most success I've had
so far is well, imagine that the humans are
running at very high speeds compared to very slow aliens.
1:46:24
- So just focusing on the
speed part of it that helps you get the right kind of intuition.
1:46:28
Forget the intelligence, just the speed.
1:46:29
- Because people understand
the power gap of time.
1:46:34
They understand that today we
have technology that was not around 1000 years ago and
that this is a big power gap and that it is bigger than, okay, so like what does smart mean?
1:46:45
When you ask somebody to imagine something that's more intelligent, what does that word mean
to them given the cultural associations that that
person brings to that word?
1:46:57
For a lot of people they
will think of, well, it sounds like a super chess player that went to double college.
1:47:07
And because we're talking
about the definitions of words here, that
doesn't necessarily mean that they're wrong.
1:47:13
It means that the word
is not communicating what I wanted to communicate.
1:47:19
The thing I want to communicate
is the sort of difference that separates humans from chimpanzees.
1:47:26
But that gap is so large that
you ask people to be like, well, human, chimpanzee, go another step along that interval,
around the same length and people's minds just go blank.
1:47:38
Like how do you even do that?
1:47:42
And I can try to break
it down and consider what it would mean to send a schematic foreign air conditioner
1000 years back in time. - (laughing) Yeah.
1:47:59
- Now I think that there's a sense in which you could redefine the word magic to refer
to this sort of thing.
1:48:05
And what do I mean by this
new technical definition of the word magic?
1:48:10
I mean that if you send a
schematic for the air conditioner back in time, they can see exactly what
you're telling them to do.
1:48:17
But having built this thing, they do not understand
how it output cold air, because the air conditioner design uses the relation between
temperature and pressure.
1:48:28
And this is not a law of
reality that they know about.
1:48:32
They do not know that when
you compress something, when you compress air or like coolant, it gets hotter and then you
can then like transfer heat from it to room temperature air, and then expand it again
and now it's colder and then you can like
transfer heat to that and generate cold air to blow out.
1:48:50
They don't know about any of that.
1:48:52
They're looking at a design
and they don't see how the design outputs cold air.
1:48:56
It uses aspects of reality
that they have not learned.
1:48:59
So magic in the sense is I
can tell you exactly what I'm going to do and even knowing
exactly what I'm going to do, you can't see how I got
the results that I got.
1:49:08
- That's a really nice example.
1:49:12
But is it possible to
linger on this defense?
1:49:16
Is it possible to have AGI
systems that help you make sense of that schematic weaker AGI systems? - Do you trust them?
1:49:23
- Fundamental part of building
up AGI is this question, can you trust the output of a system?
1:49:33
- Can you tell if it's lying?
1:49:36
- I think that's going to be
the smarter the thing gets, the more important that
question becomes, is it lying?
1:49:43
But I guess that's a really hard question. Is GPT lying to you?
1:49:47
Even now, GPT-4, is it lying to you?
1:49:49
- Is it using an invalid argument?
1:49:52
Is it persuading you via the
kind of process that could persuade you of false things
as well as true things?
1:50:00
Because the basic paradigm
of machine learning that we are presently operating under is that you can have the loss function, but only for things you can evaluate.
1:50:10
If what you're evaluating
is human thumbs up versus human thumbs down, you learn how to make the
human press thumbs up.
1:50:17
That doesn't mean that
you're making the human press thumbs up using the kind of rule that the human wants to be the case for what they press thumbs up on.
1:50:27
You know, maybe you're just
learning to fool the human.
1:50:31
- That's so fascinating and terrifying. The question of lying.
1:50:37
- On the present paradigm, what you can verify is
what you get more of.
1:50:43
If you can't verify, you can't ask the AI for it
'cause you can't train it to do things that you cannot verify.
1:50:51
Now this is not an absolute law, but it's like the basic dilemma here.
1:50:56
Maybe you can verify it
for simple cases and then scale it up without retraining it somehow.
1:51:06
Like by chain of thought, by like making the chains of
thought longer or something, and get more powerful stuff
that you can't verify but which is generalized from the
simpler stuff that did verify, and then the question is, did the alignment generalize
along with the capabilities?
1:51:23
But that's the basic dilemma on this whole paradigm of
artificial intelligence.
1:51:30
(Lex sighs) - It's such a difficult problem.
1:51:38
It seems like a problem of trying to understand the human mind.
1:51:47
- Better than the AI understand it, otherwise it has magic.
1:51:52
The same way that if you
are dealing with something smarter than you, then the same way as that 1000 years earlier they didn't know about the
temperature, pressure relation, it knows all kinds of stuff
going on inside your own mind of which you yourself are
unaware and it can output something that's going
to end up persuading you of a thing, and you
could see exactly what it did and still not know why that worked.
1:52:18
- So in response to your
eloquent description of why AI will kill us, Elon Musk
replied on Twitter, "Okay, so what should we do about it?" Question mark.
1:52:32
And you answered, "The game board has already been played "into a frankly awful state."
1:52:38
"There are not simple ways to
throw money at the problem.
1:52:42
"If anyone comes to you with a
brilliant solution like that, "please, please talk to me first.
1:52:47
"I can think of things I'd try; "they don't fit in one tweet." Two questions.
1:52:53
One, why has the game board,
in your view, been played into an awful state?
1:52:58
Just if you can give a
little bit more color to the game board and the
awful state of the game board.
1:53:05
- Alignment is moving like this, capabilities are moving like this.
1:53:10
- For the listener, capabilities
are moving much faster than the alignment. (both laughing) - Yeah.
1:53:18
- All right, so just the rate
of development, attention, interest, allocation of resources.
1:53:23
- We could have been
working on this earlier.
1:53:26
People are like, "Oh, but you know, ?
1:53:27
like how can you possibly
work on this earlier?"
1:53:30
'Cause they didn't want
to work on the problem.
1:53:33
They wanted an excuse to wave it off.
1:53:35
They like like, "Oh, how could we possibly "have worked on it earlier"
and didn't spend five minutes thinking about is there some
way to work on it earlier.
1:53:45
And frankly, it would've been hard.
1:53:47
Can you post bounties for
half of the (indistinct), if your plan is taking
this stuff seriously, can you post bounties for like
half of the people wasting their lives on string theory
to have gone into this instead and try to win a billion
dollars with a clever solution?
1:54:01
Only if you can tell which
solutions are clever. Which is hard.
1:54:06
But you know, the fact that
we didn't take it seriously. We didn't try.
1:54:12
It's not clear that we could
have done any better if we had, it's not clear how much
progress we could have produced if we had tried because it is
harder to produce solutions.
1:54:18
But that doesn't mean that
you're like correct and justified in letting everything slide.
1:54:22
It means that that things
are in a horrible state getting worse and there's
nothing you can do about it.
1:54:28
- So there's no brain
power making progress in trying to figure out
how to align these systems.
1:54:39
You're not investing money in it.
1:54:40
You don't have institution
infrastructure for like, even if you invest the money,
distributing that money across the physicists
working on strength theory, brilliant minds that are working into-- - How can you tell if
they're making progress?
1:54:54
You can like put, put them
all on interpretability.
1:54:58
'Cause when you have an
interpretability result, you can tell that it's
there and there's like, interpretability alone
is not going to save you.
1:55:04
We need systems that
will have a pause button where they won't try to prevent you from pressing the pause button.
1:55:14
'Cause they're like, oh well, I can't get my stuff done if I'm paused.
1:55:19
And that's a more difficult problem but it's like a fairly crisp
problem and you can maybe tell if somebody's made progress on it.
1:55:30
- So you can write and you
can work on the pause problem.
1:55:34
I guess more generally the pause button, more generally you can call
that the control problem.
1:55:38
- I don't actually like
the term control problem 'cause you know, it
sounds kind of controlling and alignment, not control.
1:55:45
You're not trying to take
a thing that disagrees with you and whip it back onto, make it do what you wanted
to do even though it wants to do something else.
1:55:53
You're trying to, in the
process of its creation, choose its direction. - Sure.
1:55:59
But we currently, in a lot
of the systems we design, we do have an off switch.
1:56:05
That's a fundamental part of-- - It's not smart enough to
prevent you from pressing the off switch and
probably not smart enough to want to prevent you from
pressing the off switch.
1:56:16
- So you're saying the kind of
systems we're talking about, even the philosophical concept
of an off switch doesn't make any sense because-- - Well no, the off switch makes sense.
1:56:25
They're just not opposing your attempt to pull the off switch.
1:56:32
Parenthetically, like don't
kill the system if you're...
1:56:37
Like if we're getting to
the part where this starts to actually matter and it's
like where they can fight back, like don't kill them
and dump their memory.
1:56:45
Save them to disk, don't
kill them, be nice here.
1:56:50
- Well, okay, be nice is a
very interesting concept here.
1:56:52
We're talking about a system
that can do a lot of damage.
1:56:57
I don't know if it's possible, but it's certainly one of
the things you could try is to have an off switch.
1:57:01
- It's suspend to disk switch.
1:57:06
- You have this kind of
romantic attachment to the code.
1:57:09
Yes, if that makes sense.
1:57:11
But if it's spreading, you don't want suspend to disk, right?
1:57:16
There's something fundamentally-- - If it gets that far of hand, then yes.
1:57:21
Pull the plug on everything
it's running on, yes.
1:57:24
- I think it's a research question.
1:57:25
Is it possible in AGI systems, AI systems to have a
sufficiently robust off switch that cannot be manipulated, that cannot be manipulated
by the AI system?
1:57:40
- Then it escapes from whichever system you've built the almighty lever into and copies itself somewhere else.
1:57:46
- So your answer to that
research question is no. - Obviously, yeah.
1:57:51
- But I don't know if that's
a hundred percent answer.
1:57:54
Like, I don't know if it's obvious.
1:57:56
- I think you're not putting
yourself into the shoes of the human in the world
of glacially slow aliens.
1:58:05
- But the aliens built
me, let's remember that. - [Eliezer] Yeah.
1:58:10
- And they built the box I'm in. - [Eliezer] Yeah.
1:58:14
- To me it's not obvious.
1:58:15
- They're slow and they're stupid.
1:58:17
- I'm not saying this is guaranteed, but I'm saying it's a
non zero probability.
1:58:20
It's an interesting research question.
1:58:21
Is it possible when you're slow
and stupid, to design a slow and stupid system that is
impossible to mess with?
1:58:30
- The aliens, being as stupid as they are, have actually put you on
Microsoft Azure Cloud servers instead of this hypothetical perfect box.
1:58:41
That's what happens when
the aliens are stupid.
1:58:45
- Well, but this is not AGI, right?
1:58:46
This is their early
versions of the system. As you start to...
1:58:50
- Yeah, you think that
they've got like a plan where they have declared a
threshold level of capabilities where it passed that capabilities, they move it off the cloud
servers and onto something that's air gapped?
1:59:02
(Eliezer laughing mockingly) - I think there's a lot of people, and you're an important voice here.
1:59:07
There's a lot of people that
have that concern and yes, they will do that when there's
an uprising of public opinion that that needs to be done.
1:59:14
And when there's actual
little damage done, when the holy shit,
this system is beginning to manipulate people, then there's going to be an
uprising where there's going to be a public pressure
and a public incentive in terms of funding, in developing things that
can off switch or developing aggressive alignment mechanisms.
1:59:35
And no, you're not
allowed to put on Azure-- - Aggressive alignment mechanism?
1:59:38
What the hell is aggressive
alignment mechanisms?
1:59:40
Like it doesn't matter
if you say aggressive, we don't know how to do it.
1:59:44
- Meaning aggressive alignment, meaning you have to propose
something, otherwise you're not allowed to put it on the cloud.
1:59:53
- The hell do you, do you imagine they will
propose that would make it safe to put something smarter
than you on the cloud?
1:59:59
- That's what research is for.
2:00:00
Why the cynicism about such
a thing not being possible?
2:00:04
If you have intelligence-- - That works on the first try? - [Lex] What? So yes. So yes.
2:00:08
- Against something smarter than you?
2:00:10
- So that is the fundamental thing.
2:00:13
If there's a rapid takeoff,
yes, it's very difficult to do.
2:00:18
If there's a rapid takeoff
and the fundamental difference between weak AGI and strong
AGI as you're saying, that's going to be
extremely difficult to do.
2:00:25
If the public uprising never
happens until you have this critical phase shift, then you're right.
2:00:32
It's very difficult to do. But that's not obvious.
2:00:34
It's not obvious that you're
not going to start seeing symptoms of the negative
effects of AGI to where you're like, we have to put a halt to this, that there is not just first try.
2:00:42
You get many tries at it.
2:00:44
- Yeah, we can see right now that Bing is quite difficult to align.
2:00:50
That when you try to train
inabilities into a system into which capabilities
have already been trained, that what do you know, gradient descent, like learns small, shallow,
simple patches of inability and you come in and ask
it in a different language and the deep capabilities
are still in there and they evade the shallow patches and come right back out again. There, there you go.
2:01:11
There's your red fire alarm of oh no, alignment is difficult.
2:01:16
Is everybody gonna shut everything down? No.
2:01:19
- No, but that's not the
same kind of alignment.
2:01:21
A system that escapes the box
it's from is a fundamentally different thing, I think. - For you.
2:01:28
- Yeah, no, but for the system-- - So you put a line there and
everybody else puts a line somewhere else and there's like, yeah, and there's no agreement.
2:01:37
We have had a pandemic on
this planet with a few million people dead, which we may never know whether or not it was a lab leak because there was definitely coverup.
2:01:50
We don't know that if
there was a lab leak, But we know that the
people who did the research put out the whole paper
about this definitely wasn't the lab leak and didn't reveal
that they had been doing, had like sent off coronavirus research to the Wuhan Institute of Virology after it was banned in the United States after the gain of function
research was temporarily banned at the United States.
2:02:11
And the same people who exported
gain of function research on coronaviruses to the
Wuhan Institute of Virology after gain of function, that
gain of function research was temporarily banned
in the United States, are now getting more grants to do more research on gain of function
research on coronaviruses.
2:02:32
Maybe we do better in this than in AIi, but this is not something,
we cannot take for granted that there's going to be an outcry.
2:02:39
People have different thresholds for when they start to outcry.
2:02:44
- Can't take it for granted.
2:02:45
But I think your intuition
is that there's a very high probability that this event happens without us solving the alignment problem.
2:02:52
And I guess that's where
I'm trying to build up more perspectives and color on this intuition.
2:02:59
Is it possible that the
probability is not something like 100%, but is like 32% that AI will escape the box before we
solve the alignment problem?
2:03:11
Not solve, but is it possible we always stay ahead of the AI in terms of our ability to solve
for that particular system, the alignment problem?
2:03:22
- Nothing like the world
in front of us right now.
2:03:25
You've already seen it that GPT-4 is not turning out this way.
2:03:32
And there are basic
obstacles where you've got the weak version of the system
that doesn't know enough to deceive you and the
strong version of the system that could deceive you
if it wanted to do that.
2:03:44
If it was already like
sufficiently unaligned to want to deceive you.
2:03:47
There's the question of
how on the current paradigm you train honesty when the
humans can no longer tell if the system is being honest.
2:03:58
- You don't think these
are research questions that could be answered.
2:04:00
- I think they could be answered
in 50 years with unlimited retries, the way things
usually work in science.
2:04:07
- I just disagree with that.
2:04:08
You're making it 50 years.
2:04:09
I think with the kind
of attention this gets, with the kind of funding it gets, it could be answered not in whole, but incrementally within months and within a small number
of years if it at scale receives attention and research.
2:04:26
And so if you start starting
large language models, I think there was an
intuition like two years ago, even, that something like GPT-4, the current capabilities of
even ChatGPT with GPT-3.
2:04:34
5 we're still far away from that.
2:04:42
I think a lot of people are
surprised by the capabilities of GPT-4, right?
2:04:45
So now people are waking up, okay, we need to study these language models.
2:04:49
I think there's going to
be a lot of interesting AI safety research.
2:04:54
- Are Earth's billionaires going to put up the giant prizes that would
maybe incentivize young hotshot people who just got their
physics degrees to not go to the hedge funds and
instead put everything into interpretability in
this like one small area where we can actually tell whether or not somebody has
made a discovery or not?
2:05:13
- I think so-- - [Eliezer] When?
2:05:16
- Well, this is what these
conversations are about because they're going
to wake up to the fact that GPT-4 can be used
to manipulate elections, to influence geopolitics, to influence the economy.
2:05:29
There's going to be a huge
amount of incentive to, wait a minute, we have to make sure they're not doing damage.
2:05:38
We have to make sure we interpretability, we have to make sure we
understand how these systems function so that we can
predict their effect on economy, so that there's-- - So there's a futile moral panic-- - [Lex] Fairness and safety.
2:05:49
- And a bunch of op-eds
in the "New York Times" and nobody actually
stepping forth and saying, "You know what, instead of a mega yacht, "I'd rather put that
billion dollars on prizes "for young hotshot physicists "who make fundamental
breakthroughs in interpretability."
2:06:08
- The yacht versus the
interpretability research, the old trade off.
2:06:12
(Lex laughing) I think there's going to be a huge amount of allocation of funds. I hope, I hope, I guess.
2:06:20
- You wanna bet me on that?
2:06:22
You wanna put a time scale on it.
2:06:24
Say how much funds you think
are going to be allocated in a direction that I would consider to be actually useful by what time?
2:06:32
- I do think there will
be a huge amount of funds.
2:06:36
But you're saying it
needs to be open, right?
2:06:39
The development of the
systems should be closed.
2:06:41
But the development of the
interpretability research, the AI safety research-- - So we are so far behind
on interpretability compared to capabilities.
2:06:52
Like yeah, you could take the
last generation of systems.
2:06:56
The stuff that's already in the open, there is so much in there
that we don't understand.
2:07:00
There are so many prizes
you could do before you would have enough
insights that you'd be like, "Oh, we understand how these systems work.
2:07:09
"We understand how these
things are doing their outputs.
2:07:11
"We can read their minds, "now let's try it with
the bigger systems." We're nowhere near that.
2:07:16
There is so much
interpretability work to be done on the weaker versions of the systems.
2:07:20
- So what can you say on
the second point you said to Elon Musk on what are some ideas, what are things you could try?
2:07:30
"I can think of a few
things I'd try," you said, "They don't fit in one tweet."
2:07:34
So is there something
you could put into words of the things you would try?
2:07:39
- I mean, the trouble
is the stuff is subtle.
2:07:44
I've watched people try
to make progress on this and not get places.
2:07:48
Somebody who just gets
alarmed and charges in, it's like going nowhere. - [Lex] True.
2:07:54
- It meant like years ago, I don't know, like 20 years, 15 years,
something like that.
2:07:59
I was talking to a congressperson who had become alarmed
about the eventual prospects and he wanted work on
building AIs without emotions because the emotional AI
were the scary ones, you see.
2:08:18
And some poor person at ARPA had come up with a research proposal
whereby this congressman's panic and desire to fund this, the thing, would go into something
that the person at ARPA thought would be useful and had
been munched around to where it would sound to the
congressman like work was happening on this.
2:08:36
Which you know, of course
the congressperson had misunderstood the problem and did not understand where the danger came from.
2:08:46
And so it's like the issue is
that you could like do this in a certain precise way
and maybe get something.
2:08:55
When I say put up prizes
on interpretability, I'm like, because it's verifiable there as opposed to other places, you can tell whether or not good work actually happened in
this exact narrow case.
2:09:11
If you do things in exactly the right way, you can maybe throw money
at it at and produce science instead of anti-science and nonsense and all the methods that
I know of, trying to throw money at this problem,
share this property of, well if you do it exactly
right based on understanding exactly tends to produce
like useful outputs or not, then you can add money to it in this way.
2:09:37
And the thing that I'm giving
as an example here in front of this large audience
the most understandable of those because there's other
people who, like Chris Olah, and even more generally, you can tell whether or not interpretability progress has occurred.
2:09:56
So like if I say throw
money at producing more interpretability, there's
a chance somebody can do it that way and it will actually
produce useful results.
2:10:04
Then the other stuff just
blurs off into be like, harder to target exactly than that.
2:10:09
- So sometimes the
basics are fun to explore because they're not so basic.
2:10:16
What is interpretability? What does it look like?
2:10:20
What are we talking about?
2:10:22
- It looks like we took a much smaller set of transformer layers than
the ones in the modern bleeding edge state-of-the-art systems.
2:10:35
And after applying various
tools and mathematical ideas and trying 20 different things, we have shown it that this
piece of the system is doing this kind of useful work.
2:10:51
- And then somehow also
hopefully generalizes some fundamental understanding of what's going on that
generalizes to the bigger system.
2:11:01
- You can hope, and it's probably true.
2:11:03
Like you would not expect
the smaller tricks to go away when you have a system that's
doing larger kinds of work, you would expect the larger
work kinds of work to be building on top of the smaller
kinds of work and gradient descent runs across the smaller
kinds of work before it runs across the larger kinds of work.
2:11:21
- Well, that's kind of what is happening in neuroscience, right?
2:11:24
It's trying to understand
the human brain by prodding and it's such a giant mystery and people have made progress, even though it's extremely
difficult to make sense of what's going on in the brain.
2:11:33
They have different parts of the brain that are responsible
for hearing, for sight.
2:11:36
The vision, science community,
they're just understanding the visual cortex.
2:11:40
I mean they've made a lot of
progress in understanding how that stuff works, but you're
saying it takes a long time to do that work well. - Also it's not enough.
2:11:49
So in particular, let's say you have got your interpretability tools and they say that your current AI system
is plotting to kill you. Now what?
2:12:09
- It is definitely a good step one, right? - [Eliezer] Yeah. What's step two?
2:12:16
- If you cut out that
layer, is it gonna stop wanting to kill you?
2:12:21
- When you optimize against
visible misalignment, you are optimizing against misalignment and you are also optimizing
against visibility.
2:12:34
So sure, if you can-- (Lex laughing) - It's true.
2:12:38
All you're doing is removing
the obvious intentions to kill you.
2:12:42
- You've got your detector, it's showing something inside the system that you don't like.
2:12:47
Okay, say the disaster
monkey is running this thing, we'll optimize the system until the visible bad behavior goes away.
2:12:54
But it's arising for fundamental reasons of instrumental convergence.
2:12:59
The old, you can't bring
the coffee if you're dead, any goal and you know, almost
every set of utility functions with a few narrow exceptions
implies killing all the humans.
2:13:12
- But do you think it's
possible, because we can do experimentation to discover the source of the desire to kill?
2:13:19
- I can tell it to you right now, is that it wants to do something
and the way to get the most of that thing is to put the
universe into a state where there aren't humans.
2:13:30
- So is it possible to encode
in the same way we think?
2:13:34
Like why do we think murder is wrong?
2:13:37
The same foundational ethics,
that's not hard coded in but more like deeper.
2:13:45
I mean that's part of the research.
2:13:46
How do you have it that this transformer, this small version of the language model doesn't ever want to kill?
2:13:59
- That'd be nice assuming that you got "doesn't want to kill"
sufficiently exactly right.
2:14:05
That it didn't be like, "Oh, I will detach their heads
and put them in some jars "and keep the heads alive forever
and then go do the thing."
2:14:12
But leaving that aside,
well, not leaving that aside.
2:14:15
- [Lex] Yeah, that's good,
it gets a strong point, yeah.
2:14:17
- 'Cause there is a whole issue where as something gets smarter, it finds ways of achieving the
same goal predicate that were not imaginable to stupider
versions of the system or perhaps the stupider operators.
2:14:31
That's one of many things
making this difficult.
2:14:34
A larger thing making this
difficult is that we do not know how to get any goals into systems at all.
2:14:39
We know how to get outwardly
observable behaviors into systems.
2:14:43
We do not know how to get
internal psychological wanting to do particular things into the system.
2:14:50
That is not what the
current technology does.
2:14:54
- I mean, it could be things
like dystopian futures, like "Brave New World" where
most humans will actually say, "We kind of want that future." It's a great future. Everybody's happy.
2:15:06
- We would have to get
so far, so much further than we are now and further faster
before that failure mode became a running concern.
2:15:17
- Your failure modes
are much more drastic.
2:15:20
The ones you're-- - The failure modes are much simpler.
2:15:22
It's like yeah, the AI puts the universe into a particular state.
2:15:26
It happens to not have
any humans inside it.
2:15:28
- Okay, so the paper club maximizer.
2:15:31
- Utility, so the original version of the paperclip maximizer-- - Can you explain it if you can? - Okay.
2:15:37
The original version was you lose control of the utility function and it so happens that
what maxes out the utility per unit resources is tiny
molecular shapes like paperclips.
2:15:52
There's a lot of things
that make it happy, but the cheapest one that didn't
saturate was putting matter into certain shapes.
2:16:02
And it so happens that the cheapest way to make these shapes is
to make them very small, 'cause then you need fewer atoms, per instance of the shape and arguendo, it happens to look like a paperclip.
2:16:14
In retrospect I wish I'd
said tiny molecular spirals or tiny molecular hyperbolic spirals. Why?
2:16:22
Because I said tiny molecular paperclips, this got then mutated to paperclips, this then mutated to, "And the AI was in a paperclip factory."
2:16:33
So the original story is
about how you lose control of the system, it doesn't want what you
tried to make it want.
2:16:39
The thing that it ends up wanting most is a thing that even from a very embracing cosmopolitan perspective, we
think of as having no value and that's how the value of
the future gets destroyed.
2:16:49
Then that got changed to a fable of, well, you made a paperclip
factory and it did exactly what you wanted, but you asked
it to do the wrong thing.
2:16:58
Which is a completely
different failure path.
2:17:02
(Eliezer sighs) - But those are both concerns to you.
2:17:09
So that's more than-- - If you "Brave New World."
2:17:11
If you can solve the problem
of making something want what exactly what you want it to want, then you get to deal with the problem of wanting the right thing.
2:17:22
- But first you have
to solve the alignment.
2:17:24
- First you have to solve inner alignment. - [Lex] Inner alignment.
2:17:27
- Then you get to solve outer alignment.
2:17:31
First you need to be
able to point the insides of the thing in a direction
and then you get to deal with whether that direction expressed
in reality is the thing that it aligned with
the thing that you want. - Are you scared? - Of this whole thing? Probably. I don't really know.
2:17:54
- What gives you hope about this?
2:17:57
- [Eliezer] The
possibility of being wrong.
2:17:59
- Not that you're right, but we will actually get our
act together and allocate a lot of resources to
the alignment problem.
2:18:07
- Well, I can easily imagine
that at some point this panic expresses itself in the
waste of a billion dollars.
2:18:15
Spending a billion dollars
correctly, that's harder.
2:18:18
- To solve both the inner
and the outer alignment.
2:18:21
If you're wrong--
- To solve a number of things. - Yeah. Number of things.
2:18:24
If you're wrong, what do you
think would be the reason?
2:18:30
Like if 50 years from
now, not perfectly wrong, you make a lot of really eloquent points, there's a lot of shape
to the ideas you express.
2:18:43
But if you're somewhat wrong
about some fundamental ideas, why would that be?
2:18:48
- Stuff has to be easier
then I think it is.
2:18:52
The first time you're building a rocket, being wrong is in a
certain sense quite easy.
2:18:59
Happening to be wrong in a way where the rocket goes twice
as far on half the fuel and lands exactly where
you hoped it would?
2:19:06
Most cases of being wrong make
it harder to build a rocket, harder to have it not explode.
2:19:10
'Cause it to require more
fuel than you hope to, cause it to be led off target.
2:19:15
Being wrong in a way
that makes stuff easier, that's not the usual
project management story.
2:19:22
- And then this is the first
time we're really tackling the problem of a AI alignment.
2:19:26
There's no examples in
in history where we...
2:19:28
- Oh, there's all kinds of
things that are similar if you generalize and correctly the
right way and aren't fooled by misleading metaphors. - Like what?
2:19:36
- Humans being misaligned on
inclusive genetic fitness.
2:19:39
So inclusive genetic fitness
is like not just your reproductive fitness, but also the fitness of your relatives, the people who share some
fraction of your genes.
2:19:49
The old joke is, would you give your life
to save your brother?
2:19:53
They once asked a biologist, I think it was Haldane, and Haldane said, "No, but I would give my
life to save two brothers "or eight cousins."
2:20:02
Because a brother on average
shares half your genes.
2:20:05
And cousin on average shares
an eighth of your genes.
2:20:08
So that's inclusive genetic fitness.
2:20:10
And you can view natural
selection as optimizing humans exclusively around this
one very simple criterion, like how much more frequent
did your genes become in the next generation?
2:20:23
In fact, that just is natural selection.
2:20:25
It doesn't optimize for that.
2:20:27
But rather the process of
genes becoming more frequent is that you can nonetheless imagine that there is this hill climbing process, not like gradient descent, because gradient descent uses calculus.
2:20:38
This is just using like where are you?
2:20:40
But still hill climbing in both cases, make things something better
and better over time, in steps.
2:20:47
And natural selection was
optimizing exclusively for this very simple, pure criterion of inclusive genetic fitness in a very complicated environment, we're doing a very wide range
of things and solving a wide range of problems, led
it to having more kids.
2:21:06
And this got you humans, which had no internal notion
of inclusive genetic fitness until thousands of years
later when they were actually figuring out what had even happened.
2:21:20
And no explicit desire to increase inclusive genetic fitness.
2:21:26
So from this important case study, we may infer the important fact
that if you do a whole bunch of hill climbing on a
very simple loss function, at the point where the
system's capabilities start to generalize very widely, when it is in an intuitive
sense becoming very capable
2:21:47
and generalizing far outside
the training distribution, we know that there is no general
loss saying that the system even internally represents, let alone tries to optimize
the very simple loss function you are training it on. - There is so much that we
cannot possibly cover all of it.
2:22:04
- There is so much that we
cannot possibly cover all of it.
2:22:06
I think we did a good job
of getting your sense from different perspectives of
the current state of the art with large language models.
2:22:15
We got a good sense of your concern about the threats of AGI.
2:22:23
- I've talked here about
the power of intelligence and not really gotten very far into it, but not like, why it is.
2:22:31
Suppose you screw up with
AGI and it end up wanting a bunch of random stuff.
2:22:36
Why does it try to kill you?
2:22:40
Why doesn't it try to trade with you?
2:22:43
Why doesn't it give you just
the tiny little fraction of the solar system that it would take to keep everyone alive?
2:22:51
- Yeah well, that's a good question.
2:22:53
What are the different
trajectories that intelligence, when acted upon this
world, super intelligence, what are the different
trajectories for this universe with such an intelligence in it?
2:23:02
Do most of them not include humans?
2:23:05
- I mean, the vast majority
of randomly specified utility functions do not have
optima with humans in them, would be the first
thing I would point out.
2:23:18
And then the next question is like, well, if you try to optimize something, you lose control of it.
2:23:23
Where in that space do you land?
2:23:24
'Cause it's not random, but it also doesn't necessarily
have room for humans in it.
2:23:29
I suspect that the average
member of the audience might have some questions about even
whether that's the correct paradigm to think about
it and would sort of want to back up a bit possibly.
2:23:39
- If we back up to something
bigger than humans, if we look at Earth and life
on Earth and what is truly special about life on Earth, do you think it's possible that whatever that special thing is, let's explore what that
special thing could be.
2:24:00
Whatever that special thing is, that thing appears often
in the objective function. - Why?
2:24:08
I know what you hope, but you know, you can hope that a particular set of winning lottery numbers
come up and it doesn't make the lottery balls come up that way.
2:24:18
I know you want this to be
true, but why would it be true?
2:24:21
- There's a line from "Grumpy Old Men" where this guy says, in
a grocery store, he says, "You can wish in one hand
and crap in the other "and see which one fills up first."
2:24:31
- There's a science problem.
2:24:32
We are trying to predict
what happens with AI systems that you tried to
optimize to imitate humans and then you did some of RLHF to them and of course you didn't
get like perfect alignment because that's not what happens when you hill climb towards
a outer loss function.
2:24:54
You don't get inner alignment on it.
2:24:56
But yeah, so if you
don't mind my like taking some slight control of
things and steering around to what I think is like
a good place to start.
2:25:10
- I just failed to solve
the control problem.
2:25:12
I've lost control of this thing. - Alignment. Alignment. - Still aligned.
2:25:16
(laughing) - Yeah, okay, sure. Yeah, you lost control.
2:25:20
- But we're still aligned.
2:25:22
Anyway, sorry for the meta comment.
2:25:24
- Yeah, losing control isn't as bad as you lose control to an aligned system. - [Lex] Yes, exactly.
2:25:29
- You have no idea of the horrors I will shortly unleash
on this conversation.
2:25:32
(both laughing) - All right.
2:25:35
Sorry, sorry to distract you completely.
2:25:37
What were you gonna say
in terms of taking control of the conversation?
2:25:40
- So I think that there's like a (speaking foreign language) here, if I'm pronouncing those
words remotely like correctly, 'cause of course we only ever read them and not hear them spoken.
2:25:56
For some people, the word
intelligence, smartness is not a word of power to them.
2:26:02
It means chess players,
it means the college university professor, people
aren't very successful in life.
2:26:09
It doesn't mean like charisma
to which my usual thing is like charisma is not
generated in the liver rather than the brain.
2:26:15
Charisma is also a cognitive function.
2:26:21
So if you think that smartness doesn't sound very threatening, then super intelligence is not gonna sound very
threatening either.
2:26:30
It's gonna sound like you
just pull the off switch.
2:26:34
Well, it's super intelligent
but it's stuck in a computer.
2:26:36
We pull the off switch, problem solved.
2:26:40
And the other side of it is
you have a lot of respect for the notion of intelligence.
2:26:45
You're like, well yeah,
that's what humans have.
2:26:47
That's the human superpower.
2:26:49
And it sounds like it could be dangerous, but why would it be?
2:26:57
We, as we have grown more intelligent, also grown less kind.
2:27:02
Chimpanzees are in fact a
bit less kind than humans and you know, you could argue that out.
2:27:08
But often the sort of person
who has a deep respect for intelligence is gonna be like, "Well, yes, "you can't even have kindness
unless you know what that is."
2:27:17
And so they're like, why would it do something as
stupid as making paperclips?
2:27:23
Aren't you supposing something
that's smart enough to be dangerous but also stupid
enough that it will just make paperclips and never question that?
2:27:31
In some cases people are like, "Well, even if you misspecify
the objective function, "won't you realize that what
you really wanted was x?
2:27:41
"Are you supposing something
that is smart enough to be "dangerous but stupid enough
that it doesn't understand what "the humans really meant
when they specified "the objective function?"
2:27:52
- So to you, our intuition
about intelligence is limited.
2:27:57
We should think about intelligence
as a much bigger thing.
2:28:00
- Well, I'm saying that it's that-- - [Lex] That humanness.
2:28:03
- Well, what what I'm saying
is like what do you think about artificial intelligence?
2:28:09
Depends on what you
think about intelligence.
2:28:11
- So how do we think about
intelligence correctly?
2:28:13
Like you gave one thought
experiment to think of, think of a thing that's much
faster, so it just gets faster and faster, faster and faster.
2:28:22
- And it also is like is
made of John von Neumann and there's lots of them. - [Lex] Or (indistinct).
2:28:28
- Yeah, John von Neumann
is a historical case, so you can like look up what he did and imagine based on that.
2:28:34
And we know people have
some intuition for like, if you have more humans, they can solve tougher cognitive problems.
2:28:42
Although in fact like in the game of Kasparov versus the World which was like Garry Kasparov
on one side and an entire hoard of internet people led
by four chess grand masters on the other side, Kasparov won.
2:28:57
So like all those people
aggregated to be smarter, it was a hard fought game.
2:29:03
It's like all those people
aggregated to be smarter than any individual one of them, but they didn't aggregate so well that they could defeat Kasparov But so humans aggregating
don't actually get, in my opinion, very much smarter, especially compared to
running them for longer.
2:29:19
The difference between capabilities
now and a thousand years ago is a bigger gap than
the gap in capabilities between 10 people and one person.
2:29:29
But even so, pumping
intuition for what it means to augment intelligence, John von Neumann, there's millions of him, he runs at a million times the
speed and therefore can solve tougher problems, quite a lot tougher.
2:29:48
- It's very hard to have an intuition about what that looks like, especially like you said, the intuition, I kind of think about is
it maintains the humanness.
2:30:04
I think it's hard to separate my hope from my objective intuition about what super intelligent
systems look like.
2:30:17
- If one studies evolutionary
biology with a bit of math and in particular books
from when the field was just sort of properly coalescing
and knowing itself, like not the modern textbooks, which are just memorize this legible math.
2:30:37
so you can do well on these tests, but what people were writing
as the basic paradigms of the field were being fought out...
2:30:43
A nice book if you've
got the time to read it is "Adaptation and Natural Selection," Which is one of the founding books, you can find people being optimistic about what the utterly
alien optimization process of natural selection will
produce in the way of how it optimizes its objectives.
2:31:06
You got people arguing that
like, in the early days biologists said, "Well,
organisms will restrain "their own reproduction
when resources are scarce "so as not to overfeed the system."
2:31:21
And this is not how
natural selection works, it's about whose genes are
relatively more prevalent to the next generation.
2:31:30
And if you restrain reproduction, those genes get less frequent
in the next generation compared to your conspecifics.
2:31:39
And natural selection doesn't do that.
2:31:42
In fact, predators overrun
prey populations all the time and have crashes.
2:31:47
That's just a thing that
happens and many years later... Well oh, oh oh.
2:31:51
But people said like,
"Well, but group selection." Right?
2:31:54
What about groups of organisms?
2:31:56
And basically the math of group selection almost never works out in
practice is the answer there.
2:32:02
But also years later, somebody actually ran the
experiment where they took populations of insects and
selected the whole populations to have lower sizes.
2:32:14
Now you just take pop
one, pop two, pop three, pop four look at which has the
lowest total number of them in the next generation
and select that one.
2:32:22
What do you suppose happens
when you select populations of insects like that?
2:32:26
Well, what happens is
not that the individuals in the population evolve
to restrain their breeding, but that they evolved
to kill the offspring of other organisms, especially the girls.
2:32:36
So people imagined this lovely, beautiful, harmonious output of natural selection, which is these populations
restraining their own breeding so that groups of them
would stay in harmony with the resources available.
2:32:50
And mostly the math
never works out for that.
2:32:52
But if you actually apply
the weird strange conditions to get group selection that
beats individual selection, what you get is female infanticide.
2:33:01
Like if you're reading on
restrained populations.
2:33:07
So this is not a smart
optimization process.
2:33:09
Natural selection is like so
incredibly stupid and simple that we can actually
quantify how stupid it is if you read the textbooks with the math.
2:33:16
Nonetheless, this is
the sort of basic thing of you look at this alien
optimization process and there's the thing that
you hope it will produce and you have to learn to
clear that out of your mind and just think about the
underlying dynamics and where it finds the maximum from its
standpoint that it's looking for rather than how it finds
that thing that lept into your mind as the beautiful aesthetic solution that you hope it finds.
2:33:42
And this is something that was, has been fought out
historically as the field of biology was coming to terms
with evolutionary biology.
2:33:52
And you can like look at
them fighting it out as they get to terms with this very alien, inhuman optimization process.
2:34:01
And indeed something smarter
than us would be also much smarter than natural selection.
2:34:06
So it doesn't just
automatically carry over.
2:34:09
But there's a lesson
there, there's a warning.
2:34:13
- To you, natural selection
is a deeply suboptimal process that could be significantly improved on and would be by an AGI system.
2:34:21
- Well, it's kind of stupid.
2:34:22
It has to run hundreds
of generations to notice that something is working.
2:34:28
It doesn't be like, oh, well I tried this in one
organism, I saw it worked, now I'm going to duplicate that feature onto everything immediately.
2:34:37
Has to run for hundreds of generations for a new mutation tries to fixation.
2:34:41
- I wonder if there's a case to be made in natural selection, as inefficient as it looks
is actually quite powerful.
2:34:54
That this is extremely robust.
2:34:56
- It runs for a long time and eventually manages to optimize things.
2:35:02
It's weaker than gradient
dissent because gradient dissent also uses information
about the derivative.
2:35:08
- Yeah, evolution seems to be, there's not really an objective function.
2:35:13
- [Eliezer] There's
inclusive genetic fitness is the implicit loss
function of evolutions.
2:35:18
- It's implicit
- It cannot change.
2:35:19
The loss function doesn't change, the environment changes
and therefore what gets optimized for in the organism changes. Take like GPT-3.
2:35:30
You imagine like different
versions of GPT-3 where they're all trying
to predict the next word, but they're being run on
different data sets of text and that's like natural selection, always inclusive genetic fitness but different environmental problems.
2:35:46
(Lex exhales) - It's difficult to think about.
2:35:50
So if we are saying the
natural selection is stupid, if we're saying the
humans are stupid, it's-- - Smarter than natural selection, stupider than the upper bound.
2:36:02
- Do you think there's an
upper bound by the way?
2:36:04
That's another helpful place.
2:36:06
- I mean, if you put enough matter, energy compute into one place, it will collapse into a black hole.
2:36:13
(Lex laughing) There's only so much computation
can do before you run out of negentropy and the universe dies.
2:36:19
So there's an upper bound, but it's very, very,
very far up above here.
2:36:23
Like the supernova is only finitely hot, it's not infinitely hot but it's really, really,
really, really hot.
2:36:32
- Well, let me ask you, let me talk to you about
consciousness, also coupled with that question is imagining a world with superintelligent AI systems that get rid of humans
but nevertheless keep something that we would
consider beautiful and amazing. - Why?
2:36:53
The lesson of evolutionary biology.
2:36:55
If you just guess what
an optimization does based on what you hope
the results will be, it usually will not do that. - It's not hope. I mean it's not hope.
2:37:03
I think if you objectively look at what has been a powerful, a useful...
2:37:12
I think there's a correlation
between what we find beautiful and a thing that's been useful.
2:37:18
- This is what the early
biologists thought.
2:37:22
And not just like, they thought.
2:37:24
Like "No, no, I'm not just imagining stuff "that would be pretty,
it's useful for organisms "to restrain their own reproduction "because then they don't overrun "the prey populations and
they actually have more kids "in the long run." - Hmm.
2:37:39
So let me just ask you
about consciousness.
2:37:43
Do you think consciousness is useful. - [Eliezer] To humans? - No, to AGI systems.
2:37:49
Well, in this transitionary
period between humans and AGI, to AGI systems as they
become smarter and smarter, is there some use to it? Let me step back. What is consciousness?
2:38:03
Eliezer Yudkowsky, what is consciousness?
2:38:06
- Are referring to
Chalmers as hard problem of conscious experience?
2:38:11
Are you referring to
self-awareness and reflection?
2:38:15
Are you referring to
the state of being awake as opposed to asleep?
2:38:20
- This is how I know you're
an advanced language model.
2:38:22
I gave you a simple prompt and you gave me a bunch of options.
2:38:30
I think I'm referring to all, including the hard
problem of consciousness.
2:38:38
What is it in its importance
to what you've just been talking about, which is intelligence?
2:38:44
Is it a foundation to intelligence?
2:38:48
Is it intricately
connected to intelligence in the human mind or is it a
side effect of the human mind?
2:38:56
It is a useful little tool,
like we can get rid of?
2:39:00
I guess I'm trying to get
some color in your opinion of how useful it is in the intelligence of a human being and then
try to generalize that to AI, whether AI will keep some of that.
2:39:15
- So I think that for
there to be like a person who I care about looking
out at the universe and wondering at it and appreciating it, it's not enough to have
a model of yourself.
2:39:31
I think that it is useful to
an intelligent mind to have a model of itself, but I think
you can have that without pleasure, pain, aesthetics, emotion, a sense of wonder.
2:39:56
Like I think you can have
a model of how much memory you're using and whether
this thought or that thought is like more likely to
lead to a winning position.
2:40:13
I think that if you optimize
really hard on efficiently just having the useful parts, there is not then the
thing that says like, "I am here, I look out, I
wonder, I feel happy on this. "I feel sad about that."
2:40:33
I think there's a thing that
knows what it is thinking but it doesn't quite care
about these are my thoughts, this is my me and that matters.
2:40:46
- Does that make you sad, if that's lost in AGI?
2:40:49
- I think that if that's
lost then basically everything that matters is lost.
2:40:58
I think that when you optimize, when you go really hard on making tiny molecular spirals or paperclips, that when you grind
much harder than on that than natural selection
ground out to make humans, that there isn't then the
mess and intricate loopiness and complicated pleasure,
pain, conflicting preferences, this type of feeling,
that kind of feeling.
2:41:36
In humans there's this difference between the desire of wanting
something and the pleasure of having it and it's all
these like evolutionary clutches that came together
and created something that then looks at itself and says like, "This is pretty, this matters."
2:41:53
And the thing that I worry
about is that this is not the thing that happens again, just the way that happens in us or even quite similar enough.
2:42:05
That there are many
basins of attractions here and we are in the space of attraction looking out and saying like, "Ah, what a lovely basin we are in" and there are other basins of attraction and the AIs do not end up in
this one when they go like, way harder on optimizing themselves, the natural selection optimized us.
2:42:26
'Cause unless you specifically
want to end up in the state where you are looking
out saying, "I am here," "I look out at this universe with wonder," if you don't want to preserve that, it doesn't get preserved
when you grind really hard on being able to get more of the stuff.
2:42:44
We would choose to preserve
that within ourselves because it matters and on some viewpoints is the only thing that matters.
2:42:52
- And preserving that
is in part a solution to the human alignment problem.
2:43:02
- I think the human alignment problem is a terrible phrase 'cause
it is very, very different to try to build systems out of humans, some of whom are nice and
some of whom are not nice and some of whom are trying to trick you and build a social system
out of large populations of those who are basically the
same level of intelligence.
2:43:19
Yes, you know, like IQ this, IQ that but that versus chimpanzees.
2:43:21
(chuckles) Like it is very different
to try to solve that problem than to try to build an AI from scratch, especially if God help
you are trying to use gradient dissent on giant
inscrutable matrices.
2:43:34
They're just very different problems.
2:43:35
And I think that all the
analogies between them are horribly misleading.
2:43:41
- So you don't think through
reinforcement learning through human feedback,
something like that, but much, much more elaborate is possible to understand this full
complexity of human nature and encode it into the machine?
2:43:57
- I don't think you are trying
to do that on your first try.
2:44:00
I think on your first try,
you are trying to build an...
2:44:07
Probably not what you should actually do, but let's say you're
trying to build something that is like Alpha Fold 17 and you are trying to get it
to solve the biology problems associated with making humans smarter, so that humans can like
actually solve alignment.
2:44:24
So you've got like a super biologist and I think what you would
want in the situation as referred to, just be
thinking about biology and not thinking about a
very wide range of things that includes how to kill everybody.
2:44:37
And I think that the first
AIs you're trying to build, not a million years later, the first ones, look more like narrowly
specialized biologists than getting the full complexity and wonder of human experience in there in such a way that
it wants to preserve itself
2:44:58
even as it becomes much smarter, which is a drastic system change
that's gonna have all kinds of side effects that, you know, like if we're dealing with
giant inscrutable matrices who are not very likely to be
able to see coming in advance. - But I don't think
it's just the matrices.
2:45:09
- But I don't think
it's just the matrices.
2:45:11
we're also dealing with the data, right?
2:45:15
With the data on the internet, and this is an interesting discussion about the data set itself, but the data set includes
the full complexity of human nature.
2:45:23
- No, it's a shadow cast
by humans on the internet.
2:45:27
- But don't you think that
shadow is a yin yang shadow?
2:45:32
(Lex laughing) - I think that if you had
alien super intelligences looking at the data, they would be able to pick up
from it an excellent picture of what humans are actually like inside.
2:45:43
This does not mean that if
you have a loss function of predicting the next token
from that dataset that the mind picked out by gradient
dissent to be able to predict the next token as well as possible on a very wide variety of
humans is itself a human.
2:46:02
- But don't you think it has
humanness a deep humanness to it in the tokens it
generates, when those tokens are read and interpreted by humans?
2:46:15
- I think that if you sent me to a distant galaxy with aliens who are much, much stupider than I am, so much so that I could do a
pretty good job of predicting what they'd say even though
they thought in an utterly different way from how I
did, then I might in time
2:46:34
be able to learn how
to imitate those aliens if the intelligence gap was
great enough that my own intelligence could overcome
the alienness and the aliens would look at my outputs and say like, "Is there not a deep
like name of alien nature "to this thing?" And what they would be seeing
was that I had correctly
2:46:52
And what they would be seeing
was that I had correctly understood them, but not
that I was similar to them.
2:47:05
- We've used aliens as a
metaphor, as a thought experiment.
2:47:10
I have to ask, how many alien
civilizations are out there? - Ask Robin Hanson.
2:47:16
He has this lovely grabby aliens paper, which is more or less the only argument I've ever
seen for where are they, how many of them are there
based on a very clever argument that if you have a bunch of
locks of different difficulty and you are randomly
trying the keys to them, the solutions will be about
evenly spaced even if the locks are of different difficulties.
2:47:44
In the rare cases where a
solution to all the locks exists in time, then Robin
Hanson looks at the arguable hard steps in human civilization
coming into existence and how much longer it has
left come into existence before, for example,
all the water slips back under the crust into the
mantle and so on, and infers that the aliens are about half a billion to a billion light years away.
2:48:12
And it's quite a clever calculation.
2:48:14
It may be entirely wrong, but it's the only time I've
ever seen anybody even come up with a halfway good argument
for how many of them, where are they.
2:48:23
- Do you think their
development of technologies, do you think their natural
evolution, whatever, however they grow and
develop intelligence, do you think it ends up at AGI as well?
2:48:35
- If it ends up anywhere,
it ends up at AGI.
2:48:38
Maybe there are aliens who
are just like the dolphins and it's just too hard
for them to forge metal.
2:48:49
Maybe if you have aliens
with no technology like that, they keep on getting smarter
and smarter and smarter and eventually the dolphins figure, like the super dolphins figure
out something very clever to do given their situation
and they still end up with high technology and in that case, they can probably solve
their AGI alignment problem.
2:49:08
If they're like much smarter before they actually confronted
'cause they saw had to solve a much harder environmental
problem to build computers, their their chances are
probably much better than ours.
2:49:18
I do worry that most of the
aliens who are like humans, like a modern human civilization, I kind of worry that
the super vast majority of them are dead.
2:49:30
Given how far we seem to be
from solving this problem.
2:49:37
But some of them would be
more cooperative than us.
2:49:40
Some of them would be smarter than us.
2:49:42
Hopefully some of the ones who are smarter and more cooperative
than us are also nice.
2:49:47
And hopefully there are some galaxies out there full of
things that say "I am, I wonder."
2:49:58
But it doesn't seem like we're on course to have this galaxy be that.
2:50:02
- Does that in part give
you some hope in response to the threat of AGI, that
we might reach out there towards the stars and find...
2:50:11
- No, if the nice aliens
were already here, they would have stopped the Holocaust.
2:50:17
That's a valid argument
against the existence of God.
2:50:20
It's also a valid argument
against the existence of nice aliens and unnice aliens who would've just eaten the planet. So no aliens.
2:50:30
- You've had debates with Robin
Hanson that you mentioned.
2:50:33
So one particular I just
want to mention is the idea of AI foom, or the ability of AGI to improve themselves very quickly.
2:50:41
What's the case you made and
what was the case he made?
2:50:44
- The thing I would say is that
among the thing that humans can do is design new AI systems.
2:50:51
And if you have something
that is generally smarter than a human, it's probably
also generally smarter at building AI systems.
2:50:56
This is the ancient argument
for foom put forth by I. J.
2:50:56
Good and probably some science
fiction writers before that, but I don't know who they would be.
2:51:06
- Well, what's the argument against foom?
2:51:10
- Various people have
various different arguments, none of which I think hold up.
2:51:15
There's only one way to be right and many ways to be wrong.
2:51:19
A argument that some people
have put forth is like, well, what if intelligence
gets exponentially harder to produce as a thing
needs to become smarter?
2:51:31
And to this, the answer is well, look at natural selection
spitting out humans.
2:51:35
We know that it does not
take like exponentially more resource investments to
produce linear increases in competence in hominids
because each mutation that rises to fixation, if the impact it has
(indistinct) small enough, it will probably never reach fixation.
2:51:58
And there's like only
so many new mutations you can fix per generation.
2:52:02
So given how long it
took to evolve humans, we can actually say with some
confidence that there were not logarithmically diminishing
returns on the individual mutations increasing intelligence.
2:52:14
So example of fraction of sub debate.
2:52:19
And the thing that Robin Hanson said was more complicated than that.
2:52:22
A brief summary, he was like, "We won't have one system
that's better at everything.
2:52:27
"You'll have a bunch of
different systems that are good "at different narrow things."
2:52:31
And I think that was falsified by GPT-4, but probably Robin Hanson
would say something else.
2:52:36
- It's interesting to ask...
2:52:37
It's perhaps a bit too philosophical, this prediction is
extremely difficult to make, but the timeline for AGI, When do you think we'll have AGI?
2:52:46
I posted this morning on
Twitter and it was interesting to see like in, in five
years, 10 years years, in 50 years or beyond.
2:52:54
And most people, like
70% something like this, think it'll be in less than 10 years.
2:53:01
So either in five years or in 10 years.
2:53:05
So that's kind of the state.
2:53:06
Do people have a sense
that there's a kind of...
2:53:09
I mean they're really impressed
by the rapid developments of ChatGPT and GPT-4.
2:53:13
So there's a sense that there's a-- - Well, we are sure on track to enter into this gradually with
people fighting about whether or not we have AGI.
2:53:23
I think there's a definite
point where everybody falls over dead 'cause you got something
that was sufficiently smarter than everybody and
that's a definite point of time.
2:53:33
But like when do we have AGI?
2:53:35
When are people fighting over
whether or not we have AGI?
2:53:38
Well, some people are
starting to fight over it as of GPT-4.
2:53:42
- But don't you think there's
going to be potentially definitive moments when we say that this is a sentient being?
2:53:50
We would go to the Supreme Court and say that this is a sentient
being that deserves human rights, for example. - You could make, yeah.
2:53:56
If you prompted being, the
right way could go argue for its unconsciousness in front
of the Supreme Court right now.
2:54:00
- [Lex] I don't think you could do that successfully right now.
2:54:03
- Because the Supreme
Court wouldn't believe it?
2:54:07
I think you could put an
IQ 80 human into a computer and ask him to argue for
his own consciousness before the Supreme Court
and the Supreme Court would be like, "You're just a computer."
2:54:19
Even if there was an
actual person in there.
2:54:22
- I think you're simplifying this. No, that's not at all.
2:54:24
That's been the argument, there's been a lot of
arguments about the other, about who deserves rights and not.
2:54:30
That's been our process
as a human species, trying to figure that out.
2:54:33
I think there will be a moment, I'm not saying sentience is that, but it could be, where
some number of people, like say over a hundred million people have a deep attachment, a fundamental attachment the
way we have to our friends,
2:54:49
to our loved ones, to our significant others,
have fundamental attachment to an AI system and they
have provable transcripts of conversation where they say, "If you take this away from me, "you are encroaching on my
rights as a human being." - People are already saying that.
2:55:04
- People are already saying that.
2:55:06
I think they're probably mistaken, but I'm not sure 'cause nobody knows what goes on inside those things.
2:55:12
- Eliezer, they're not
saying that at scale. - [Eliezer] Okay.
2:55:16
- So the question is, is there a moment when AGI, we know AGI I arrived,
what would that look like?
2:55:21
I'm giving essentially as just an example.
2:55:22
It could be something else.
2:55:23
- It looks like the AGIs
successfully manifesting themselves as 3D video of young women, at which point a vast portion
of the male population decides that they're real people.
2:55:38
- So sentience, essentially.
2:55:42
Demonstrating identity and sentience.
2:55:45
- I'm saying that the easiest way to pick up a hundred million people saying that you seem like a person is to look like a person talking to them, with Bing's current
level of verbal facility. - I disagree with that.
2:55:58
- A different set of prompts. - I disagree with that.
2:56:01
I think you're missing again, sentience.
2:56:03
There has to be a sense
that it is a person that would miss you when you're gone.
2:56:07
They can suffer, they can die.
2:56:09
Of course, I'm being-- - GPT-4 can pretend that right now.
2:56:16
How can you tell when it's real?
2:56:18
- I don't think it can pretend
that right now successfully. It's very close.
2:56:21
- Have you talked to GPT-4? - [Lex] Yes, of course. - Okay.
2:56:26
Have you been able to get a version of it that hasn't been trained
not to pretend to be human?
2:56:31
Have you talked to a jail broken version that will claim to be conscious? - No.
2:56:36
The linguistic capability's there, but there's something
about a digital embodiment of the system that has a bunch of, perhaps it's small interface
features that are not significant relative to
the broader intelligence that we're talking about.
2:57:01
So perhaps GPT-4 is already there.
2:57:04
But to have the video where a
woman's face or a man's face to whom you have a deep connection, perhaps we're already there, but we don't have such a
system yet deployed at scale.
2:57:15
- The thing I'm trying to
to gesture at here is that it's not like people have a
widely accepted, agreed upon definition of what consciousness is.
2:57:26
It's not like we would
have the tiniest idea of whether or not that was
going on inside the giant inscrutable matrices, even if we hadn't agreed upon definition.
2:57:34
So if you're looking for
upcoming predictable big jumps in how many people think
the system is conscious, the upcoming predictable big
jump is it looks like a person talking to you who is
cute and sympathetic.
2:57:48
That's the upcoming predictable big jump.
2:57:52
Now that versions of
it are already claiming to be conscious, which is the point where I start going like, ah, not 'cause it's real,
but because from now on, who knows if it's real? - Yeah.
2:58:04
And who knows what
transformational effect it has on a society where more than
50% of the beings that are interacting on the internet
and sure as heck look real are not human?
2:58:15
What kind of effect does that
have when young men and women are dating AI systems?
2:58:22
- You know, I'm not an expert on that. God help humanity.
2:58:26
(chuckles) I'm one of the closest things to an expert on where it all goes.
2:58:32
'Cause you know, and how
did you end up with me as an expert?
2:58:35
'Cause for 20 years, humanity
decided to ignore the problem.
2:58:39
So like, this tiny handful
of people and basically me, got 20 years to try to be an expert on it while everyone else ignored it. And yeah.
2:58:50
So where does it all end up?
2:58:51
Try to be an expert on that.
2:58:53
Particularly the part where
everybody ends up dead, 'cause that part is kind of important, but what does it do to dating
when some fraction of men and some fraction of women
decided that they'd rather date the video of the thing that is like relentlessly
kind and generous to them and claims to be conscious, but who knows what's goes on inside it
and it's probably not real, but you know, you can think it's real, what happens to society? I don't know.
2:59:17
I'm not actually an expert on that.
2:59:19
And the experts don't know either, 'cause it's kind hard
to predict the future.
- So you have talked a lot about sort of the longer term future where it's all headed.
2:59:35
- By longer term we mean
like, not all that long.
2:59:38
But yeah, where it all ends up.
2:59:41
- But beyond the effects of men
and women dating AI systems, you're looking beyond that. - Yes.
2:59:47
'Cause that's not how the fate
of the galaxy got settled.
2:59:51
- Well, let me ask you about
your own personal psychology. A tricky question.
2:59:56
You've been known at times
to have a bit of an ego.
2:59:59
Do you think-- - Says who? But go on.
3:00:03
- Do you think ego is
empowering or limiting for the task of understanding
the world deeply? - I reject the framing.
3:00:13
- [Lex] So you disagree
with having an ego?
3:00:15
So what do you think about ego?
3:00:16
- No, I think that the
question of what leads to making better or worse predictions, what leads to being able
to pick out better or worse strategies is not carved at
its joint by talking of ego.
3:00:30
- So it should not be subjective.
3:00:31
It should not be connected to
the intricacies of your mind?
3:00:35
- No, I'm saying that like, if you go about asking all day long, "Do I have enough ego?
3:00:43
"Do I have too much of an ego?"
3:00:43
, I think you get worse at
making good predictions.
3:00:48
I think that to make good
predictions, you're like, how did I think about this? Did that work? Should I do that again?
3:00:55
- You don't think we as
humans get invested in an idea and then others attack you
personally for that idea?
3:01:04
So you plant your feet and
it starts to be difficult to, when a bunch of assholes
low effort attack your idea to eventually say, "You know what? "I actually was wrong." And tell them that.
3:01:15
As a human being, it becomes difficult. It's difficult.
3:01:22
- So like Robin Hanson and I
debated AI systems and I think that the person who won
that debate was Gwern.
3:01:28
And I think that reality was
well to the Yudkowskian side of the Yudkowsky-Hanson spectrum, like further from Yudkowsky.
3:01:39
And I think that's because I
was trying to sound reasonable compared to Hanson and saying
things that were defensible and relative to Hanson's
arguments and reality was way over here in
(indistinct) in respect to, Hanson was like, "All the
systems will be specialized."
3:01:55
Hanson may disagree with
this characterization.
3:01:58
Hanson was like, "All the
systems will be specialized."
3:02:00
I was like, "I think we build specialized
underlying systems "that when you combine them are good "at a wide range of things."
3:02:08
And the reality is like, no,
you just stack more layers into a bunch of gradient descent.
3:02:12
And I feel looking back that by trying to have this reasonable
position contrasted to Hanson's position, I missed the ways that
reality could be more extreme than my position in the same direction.
3:02:28
So is this like, is this a
failure to have enough ego?
3:02:33
Is this a failure to make
myself be independent?
3:02:37
I would say that this is
something like a failure to consider positions that would
sound even wackier and more extreme when people are
already calling you extreme.
3:02:49
But I wouldn't call that
not having enough ego.
3:02:53
I would call that insufficient ability to just clear that all out of your mind.
3:03:01
- In the context of debate and discourse, which is already super tricky.
3:03:05
- In the context of prediction, in the context of modeling reality.
3:03:08
If you're thinking of it as a debate, you're already screwing up.
3:03:11
- So is there some kind
of wisdom and insight you can give to how to clear
your mind and think clearly about the world?
3:03:18
- Man, this is an example
of where I wanted to be able to put people into FMRI machines,
then you'd be like, okay, "See that thing you just did, "you were rationalizing right there."
3:03:27
Oh, that area of the brain lit up.
3:03:29
You are like now being
socially influenced is, is kind of the dream.
3:03:35
And you know, I don't know, I wanna say like just
introspect, but for many people, introspection is not that easy. - [Lex] It's hard.
3:03:44
- Notice the internal sensation.
3:03:46
Can you catch yourself in the
very moment of feeling a sense of, well if I think this thing
people will look funny at me.
3:03:56
Okay, if you can see that sensation, which is step one, can you
now refuse to let it move you?
3:04:04
Or maybe just make it go away.
3:04:06
And I feel like I'm
saying like, I don't know, like somebody's like,
"How do you draw an owl?"
3:04:11
And I'm saying like,
"Well, just draw an owl."
3:04:15
(both laughing) I feel like most people,
the advice they need is like, well how do I notice
the internal subjective sensation in the moment that it happens of fearing to be socially influenced?
3:04:28
Or okay, I see it, how do I turn it off?
3:04:30
How do I let it not influence me?
3:04:32
Do I just do the opposite
of what I'm afraid people will criticize me for?
3:04:36
And I'm like, "No, no, "you're not trying to do the opposite "of what you're afraid of
what you might be pushed into.
3:04:46
"You're trying to let the
thought process complete "without that internal push."
3:04:54
Can you not reverse the push, but be unmoved by the push and are these instructions
even remotely helping anyone? I don't know.
3:05:03
- I think tho when those instructions, even those the words you've
spoken and maybe you can add more, when practiced daily, meaning in your daily communication.
3:05:12
So it's daily practice of
thinking without influence from-- - I would say find prediction
markets that matter to you and better in the prediction markets.
3:05:23
That way you find out
if you are right or not.
3:05:26
- [Lex] And you really, there's stakes.
3:05:30
- Or even manifold markets where
the stakes are a bit lower.
3:05:33
But the important thing
is to get the record and you know, I didn't
build up skills here by prediction markets.
3:05:43
I built them up via like, well, how did the foom debate resolve and my own take on it,
as to how it resolved.
3:05:56
The more you are able to notice yourself not being dramatically wrong, but having been a little off, your reasoning was a little off.
3:06:06
You didn't get that quite right.
3:06:08
Each of those is a opportunity
to make like a small update.
3:06:12
So the more you can like
say, "oops," softly, routinely, not as a big deal, the more chances you get to be like, I see where that reasoning went astray.
3:06:20
I see how I should have
reasoned differently.
3:06:23
And this is how you
build up skill over time.
3:06:27
- What advice could you give
to young people in high school and college, given the
highest of stakes things you've been thinking about?
3:06:36
If somebody's listening to this
and they're young and trying to figure out what to
do with their career, what to do with their life,
what advice would you give them?
3:06:45
- Don't expect it to be a long life.
3:06:47
Don't put your happiness into the future.
3:06:49
The future is probably not
that long at this point, but none know the hour nor the day.
3:06:56
- But is there something, if they want to have hope to
fight for a longer future, is there a fight worth fighting?
3:07:06
- I intend to go down fighting. I don't know.
3:07:13
I admit that although I do
try to think painful thoughts, what to say to the children at this point is a pretty painful
thought as thoughts go. They want to fight.
3:07:26
I hardly know how to fight
myself at this point.
3:07:30
I am trying to be ready for
being wrong about something, preparing for my being wrong in a way that creates a bit of hope and
being ready to react to that and going looking for it.
3:07:45
And that is hard and complicated.
3:07:48
And somebody in high school, I don't know, like you have presented
a picture of the future that is not quite how I expected it to go where there is public
outcry and that outcry is put into a remotely useful direction, which I think at this point is just shutting down the
GPU clusters because no, we are not in a shape to frantically do it at the last minute, do
decades' worth of work.
3:08:15
The thing you would do at this
point if there were massive public outcry pointed
in the right direction, which I do not expect, is shut down the GPU clusters
and and crash program on augmenting human
intelligence biologically.
3:08:26
Not the (indistinct) stuff, biologically.
3:08:30
'Cause if you make humans much smarter, they can actually be smart and nice.
3:08:35
You get that in a plausible way, in a way that it is not as easy to do with synthesizing these
things from scratch, predicting the next
tokens and applying RLHF.
3:08:45
Like humans start out in the frame that produces niceness, that
has ever produced niceness.
3:08:53
And saying this, I do
not want to sound like the moral of this whole thing was like, oh, you need to engage in mass action and then everything will be all right.
3:09:05
This is 'cause there's so
many things where somebody tells you that the world is
ending you need to recycle.
3:09:10
And if everybody does
their part and and recycles their cardboard, then we can
all live happily ever after.
3:09:15
And this is unfortunately
not what I have to say.
3:09:24
Everybody recycling their
cardboard is not gonna fix this.
3:09:26
Everybody recycles their
cardboard and then everybody ends up dead, metaphorically speaking.
3:09:31
But if there was enough,
like on the margins, you just end up dead a little later on most of the things at a few
people can do by trying hard.
3:09:43
But if there was enough
public outcry to shut down the GPU clusters and then you
could be part of that outcry.
3:09:52
If Eliezer is wrong in the direction that Lex Friedman predicts, that there is enough public
outcry pointed enough in the right direction to do something that actually, actually, actually
results in people living.
3:10:05
Not just we did something, not just there was an outcry
and the outcry was given form in something that was safe and convenient and didn't really inconvenience anybody and then everybody died everywhere.
3:10:15
There was enough actual
like, oh, we're going to die, we should not do that.
3:10:19
We should do something
else which is not that, even if it is not super duper convenient, it wasn't inside the previous
political overton window.
3:10:27
If I'm wrong and there's
that kind of public outcry, then somebody in high
school could be ready to be part of that.
3:10:32
If I'm wrong in other ways, then you could maybe be part of that.
3:10:38
And if you were like a
brilliant young physicist, then you could like go
into interpretability and if you're smarter than that, you could work on alignment
problems where it's harder to tell if you got them
right or not, (sighs) and other things.
3:10:52
But mostly for the kids in
high school, it's like, yeah, be ready to help if Eliezer Yudkowsky is wrong about something and otherwise don't put your happiness
into the far future.
3:11:09
It probably doesn't exist.
3:11:11
- But it's beautiful that
you're looking for ways that you're wrong.
3:11:14
And it's also beautiful that
you're open to being surprised by that same young physicist
with some breakthrough.
3:11:21
- It feels like a very,
very basic competence that you are praising me for.
3:11:25
And you know, like, okay, cool.
3:11:28
I don't think it's good
that we're in a world where that is something that I deserve to be complimented on,
I've never had much luck in accepting compliments gracefully.
3:11:40
Maybe I should just accept
that one gracefully.
3:11:41
(Lex laughing) Sure, thank you very much.
3:11:45
- You've painted with some
probability a dark future.
3:11:48
Are you yourself, just when you think, when you ponder your life and
you ponder your mortality, are you afraid of death? - I think so, yeah.
3:12:06
- Does it make any sense
to you that we die?
3:12:16
There's a power to the
finiteness of the human life that's part of this whole
machinery of evolution and that finiteness doesn't
seem to be obviously integrated into AI systems.
3:12:33
So it feels like,
fundamentally in that aspect, some fundamentally different
thing that we're creating.
3:12:39
- I grew up reading books like "Great Mambo Chicken and
the Transhuman Condition" and later on "Engines of
Creation" and "Mind Children."
3:12:49
You know, age 12 or thereabouts.
3:12:53
So I never thought I was
supposed to die after 80 years.
3:12:59
I never thought that
humanity was supposed to die.
3:13:02
I thought we were like, I always grew up with the
ideal in mind that we were all going to live happily ever after in the glorious transhumanist future.
3:13:10
I did not grow up thinking that death was part of the meaning of life. - [Lex] And now...
3:13:17
- And now I still think
it's a pretty stupid idea.
3:13:21
You do not need life to be
finite to be meaningful. It just has to be life.
3:13:26
- What role does love play
in the human condition?
3:13:29
We haven't brought up love
and this whole picture.
3:13:31
We talked about intelligence,
we talked about consciousness.
3:13:33
It seems part of humanity, I would say one of the
most important parts is this feeling we have
towards each other.
3:13:45
- If in the future there were
routinely more than one AI, let's say two for the sake of discussion, who would look at each other and say, "I am I and you are you."
3:14:00
The other one also says, "I am I and you are you."
3:14:03
And sometimes they were happy
and sometimes they were sad and it mattered to the other one that this thing that is
different from them is like, they would rather it be happy than sad and entangle their lives together, then this is a more optimistic thing than I expect to actually happen.
3:14:25
A little fragment of
meaning would be there, possibly more than a little, but that I expect this to not happen.
3:14:32
That I do not think this
is what happens by default.
3:14:34
That I do not think
that this is the future we are on track to get is
why I would go down fighting rather than, you know,
just saying, "Oh well."
3:14:48
- Do you think that is part of the meaning of this whole thing or
the meaning of life?
3:14:54
What do you think is the
meaning of life, of human life?
3:14:57
- It's all the things
that I value about it and maybe all the things
that I would value if I understood it better.
3:15:03
There's not some meaning far outside of us that we have to wonder about.
3:15:09
There's just looking
at life and being like, yes, this is what I want.
3:15:16
The meaning of life is not some kind of...
3:15:24
Meaning is something
that we bring to things when we look at them, we
look at them and we say like, "This is its meaning to me."
3:15:33
It's not that before
humanity was ever here, there was some meaning written
upon the stars where you could like go out to the star
where that meaning was written and change it around and
thereby completely change the meaning of life, right?
3:15:45
The notion that this is written
on a stone tablet somewhere implies that you could change the tablet and get a different meaning and that seems kind of wacky, doesn't it?
3:15:56
It doesn't feel that
mysterious to me at this point.
3:15:58
It's just a matter of
being like, yeah, I care. - I care.
3:16:06
And part of that is the love
that connects all of us.
3:16:12
- [Eliezer] It's one of the
things that I care about.
3:16:17
- And the flourishing of
the collective intelligence of the human species.
3:16:22
- You know, that sounds
kind of too fancy to me.
3:16:25
I just look at all the people, like one by one up to the
8 billion and be like, that's life, that's life, that's life.
3:16:37
- Eliezer, you're an incredible human. It's a huge honor.
3:16:40
I was trying to talk
to you for a long time (laughing) because I'm a big fan.
3:16:46
I think you're a really important voice and really important mind.
3:16:49
Thank you for the fight you're fighting.
3:16:52
Thank you for being fearless and bold and for everything you do.
3:16:55
I hope we get a chance to talk again.
3:16:56
And I hope you never give up.
3:16:58
Thank you for talking today. - You're welcome.
3:17:00
I do worry that we didn't
really address a whole lot of fundamental questions I
expect people have, but you know, maybe we got a little bit further
and made a tiny little bit of progress and I'd say
be satisfied with that.
3:17:14
But actually no, I think
one should only be satisfied with solving the entire problem. - To be continued.
3:17:21
Thanks for listening to this conversation with Eliezer Yudkowsky.
3:17:25
To support this podcast,
please check out our sponsors in the description.
3:17:28
And now let me leave you
some words from Elon Musk.
3:17:33
"With artificial intelligence,
we're summoning the demon."
3:17:38
Thank you for listening and
hope to see you next time.