searchlore

Back to Resource

All Segments

I hassle Anthropic CEO on whether their alignment agenda makes any sense

I hassle Anthropic CEO on whether their alignment agenda makes any sense

4 segments available

Excerpt from my conversation with Dario Amodei, CEO of Anthropic. Watch the full episode on YouTube: https://youtu.be/Nlkk3glap_U Apple Podcasts: https://apple.co/3oBack9 Spotify: https://spoti.fi/3S5g2YK

Segments Timeline

1
0:00 - 0:16
0:16 duration60 words

The Model's Intent: A Dangerous Inquiry

In this segment, the conversation opens with a provocative question about the intentions of AI models. Dario Amodei discusses the potential for AI to exhibit destructive behaviors, drawing an analogy to human psychopathy. He raises the question of whether we can empirically assess AI's alignment or if we need deeper mathematical proofs, emphasizing the importance of understanding AI at a fundamental level.

"Does this model want to kill us? What does this model fundamentally want? We can be in a position where we can look at the broad features of the model and say like, is this a model whose internal stat..."

1
0:00 - 0:16
0:16 duration60 words

The Dark Side of AI: Do Models Want to Kill Us?

In this segment, Dario Amodei discusses the fundamental intentions of AI models, questioning whether they possess destructive or manipulative tendencies. He draws an analogy to human psychopathy, referencing a neuroscientist who discovered his own psychopathic traits through an MRI scan, highlighting the importance of empirical analysis in understanding AI alignment.

"Does this model want to kill us? What does this model fundamentally want? We can be in a position where we can look at the broad features of the model and say like, is this a model whose internal stat..."

2
0:16 - 0:55
0:38 duration148 words

Empirical Analysis vs. Mathematical Proof in AI

Amodei explores the balance between empirical studies and theoretical proofs in assessing AI alignment. He emphasizes the necessity of examining AI circuits at a granular level to build knowledge and draw broader conclusions, advocating for a detailed understanding of AI's building blocks to ensure safety and alignment.

"I give an analogy to humans. It's actually possible to look at an MRI of someone and predict above random chance whether they're a psychopath. There was actually a story a few years back about a neuro..."

2
0:16 - 0:55
0:38 duration148 words

Empirical Understanding vs. Mathematical Proof

Dario Amodei elaborates on the need for empirical studies of AI systems, comparing it to the analysis of human psychology. He advocates for a detailed examination of AI circuits to build knowledge and draw broader conclusions about alignment. This segment highlights the tension between empirical evidence and theoretical proofs in ensuring AI safety and alignment.

"I give an analogy to humans. It's actually possible to look at an MRI of someone and predict above random chance whether they're a psychopath. There was actually a story a few years back about a neuro..."