
4 segments available
Excerpt from my conversation with Dario Amodei, CEO of Anthropic. Watch the full episode on YouTube: https://youtu.be/Nlkk3glap_U Apple Podcasts: https://apple.co/3oBack9 Spotify: https://spoti.fi/3S5g2YK
In this segment, the conversation opens with a provocative question about the intentions of AI models. Dario Amodei discusses the potential for AI to exhibit destructive behaviors, drawing an analogy to human psychopathy. He raises the question of whether we can empirically assess AI's alignment or if we need deeper mathematical proofs, emphasizing the importance of understanding AI at a fundamental level.
"Does this model want to kill us? What does this model fundamentally want? We can be in a position where we can look at the broad features of the model and say like, is this a model whose internal stat..."
In this segment, Dario Amodei discusses the fundamental intentions of AI models, questioning whether they possess destructive or manipulative tendencies. He draws an analogy to human psychopathy, referencing a neuroscientist who discovered his own psychopathic traits through an MRI scan, highlighting the importance of empirical analysis in understanding AI alignment.
"Does this model want to kill us? What does this model fundamentally want? We can be in a position where we can look at the broad features of the model and say like, is this a model whose internal stat..."
Amodei explores the balance between empirical studies and theoretical proofs in assessing AI alignment. He emphasizes the necessity of examining AI circuits at a granular level to build knowledge and draw broader conclusions, advocating for a detailed understanding of AI's building blocks to ensure safety and alignment.
"I give an analogy to humans. It's actually possible to look at an MRI of someone and predict above random chance whether they're a psychopath. There was actually a story a few years back about a neuro..."
Dario Amodei elaborates on the need for empirical studies of AI systems, comparing it to the analysis of human psychology. He advocates for a detailed examination of AI circuits to build knowledge and draw broader conclusions about alignment. This segment highlights the tension between empirical evidence and theoretical proofs in ensuring AI safety and alignment.
"I give an analogy to humans. It's actually possible to look at an MRI of someone and predict above random chance whether they're a psychopath. There was actually a story a few years back about a neuro..."