
4 segments available
Excerpt from my conversation with Dario Amodei, CEO of Anthropic. Watch the full episode on YouTube: https://youtu.be/Nlkk3glap_U Apple Podcasts: https://apple.co/3oBack9 Spotify: https://spoti.fi/3S5g2YK
Dario Amodei discusses the potential reasons why large language models (LLMs) might not achieve human-level intelligence. He explores practical issues such as running out of data or compute resources, while also addressing the fundamental scaling laws that govern LLM performance. Amodei emphasizes that while hitting a plateau is possible, it is unlikely from a theoretical standpoint.
"if it turns out that scaling plateaus before we reach human level intelligence looking back on it what would be your explanation if I would distinguish some problem with the fundamental Theory with so..."
Dario Amodei discusses the potential reasons why scaling might plateau before achieving human-level intelligence in LLMs. He considers practical issues such as running out of data or compute resources, but finds it unlikely that scaling laws will simply stop. This segment explores the fundamental challenges in AI development and the implications of hitting a wall in progress.
"if it turns out that scaling plateaus before we reach human level intelligence looking back on it what would be your explanation if I would distinguish some problem with the fundamental Theory with so..."
In this segment, Dario Amodei delves into the complexities of training LLMs for high-level programming tasks. He highlights the importance of the loss function in next word prediction and how it may lead to an overemphasis on certain tokens, potentially hindering the model's ability to learn essential programming concepts effectively.
"that happens my explanation would be there's something wrong with the loss when you train on next word prediction like if you really want to learn to program at a really high level it means you care a..."
In this segment, Amodei delves into the complexities of training LLMs, particularly in the context of next-word prediction. He highlights the importance of focusing on rare tokens that are essential for high-level programming skills. The discussion emphasizes the potential shortcomings of current loss functions in AI training and their impact on the learning process.
"I think from a fundamental perspective it's very unlikely that the scaling laws will just stop you could have made a case a few years ago that they can't reason they can't program I think it's a less ..."