
5 segments available
I have a much better understanding of Sutton’s perspective now. I wanted to reflect on it a bit. Read the transcript here: https://www.dwarkesh.com/p/thoughts-on-sutton TIMESTAMPS 00:00:00 The steelman 00:02:42 TLDR of my current thoughts 00:03:22 Imitation learning is continuous with and complementary to RL 00:08:26 Continual learning 00:10:31 Concluding thoughts
In this segment, the speaker reflects on Richard Sutton's perspective, particularly focusing on his essay 'The Bitter Lesson.' The discussion emphasizes the importance of leveraging compute effectively in AI training, critiquing the inefficiencies of current LLM training methods, and highlighting the need for a new architecture that allows for continual learning. The speaker argues that LLMs are limited by their reliance on human data and lack the ability to learn organically from their environment.
"Boy do you guys have a lot of thoughts about the Sutton interview. I’ve been thinking about it myself and I think I have a much better understanding now of Sutton’s perspective than I did during ..."
The speaker discusses the relationship between imitation learning and reinforcement learning (RL), arguing that they are complementary rather than mutually exclusive. This segment explores how pretrained LLMs can serve as a foundation for accumulating experiential learning, drawing parallels to historical advancements in technology and the importance of prior knowledge in developing AGI.
"That's my understanding of Richard's position. My main difference with Rich is just that I don't think the concepts he's using to distinguish LLMs from true intelligence are actually that mutual..."
This segment delves into the significance of human data in training AI models, using the analogy of fossil fuels to illustrate how initial resources can facilitate progress. The speaker argues that while human data is crucial for early learning, it should not be seen as a dead-end but rather as a stepping stone towards more advanced learning techniques. The discussion includes examples like AlphaGo and AlphaZero to highlight the evolution of AI learning methods.
"to RL. I tried to ask Richard a couple of times whether pretrained LLMs can serve as a good prior on which we can accumulate the experiential learning (aka do the RL) which will lead to AGI. Ilya..."
In this segment, the speaker emphasizes the challenges LLMs face in achieving true continual learning. They discuss the limitations of current models in extracting information from their environment and propose potential methods for integrating continual learning into LLMs. The speaker expresses optimism about the possibility of replicating human-like learning flexibility in future AI models.
"Continual learning. Sorry to bring up my hobby horse again. I'm like a comedian who's only come up with one good bit, but I'm gonna milk it for all it's worth. An LLM being RLed on outcome-based ..."
The speaker concludes by reflecting on the evolution of AI learning paradigms, contrasting the development of LLMs with natural learning processes in humans and animals. They acknowledge the critiques posed by Sutton regarding the limitations of current models and suggest that future AGI systems may emerge from the foundational ideas presented by Sutton, despite the challenges faced by existing LLMs.
"Maybe this won't work! But I don't think these super first-principle arguments (for example, about how these LLM don't have a true world model) are actually proving much. I also don't think they’..."