
4 segments available
Full Episode: https://youtu.be/Wo95ob_s_NI Apple Podcasts: https://podcasts.apple.com/us/podcast/john-schulman-openai-cofounder-reasoning-rlhf-plan/id1516093381?i=1000655679622 Spotify: https://open.spotify.com/episode/1ivzHH9RWciXe4O1rKtldf?si=53503781e05f4d8f Transcript: https://www.dwarkeshpatel.com/p/john-schulman/ Me on Twitter: https://twitter.com/dwarkesh_sp/
John Schulman discusses the evolution of AI models from simple search engines to more complex collaborators capable of handling entire coding projects. He envisions a future where users can provide high-level instructions, and the models will autonomously write, test, and iterate on code, significantly enhancing productivity and collaboration.
"in one or two years we'll find that you can use them for a lot of um more like involved tasks than they can do now so you could um you could imagine having the models do carry out a whole coding proje..."
Schulman explains how future AI models will be better at recovering from errors and handling edge cases. He emphasizes the importance of sample efficiency, suggesting that improved generalization will allow models to learn from minimal data, thus enhancing their ability to stay on track during complex tasks.
"any any kind of training uh any like doing RL uh to learn how to do these tasks uh however you do it whether it's whether you're supervising the final output or supervising it like each step um I thin..."
In this segment, Schulman speculates on the potential for AI models to achieve human-level coherence over extended periods. He discusses the implications of long-horizon reinforcement learning training and its potential to enable models to plan and execute complex projects, akin to collaborating with a human colleague.
"from uh from other um abilities will allow them to get back on the track on track whereas current models might just get stuck and get lost I'm not sure I'm not sure actually how uh understand more sup..."
Schulman reflects on the challenges that remain in achieving Artificial General Intelligence (AGI). He acknowledges that while improving coherence is crucial, there are other weaknesses in current models that need to be addressed. He emphasizes the complexity of developing AI that can function as a fully capable colleague, highlighting the ongoing journey toward AGI.
"if it's the case that once you start this long Horizon RL training regime it immediately unlocks your ability to be coherent for longer periods of time should we be predicting something that is human ..."