
6 segments available
Full Episode: https://youtu.be/Wo95ob_s_NI Apple Podcasts: https://podcasts.apple.com/us/podcast/john-schulman-openai-cofounder-reasoning-rlhf-plan/id1516093381?i=1000655679622 Spotify: https://open.spotify.com/episode/1ivzHH9RWciXe4O1rKtldf?si=53503781e05f4d8f Transcript: https://www.dwarkeshpatel.com/p/john-schulman/ Me on Twitter: https://twitter.com/dwarkesh_sp/
John Schulman discusses the initial realization at OpenAI that large language models (LLMs) could be effectively utilized through chatbots. He explains the evolution from basic models to instruction-following models, highlighting the challenges of prompting and the need for user-friendly interactions.
"you let the creation of Chad jbt at what point do you did you realize first of all these llms are the pat to go and then a Chad bot would be or some way to instruct them would be a useful thing to do ..."
Schulman elaborates on the development of conversational AI, referencing Google's chat models like LaMDA and the importance of follow-up questions in chat interactions. He shares insights from his previous project, WebGPT, which focused on question answering and the necessity for a chat-based approach.
"thinking about um chat so uh so Google had some papers uh like they had uh Lambda and um earlier Mina so they had these chat Bots and it was more like um uh like you had a it was more like a base mode..."
The conversation shifts to the advancements made with GPT-3.5, which was trained to excel in language and coding tasks. Schulman discusses the features considered during development, including browsing capabilities, and the decision to prioritize the model's internal knowledge over external browsing.
"because um you always want to ask follow-up questions or sometimes you need a clar the the model should ask a clarifying question because the question is ambiguous so it was kind of clear after we did..."
Schulman reflects on the challenges faced with GPT-4, including its impressive outputs and occasional unreliability. He discusses the excitement surrounding instruction-following models and the need for further refinement before public release.
"um and then uh we were thinking about we had it out for beta testing or to friends and family for a while and we were thinking about doing a public release um but um at that time uh actually GPD 4 fin..."
In this segment, Schulman explains the strategy of combining instruction and chat datasets to enhance model performance. He emphasizes the intuitive understanding users have of chatbots, which led to more coherent and sensible behavior in the models.
"outputs so it was clearly not quite ready for prime time but it was like obviously very good um and uh yeah so I guess that um people forgot about chat for a little while after thatc about this like a..."
Schulman addresses the complexities involved in fine-tuning models like ChatGPT. He discusses the iterative process required for effective training, the challenges of using human-generated data, and the potential for creating a close approximation of ChatGPT using publicly available APIs.
"um I think people had an intuitive sense of uh like what a helpful robot should be like so I think it was uh just much easier to tell people uh like uh to to get for people to get the idea of what wha..."