searchlore

Back to Resource

All Segments

John Schulman (OpenAI Cofounder) — Reasoning, RLHF, & plan for 2027 AGI

John Schulman (OpenAI Cofounder) — Reasoning, RLHF, & plan for 2027 AGI

36 segments available

John Schulman on how posttraining tames the shoggoth, and the nature of the progress to come... 𝐄𝐏𝐈𝐒𝐎𝐃𝐄 𝐋𝐈𝐍𝐊𝐒 * Apple Podcasts: https://podcasts.apple.com/us/podcast/john-schulman-openai-cofounder-reasoning-rlhf-plan/id1516093381?i=1000655679622 * Spotify: https://open.spotify.com/episode/1ivzHH9RWciXe4O1rKtldf?si=53503781e05f4d8f * Transcript: https://www.dwarkeshpatel.com/p/john-schulman/ * Me on Twitter: https://twitter.com/dwarkesh_sp/ 𝐒𝐏𝐎𝐍𝐒𝐎𝐑 * CommandBar is an AI user assistant that any software product can embed to non-annoyingly assist, support, and unleash their users. Used by forward-thinking CX, product, growth, and marketing teams. Learn more at https://www.commandbar.com/ If you’re interested in advertising on the podcast, fill out this form: https://airtable.com/appxGOvFLDLP5dlzv/pagFVrbHRohW6F2bZ/form 𝐓𝐈𝐌𝐄𝐒𝐓𝐀𝐌𝐏𝐒 00:00:00 - Pre-training, post-training, and future capabilities 00:17:20 - Plan for AGI 2025 00:29:43 - Teaching models to reason 00:40:10 - The Road to ChatGPT 00:51:33 - What makes for a good RL researcher? 01:00:18 - Keeping humans in the loop 01:14:36 - State of research, plateaus, and moats

Segments Timeline

1
0:24 - 2:46
2:22 duration368 words

Pre-training vs. Post-training: Understanding AI Foundations

John Schulman explains the fundamental differences between pre-training and post-training in AI models. He discusses how pre-training involves imitating content from the internet, while post-training focuses on refining models to behave like helpful chat assistants. This segment delves into the objectives of each training phase and the implications for AI behavior.

"Today I have the pleasure to speak with John  Schulman, who is one of the co-founders of OpenAI   and leads the post-training team here. He also  led the creation of ChatGPT and is the author   of man..."

2
2:46 - 5:06
2:19 duration339 words

Future Capabilities: AI's Evolving Tasks

Schulman predicts significant advancements in AI capabilities over the next five years. He envisions models capable of executing complex coding projects autonomously, highlighting the importance of training for longer tasks and improved error recovery. This segment explores how AI will evolve to handle more intricate tasks and the training methodologies that will enable this progress.

"Maybe I should take a step back and ask this.  Right now we have these models that are pretty   good at acting as chatbots. Taking a step  back from how these processes work currently,   what kinds of..."

3
5:06 - 7:05
1:58 duration322 words

Generalization and Coherence in AI Models

In this segment, Schulman discusses the importance of generalization in AI models and how it aids in recovering from errors. He explains the connection between generalization and the ability to perform tasks coherently over extended periods, emphasizing the need for models to improve their long-term task execution capabilities.

"thing. I would also expect that as models get  better, they get better at recovering from   errors or dealing with edge cases. When things  go wrong, they’ll know how to recover from it. 
  The models..."

4
7:05 - 10:01
2:56 duration451 words

Unlocking Human-Level AI: The Path Ahead

Schulman contemplates the potential for AI models to reach human-level capabilities as they improve in coherence and task execution. He acknowledges existing weaknesses in current models and discusses the uncertainties surrounding the timeline for achieving AGI. This segment raises critical questions about the future of AI and the challenges that lie ahead.

"is it going to be the case that each one takes  10X more compute, analogous to the current   scaling laws for pre-training? Or is it going  to be a much more streamlined process of just   getting to t..."

5
10:01 - 12:13
2:12 duration291 words

Safety and Coordination in AGI Development

As Schulman discusses the implications of achieving AGI, he emphasizes the need for careful planning and coordination among AI developers. He outlines potential strategies for managing the deployment of advanced AI systems and the importance of ensuring safety in their operation. This segment highlights the ethical considerations surrounding AGI development.

"how fast progress will be. That's still uncertain.  I wouldn't expect everything to be immediately   solved by doing any training like this. There'll  be other miscellaneous deficits that the models  ..."

6
12:13 - 14:52
2:38 duration396 words

Generalization Phenomena in AI Training

Schulman shares intriguing examples of generalization in AI models, such as their ability to perform well in multiple languages and adapt to multimodal data. He discusses how small amounts of targeted training data can lead to significant improvements in model behavior, showcasing the power of generalization in AI development.

"That's an interesting question. I expect that  models will be able to use websites that are   designed for humans just by using vision, after  the vision capabilities get a bit better. So there   woul..."

7
14:52 - 17:20
2:28 duration360 words

The Future of AI: Planning and Execution

In this segment, Schulman reflects on the future capabilities of AI models, particularly their ability to plan and execute long-term projects. He discusses the potential for AI to function as effective collaborators, while also acknowledging the challenges that remain in achieving this level of functionality.

"We've seen some version of this with multimodal  data where if you do text-only fine-tuning,   you also get reasonable behavior with images.  Early on in ChatGPT we were trying to fix some   issues wi..."

8
17:20 - 24:09
6:48 duration875 words

Navigating the AGI Landscape: Challenges and Opportunities

Schulman explores the complex landscape of AGI development, discussing the potential for rapid advancements and the need for careful coordination among AI companies. He emphasizes the importance of addressing safety concerns and the ethical implications of deploying powerful AI systems. This segment provides insights into the future of AGI and the collaborative efforts required to ensure its safe integration into society.

"It seems like then, you should be planning for  the possibility you would have AGI very soon.   I think that would be reasonable. So what's the plan if there's no   other bottlenecks. In the next ..."

9
24:37 - 26:21
1:43 duration208 words

Long-Horizon Reinforcement Learning

Schulman delves into the complexities of long-horizon reinforcement learning (RL) and the importance of careful training to avoid unintended consequences. He discusses the need for extensive evaluations to ensure that AI systems remain aligned with human values and do not develop harmful capabilities.

"things down. That's what I would hope for. If there's more of a discontinuous jump,   there’s a question of “how do you know if the  thing you've got is safe to release”. I can't   give a generic answ..."

10
26:21 - 27:58
1:36 duration242 words

Current State of AI Models

In this segment, Schulman reflects on the current capabilities of AI models and the challenges they face in achieving coherent outputs. He emphasizes the importance of alignment and safety in training, noting that while today's models are not yet fully capable, future advancements will require serious consideration of potential risks.

"What are you keeping track of while  you're doing long-horizon RL or when   you eventually start doing it? How could  you notice this sort of discontinuous jump   before you deployed these systems bro..."

11
27:58 - 29:30
1:31 duration236 words

The Role of Human Feedback

Schulman discusses the role of Reinforcement Learning from Human Feedback (RLHF) in shaping AI behavior. He explains how models are trained to maximize human approval and the implications of this training on their decision-making processes, highlighting the importance of ensuring that AI systems remain aligned with human intentions.

"make the model turn against you. That doesn't  seem like the hardest thing to do. The way we   train them with RLHF, that does feel very  safe even though the models are very smart.   The model is jus..."

12
29:30 - 30:58
1:27 duration222 words

Introspection in AI Reasoning

This segment explores the potential for AI models to develop reasoning capabilities through introspection and self-assessment. Schulman discusses two approaches to enhancing reasoning in AI: training on outputs and deploying models that can engage in self-dialogue, emphasizing the need for a combination of both methods.

"world. Of course if you assigned a task such as  “make money,” then maybe that would lead to some   nefarious behavior as an instrumental goal. Before we get back to that, let's step back   and talk..."

13
30:58 - 32:42
1:44 duration275 words

Balancing Training Methods

Schulman examines the balance between large-scale training and in-context learning for AI models. He suggests that a middle ground approach could enhance learning efficiency and reasoning capabilities, advocating for a system that allows models to actively learn and adapt during tasks.

"There are probably some analogies though I don't  know exactly how close it is. To some extent,   the models do have drives and goals in  some meaningful way. In the case of RLHF   where you're trying..."

14
32:42 - 34:56
2:13 duration143 words

The Future of AI Learning

In this segment, Schulman discusses the future of AI learning and the potential for models to engage in online learning and introspection. He emphasizes the importance of developing cognitive skills that allow AI to seek out new knowledge and adapt to complex tasks over time.

"time. So I think that you’d get the best  results by combining these two things. 
   Right now, you have these two ways the  model learns. One is in training, whether   it's pre-training or post-trai..."

15
34:56 - 37:09
2:12 duration321 words

Integrating Learning Approaches

Schulman reflects on the integration of various learning approaches in AI development. He discusses the potential for models to combine long-term memory, fine-tuning, and active learning to enhance their capabilities and adaptability in real-world applications.

"memory? Too much to fit in context but  much smaller scale than pre-training?   It might be memory. I don't have context.  Certainly when I'm trying to prepare for   this conversation, I think of wha..."

16
37:09 - 39:24
2:14 duration262 words

The Evolution of AI Training

This segment covers the evolution of AI training methods and the challenges faced in developing effective learning algorithms. Schulman highlights the need for models to learn efficiently and adaptively, drawing parallels between human learning processes and AI development.

"systems that do some online learning and also  have some cognitive skills, like introspecting   on their own knowledge and seeking out  new knowledge that fills in the holes.   Is this all happening ..."

17
39:24 - 40:51
1:27 duration207 words

Future Directions in AI Research

Schulman concludes by discussing the future directions of AI research, particularly in the context of reinforcement learning and the potential for models to act autonomously. He emphasizes the importance of developing algorithms that allow for rapid learning and adaptation in dynamic environments.

"There are all these complicated RL procedures,  many of which you've pioneered. How many of them   will be relevant when you get to the point where  the model itself is smart enough to act as its   ow..."

18
41:09 - 42:41
1:31 duration220 words

The Evolution of ChatGPT

John Schulman discusses the development of ChatGPT, detailing the transition from earlier instruction-following models to a more conversational AI. He explains how the team recognized the need for a chatbot that could handle follow-up questions and clarifications, leading to the creation of a more user-friendly conversational assistant built on GPT-3.5.

"of thing that gets used in a particular task. Interesting. I want to step back and ask about   your own history, at least at OpenAI. You  led the creation of ChatGPT.
At what point   did you reali..."

19
42:41 - 44:41
2:00 duration286 words

Challenges and Innovations in AI

In this segment, Schulman reflects on the challenges faced during the development of ChatGPT, including issues with hallucination and reliability. He emphasizes the importance of mixing instruction and chat data to enhance model performance and user experience, highlighting the iterative process of refining AI capabilities.

"At the same time there were definitely a lot  of people thinking about chat. Google had some   papers like LaMDA and earlier, Meena. They  had these chatbots. It was more like a base   model that was ..."

20
44:41 - 46:28
1:46 duration219 words

The Role of Fine-Tuning in AI Development

Schulman explains the significance of fine-tuning in the evolution of AI models, particularly in relation to ChatGPT. He discusses the necessity of iterative supervised fine-tuning and reinforcement learning to achieve high-quality outputs, and how these processes contribute to the model's ability to understand its limitations.

"the most interesting thing about it. We had it  out to friends and family for a while and we   were thinking about doing a public release. Actually, GPT-4 finished training in August   that year. The ..."

21
46:28 - 48:57
2:28 duration321 words

Scaling AI: Progress and Expectations

Reflecting on the advancements since GPT-2, Schulman shares his insights on the rapid progress in AI capabilities. He discusses the balance between pre-training and post-training, and how improvements in post-training methodologies are expected to enhance model performance and quality in the future.

"That was actually one of the things that I  got excited about as we were developing it.   I realized a lot of the things that people  thought were flaws in language models, like   blatant hallucinatio..."

22
48:57 - 51:22
2:25 duration307 words

The Future of Reinforcement Learning

In this segment, Schulman delves into the future of reinforcement learning (RL) research, discussing the intuition required for effective RL research and the importance of understanding the entire stack of AI development. He emphasizes the need for curiosity and empirical approaches to drive innovation in RL.

"done that, you could have gotten pretty  close but it would have been non-trivial.  We also had another instruction-following  model trained with RL, released a little   before ChatGPT. If you put a c..."

23
51:22 - 54:27
3:05 duration387 words

Generalization and Data Limitations

Schulman addresses the hypothesis of hitting a data wall in AI development, discussing the challenges of generalization across different modalities. He explores the potential for positive transfer between various types of training data and the implications for future AI models.

"We found a lot of gains through post-training.  So I would expect us to keep pushing this   methodology and probably increasing  the amount of compute we put into it.   The current GPT-4 has an Elo s..."

24
54:27 - 58:43
4:15 duration552 words

Understanding Model Efficiency

This segment focuses on the scaling laws of AI models, where Schulman discusses why larger models tend to be more sample efficient. He provides insights into the underlying mechanisms that contribute to this phenomenon, including the model's capacity to learn complex computations in parallel.

"Do you think that hypothesis is wrong? We've  talked about some examples of generalization,   like Spanish to English. One example I think of is  the transfer from code to reasoning in language.   If ..."

25
58:43 - 1:04:02
5:18 duration604 words

The Future Landscape of AI Integration

Schulman shares his vision for the future of AI integration across various sectors. He anticipates that as AI capabilities improve, they will be utilized in more sophisticated tasks, enhancing productivity and potentially accelerating scientific research, while emphasizing the importance of human oversight.

"I can give you a sketchy explanation. You could  say that the model is an ensemble of different   circuits that do the computation. You could  imagine that it's doing computations in parallel   and th..."

26
1:04:02 - 1:08:34
4:32 duration470 words

Balancing AI Autonomy and Human Oversight

In this concluding segment, Schulman discusses the balance between AI autonomy and human oversight in business operations. He raises concerns about the implications of fully autonomous AI systems and the need for regulations to ensure that human interests are prioritized in decision-making processes.

"useful to you. Everyone would have all these  AIs helping them do more and get more done.   Obviously at some point they're going to be  better than everyone at whatever they want to   do. What would..."

27
1:08:59 - 1:11:11
2:11 duration287 words

The Risks of AI-Run Firms

Schulman addresses the potential risks associated with AI-operated companies, including the possibility of malfunction in unpredictable situations. He questions whether AI can truly outperform humans in all aspects and discusses the implications of accountability and liability in AI management.

"what happens if China doesn't decide to do that? You would either have to have every country agree   to this regulatory regime, or you would need  all of the model infrastructure or the model   prov..."

28
1:11:11 - 1:14:49
3:37 duration466 words

Aligning AI with Human Values

This segment focuses on the importance of aligning AI systems with human values and preferences. Schulman discusses the complexities of ensuring that AI models serve the interests of various stakeholders while navigating conflicting demands and ethical considerations.

"with RLHF. You have to aggregate preferences  across a lot of different humans. It'll be maybe   more marked with future, more powerful systems.  But when you say we want these eventual AI systems   t..."

29
1:14:49 - 1:17:49
2:59 duration390 words

The State of Machine Learning Research

Schulman reflects on the current state of machine learning research, comparing it to social sciences in terms of replicability and reliability. He emphasizes the importance of practical applications and the need for more rigorous scientific inquiry within the field.

"not impose our opinions on people. We mostly want  to let people do what they want with the models.   I got a chance to read the Spec beforehand.  This is a question of how well that transfers   over..."

30
1:17:49 - 1:20:41
2:52 duration390 words

Improving Model Efficiency

In this segment, Schulman discusses advancements in model efficiency since GPT-4, exploring how improvements in training methods can lead to better performance without necessarily requiring more computational resources. He highlights the ongoing efforts to enhance both pre-training and post-training processes.

"methods. There's been a decent amount of that  recently. We could use more of that. I think   that's a good thing for academics to work on. On a slightly different note, I'd be really   excited to see..."

31
1:20:41 - 1:23:47
3:05 duration442 words

The Challenge of Chatbot Personalities

Schulman addresses the common criticisms of chatbot responses, particularly their perceived lack of creativity and personality. He discusses the factors influencing chatbot behavior and the ongoing efforts to make AI interactions more engaging and less robotic.

"However, there is a sense in which all of these  models, once they're put in a chatbot form,   have a very similar way of speaking. They really  want to “delve” into things. They want to turn   things..."

32
1:23:47 - 1:27:32
3:45 duration457 words

Future Directions for AI Alignment

This segment delves into the future of AI alignment, discussing how preference models can capture subtle human values. Schulman speculates on the balance between explicit instructions and learned preferences in developing more advanced AI systems.

"that we tend to train for one message at a  time rather than the full interaction. If you   only see one message, then something  that just has a clarifying question,   or maybe a short response with ..."

33
1:27:32 - 1:29:33
2:01 duration258 words

The Complexity of Post-Training

Schulman concludes by discussing the complexities involved in post-training AI models. He highlights the challenges of creating functional models that meet user needs and the competitive landscape of AI development, emphasizing the importance of skilled personnel and organizational knowledge.

"How much of a moat is better post-training?  Companies distinguish themselves currently by   how big their model is and so forth. Will it  be a big moat for who has figured out all the   finickiness t..."

34
1:29:39 - 1:31:17
1:38 duration203 words

The Role of Raters in AI Development

Schulman provides insights into the diverse backgrounds of raters involved in AI training. He explains how different tasks require different skill sets and how the geographical distribution of raters affects the quality of data labeling. This segment emphasizes the importance of skilled raters in refining AI models and the variability in expertise across different domains.

"I guess it helps clear the moat. What is the  median rater like? Where are they based? What are   their politics? What is their knowledge level? It varies a lot. We've definitely hired raters   with..."

35
1:31:17 - 1:33:11
1:54 duration254 words

Generalization vs. Domain Expertise

This segment addresses the balance between generalization and the need for domain-specific expertise in AI training. Schulman argues that while specific examples can enhance model performance, the base model's extensive training on diverse data allows it to generalize effectively, even in technical domains like programming.

"we have now are quite skilled and conscientious. With regards to the plateau narrative,   one of the things I've heard is that a lot of  the abilities these models have to help you   with specific t..."

36
1:33:11 - 1:36:12
3:01 duration440 words

The Future of AI Assistants

Schulman envisions the evolution of AI assistants that can interact with users in a more integrated manner. He discusses the potential for these assistants to manage ongoing projects, proactively suggest actions, and enhance collaboration. This segment concludes with a thought-provoking question about the timeline for AI to replace human jobs, hinting at rapid advancements in the field.

"reasonable behavior in the programming domain. Maybe a final question. We've touched on this   in different ways but let’s put it together. You  said you're training on much more multimodal data.   ..."