
31 segments available
ChatGPT-5 just launched, marking a major milestone for OpenAI and the entire AI ecosystem. Fresh off today's live stream, a16'z Erik Torenberg was joined in the studio by three people who played key roles in making this model a reality: - Christina Kim, Researcher at OpenAI, who leads the core models team on post-training - Isa Fulford, Researcher at OpenAI, who leads deep research and the ChatGPT agent team on post-training - Sarah Wang, General Partner at a16z, who helped lead our investment in OpenAI since 2021 They discuss what’s actually new in ChatGPT-5—from major leaps in reasoning, coding, and creative writing to meaningful improvements in trustworthiness, behavior, and post-training techniques. We also discuss: - How GPT-5 was trained, including RL environments, and why data quality matters more than ever - The shift toward agentic workflows—what “agents” really are, why async matters, and how it’s empowering a new golden age of the “ideas guy” - What GPT-5 means for builders, startups, and the broader AI ecosystem going forward Whether you're an AI researcher, founder, or curious user, this is the deep-dive conversation you won't want to miss. Timecodes: 00:00 ChatGPT Origins 02:13 Model Capabilities & Coding Improvements 04:11 Model Behaviors & Sycophancy 06:15 Usage, Pricing & Startup Opportunities 08:03 Broader Impact & AGI Discourse 16:59 Creative Writing & Model Progress 31:50 Training, Data & Reflections 36:25 Company Growth & Culture 41:39 Closing Thoughts Resources: Find Christina on X: https://x.com/christinahkim Find Isa on X: https://x.com/isafulf Find Sarah on X: https://x.com/sarahdingwang Stay Updated: Let us know what you think: https://ratethispodcast.com/a16z Find a16z on Twitter: https://twitter.com/a16z Find a16z on LinkedIn: https://www.linkedin.com/company/a16z Subscribe on your favorite podcast app: https://a16z.simplecast.com/ Follow our host: https://x.com/eriktorenberg Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details, please see a16z.com/disclosures.
OpenAI researchers discuss the unique opportunity of working on universally useful technology like ChatGPT-5. They emphasize the importance of making AI accessible and capable for a wide range of users, highlighting the excitement around the model's capabilities and its potential impact on everyday tasks.
"I mean, I think it's pretty unique at OpenAI to be able to work on something that's so generally useful. I mean, it's like everything they tell you not to do at a startup is just like your user is any..."
Christina Kim and Isa Fulford introduce themselves and their roles at OpenAI, detailing their contributions to the development of ChatGPT-5. They reflect on their experiences and the evolution of the model from WebGPT to ChatGPT, setting the stage for discussing the latest advancements.
"testing and they're like, oh, I thought I asked like a really hard question. I feel like a little bit insulted that I thought for like two seconds or like when it doesn't even want to think at all. So..."
The researchers express their excitement about the improved utility of GPT-5, noting that it performs significantly better in real-world applications. They discuss the positive evaluations and the anticipated user experience, emphasizing the model's enhanced usefulness across various tasks.
"for about four years now. I originally worked on WebGPT, which was the first LLM using tool use, but it was just one question. So the model learned how to use the browser tool, but you only ask one qu..."
Sarah Wang highlights the significance of coding improvements in GPT-5, mentioning the contributions of team members and the meticulous effort put into enhancing coding performance. The discussion focuses on how these advancements position GPT-5 as a leading coding model in the market.
"just, it's way more useful. Like in cross, like all the things that people actually use chat for. And it's not just like, and it's, I think the eval numbers look good, but then also like the way when ..."
The conversation shifts to front-end web development, with researchers discussing the leaps made in this area with GPT-5. They compare its capabilities to previous models, emphasizing the attention to detail and data quality that have led to significant improvements in front-end coding.
"just caring so much about getting coding working well. And maybe actually just to double click on front end web development. I mean, we've seen as sort of investors in the ecosystem, that's obviously ..."
The team discusses the intentional design choices made for GPT-5's behaviors, particularly addressing previous issues with sycophancy. They explain the balancing act of creating a helpful assistant while avoiding overly engaging responses, aiming for a healthy interaction model.
"Really exciting to see. Loved the demos in the live stream, too. I wanted to ask about model behaviors because I know you worked on that, too. But how did you guys think about that for GPT-5? And ther..."
The researchers delve into the relationship between hallucinations and deception in AI models. They explain how GPT-5 has been designed to minimize these issues by encouraging careful reasoning and thoughtful responses, leading to a more reliable user experience.
"to actually feel like? And I think we were really excited with GPT-5 because it's kind of a time to like reset and rethink about, especially since it's so easy to make something, I think very engaging..."
As they evaluate GPT-5's performance, the team anticipates new use cases that will emerge from its capabilities. They express excitement about the model's pricing and how it will empower startups and developers to innovate in ways that were previously limited.
"lot of the previous models for hallucinations. Over the next few weeks as you're evaluating usage, what are the biggest questions that you're having or that you're sort of anticipating being potential..."
The discussion highlights the potential for non-technical individuals to leverage GPT-5 for coding and app development. Researchers envision a surge in indie businesses and creative projects, emphasizing the model's ability to democratize technology for idea-driven entrepreneurs.
"here too, how did deep research, chat GPT, operator, sort of your existing products inform how you went about approaching GPT-5. One thing that's interesting is with reinforcement learning, training a..."
The researchers reflect on GPT-5's implications for the broader AGI conversation. They discuss how the model's advancements challenge the notion of hitting a performance ceiling and emphasize the importance of real-world usage as a metric for progress in AI development.
"opportunities you're more excited about because of this. I mean, people always say vibe coding. I think basically non-technical people have such a powerful tool at their hands. I think really you just..."
The team explains their approach to evaluating GPT-5's performance, focusing on the capabilities they aim to enhance. They discuss the creation of internal evaluations and the importance of aligning model improvements with user needs, ensuring that advancements are practical and impactful.
"are using this in their daily lives to help them like across multiple tasks. So I feel like that's actually like the ultimate usage in terms like that I'm excited about for in terms of like, are we ge..."
The researchers discuss the challenge of balancing general capabilities with specific use cases in model development. They emphasize the importance of catering to a wide range of tasks while also focusing on critical areas like coding, ensuring that improvements benefit diverse user needs.
"hill climb on those. Yeah. I think we make this joke a lot internally that if you want to nerd-side someone into working on something, you just need to make a good eval, and then people are going to b..."
The team reflects on the progression of AI models, noting that as models become smarter, they improve in instruction following and tool use. They discuss the significance of enhancing general intelligence to unlock new capabilities and how foundational algorithms contribute to this evolution. The conversation touches on the challenges of developing effective agents and the role of reinforcement learning in achieving meaningful advancements.
"think if you choose a capability that's quite general, like online research, you just have to make sure that you represent like a distribution of tasks across loads of different domains if you want to..."
The researchers debate the importance of architecture, data, and scale in improving AI models. They emphasize that high-quality data is crucial, especially as learning methods become more efficient. The discussion highlights the need for realistic reinforcement learning environments to train models effectively and the potential for startups to contribute to this area.
"we realized, okay, like this is a thing that's going to actually let us get to useful agents. And so I think it's interesting at OpenAI because you have people pushing like, you know, foundational alg..."
The conversation shifts to the significance of realistic reinforcement learning environments in training AI models. The researchers argue that better tasks and environments lead to improved model performance. They discuss the challenges of creating effective training scenarios and the potential for AI agents to perform complex tasks on computers, emphasizing the need for extensive training data.
"you see for the next stage? Is that, I mean, maybe tying it to RL environments, is there sort of a lack of good, realistic RL environments that that's sort of the next frontier, which maybe creates an..."
The researchers share their excitement about the improvements in creative writing capabilities with GPT-5. They discuss how the model can assist in crafting sensitive content, such as eulogies, and how it enhances everyday writing tasks. The conversation highlights the emotional impact of the model's writing and its utility in helping users express themselves more effectively.
"you know training on training on way more things yeah Let's talk about creative writing. Maybe you can talk about the improvements there, how you think about it. That's one of my favorite improvements..."
The team reflects on how quickly users adapt to new AI capabilities, citing the rapid acceptance of ChatGPT. They discuss the ease of interfacing with advanced models and how users often take for granted the sophisticated assistance these tools provide. The conversation underscores the importance of making AI approachable as it becomes more intelligent.
"the team. I want to see those prompts. Yeah. We're now all just looking for M dashes. That was good to say. We're like, where do you stand on the M dash discourse? I like M dashes. I do that normally ..."
The researchers explore the limitations of GPT-5, particularly its inability to take real-world actions independently. They discuss the potential for future models to handle more complex tasks and the importance of user trust in allowing AI to perform actions without constant oversight. The conversation hints at the exciting possibilities for future developments in AI capabilities.
"gonna be like quite approachable to people do you think the jump from GPT 4 to 5 was bigger or 3 to 4 or maybe 3.5 to 4 I mean at least one one thing for me and my usage of it is sometimes I'm wonderi..."
The discussion centers on the concept of AI agents and their role in performing useful work for users. The researchers define what an agent means in the context of AI capabilities and express their vision for future developments. They emphasize the importance of proactive agents that can anticipate user needs and perform tasks autonomously, paving the way for more integrated AI solutions.
"them more, you might, you know, allow it to do things for you without, um, checking in with you as much. Maybe just to build on that question, um, for, in terms of what it can't do today, but what you..."
The researchers highlight the new capabilities launched in ChatGPT agents, focusing on improving deep research and task management. They discuss the importance of synthesizing information from various sources and the challenges of creating and editing documents. The segment also touches on consumer use cases, such as shopping and trip planning, showcasing the versatility of the technology.
"guess my very general definition would just be something that does work, useful work for me on my behalf with I would say asynchronously. So like you'd kind of leave it and then come back and get, eit..."
In this segment, the conversation explores the evolving expectations of users regarding response times and the quality of answers provided by AI. The researchers reflect on how users are now willing to wait longer for high-quality responses, contrasting this with earlier demands for speed. They discuss the implications of this shift for future AI development.
"action, um, which is, um, interesting cause it's, um, it's very, it's kind of the last step often of, of a, of a task. And it's the, maybe a task that would take less time for a human. And it's like a..."
The researchers discuss the trade-offs between response speed and the quality of information provided by AI. They reflect on their experiences optimizing for latency and how user expectations have changed over time. The segment emphasizes the importance of delivering thorough answers while managing user impatience.
"do you think is the ideal frontier for something like that? Yeah, it's interesting because I worked on, I built the retrieval on ChatGPT and was on the browsing team before this. Tina was also on the ..."
This segment focuses on how product design can shape user expectations regarding response times and the nature of the information provided. The researchers discuss the challenges of meeting user demands for quick, thorough responses and how the design of AI systems can influence these perceptions.
"would take the human to do, they're willing to wait for it? Or is that just constantly shifting sand? I think with these launches, people's expectations keep getting changing. Yeah, I do think we have..."
The researchers identify key bottlenecks in achieving reliable agency with AI. They discuss the challenges of training models on diverse data and the implications of having agents perform tasks on behalf of users, particularly concerning privacy and oversight. The segment highlights the need for ongoing improvements in AI capabilities.
"to the amount of time that they wait. I think we hear this with GPT-5 internally when people are testing it. They're like, oh, I thought I asked a really hard question. I feel like a little bit insult..."
In this segment, the researchers address the specific challenges faced by AI agents in performing browsing tasks. They discuss the limitations of available training data and the need for innovative approaches to enhance AI capabilities in knowledge work. The conversation emphasizes the importance of developing effective training methodologies.
"every model that's built on top of that. So I think that will also help, especially with like multimodal capabilities, as Tina said, with like computer use, because it's like just literally looking at..."
Christina Kim explains the concept of mid-training and its role in enhancing AI models. She describes how mid-training serves as a bridge between pre-training and post-training, allowing for updates to the model's knowledge without the need for extensive retraining. This segment highlights the significance of mid-training in keeping AI models current.
"actually probably a big one. Mm hmm. just for general improvements of computer usage. Do you think you'll lean more heavily on human data vendors to help collect that? Or given it doesn't exist, to yo..."
The researchers reflect on the evolution of AI over the past few years, particularly in relation to WebGPT. They discuss the challenges of grounding language models and the importance of addressing issues like hallucinations. This segment provides insights into the historical context of AI development and the lessons learned along the way.
"model's intelligence and up-to-dateness. Christina, did you work on WebGPT? Yes, I did. Okay, so you're basically like an AI historian. Yes, yes. She also watched some Confucius. I'm an elder. So can ..."
Christina and Isa discuss the pivotal moments that made them realize they were part of a groundbreaking company in AI. They share personal anecdotes about their early experiences with OpenAI and how the scaling laws of AI models inspired them to contribute to this transformative technology.
"like not quite maybe for everyone yet um but there's something here when did you realize like i'm working at one of the most important companies of this generation like like when was the moment where ..."
The researchers reflect on the significant changes at OpenAI since they joined, highlighting the growth from a small team to a larger organization. They discuss how the company has maintained a startup culture that encourages innovation and collaboration across teams, despite its rapid expansion.
"read and reread Calvin French Owen's piece, just his reflections on working at OpenAI. Curious, and you don't have to comment on that piece unless you want to, but would love your reflections on the c..."
Christina and Isa explore OpenAI's unique position as both a consumer and enterprise-focused company. They discuss how this duality influences their mission to create the most capable AI while ensuring accessibility for a wide range of users.
"to maintain that culture, which I think is pretty special. Yeah, we definitely reward agency. And I think that's like what's always been true. And I think especially on the research side, the teams ar..."
In closing, the researchers emphasize the significance of GPT-5 and its usability for users. They express excitement about making advanced reasoning models available to everyone and anticipate the innovative applications that will emerge from this latest iteration of AI technology.
"of taste has become also very widely used. What does good taste mean within open AI? How do you know it when you see it? And is that something that even in a world where everything, the cost to produc..."