OpenAI researcher on why soft skills are the future of work | Karina Nguyen

0:00

Not only are you working at the  cutting edge of AI and LLMs, you're actually building the cutting edge.

0:06

When I first came to Anthropic and I was like,  "Oh my God, I really love front-end engineering."

0:10

And then the reason why I switched to research is  because I realized, "Oh my God, Claude is getting better at front-end.

0:14

Claude is getting better  at coding.

0:14

I think Claude can develop new apps."

0:20

What skills do you think will be most valuable  going forward for product teams, in particular?

0:26

Creative thinking and you kind of  want to generate a bunch of ideas and filter through them and not just  build the best product experience.

0:30

I think it's actually really, really hard  to teach the model how to be aesthetic or really good visual design or how to be  extremely creative in the way they write.

0:42

What do you think people most  misunderstand about how models are created?

0:45

When you taught the model, some of the  self-knowledge of you actually don't have a physical body to operate in the physical  world, the model would get extremely confused.

0:58

Today my guest is Karina Nguyen.

0:58

Karina is an  AI researcher at OpenAI where she helped build Canvas, tasks, the o1 chain-of-thought model  and more.

1:04

Prior to OpenAI, she was at Anthropic where she led work on post-training and evaluation  for the Claude 3 models, built a document upload feature with 100K context windows and so much  more.

1:15

She was also an engineer at New York Times, was a designer at Dropbox and at Square.

1:20

It's very  rare to get a glimpse into how someone working on the bleeding edge of AI and LLMs operates and  how they think about where things are heading.

1:31

In our conversation, we talk about how teams  that OpenAI operate and build products, what skills she thinks you should be building  as AI gets smarter, how models are created, why synthetic data will allow models to  keep getting smarter and why she moved from engineering to research after realizing  how good LLMs are going to be at coding.

1:44

If you enjoy this podcast, don't forget to  subscribe and follow it in your favorite podcasting app or YouTube.

1:52

It's the best  way to avoid missing feature episodes and it helps the podcast tremendously.

1:56

With that, I bring you Karina Nguyen.

2:02

This episode is brought to you by Enterpret.

2:02

Enterpret unifies all your customer interactions from Gong calls to Zendesk tickets to Twitter  threads to app store reviews, and makes it available for analysis.

2:11

It's trusted by leading  product orgs like Canva, Notion, Loom, Linear, monday.

2:17

com, and Strava, to bring the voice of the  customer into the product development process, helping you build best-in-class products faster.

2:22

helping you build best-in-class products faster.  What makes Enterpret special is its ability to build and update customer-specific AI models that  provide the most granular and accurate insights into your business, connect customer insights  to revenue and operational data in your CRM or

2:37

data warehouse to map the business impact of  each customer need and prioritize confidently, and empower your entire team to easily take  action on use cases like win-loss analysis, critical bug detection and identifying drivers  of churn with Enterpret's AI system, Wisdom. Looking to automate your feedback loops and  prioritize your roadmap with confidence,

2:53

Looking to automate your feedback loops and  prioritize your roadmap with confidence, like Notion, Canva and Linear?

2:57

Visit  E-N-T-E-R-P-R-E-T.

2:57

com/Lenny to connect with the team and to get two free months when you  sign up for an annual plan.

3:04

This is a limited time offer. That's Enterpret. com/Lenny.

3:09

This episode is  brought to you by Vanta.

3:09

And I am very excited to have Christina Cacioppo, CEO and co-founder Vanta,  joining me for this very short conversation. Great to be here.

3:22

Big fan of  the podcast and the newsletter.

3:24

Vanta is a longtime sponsor of the show, but for some of our newer listeners,  what does Vanta do and who is it for? Sure.

3:32

So we started Vanta in  2018.

3:32

Sure. So we started Vanta in  2018. Focused on founders, helping them start to build out their  security programs and get credit for all of that hard security work with compliance  certifications like SOC 2 or ISO 27001 today,

3:46

we currently help over 9,000 companies, including  some startup household names like Atlassian, Ramp and LangChain start and scale their  security programs and ultimately build trust by automating compliance, centralizing  GRC, and accelerating security reviews. That is awesome. I know from experience  that these things take a lot of time and That is awesome.

4:02

I know from experience  that these things take a lot of time and a lot of resources and nobody  wants to spend time doing this.

4:10

That is very much our experience,  but before the company and to some extent during it.

4:13

But the idea  is with automation, with AI, with software, we are helping customers  build trust with prospects and customers in an efficient way.

4:21

And our joke, we started  this compliance company so you don't have to.

4:26

We appreciate you for doing that.

4:26

And you  have a special discount for listeners, they can get $1,000 off Vanta at Vanta.

4:30

com/Lenny, that's V-A-N-T-A.

4:34

com/Lenny for $1,000  off Vanta.

4:34

Thanks for that, Christina. Thank you.

4:45

Karina, thank you so much for  being here. Welcome to the podcast.

4:48

Thank you so much, Lenny, for inviting me.

4:50

I'm very excited to have you here because not  only are you working at the cutting edge of AI and LLMs, you're actually building the cutting edge of  AI and LLMs.

4:56

You recently launched this feature, which basically...

5:02

the first agent feature  of OpenAI.

5:02

I also just did this survey, I don't know if you know about this.

5:08

I did  a survey of my readers and asked them what tools do you use every day in your work  and most use?

5:11

And ChatGPT was number one, above Gmail, above Slack, above anything else.

5:16

90% of people said they use ChatGPT regularly. That's quite good. It's absurd.

5:23

It wasn't around two years ago. Yeah.

5:26

Also, we're recording this the week that OpenAI  announced Stargate, which is this half trillion dollar investment in AI infrastructure.

5:31

So  there's just a lot happening constantly in AI and you have a really unique glimpse into  how things are working, where things are going, how work gets done.

5:40

So I have a lot of questions  for you.

5:40

I want to talk about how you operate and how you work at OpenAI, where you think things  are going, what skills are going to matter more and less in the future, and also just where  things are going broadly. So how does that sound? Sounds great. Thank you so much.

5:54

Yeah, I was extremely lucky to join early days Anthropic and learned  a lot of things there.

6:00

And I joined OpenAI around eight months ago.

6:07

So, yeah,  I'm excited to dive more in into- Okay, I'm going to definitely ask you about  the differences between those, but I want to start more technical and just dive right in.

6:14

I  want to talk about model training.

6:14

People always hear about models being trained, these big models,  how much data takes, how long it takes, how much money toss it takes, how we're running out of  data, which I want to talk about.

6:25

Let me just ask you this question.

6:30

What do you think people  most misunderstand about how models are created?

6:36

Model training is more an art than a science.

6:36

And in a lot of ways we, as model trainers, think a lot about data quality.

6:45

It's one of  the most important things in model training is like how do you ensure the highest quality  data for certain interaction model behavior that you want to create?

6:57

But the way you debug  models is actually very similar the way you debug software.

7:03

So one of the things that I've learned  early days at Anthropic was we've discovered especially this Claude 3 training, when you  taught the model some of the soft-knowledge of, "Hey, you actually don't have a physical  body to operate in the physical world."

7:21

But then at the same time you had data that  taught the model some of the function calls, which is like, "This is how you set the alarm."

7:27

And so the model would get extremely confused about whether it can set an alarm, but it doesn't  have a body in the physical world.

7:33

So it's like the model gets confused and sometimes it'll over  accuse.

7:41

So sometimes it says, "Look, I don't know.

7:46

Sorry, I cannot help you."

7:46

And so there is always  a balance trade off between how do you make the model to be more helpful for users, but also not  being harmful in other scenarios.

7:53

And so it's always about how do you make the model more robust  and operate across a variety of diverse scenarios. That is so funny.

8:09

I never thought about that.

8:09

Most of the data that it's trained on is kind of assuming it's like a human describing  the world and how they operate.

8:12

It assumes there's a body and you could do things, and  the model is told you don't have a body. Yeah. Okay.

8:21

I want to talk a little bit about  data while we're on this topic.

8:21

I know you have strong opinions here.

8:26

There's this meme  that models are going to stop getting smarter because they're running out of data.

8:30

They're  trained in a large part on the internet and there's only one internet and they've already  been trained on it, what more can you show them about the world?

8:38

And there's this trend  of synthetic data, this term synthetic data. What is synthetic data?

8:43

Why do you think it's  important?

8:43

Do you think it's going to work?

8:46

I think there are two questions here.

8:46

We can  unpack one at a time.

8:46

But people say we are hitting the data wall.

8:54

I think people think more  in the terms of pre-trained large models that are trained on the entire internet to predict  the next token.

9:01

But what actually the model is learning during that process is actually how do  you compress the compression algorithm here?

9:09

The model learns to compress a lot of knowledge  and it learns how to model the world.

9:15

So the next prediction of the word, like, "Teach me  how to drive," basically.

9:22

And you only have a few words that will match that, a car.

9:29

So the  model actually learns about the world in itself.

9:38

So it's like it's modeling human behavior,  sometimes it's modeling...

9:38

And when you talk to pre-trained models which are very, very large,  they're actually extremely diverse and extremely creative because you can talk to almost any  Reddit user through a pre-trained model.

9:56

But I think what's happening right now with  new paradigm of o1 series is that the scaling in post-training itself is not hitting the wall.

10:04

And that's because basically we went from raw data sets from pre-trained models to infinite amount  of tasks that you can teach the model in the post-training world via reinforcement learning.

10:21

So any task, for example, how to search the web, how to use the computer, how to write, wow,  all sorts of tasks that you trying to teach the model all the different skills.

10:36

And that's why  we're saying there's no data wall or whatever, because there will be infinite amount of tasks  and that's how the model becomes extremely super intelligent.

10:47

And we are actually  getting saturated in all benchmarks.

10:52

So I think the bottleneck is actually in  evaluations that we don't have all the frontier, like evals like, I don't know, GPQA,  which is a Google-proof question answering, PhD level intelligence.

11:08

The benchmark is  getting to, I don't know, more than 60, 70%, which is what PhD gets.

11:13

So it's  literally hitting the wall in like evals.

11:20

I want to follow both those threads.

11:20

So the  first is on this idea of synthetic data.

11:20

Is a simple way to understand it, that the models are  generating the data that future models are trained on and you ask it to generate all these ways of  doing stuff, all these tasks as you described, and then the newer models trained on this  data that the previous model generated?

11:39

Some tasks are synthetically curated.

11:39

So this  is an active research area is how can you synthetically construct new tasks with models  to learn.

11:46

Sometimes when you develop products, you get a lot of data from the product and user  feedback and you can use that data too in this cross-training world.

12:01

Sometimes you still want  to use human data because actually some of the tasks can be really, really hard to teach.

12:10

Experts  only know certain knowledge about some chemicals or biological knowledge, so you actually need to  tap into the experts' knowledge a lot.

12:20

So yeah, I think to me synthetic data training is  more for product...

12:28

It's a rapid model iteration for similar product outcomes.

12:35

And  we can dive more into it, but the way we made Canvas and tasks and new product features for  ChatGPT was mostly done by synthetic training.

12:49

Let's actually get into that.

12:49

That's really  interesting.

12:49

I want to talk about evals, but let's follow that thread.

12:52

So talk  about how this helped you create Canvas.

12:56

So when I first came to OpenAI,  I really had this idea of, "Okay, it would be really cool for ChatGPT to actually  change the visual interface but also change the way it is with people."

13:11

So going from being  a chatbot to more of a collaborative agent, and the collaborator is a step towards more  genetic systems that become innovators ultimately.

13:28

And so the entire team of applied engineers,  designers, products, research got formed in the air almost out of nothing.

13:37

It's just like a  collection of people who just got together and we rapidly started iterating with each other.

13:43

Actually Canvas is one of the...

13:43

I would say the first project at OpenAI, where researchers  and applying engineers started working together from the very beginning of the product  development cycle.

13:55

And I think there's a lot of things that we have learned on the  way, but I definitely came with the mindset of, "We need to do a really rapid model situation  such that it would be much easier for engineers to work with the latest model possible, but also  learn from user feedback or early internal dog food.

14:24

How do we improve the model very rapidly?"

14:24

And it's really hard to kind of like figure out how people...

14:34

when you deploy a product, how  people would be able to use it.

14:34

And so the way you synthetically train the model is physically  figuring out what are the most core behaviors that you wanted the product feature to do.

14:47

And  for Canvas, for example, it came down to three main behaviors.

14:55

It was how do you trigger Canvas  for prompts like, "Write me a long essay," when the user intention is mostly iterating over long  documents?

15:02

Or, "Write me a piece of code," or when to not trigger Canvas for prompts like, "Can you  tell me more about President..."

15:09

I don't know, some of the general questions.

15:18

So you  don't want to trigger Canvas because the user intention is mostly getting answer, not  necessarily iterate over the long document.

15:28

The second behavior is how do we teach the  model to update the document when the user asks?

15:36

So one of the behaviors that we taught  the model is actually have some agency and autonomy to literally go to the document and  select specific sections and either delete it or edit, so highlight it and rewrite certain  sections.

15:51

Sometimes the user would just say, "Change the second paragraph to be something  friendlier," and we would have to teach the model to literally find the second paragraph  in the document and change it to a friendly tone.

16:09

So basically you teach both how to trigger  edit itself, but also how do you teach the model to get higher quality edit for the document?

16:17

In case of coding, for example, there's also the question of how good the model is  of completely rewriting the document, versus having a very specific target edits.

16:30

So that's another layer of decision boundary within edit itself is, "Let's select the entire  document and rewrite completely, or do you want to have a very targeted custom behavior."

16:42

And  when we first launched the model, we would bias the model towards more rewrites because we saw the  quality of the rewrites were much higher.

16:48

But over time you are shifting based on user feedback and  what you're learning from iterative deployment.

17:02

Lastly, the third behavior that we taught  synthetically the model is how to make comments on any document.

17:08

So the way we used that is we would  use o1 model to seem a way of user conversation, let's say like, "Write me a document  about XYZ."

17:19

But then we used o1 to produce the document and then we injected user  prompt to be like, "Oh, make some comments, critique my piece of writing or critique this  piece of writing that you just made."

17:33

And then we taught the model to make comments on the  document on very specific [inaudible 00:17:45] So it's also what kind of comments you want the  model to make.

17:45

Do they make sense or not?

17:45

How do you teach the quality of that?

17:51

And it all  came down to measuring progress via very robust evals.

18:00

But, yeah, this is how you used o1 and  a synthetic data generation for the training.

18:07

Okay, that's so interesting.

18:07

So you talk about  this idea of teaching the model and you mentioned how it's using synthetic data to teach the model  different behaviors is a simple way to think about it.

18:17

Basically that's where you do that by showing  it what success looks like using basically evals.

18:23

Is that the simple way to think about it?

18:23

Like,  "Here's what you doing this successfully would look like," and that teaches it, "Okay, I see this  is what I should be doing [inaudible 00:18:31]" Yeah, great. Yeah, amazing. Yeah, you got it. Okay, got it.

18:32

I want to start unpacking  what your day-to-day looks like as you're building these sort of things.

18:36

Is  it like you sitting there talking to some version of ChatGPT, crafting these evals? Sometimes I do that.

18:44

Sometimes I do sit with  ChatGPT.

18:44

Actually, I think I learned this so much from Anthropic, is people spend so much time  prompting models and where quality's a really bad batch all the time, and you actually get a lot  of new ideas of how do you make the model better?

19:05

It's like, "This response is kind of weird. Why's it doing this?"

19:05

And you start debugging or something, or you start figuring out new methods  of how do you teach the model to respond in the different way, have better personality, let's say.

19:16

So it's the same thing of how personality is made in the models with those.

19:25

It's  very similar methods.

19:25

But, yes, I think my time at OpenAI have changed.

19:29

I think  when I first came, I was mostly research IC work so I was like building a lot of...

19:36

I was  running code, training models, write evals, working with PMs and designers to learn, teach  them how to even think about evaluation.

19:43

I think that was really cool experience and I think it  was just like an adoption of, "How do we do this product management of AI feature for our AI  models?"

19:59

Yeah, but now it's mostly management and mentorship.

20:10

I'm still doing IC research code up  to 4:00 PM, although.

20:10

But I just kind of changed.

20:21

All right, don't talk too  much about being a manager. Okay.

20:22

Because everyone's in firing their managers.

20:22

"Who needs managers anymore?" That's what I hear now. Just kidding.

20:26

It's interesting that so  much of your time was spent on teaching product teams how evals integrate and how important  it is.

20:32

And I've heard this a few times and I haven't personally experienced it yet, so I  think it's an important thread to follow is just how writing these evaluations is going  to become increasingly an important part of the job of product teams, especially when  they're building AI features and working with LLMs.

20:50

So can you just talk a bit more  about what that looks like?

20:50

Is it sitting there with an Excel spreadsheet basically showing,  "Here's the input, here's the output, here's how good the result was"?

20:58

Talk about what  that actually looks like very practically.

21:02

It certainly depends on what you're developing,  but there are various types of evaluations.

21:09

Sometimes I do ask product managers, or there's  also new roles that we have, model designers, to go through some of the user feedback maybe or  think of various user conversations that should have triggered...

21:26

Under these circumstances,  it should trigger Canvas.

21:26

And then you have this ground truth label of, "Okay with this  conversation it should look trigger Canvas, under this conversation it should not trigger Canvas."

21:36

And you have this very deterministic kind of eval that for decision-making behaviors is like this.

21:43

When we were launching tasks, for example, how do you make correct schedules is actually  really hard for the model.

21:49

But we built out some of the deterministic evaluations that  is like, "Okay, if the user says 7:00 PM, the model should say 7:00 PM."

22:04

So if you can  have deterministic evals whether it's pass or fail.

22:09

And the way it works is all the...

22:09

Sometimes I ask product managers to just go create a double sheet, have different  tabs and what's the current behavior, what's the ideal behavior and why, and some notes.

22:22

And sometimes they usually use it with evals, sometimes we use it for training.

22:30

Because  if you give the spreadsheet to o1 model, it can probably figure out how to teach itself  a good behavior.

22:35

And I think there are second type of evals that is more prevalent is human  evaluations.

22:44

And you can have specific trainers or you can have internal people to when you have  a conversation of the prompt and then you have various completion of models, you choose the win  rate. Which model is the best?

23:00

Which model produce the highest quality comment or edit?

23:06

And then you  can have continuous win rates.

23:06

And as you develop new models it should always win over the previous  models.

23:14

So it depends on what you want to measure. So interesting.

23:22

Basically what I'm hearing,  and there's something I'm learning about as I talk to people, is product development  might move from this, "Here's a spec PRD, let's build it together and then cool, let's  review it. Are we happy with this?"

23:32

From that to, "Hey, AI, build this thing for me and  here's what correct looks like," and I'm spending all my time on what does  correct look like on evals essentially.

23:47

You definitely want to measure progress  of your model and this is where evals is, is because you can have prompted model as  a baseline already.

23:52

And the most robust evals is the one where prompted baselines get the  lowest score or something.

23:59

And then because then you know if you're trained a good model, then it  should just hill climb on that eval all the time, while not also regressing on other intelligence  evals.

24:12

That's what I'm saying, it's more of an art than science.

24:19

It's like, "Okay, if  you optimize the model for this behavior, you don't want to brain damage in other areas  of intelligence or..."

24:23

This is happening all the time in every lab, in every research team.

24:30

I would say prompting is also a way to prototype new product ideas.

24:39

Early days at Anthropic  when I was working file uploads feature, I remember I was just prompting the model to  just...

24:46

I remember we were launching a hundred key contexts.

24:54

I was just prototyping this in their  local browser. I did the demo.

24:54

People really, really loved it.

24:59

And they just wanted API for  file uploads or something.

24:59

And then that's when it clicked to me, and also one of the blog posts a  long time ago, it clicked on me prompting is a new way of product development or prototyping  for designers and for product managers.

25:20

For example, one of the features that I  want to do is have a personalized starter prompts.

25:27

So whenever you come to Claude, it  should recommend you starter prompts based on what your interests are.

25:35

And so you  can literally do it prompting for that. Mm-hmm. To experiment with that.

25:44

Another feature was generating titles  for the conversations.

25:44

It's a very small micro experience but I'm really  proud of.

25:50

The way we did that was we took five latest conversation  from the model, asked the model, "What's the style of the user?"

25:58

And then for  the next new conversation, the generated title will be of the same style.

26:04

It's just like  really little micro experiences like this. That's so cool.

26:11

Did you do  that at Anthropic or at OpenAI? At Anthropic. Okay, cool.

26:15

I love the file upload  feature that Claude has by the way.

26:18

ChatGPT doesn't have that yet, is that right? I think has the way.

26:22

[inaudible 00:26:23] I think the way it's implement  is very different though. Okay.

26:25

Maybe it's the PDF feature, because  I use it all the time with Claude. Yeah. Okay. That's cool.

26:29

Somebody needs to get on that.

26:29

Main, it's wild  how many features you built that I use every day and that many people use every day.

26:33

This  prototyping point you made is really important.

26:37

It's something that comes up a ton on this  podcast also of how that...

26:37

is maybe the way that AI has most impacted the job of product  builders recently is just prototyping instead of going from showing just like, "Here's a PRD,  here's a design."

26:46

PMs are more and more just, "Here's the prototype with the idea that I  have," and it's working. You can play with it. Yeah. Yeah.

26:55

Okay, I want to spend a little more time  on how you operate.

26:55

So you talked about you built this in launch of this tasks feature,  is that the way to describe your tasks? Yeah.

27:06

So talk about how that emerged and let's  better understand just how you collaborate with product teams and how OpenAI works  in that way, whatever you can share there.

27:14

I think Canvas and tasks are going into the  bucket of projects where it's more short or medium terms.

27:21

And actually the way Canvas  and tasks came about to be was it started with one person prototyping and creating  a spec. It's kind of like PRD.

27:30

It's like creating a spec of the behavior of the model.

27:39

I don't think tasks is extremely groundbreaking feature necessarily.

27:49

What makes it really cool  is because the models are so general...

27:49

Model can now search, they can write sci-fi  stories, they can search for stocks, they can summarize the news every day.

28:02

Because  the models are so general giving something familiar to people that notifications is  very familiar, having reminders is very familiar.

28:13

So feeling like a form factor for the  people who are very familiar, same as Canvas, Google Docs is very familiar, but then you add  magical AI moment and it becomes very powerful.

28:26

But the way it comes usually operationally...

28:26

Yeah, size is like a prototype, literally prompted prototype of how you would want  the model to behave.

28:31

For tasks, for example, you need to design...

28:38

Literally design thinking  is like okay, well, if the user says, "Remind me to go to lunch at 8:00 AM tomorrow," what  information does the model need to extract from that prompt in order to create a reminder?

28:55

And so  this is how you design a spec for a new feature, like a tool.

29:04

Canvas and tasks are all tools.

29:04

So it's like how do you create the tool stack?

29:09

And then it's mostly like developing  JSON schema.

29:09

It was like, "Okay, from this problem maybe the model should extract  the time that the user requested."

29:16

And then you think about which format do you want the time to  be?

29:23

And then how do you want the model to notify you is basically the user should give instruction  to the model.

29:30

And then this instruction would fire off every day or something at that particular  time.

29:39

So, for example, if you say, "Every day I want to learn know about the latest AI news,"  the model should rewrite into, "Okay search for the latest AI news and this task will get fired  at that particular type that the user requested."

30:02

And then your design is like tool spec. Actually,  I don't know.

30:02

I feel like sometimes it's through conversations I...

30:09

Either people ask me to join  the [inaudible 00:30:15] team and they're like, "Oh my god, we need researchers."

30:17

Or like,  "We need some support.

30:17

We need to train the models," or sometimes.

30:22

Canvas was mostly  like I just pitched the idea of...

30:22

It got staffed quite immediately during the break, so  it's dependent on the project.

30:28

And then usually with staffing is mostly a product manager,  model designer, actual product designer, a couple of researchers and a bunch of applied  engineers.

30:42

Depends on the complexity of a project.

30:48

And then for tasks it took, I don't know, like  two months or so to go from zero to one basically. Oh wow.

30:58

For Canvas this was like four, five months,  I guess, to go from zero to one.

30:58

And then you teach product managers how to build evals and  maybe how do we not only ship the better feature, but how do we think longer term?

31:16

What kind of  cool features did you want tasks to have?

31:16

I think it would be nice for tasks to be a little  bit more personalized.

31:22

It'd be nice to have to create tasks via voice on a mobile,  right?

31:28

This is how you get research roadmap right here is thinking how the  feature will be developed in the future.

31:39

And then from there it's like you  start getting data sets.

31:39

With evals, you want to make sure that goes well.

31:46

And then you  need to have a trade-off between what methods you want to use.

31:53

And the reason why I really love  relying purely on synthetic data instead of collecting data from humans is because it's much  more scalable, it's cheap, less than half.

31:59

You literally sample from the model and you teach  the core behaviors of the models and that will generalize to all sorts of diverse coverage.

32:10

And when you launch the beta feature, you learn so much from the users that  you can...

32:17

All your synthetic sets can be shifted in the distribution and how the users  behave on the product behavior.

32:24

And this is how we improve.

32:28

And this is what happened with  Canvass too when we launched from beta to GA. Okay.

32:34

This episode is brought to you by Loom.

32:36

Loom lets you your  screen, your camera and your voice to share video messages easily.

32:42

Record a Loom and send  it out with just a link to gather feedback, add context or share an update.

32:47

So now you  can delete that novel link email that you were writing.

32:52

Instead, you can record your  screen and share your message faster.

32:52

Loom can help you have fewer meetings and make the  meetings that you do have much more productive.

33:02

Meetings start with everyone on the same page  and end early.

33:02

Problem solved, time saved.

33:02

We know that everyone isn't a one-take wonder when it  comes to recording videos.

33:08

So Loom comes with easy editing and AI features to help you record once  and get back to the work that counts.

33:13

Save time, align your team, stay connected and get  more done with Loom.

33:19

Now part of Atlassian, the makers of Jira.

33:24

Try Loom for free today  at Loom. com/Lenny. That's L-O-O-M. com/Lenny.

33:34

Something that I want to help people understand,  and I don't even 100% understand this, is what's the simplest way to understand the job  of a researcher versus say a model designer and other folks involved?

33:43

What's the simplest way  to understand what researchers do at OpenAI?

33:48

So the project that I described are mostly  product-oriented.

33:48

Research is mostly product research.

33:52

Another component of my team is actually  more longer term exploratory projects.

33:52

And it's more about developing new methods, understanding  those methods under a variety of circumstances.

33:59

So basically developing methods, you need to follow  very similar recipe of building evals but it's much more sophisticated evals.

34:17

You want to have  outer distribution or if you want to measure generalization, you need to capture that.

34:22

But it is basically more sciencey in a way where...

34:29

If we talk about synthetic data, one of  the hardest things about synthetic data is how do you make it more diverse?

34:36

Diversity in synthetic  data is one of the most important questions right now.

34:41

And so it's like exploring ways to inject  diversity as a general method that will work for all is one of the research explorations.

34:48

Other  ones is more developing new capabilities.

34:48

I feel like it's always about you work on this new method  and you have signs of life that it's working, either you think of how do you make it more  general or you think of how do you make it very useful?

35:09

And this is how the longer-term  projects become more medium, short-term project. That makes sense.

35:15

Essentially working on  developing ways to make the model smarter, o4, o5, o6. New ways to...

35:20

o1  was a big breakthrough, right? Yeah.

35:25

The way it operates where it's not  just, "Here's your answer," it actually thinks and takes time to think through the  process of coming up with an answer. Okay. Yeah. Very helpful.

35:34

Speaking of that, of thinking  about the future, where things are going, I want to spend some time on just this insight that  basically you are building the cutting edge of AI, at the very bleeding edge of where AI is going  and where it is.

35:45

And so I'm very curious to hear just your take on how you think things are going  to change in the world and how people work based on where you see things are going.

35:58

And I know  it's a broad question, but let's say in the next three years, how do you see the world changing?

36:02

How do you see people's way of working changing?

36:08

It's a very humbling experience to be in  both labs, I guess.

36:08

To me when I first came to Anthropic and I was like, "Oh no, I  really love front-end engineering."

36:12

And then the reason why I switched to research is because  I realized at that time it's like, "Oh my god, Claude is getting better at front-end.

36:22

Claude  is getting better at coding.

36:22

I think Claude can develop new apps or something and so it  can develop new features for the thing that I'm working."

36:33

So it was kind of like this meta  realization where it's like, "Oh my god, the world is actually changing."

36:40

And when we first launched  100K context at that time, obviously I'm thinking about form factors that's like file uploads were  very natural, very familiar to people.

36:48

But you can imagine we could just make infinite chats  in the Claude.

36:55

ai app, as if it's 100K context.

37:04

But because file uploads...

37:04

It's like form  follows function.

37:04

It's like the form factor, the file uploads can enable people to just  literally upload anything, the books, any reports, financial and ask any task to the model.

37:18

And then  I remember it was either enterprise customers, financial customers were really interested  in that. It's like, "Oh wow."

37:27

It's actually one of the very common tasks that people do  in that setting.

37:33

It's kind of crazy to see how some of the redundant tasks are getting  automated basically by these smart models.

37:48

And they're entering the era where, I actually  don't know for example sometimes if o1 gives me the correct answer or not because I'm not  an expert in that field.

37:55

And it's like, "I don't even know how to verify the outputs  of the models."

38:01

It's because all my experts know they can verify this.

38:07

So, yes, so  basically there are trends that are going on.

38:14

The first trend is the cost of reasoning  and intelligence is drastically going down.

38:22

I had a blog post about this.

38:22

Maybe  I should update on latest benchmarks, because at that time everybody was doing one  benchmark and they'd be...

38:27

quickly saturated the benchmarks.

38:34

So I'm like, "Now we need to do  the same plot but with another frontier eval."

38:34

But the cost of intelligence is going down because  it becomes that much cheaper.

38:41

Small models are becoming even smarter than large models and  that's because of the distillation research.

38:56

This happened with Claude 3 Haiku.

38:56

I was working  with the training on the Claude 3 Haiku and I realized it was much smarter than Claude 2, which  was way bigger, lots [inaudible 00:39:08].

39:02

But the power of small models become very intelligent  and fast and cheap.

39:10

We are moving towards that world.

39:16

That has multiple implications,  but the news is that people will have more access AI and that's really good.

39:23

Builders and  developers will have much better access to AI, but also it means all the work that has been  bottlenecked by intelligence will be unblocked.

39:40

I'm thinking about healthcare, right?

39:40

Instead  of going to a doctor, I can ask ChatGPT or give ChatGPT a list of symptoms and ask me, "Would  I have a cold, flu, something else?"

39:47

I can literally get the access to doctor almost.

39:58

And  there's been some research studies around that.

40:05

There was a New York Times story about that  where they compared doctors to doctors using ChatGPT to just ChatGPT and just ChatGPT was  the best of them.

40:10

All doctors made it worse. Yeah, that's crazy. Yeah.

40:18

Yeah, that's crazy,  right?

40:18

Education I think I would have dreamt if I had the tool like ChatGPT when I was  young and would learn so much.

40:25

But it's like people can now learn almost anything from  these models.

40:30

So they can learn new language, they can learn how to build new look apps and  write anything they do want. It's humbling to have...

40:46

launch Canvas and bring that thing to  the people, enable them to do something else that they couldn't have ever before.

40:52

There's  something magical around this experience.

40:57

Education will have massive implications.

40:57

I guess  like scientific research, I think it's the dream of any AI research is to automate AI research.

41:03

It's kind of scary, I'd say, which makes me think that people management will stay.

41:11

It's one of  the hardest thing to...

41:11

Emotional intelligence with the models, creativity in itself is one of  the hardest things.

41:18

So writers, I don't think people should be worried as much.

41:26

I think will  alleviate a lot of redundant tasks for people. This is awesome.

41:34

Okay, I want to follow this  thread for sure.

41:34

And it's funny that what you described as you were an engineer at Anthropic  and you're like, "Okay, Claude is going to be very good at engineering.

41:42

This isn't going to  be a potentially career long term, so I'm going to move into research and AI is going to need me  for a long time to build it, to make it smarter."

41:53

I would say we still have...

41:53

I think  Canvas team has still have really cool front engineers that are really people  who really care about interaction, design, interacting experience.

42:05

I don't  think models are there yet I think if...

42:10

But we can get the models to this top  1% of front-ends and things for sure.

42:15

So what I want to move on to next along these  lines is just, and this is just speculation, but what skills do you think will  be most valuable going forward for product teams in particular?

42:26

So folks  are listening and they're like, "Okay, this is scary.

42:29

What should I be building now  to help me stay ahead and not be in trouble down the road?"

42:37

What skills do you think are  going to be more and more important to build?

42:42

Yeah, I think creative thinking.

42:42

You want to  generate a bunch of ideas and filter through them and not just build the best product experience. Listening.

42:52

You want to build something that the most general model will not replace you.

42:59

And  oftentimes you build something and you make it really, really good for specific set of users and  actually the mode is now in your user feedback.

43:17

The mode is more in whether you listen to them,  whether you can rapidly iterate. The mode is in here.

43:26

I don't think we are yet to...

43:26

There are  so many ideas, I think there's an abundance of ideas that you can work on. I wouldn't be  worried.

43:32

I feel like in fact I just think people in AI field are like...

43:37

I wish they were a  little bit more creative and connecting the dots across the print fields or something like  that to develop really cool new generation and new paradigms of interactions with this AI.

43:50

I don't think we've cracked this problem at all.

43:56

A couple of years ago I was telling some people,  I was like, "You want to build for the future."

44:03

So it's like it doesn't necessarily matter  whether the model is good or not, good right now, but you can build product ideas such that  by the time the models will be really good, it'll work really well.

44:17

I think it just happened  naturally.

44:17

For example, at Anthropic the Claude artifacts...

44:26

And I feel early days of  Canvas was, back in 2022 before ChatGPT, writing ideas was our knowledge [inaudible  00:44:36].

44:33

But I feel like Claude 1.

44:33

3 model itself was not there to have made really extreme  good high quality edits.

44:38

For example, like coding.

44:47

And I feel like I see startups like Kaeser  was doing super well.

44:47

And that's because they iterate so fast.

44:53

They invent new ways  of training models. They move really fast.

45:01

They listen to what users like, massive  distributions. Yeah, it's kind of cool.

45:08

That's really helpful actually.

45:08

So what I'm  hearing is that soft skills essentially are going to be more and more important, powerful.

45:12

You just talked about management, leading people, being creative and coming up with innovative  insights, listening.

45:16

There's a post I wrote that I'll link to where I try to analyze  how AI will impact product management.

45:22

And we're actually very aligned, and my sense was the  same thing, that soft skills are going to become more and more important.

45:32

And the things that  are going to be replaced is the hard skills, which is interesting because usually people  value the hard skills like coding, design, writing really well.

45:41

And it's interesting that  AI is actually really good at that because it's taking a bunch of data, synthesizing it and  writing, creating a thing, versus all these fuzzy things around of what influences, convinces  people to do things and aligning and listening, like you said, creativity, anything  along those lines come up as I say that.

46:01

I think it's actually a really, really hard to  teach the model how to be aesthetic or do really good visual design or how to be extremely creative  in the way they write.

46:08

I still think ChatGPT kind of sucks at writing and that's because it's  bottlenecked by this creative reasoning.

46:16

I think characterization is one of the most  important...

46:22

I think for a manager, I feel like...

46:28

Actually, AI research progress is bottlenecked by  management, research management.

46:28

It's because you have constrained set of compute and you need to  allocate the compute to the research paths that you feel the most convinced about.

46:42

It was like  you need to have a really high conviction in the research paths to put the compute, and it's more  return on investment kind of situation.

46:50

It's like, "Okay, I'm thinking a lot about across all my  projects, which projects are higher priority?"

47:04

Prioritization and also on the lower level,  "Which experiments are really important to run right now and which are not?"

47:09

and cut through  the line.

47:09

So I was thinking prioritization, communication, management.

47:14

People skills like  empathy, understanding people, collaboration.

47:23

I think Canvas wouldn't be an amazing launch  if it wasn't about people and I think it's a wonderful group of people.

47:31

And I get a chance  to work with people like Lee Byron who's a co-creator at GraphQL and some of the best Apple  designers. It's so cool to see...

47:36

and how do you create this collaboration between people.

47:45

It's  just something that's still humane, I think.

47:50

Let me just follow through a little bit.

47:50

I  imagine people listening are like, "Okay, but once we have AGI or SGI it's like it'll  do all this."

47:54

There's a world where like, "Why isn't all this done?"

47:59

I think it's  easy to just assume all that.

47:59

I'm curious this idea of creativity and listening,  why you think AI isn't good at it, other than it's just very hard to train it to  do this well.

48:10

Is there anything there of just why this is especially difficult  for AI and LLMs to get good at?

48:20

I think currently it's difficult  for many reasons.

48:20

I think it's still an active research area and it's something that  I think my team is working on.

48:26

It's like, "Okay, how do we teach the models to be more creative  in the writing?"

48:32

And so I'm thinking this new paradigm of wise that the models think more  should actually lead to better writing in itself.

48:45

But when it comes down to idea generation  or discriminating of what is a good visual design or not, I feel like it hasn't had learned  examples from people to discriminate it very well.

49:02

I do think it's because there are not that  many people who are actually really...

49:02

It's not accessible to models to learn from these people I  guess.

49:12

So I definitely think that's why it sucks. Yeah, that makes sense.

49:19

Basically  there's not enough of you yet, researchers teaching it to do these  things, slash people that have incredible taste and creativity that can teach  these things.

49:26

You could argue this will come. Right.

49:31

But we don't need to keep going down that thread.

49:31

Let me ask you a specific question.

49:31

In this post I wrote, I made this argument that a lot of people  disagreed with that strategy is something that AI tooling will become increasingly great at and  take over.

49:42

There's the sense that that's the thing that people will continue to be much better at and  you can't offload to AI basically developing your strategy, telling you what to do to win.

49:53

My case  is, "Isn't strategy, just take all the inputs, all the data you have available, understand  the world around you and come up with a plan to win?"

50:03

It feels like AI and LLM would be  incredibly smart at this. What's your take? I think so too.

50:09

I think again, you teach the model  all sorts of tools and capabilities and reasoning and it's like when it comes down to...

50:17

For Canvas  right now, it would be very cool for the model just aggregate all the feedback from users,  summarize me the top five most painful flows on user experiences.

50:30

And then the model itself  is very capable of thinking of knowing how it's been made, figure out how to create a dataset for  itself to train on it.

50:38

And I don't think that we are far away from that self-improvement,  models becoming self-improved by...

50:55

That, and the part of development, is basically  self-improving.

50:55

It's kind of like its own organism or something.

51:00

Again, like strategies, it's more  like data analysis and coming up with...

51:00

I think what models are really good at is connecting the  dots, I think.

51:12

It's like if you have user feedback from this source, but you also have an internal  dashboard with metrics and then you have other feedback or input and then it can create a plan  for you, recommendations even.

51:31

And I think this is one of the most common use cases for ChatGPT  too, is coming up with these sort of things.

51:46

That makes sense essentially a human can only  comprehend so much information at once and look at so much data at once to synthesize takeaways.

51:51

And as you said, these context windows are huge now.

51:56

Here's all the information, what's  the most important thing I should do?

51:59

Yeah, same as scientific research.

51:59

Ideally  the model would be able to suggest ideas, new ideas, or iterate on the experimental  given the empirical results of the previous experiments like how do you come  up with new ideas or the methods? Yeah. Oh, man.

52:18

Okay, so just to close the loop on  this conversation, this part of the thread is the skills you're suggesting people focus on building  and leaning into is soft skills like creativity, managing influence, collaboration, looking for  patterns.

52:32

Is that generally where your mind is at?

52:40

Yeah, I'm thinking a lot about how do we  make organizations more effectively and I think this is mostly management, I guess.

52:43

It's like how do you organize research teams or generally teams combined...

52:49

Compose teams  such that they will be at their maximally succeed or at the maximal performance of what  can possibly...

52:56

We can literally create the next generation of computers.

53:04

It's just the  matter of conviction and the way you manage through that.

53:09

It's scaling organizations  or scaling product research, I guess.

53:16

Yeah, I think you're basically building  this thing and not efficiently doing it is limiting the potential of  the human species right now. Right.

53:26

It's mismanagement within the  research team in OpenAI and Anthropic and some of these other models.

53:32

Yeah, it's kind of crazy to think about it. Holy moly.

53:33

Okay, so speaking of Anthropic  and OpenAI, you've worked at both.

53:33

Very few people have worked at both companies and have  seen how they operate.

53:38

I'm curious just what you've noticed about the differences  between these two, how they operate, how they think, how they approach stuff.

53:44

What can you share along those lines?

53:48

It's more similar than different.

53:48

Obviously there  was a lot of...

53:48

There are some differences always comes to nuances. I would say culture.

53:55

I really  love Anthropic and I have a lot of friends there.

54:02

And I also love OpenAI and they still have a lot  of friends though.

54:02

So it's not about enemies.

54:02

I feel like there's in AI, it's all like, "Yeah,  they're competitors. There's enemies."

54:07

It's actually like one big community of people doing  the same thing.

54:11

I would say what I've learned from Anthropic is this real care and craft towards  model behavior, model craft, model training.

54:32

And I've been thinking a lot about, "Okay,  what makes Claude Claude and what makes ChatGPT ChatGPT?"

54:36

And it's like I still have some  sense of operational processes that leads to the outputs, to the model. It's the outputed model.

54:43

And it's like the reason why Claude has so much more personality and is more like a librarian... I  don't know. I don't know.

54:49

I am visualizing Claude being like a librarian at some point, very nerdy  or something. ...

54:59

is because I feel like it's the reflection of the creators who are making this  model.

55:07

And a lot of details around the character and the personality and whether the model  should follow up on this question or not.

55:18

What's the correct ethical behavior for the  model in these scenarios?

55:18

A lot of crafts and curated datasets.

55:26

This is where I learned  that part of art, I guess, at Anthropic.

55:35

I would say Anthropic is much smaller.

55:35

When  I joined it was, what, like 70 people?

55:35

When I left it was tons of people.

55:40

And obviously  the culture changed so much.

55:40

I really enjoyed being early days startup lives, and people knew  each other as a family. But the culture shifted.

55:53

I would say that I learned from Anthropic  that they're much better at focusing and prioritization of...

55:58

Very hardcore prioritization,  I guess. And they need to do it.

55:58

But I think OpenAI's much more innovative and much more  risk-takers in terms of product or research.

56:14

Actually, in way your full-time job can be  just teaching the model how to be creative writers.

56:21

And it's like there's some luxury in this  research freedom that comes with scale, maybe. I don't know.

56:28

I'd say I have much more creative  product freedom to do almost anything, I guess, within OpenAI, evolve ChatGPT into the vision that  we want.

56:39

It's more probably bottoms-up, I guess.

56:47

Yeah, that's how I was thinking about it.

56:47

It feels like OpenAI is more bottoms-up, distributed, people bubble up ideas, try stuff.

56:52

And that leads to more products launching, I imagine more things just kind of being  tried versus more of a, "Let's just make sure everything we do is awesome and great and  craft and thinking deeply about every investment." Right.

57:08

That's really interesting.

57:08

I've never  heard it described this way.

57:08

Karina, we've covered so much ground.

57:12

This is going  to help a lot of people with so many ways of thinking about where the future's going.

57:16

Before  we get to our very exciting lightning round, I'm curious if there's anything else that you  think might be helpful to share or get into?

57:23

One of my regrets, I guess, when I was early days  at Anthropic was that...

57:23

I think there was some luxury of the time, because pre-ChatGPT,  to actually come in with a bunch of ideas and prototype almost every day.

57:36

And I think  that we did a lot of cool ideas like Claude, and Slack was actually one of the first tool-usey  products.

57:44

It's like Claude could operate in your workplace now.

57:53

It's kind of cool because you  can add Claude to summarize the thread.

57:53

So maybe you have an entire conversation with  someone and then you want a summary of what happened you can ask Claude, "Summarize this."

58:04

Also, it was really fun to iterate on the model itself.

58:10

It's like when you just talk to the model  in Slack forever.

58:10

It created some social element, it was kind like [inaudible 00:58:19] and this  Discord, people learned so much about prompting and how to work with Claude.

58:23

Actually, one of  the features that was early tasks prototype is every Monday Claude would just summarize  the entire channel.

58:30

Or every Friday we'd just summarize a bunch of channels and give the  news about the organization, or something.

58:47

And it's kind of like really cool form factor.

58:47

I think thinking about form factor's a really important question in AI, especially we haven't  even figured out how do we create an awesome product experience with o-series models.

59:01

It's like  the paradigm between synchronous real time give an answer paradigm into more asynchronous paradigm  of agents working on the background.

59:08

But then now the question is the agents should build trust  with you, right?

59:16

And trust builds over time, which is like with humans.

59:20

And you start this  collaboration which is why this collaboration model with you and the model is so important  because you build trust and the model learns from your preferences so that it can become  more personalized and it will start predicting the next action that you want to take on  the computer or something.

59:40

And it's more predictive, much more...

59:45

We went from personal  computers to personal model basically here. Why is it not a thing?

59:54

That  seems like such an obvious feature that every LLM should have as  a Slack bot version of them.

59:56

Is that a thing I can help you install?

1:00:00

Or is that not a thing right now?

1:00:02

I know that Claude and Slack was  sunset in 2023 or something.

1:00:02

I think it was after ChatGPT was mostly the focus on  customer use cases or enterprise use cases. Mm-hmm. Bummer.

1:00:19

I think the form factor of Claude and  Slack was kind of constrained a little bit when you want to talk about new features. Bummer. I want that.

1:00:30

I know that ChatGPT had Slackbar tools.

1:00:30

I  don't know, maybe it will come back sometime.

1:00:35

All right, I would pay for that.

1:00:35

Any other  memories from that time of early days?

1:00:39

Because that's a really special place  to have been is early days Anthropic.

1:00:43

Any other memories or stories from that  time that might be interesting to share?

1:00:48

I think the very first launch when we felt...

1:00:48

When click from use, again, was 100K context launch is when the models could input the  entire book and give you a summary of the book or something. Or the financial...

1:01:02

or  catalog multi files financial reports and then give you an answer to the question, to very  specific questions.

1:01:09

I think there was something in there that was kind like, "Oh my god, this is a  really cool new capability."

1:01:16

Not model capability, but more like the capabilities that  came from the product form factor itself rather than the model capability as much.

1:01:28

I think other prototypes that we were thinking about...

1:01:38

There's one part having a Claude  workspaces and it's kind of the same idea of Claude and I would have this shared workspace  and that share workspace is like a document and we can iterate on the document.

1:01:51

And I feel  like sometimes the ideas, [inaudible 01:01:55] and they're locked for two  years, just like in this case.

1:02:00

It's interesting, there's these milestones  that kind of open up our view of what is happening and where things are going.

1:02:04

ChatGPT think was the first of just like, "Wow, this is much better than I would've  thought."

1:02:08

You talked about 100K context windows where you could upload a book and  ask it questions and have it summarize.

1:02:14

I actually use that all the time.

1:02:17

When I have  interview guests and they wrote a book, I sometimes don't have time to read the whole  book.

1:02:20

So I use it to help me understand what the most interesting parts are.

1:02:23

And then I actually  dive into the book, just to be clear.

1:02:23

And then, I don't know, maybe voice was another one where  you could talk to say ChatGPT.

1:02:29

Is there any other moments there that you're like, "Wow, this is  much better than I thought it was going to be?"

1:02:40

Yeah, I think the computer use agents,  like the model operating the desktop.

1:02:40

And you can essentially think of new kind of  experience where the model can learn the way you browse.

1:02:57

And from that preference it  can just browse as just like you.

1:02:57

It's kind of simulated persona.

1:03:04

And it's actually  very similar to the idea of like, "Okay, maybe Sam Altman doesn't have a lot of time.

1:03:11

Maybe  I want to talk to his simulation and ask..."

1:03:11

Or, for example, I really appreciate some  of the technical mentorship. Yeah, cool.

1:03:27

But he doesn't have a lot of time  so it's like I really want to ask him this questions.

1:03:30

How do you respond with simulated  environments like this would be really cool.

1:03:37

That's a great place to plug Lennybot, have one of those.

1:03:38

It's trained on  all of my podcasts and newsletters. Oh, cool. It sits on many models.

1:03:43

I don't  know which exactly they use, but it's exactly that.

1:03:46

And it's not even  me, it's all the guests that have been on the podcast and on newsletter as  I wrote.

1:03:51

And you could just ask it, "How do I grow my product?

1:03:53

How do I develop a  strategy?"

1:03:53

And it's actually shockingly good.

1:03:58

Do you feel like it reflects who you are? Yeah. Or would it be... Okay.

1:04:01

The best part of it is you can talk  to it.

1:04:01

There's an ElevenLabs voice version that's trained on  my voice from this podcast, and it's actually very good and people have  told me they sit there for hours talking to it. Wow.

1:04:15

And somebody told it, "Interview  me like I am on Lenny's podcast, ask me questions about my career."

1:04:21

And he did  a half hour podcast episode with Lennybot.

1:04:25

Oh my god, that's so fun. It's incredible. Future is wild. Yeah.

1:04:29

I think content transformation is...

1:04:29

I would  imagine sometime when you generate a sci-fi story in Canvas, you can transform this into audiobook  where you have very natural content transformation of one media to another media.

1:04:47

I think one of my  earliest inspiration is one of the last episodes of Westworld where, I don't want to spoil, but  where Dolores comes to her work at that time and she comes to this new workspace and she starts  writing a story.

1:05:04

And then as she writes a story, a 3D, virtual reality, starts creating on the fly.

1:05:13

So I kind of want to create that. Kind of cool. Wow.

1:05:23

Speaking of medium, I guess I was wondering if I should go in  this direct or not, but real quick.

1:05:26

Kevin Weill/Kevin Weill, I don't know exactly how  to pronounce his last name, the CPO of OpenAI. Kevin Weill, uh-huh. Is it Weill or Weill? I think Weill. Weill. Okay. Okay. Let's just  say that. We'll go with that. I hope, yeah.

1:05:43

He did a panel at the Lenny and Friends Summit  last year and he made this really fascinating point that chat is a really interesting interface  for these tools because they're just getting smarter and smarter and smarter and smarter  and smarter.

1:05:54

And chat continues to work as a paradigm to just interact with them, similar to  a human.

1:05:57

You could talk to Albert Einstein.

1:05:57

You could talk to someone not very smart and it's  all conversation still.

1:06:02

And so it's a really flexible way to interact with increasingly good  intelligence.

1:06:06

At some point it'll not be so great, and you were talking about all these ways that  you're adding additional ways to interact.

1:06:13

But it's interesting chat proved to be a really  powerful layer on top of all this stuff. Yeah, that's real cool.

1:06:22

I feel like chat also has  social element which is very humane.

1:06:22

It's like, yeah, you sometimes want to get into group chat.

1:06:28

And having conversations with AI is kind of like a group chat in itself, as messaging.

1:06:32

Actually,  this idea of how do you build features like this?

1:06:40

I see tasks as this general feature that will  scale very nicely as the models would develop new capabilities themselves.

1:06:50

The models will be  able to do better searches and create new...

1:06:50

come up with more creative writing on render, react  apps and like HTML apps.

1:06:58

And you can have everyday new puzzle for you, every day continue the story  from the previous days. It scales very nicely.

1:07:13

You mentioned something as we were getting into  this extra section that we ended up going down is this idea of the agents using a computer.

1:07:18

I  know this is actually something you are going to launch today, the day we're recording it,  which will be out by the time this comes out, called Operator, can you talk about this very  cool feature that people will have access to?

1:07:33

Yeah, so I unfortunately did not work on  that, but I'm really, really excited about this launch.

1:07:39

It's basically an agent that can  complete the task in its own virtual computer, in its own virtual environment.

1:07:49

You can do any  literally task like order me a book on Amazon.

1:07:56

And then ideally the model will either  follow up with you which book do you want, or know you so well that it start recommending,  "Oh, here is the five books that I might recommend you to buy."

1:08:07

And then you hit, "Yeah, help  me buy."

1:08:07

And then the model goes off into its own virtual little browser and complete the  task and buy the book on the Amazon.

1:08:16

And then if you give the model credentials, credit cards,  obviously it comes with a lot of trust and safety, then it will just complete the thing  for you.

1:08:30

It's a virtual assistant.

1:08:37

It's interesting how this just sounds  like obviously this should happen.

1:08:37

Why is this not yet a thing?

1:08:39

Which is also  mind-blowing that we're just assuming this should exist.

1:08:44

Just some AI doing things  for you on a computer we just ask it to do. Yeah. It's absurd.

1:08:51

It's actually really hard.

1:08:51

And I  think you're still cracking this, but feel like...

1:08:58

I don't know if you use  Tuple like a pair programming product. No.

1:09:04

But at Anthropic we loved pair  programming, so if you used- Oh yeah, Shopify uses this.

1:09:07

I remember  it came up on a podcast episode. Oh, nice.

1:09:10

Yeah, so it is a very cool product  where you can just call anyone at any time and then share screen and the other person  can have access to the screen or start literally operating your computer.

1:09:21

And it's  very realtime...

1:09:21

The allegiance is very... it's very high quality.

1:09:30

And it's just like I  kind of want the same.

1:09:30

I want to pair program with my model and the model should even talk to  me.

1:09:37

Draw very specific section in my code and just go to tell me...

1:09:45

Obviously teach me and  we can have different modes.

1:09:45

It's like right, this is a product right here for you. I  don't know.

1:09:49

Some people should build that.

1:09:57

It sounds like a startup just got birthed- Yes. ...

1:10:00

from someone listening to this.

1:10:00

You mentioned that it's very hard to do this agent controlling a computer  as you and helping out.

1:10:02

What makes it so hard for whatever, however  much you can explain briefly?

1:10:11

Much of it is because right now the model's  operating on pixels instead of language or whatnot.

1:10:21

Pixels is actually really, really  hard.

1:10:21

The models [inaudible 01:10:25] perception, or visual perception.

1:10:24

I think  there's still a lot of multimodal research that's going on, but I think language scaled so much  easier compared to multimodal because of that.

1:10:38

Another thing that I guess my team is working  that is how do you derive human intent very correctly?

1:10:46

It's like sometimes does the model  know enough information to ask a follow-up question or to complete the task?

1:10:52

You don't  want an agent to go off for 10 minutes and then come back with an answer that you didn't  even want.

1:10:58

That actually creates much more worse user experience.

1:11:05

And this comes with  teaching the model people skills.

1:11:05

It's like, "What do people like?

1:11:14

Kind of like creating  the mental model of the user and care about the user in order to ask certain questions.

1:11:21

Actually, that part is hard to do for the models.

1:11:28

That relates to what we talked about  earlier where this kind of the soft skill, people skills piece is not where  these models are strong yet. Yeah. Okay.

1:11:35

I'm going to skip the lightning round.

1:11:35

I want to ask just one question from  the lightning round, something fun. Yeah.

1:11:43

Okay, so when AI replaces your job, Karina, I'm  curious what you're...

1:11:43

And it gives you a stipend, gives you a monthly stipend.

1:11:48

Here's your  salary for the month.

1:11:48

What would you want to do?

1:11:53

What do you want to spend your time on?

1:11:53

What will you be doing in this future world?

1:11:57

I've been thinking about this a lot times.

1:11:57

I  feel like I have a lot of jobs options.

1:11:57

I would love to be a writer, I think.

1:12:05

I think that would  be super cool.

1:12:05

You should write short stories, sci-fi stories, novels.

1:12:11

I really like art  history, so you know those conservationists in the museums who just try to preserve art  paintings, but just painting through a long day? Mm-hmm.

1:12:29

I think that would be really cool to do. Yeah. That sounds beautiful. I don't know.

1:12:39

What I'm hearing is you need to Nerf these models  to not get very good at writing so that you can continue...

1:12:44

Although at that point you don't need  to do it from...

1:12:44

You don't need people to buy it, you're just doing it for fun, so it doesn't even  matter if they're incredibly good at writing or art conservation.

1:12:51

Oh man, what an episode of  our conversation.

1:12:51

What a wild time we're living in.

1:12:57

Karina, thank you so much for being here. Two final questions.

1:12:57

Where can folks find you online if they want to reach out and follow up on  anything?

1:13:01

And how can listeners be useful to you?

1:13:06

You can find me, I'm on Twitter it's KarinaNguyen.

1:13:06

You can also shoot me an email on my website.

1:13:06

And my team is hiring and so I'm looking for  research engineers, research scientists, as well as machine learning engineers, people  who come from product engineers who want to learn model training.

1:13:26

I'm actually hiring for my  team.

1:13:26

My team is called Frontier Product Research, and we train models, we develop new  methods but for product oriented outcomes. What a place to work. Holy moly.

1:13:39

What's the best way for people to apply for  these very lucrative roles?

1:13:46

I think you can shoot me a DM on Twitter. Okay.

1:13:49

Or I'm yet to create a job description for them. Okay.

1:13:53

This is the job description.

1:13:55

Or you can apply into post training team. Yeah. Okay.

1:13:58

You're going to get a flood of  DMs. I hope you're prepared.

1:13:58

Karina, thank you so much for being  here. This was incredible.

1:14:03

Thank you so much, Lenny. Bye, everyone. It was fun.

1:14:08

Thank you so much for listening.

1:14:08

If you found  this valuable, you can subscribe to the show on Apple Podcasts, Spotify, or your favorite  podcast app.

1:14:12

Also, please consider giving us a rating or leaving a review as that really helps  other listeners find the podcast.

1:14:18

You can find all past episodes or learn more about the show at  LennysPodcasts. com.

1:14:23

See you in the next episode.