Al Engineering 101 with Chip Huyen (Nvidia, Stanford, Netflix)

0:00

One question that get asked a lot and a lot is how do we keep up to date with the latest AI news?

0:03

Why why do you need to keep up to date with the latest AI news?

0:07

If you talk to the users and understand what they want, what they don't want, look into the feedback, then you can actually improve the application way way way more.

0:14

>> A lot of companies are building AI products.

0:16

A lot of companies are not having a good time building AI products.

0:19

>> We are in an ideal crisis now.

0:19

We have all this really cool tools.

0:21

You have do everything from scratch. It have your design.

0:24

It can have your right code. You have your website.

0:26

So in theory, we should see a lot more.

0:28

But at the same time, it more or less somehow stuck.

0:31

They don't know what to build.

0:33

>> All this AI hype, the data is actually showing most companies try it, doesn't do a lot, they stop.

0:36

What do you think is the gap here?

0:38

>> It's really hard to measure productivity.

0:39

So I do ask people to ask their managers, would you rather have give everyone on the team very expensive coding agent subscriptions or you get an extra headcount.

0:49

Almost everyone managers would say headcount.

0:51

But if you ask VP level or someone who manage a lot of teams they would say one AI assistant because as managers you are still growing.

0:59

So for you having one HR head is big whereas for executive maybe you have more business metrics that you you care about.

1:05

So you actually think about what actually drive productivity metrics for you.

1:11

>> Today my guest is Chip Hen.

1:11

Unlike a lot of people who share insights into building great AI products and where things are heading.

1:17

Chip has built multiple successful AI products, platforms, tools.

1:21

Chip was a core developer on NVIDIA's Nemo platform, an AI researcher at Netflix.

1:26

She taught machine learning at Stanford.

1:28

She's also a two-time founder and the author of two of the most popular books in the world of AI, including her most recent book called AI Engineering, which has been the most read book on the O'Reilly platform since its launch.

1:40

She's also gotten to work with a lot of enterprises on their AI strategies and so she gets to see what's actually happening on the ground inside a lot of different companies.

1:49

In our conversation, Chip explains a lot of the basics like what exactly does pre-training and post-training look like? What is RAG?

1:57

What is reinforcement learning? What is RHF?

1:58

We also get into everything she's learned about how to build great AI products, including what people think it takes and what it actually takes.

2:04

We talk about the most common pitfalls that companies run into, where she's seeing the most productivity gains, and so much more.

2:11

This episode is quite technical, more technical than most conversations I've had, and is meant for anyone looking for a more in-depth conversation about AI.

2:18

If you enjoy this podcast, don't forget to subscribe and follow it in your favorite podcasting app or YouTube.

2:24

And if you become an annual subscriber of my newsletter, you get a year free of 16 incredible products, including Devon, Lovable, Replet, Bolt, NAN, Linear, Superhum, Dcript, Whisper Flow, Gamma, Perplexity, Warp, Granola, Magic Patterns, Rickcast, JPRD, and Mobin. Head on over to Lenny's.

2:38

com and click product pass.

2:40

With that, I bring you Chip when after a short word from our sponsors.

2:45

This episode is brought to you by Dcout.

2:48

Design teams today are expected to move fast, but also to get it right.

2:53

That's where Dout comes in.

2:56

Dout is the all-in-one research platform built for modern product and design teams.

3:00

Whether you're running usability tests, interviews, surveys, or in the wild fieldwork, Dout makes it easy to connect with real users and get real insights fast.

3:08

You can even test your Figma prototypes directly inside the platform.

3:13

No juggling tools, no chasing ghost participants.

3:16

And with the industry's most trusted panel, plus AI powered analysis, your team gets clarity and confidence to build better without slowing down.

3:25

So if you're ready to streamline your research, speed up decisions, and design with impact. Head to dscout. com to learn more. That's dsc. com.

3:36

The answers you need to move confidently.

3:37

Did you know that I have a whole team that helps me with my podcast and with my newsletter?

3:42

I want everyone on that team to be super happy and thrive in their roles.

3:46

Just Works knows that your employees are more than just your employees. They're your people.

3:49

My team is spread out across Colorado, Australia, Nepal, West Africa, and San Francisco.

3:56

My life would be so incredibly complicated to hire people internationally, to pay people on time and in their local currencies, and to answer their HR questions 24/7.

4:04

But with Just Works, it's super easy.

4:06

Whether you're setting up your own automated payroll, offering premium benefits, or hiring internationally, JustWorks offers simple software and 24/7 human support from small business experts for you and your people.

4:19

They do your human resources right so that you can do right by your people.

4:23

Just works for your people.

4:30

>> Chip, thank you so much for being here and welcome to the podcast. >> Hi, Lenny.

4:34

I've been a big fan of a podcast for a while.

4:36

So, I'm really excited to be here. Thank you for having me.

4:39

>> I want to start with this table slashchart that you shared on LinkedIn a while ago that went super viral.

4:44

And I think it went super viral because it hit a nerve with a lot of people.

4:47

And let me just read this and we'll show this on YouTube for people that are watching.

4:52

So, it's this very simple table you shared of what people think will improve AI apps and what actually improves AI apps.

4:58

what people think will improve AI apps, staying up to date with the latest AI news, adopting the newest agentic framework, agonizing over vector databases to use, constantly evaluating what model is smarter, fine-tuning a model, and then you have what actually improves AI apps, talking to users, building more reliable platforms, preparing better data, optimizing end-to-end workflows, writing better prompts.

5:21

Why do you think this such a nerve with people?

5:23

And just what if you had to boil it down, what do you think is what do you think people are missing about building successful AI apps?

5:28

>> One one question that get asked a lot and a lot is that how do we keep up to date with the latest AI news?

5:33

And I'm like why why do you need to keep up to date with the latest AI news?

5:38

I know it sound very counter counter intuitive but there's so much news out there.

5:42

A lot of people also ask me questions like how do I choose between two different technologies like maybe like recently like MCP versus like Asian Asians, right?

5:53

right? like protocol and it was like which one is better or like this or that and a ser question I used to ask them is like first like if how much of the improvements could you get like from like optimal solutions versus nonoptimal solutions right and sometimes they were like actually it's not much right and I

6:10

was like okay if it's not much improvement why do you want to spend so much time debating something that doesn't uh make that much difference to your performance and another question they asked is like if you adopt a new technology like how hard it would be to switch that out to another. And

6:25

And sometimes it were like, oh, I think it would be like a lot of work switching it out.

6:31

And I was just like, hm, let's say here's a new technology.

6:33

It hasn't been tested by a lot of people.

6:35

And if you adopt it, you would be like stuck with it forever.

6:38

Like, do you actually want to adopt it, right?

6:40

And maybe you want to think twice about about like overcommit to like a new technology that hasn't been battle tested.

6:47

hasn't been battle tested. I love your just broader advice is just simple like talk to to build successful apps talk to users build better data uh write better prompts optimize the user experience versus just like what is the latest and greatest what's the best model to use

7:03

right now what's happening in AI let me follow this thread of this idea of fine-tuning and basically post-training there's all these terms that people hear in AI and I think this is going to be a really good opportunity for people to learn what we're actually talking about since you actually do these things. You You build these things.

7:19

You work with companies doing these things.

7:20

And um there's a few terms I want to sprinkle in through the conversation, but let's start with this one.

7:23

What what's the simplest way for someone to understand what is the difference between pre-training and post- training and then just how fine-tuning fits into that?

7:32

Just what fine-tuning actually is.

7:34

>> Disclaimer, I don't have like full visibility into like on what like this big secretive like frontier labs are doing.

7:40

Uh but right from what I heard, right?

7:42

right? So, so I think it's like one is um like supervised fetuning when you have demonstration data and you have like a bunch of like um experts like okay here's a prop right and here is what the answer should be like and like you you you just train it like on like to like stim like uh simulate um like emulate what the human expert would be

8:01

like and that's also like what a lot of people would like um the the open source models are doing they do it by distillation so instead of having human experts should like write really good starting a great answers to like prompts that get like very popular famous good models to like generate a respond to it and like getting this train smaller model to emulate. So, so sometimes you

8:22

So, so sometimes you see people just like so that's because like some I I really appreciate open source community by the way but like going from like have been able to train a models that can emulate a existing good model is very different from like being able to train a good models like an outperform existing good model.

8:39

So it's a big step there.

8:41

Uh so yeah so like we have supervised fetuning and another things that's like very big uh I'm not sure you have guessed talking about it already but like reinforcement learning is like everywhere.

8:50

Okay, let's pause on that because I definitely want to spend time on that and that's such a cool topic that I that's emerging more and more in my conversations.

8:57

But just to even summarize the things you just shared which I think is really really important stuff.

9:02

important stuff. So the idea here is a model is essentially this algorithm piece of code that someone writes and say the frontier models are feeding it just like the entire internet of content and basically it's trying to test itself on predicting in all of the in across all that data the next word essentially

9:19

token is a simpler way isn't the correct way to think about it but a simpler way to think about it is like the next word in a in text and as it gets it wrong it adjusts these things called weights essentially uh just like is that a simple way to think about it even though that's that even that's just like very surface level. So I think of language

9:34

So I think of language modeling as a way of encoding statistical information about a language, right?

9:40

So so let's say that um we we both speak English.

9:43

we we both speak English. So we kind get a sense of like what is more statistically likely like if I say my favorite color is then you would like okay that should be another color like the word blue would be much more likely to appear than the word like uh of table

9:58

right because statistically blue is more likely to come back to every color is so so it's just like get is um it's is it's a way of encoding statical information so like when language modeling when you train a large amount of data like it you see a lot of languages a lot of domains

10:13

so it can tell like okay you bas say this standards then if user do the prompt then it would come like with the next uh most likely token uh so by the way it's not a new idea actually every so it's the idea comes very very old like from the 1951 papers um um the like English entropy I think it's like called

10:32

Shannon it's a great paper and I think every a story I really like um it's from did you read Sherlock Holm by the way >> uh yeah I read a few Sherlock Holmes books yeah >> yeah so so this is story of like when Sherlock Holmes was using this statical information to like have sewn a case. So

10:44

So he was getting um so this in this story uh there is uh somebody left message uh with a lot of like stick figures.

10:52

with a lot of like stick figures. So Shhol was like okay he knows that in English the most common letter is E then the most common stick figure must be E right and then he goes he start like that it was just s uh so the uh the code so I think that's language so in a way

11:11

that's like simple language modeling right but instead of like at a word level he does it at like tok like character level and token is something in between right token is not quite a word uh but it's bigger than a character so let's say uh we we say token because uh it helps us like read what help us

11:27

reduce vocabulary because with character is like smallest amount of like vocabulary right so the five has like 26 character but words can have like millions and millions right uh whereas um um tokens you can like be able to like get like the sweet spot between the two so let's say that we

11:44

have like uh the new word like uh uh how to say like podcasting right let's say it's a it's a new word but it can divide podcast and ink So people understand okay podcast we know the meaning we know that ink is like uh like a verb like girant whatever it is so we we know the word like podcasting. So that's why the

12:02

So that's why the token comes in but yeah uh that's like the pre-tuning is basically like encoding statistical informations of language to have you predict what is most likely um I think like most likely is the most simple way of doing it.

12:16

is the most simple way of doing it. uh because it's more like building a distributions of like okay so the next token could be like more like like 90% of the time it could be like a color 90 like 10% of the time could be like something else right so based on distributions the language could like

12:30

pick like depending on your sampling strategy like do you want it to always pick the most likely token or you wanted to pick something more creative you know so so so I think sampling strategy I think is something extremely important can have you boost up performance in a in a huge way and very very underrated. >> Okay, awesome. So essentially uh a model >> Okay, awesome.

12:49

So essentially uh a model is just code with this whole uh set of weights essentially the statistical model that has learned to predict what comes next after certain words and phrases. >> Yeah.

13:03

>> And then post-raining and fine-tuning specifically is doing that same thing.

13:08

So pre-training you get like GPT5 fine-tuning is someone taking GPT5 and doing the same sort of thing.

13:14

uh adjusting these awaits a little bit for specific use cases on data that they uh find is necessary to do their very specific use case.

13:22

Is that a simple way to think about it?

13:24

>> Yeah, I think you ws as like functions, right?

13:27

So let's say just like you you have uh maybe has a functions of like maybe Lenny's height is maybe like 1x like 1x plus something or like 2x like one and plus something is is a waist, right?

13:40

So you change it until you fit the uh the the correct data which is like my height and your height, right?

13:46

So so you can think of the weight is just like a way like they function.

13:48

So so you like chain adjust the weight so they can fit the data which is a training data. >> Awesome. Okay.

13:54

So so we're talking about pre-training, post- training, finetuning.

13:58

Is there anything else here that's important to share about just like what this is exactly?

14:00

What people need to understand about these parts of training?

14:05

So the vast majority of time we don't touch on like pre-union model like as users we don't use already done for us. >> Yeah.

14:13

So so I think actually it's bit of fun like uh process like when my friends training models like I try to play with their pre-shing model and they're horrendous.

14:20

They like saying things like oh my god this is like yeah it's crazy.

14:25

Um so so it's it's really interesting to look at like how much of like post training can change the motor behavior.

14:32

Um yeah and I think that's where like a lot of time that a lot of people are spending energy on nowadays in front lab is on like post training because uh pre-training uh I think um so preing have been used to like increase the general c capacity of of of a model capabilities of a model and it depends on a lot of data and like model size to increase um to increase the model capabilities and

14:58

at some point we are actually like have max out like in data right and then people like texted that pretty much I think a lot of people are doing like with other data like audios and videos and everyone's trying to think of like what is the new source of data but where like post trading but like of course like this more of like everyone can have very similar pre-training data is like post trading is where they make a big

15:19

difference nowadays >> this is a good segue to you talked about supervised learning versus unsupervised learning I love we're getting into this by the way this is super interesting so you talked about labeled data basically supervised learning is AI learning on data that somebody has already labeled and told it here's correct versus incorrect for example this is spam versus not spam this is uh a good short

15:40

story this is not a good short story we've had u the cos of a lot of these companies that do this for labs uh mercor and scale handshake uh there's micro uh there's a few others so is is that essentially what these companies are doing for labs giving them labeled data high quality data to train on >> it is in a way but I think it's more like a product of big equations. So

16:00

So there are a lot more different components than that.

16:03

So that's why I was talking about reinforcement learning.

16:06

I'm not sure if your CEOs that you interview bring up like that term.

16:10

Uh so so the idea is that um you want people to like so like let's say you have a model give the model like a prop right and it produce an output right you want to buy like you want to reinforce or encourage the model to produce an

16:25

output that is better right so so like how like now time to like how do we know that the answer is good or bad right so people realize on like um signals so one way to get like a first one good or bad is like human feedback Right? It happen

16:39

It happen we have two responses.

16:41

You can okay this one is better than the other.

16:42

Um and we do that is because like as humans uh we tend to it's very hard to give like concrete score but it's easier to do comparisons right like if you ask me okay give this song a score.

16:53

okay give this song a score. I'm not a musicians like and don't know like how hard it is like it's like yeah I don't know like what like now 10 I go six you know and like if you ask me again a month from now and I completely forgot okay maybe now seven or like four I

17:06

don't know but they need to ask me okay here are two songs and which one would you prefer to play for the birthday party I was like okay I can prefer this song so like comparison is a lot easier um so say yes so we have a humans um you have human feedback uh and then you use

17:21

this human feedback to train a reward model so like tell like which like so and then Free root model help you like okay the model now produce this response is robot model can score is this good or bad you charge some bias toward like producing better model the better responses another ways like you can

17:36

instead of using a humans you can use like AI right like a response say yes or good good or bad right or the thing is that people are very big on nowadays like verifiable rewards which is like natural um so basically they give it a math problems and then math solutions uh

17:51

like is a model output a solution is you know that okay expected response should in a 42 and it doesn't provide 42 then it's then it's wrong right it's not a good response um so so yes so like uh a lot of time people like using this um human labor like human um human laborers should like produce like m like um how

18:10

to say expert questions and like the expected answers and in a way that like design a system that like verifiable so that the the models can can be trained on yeah >> okay I'm really glad you went there this is essentially RHF reinforcement learning with human feedback which is

18:26

exactly what I wanted to also talk about right >> yeah so um I think it's like it's general it's like it's a way of learning it's like training is contextual learning and whether it learned from human feedback or like AI feedback or like verifiable rewards uh I think I say you say just different way of like um

18:43

collei signals >> awesome yeah that's uh I we had the c of anthropic on the podcast and he talked about their version of RHF which is AIdriven reinforcement learning I love the way you phrased it where you basically you want to help the model. You want to reinforce correct behavior

18:55

You want to reinforce correct behavior and correct answers and this is the method to do it.

18:59

Whether it's say an engineer seeing an output from a model being like no here's how I would code it differently and then training and it's training a different model that the original model works with to tell it am I correct or not correct. Is that right? Yeah. >> Yeah.

19:16

>> Yeah. I I think I think that's a way of of looking into it and I think that's a space is so exciting nowadays because there's so many like domain expert task that the model like that model developers want model to do well on right let's say you're like accountant

19:31

right like maybe you want to use a model to have accounting task so I need a lot of like accounting data like examples from like accountants so you need to hire a lot of them should I do it or if you want to do physics problems or you want to do um I don't know like legal

19:45

questions and stuff or like engineering questions or like somebody was telling me they want to do like uh using like coding for to solve scientific problems and not just like coding to build product which is another different whole realm of things and I also like using

19:58

very specific toolings like uh yeah like I'm not sure what apps you use but maybe like um for editing app or like Quickbooks or like Google Excel like they have very specific like tune specific um expert expertise that you want the models to learn. So like they

20:12

So like they need a lot of like humans expert in this area to like create data to train them.

20:19

Um and it's a massive things.

20:19

It's like people because uh everyone wants a lot of data and like want slaps like unlimited budget.

20:24

Uh but uh whether I think this is also like a little bit of lowkey interesting economics.

20:29

I'm not sure you talk to to like the guest about I thought it's very interesting to think about because it's very lopsided, right?

20:38

because like there only like a very small numbers of frontier labs right and they want a lot of data and there's like a massive amount of like startups or companies are providing related data so like you can see this companies like this startup like doing later but they

20:52

have like maybe have like massive AR but you ask them like okay so how many customers you have and they could be like oh a very small numbers I'm not sure I'm not sure you you I saw you smiling >> yeah we chat we chat about that >> yeah so so I'm like a bit like leave me unease Right? They have like a companies

21:07

unease Right? They have like a companies growing like crazy but it's like heavily dependent on like >> two or three companies and at the same time like if I if I was this company Frontier Labs what could be the right

21:20

economical things for me to do right now I want a lot of startups I want to have a lot of providers so I can pick and choose and then this providers can also like to compete each other to lower the price and it's so dependent on me they will sell to me regardless. So, so I So, so I feel like Yeah.

21:34

So, so this economics, the whole economics is very interesting to me and I'm curious to see how it plays out.

21:41

>> What I'm hearing is you're uh you're bearish on the future of these data labeling companies because as you said, they don't have a lot of leverage over pricing because they have so few customers and there's so many people getting into the space.

21:52

So, basically, even though they're some of the fastest growing companies in the world, you're feeling like there's there's a challenge up ahead.

21:59

>> I'm not sure if I'm bearish on it.

21:59

Um I think I'm curious because I think things have has a way of work out in ways that I don't expect.

22:09

So I think that maybe these companies they have a lot of data.

22:14

Maybe they wouldn't be able to use that to like have some insight that helps them like stay ahead of the curve, you know. So so I don't know.

22:19

Uh >> a very fair answer. Okay.

22:22

While we're on this topic, I want to chat about evals, which is a very recurring topic in this podcast.

22:28

This is the other piece of data content these companies share that AI labs really need.

22:33

Can you just talk about what an eval is the simplest way to understand it and then how this helps models get smarter.

22:41

>> So I think if people approach eval I think there like two very different problems.

22:46

problems. ones is a app builder right like can I say I have an app uh that do like uh maybe a chatbot very simple and I know it's first thing that came to my mind um and I want you to know is chatbot is good or bad right so I need to come up way with like evaluate the

23:02

chatbot um another thing is uh I think of this as a uh task specific evol design so let's say I'm a model developer and I want to make my model better at curve writing right and I was like okay but how How how do I even measure cerwiting right? So I would need

23:18

So I would need someone to like okay understand cer writing and think about like what makes good story like what makes a story good and then designed the whole data set and then criteria to evaluate creative writing.

23:31

writing. Um so yeah so so I think there's that I think it's like more like eval design that is very interesting uh come criteria um come guideline how to do it and then also like train people like how to do it effectively um so I guess uh in a course I think Evar is really really fun because it's extremely

23:53

creative uh I was looking at like different avons people build and was like wow like it's not dry at all this is like super super super fun >> we had a whole podcast any vals with HML Haml and Shrea and uh and that's exactly what they talked about is just it's actually really fun to create evals for for companies especially. So let's still

24:09

So let's still dig into that one a little bit more.

24:13

There's this kind of debate online that I don't know how big of a deal this debate is, but it feels like people uh spend a lot of time thinking about this this idea of do we need evals for AI products?

24:23

Some of the best companies say they don't really do evals. They just go on vibes.

24:26

They're just like is this working well? Can I feel it or not?

24:30

What's your take on just the importance of building evals and the skill of eval for app AI apps not the model companies?

24:38

>> You don't have to be like absolutely perfect at things to win.

24:41

You just need to be like good enough and being consistent about it.

24:45

Okay, this is not the philosophy I follow but like I have worked with enough companies to see that play out.

24:52

So when I say like why company don't need evaluate let's say you are like an executive right and you want to have a new use case.

24:57

have a new use case. So here's a use case you you started out we built and it's like it works well right the customers are somewhat happy don't you don't have the exact metric for it but like the traffic keeps increasing like people seem happy people keep buying stuff right and now here's our engineer coming like okay we need eva for it and so I think it was like okay how much effort do we need to put into eva and

25:17

they were like okay uh maybe like two engineers as much as much and it could maybe would improve so was like okay so how much expected gain can I get from it and the engineer would be like oh maybe you can improve it from like 80% to like 82% 85% right and I was like okay but we take like that two engineers and we launch a new feature then it could give me like so much more like improvement right so so I think it's like one of

25:42

them is like sometime people think of evol okay this is good enough just don't touch it like if you do spend a lot of energy on eva I would like only incremental improvement where it spends the energy on like another use case and maybe it's good enough that you vive check it right so so I I do think it's like maybe like that's a debate is about um I do think that's like a lot of time people just like get things to the to

26:03

the place when it's like okay good enough people run but and then but of course it's like there's a lot of risk associated with it because if you don't have a clear metric you uh you have a good visibility to how the application models are performing it might do something very dumb or it can cost you like I know some something like crazy can happen so so yeah so um so so I do think evol is very very important. If

26:27

think evol is very very important. If you have if you operate at scale and where like failures can have like catastrophic consequences then you do need to be very tyrannical about like what you put in front of the users understand different failure modes like

26:45

what could go wrong and also maybe in a space that like is is a feature of the product is as a competitive advantage right you want to be the best at it you want to have like a very strong understanding of like where you are and like where you are with the competitors but it's just something that's like more like a low key. Okay, this like

26:59

like a low key. Okay, this like something is like okay that's not the core or like it helps with our users then maybe you don't need to be so so obsessed or like tyical about it is okay that's good enough for now and if it fails then it fails like okay I know

27:13

it's like it's so terrifying but like yeah um yeah I think it's all about like the question of like return investment um I'm a big fan of I love writing Eva say it's like I understand why some people would choose to not focus on Eva right away and choose like bringing on new functionalities instead. >> Awesome. That is a really pragmatic >> Awesome.

27:32

That is a really pragmatic answer.

27:33

What I'm hearing is eval are great, very important, especially if you're operating at scale, but pick your battles.

27:39

You don't need to write evals for every little feature.

27:41

Something that Hamlin Shrea shared is that people need just like I don't know five or seven evals for the most important elements of their product.

27:48

Is that is that what you see or do you see a lot more in production that people build and need?

27:53

Um I I don't think of like just a fixed number unlike the evol like what was the goal of evol right the go to evol is to guys of product development um so so like you see evol um because I think I'm a big fan of evol is is that it helps you uncover opportunities where the products are

28:11

doing well so I sometimes seen it very often okay we look at the AA we realize it's like okay it perform really poorly on this like specific segment of users and then we look into it's like what what what what's what's wrong with it and it turns out it's like we just like don't have a good messaging to it. So

28:27

So like we should like just focus on the things that we doing can improve significantly. Yeah.

28:32

significantly. Yeah. So I kind of the number of evol is really depends like we have seen product with like hundreds of different metrics right like people like going crazy this is because like product is like general right have different have like one avar for like I don't know

28:46

like uh verity have like one evolve for like user sensitive data um and like another is like is um for length but like has a number of like um okay let's just example complex simple like div research so so you have the application you have like build a to like do deep research for you, right? like okay like

29:04

research for you, right? like okay like have a prompt like me say okay do me a com comprehensive research on Lenny's podcast and help me like some like uh propose like show me report on what kind of topics he's interested in what kind of videos could get the most views or

29:21

like what topics that he's missing on that he should be covering right like have that kind of like prompt then how do you evaluate the the result right I don't think this like one like metrics that would help maybe just like maybe you have like a 100 I think somebody has

29:36

a benchmark and is they get like a 100 expert like write a bunch of prompts and they go through like all the all the answers on AI and I do it and it's like it's extremely costly and slow right but if I might have something else for like

29:48

one way was thinking about it um I was talking to a friend about it and and one way is like how to produce the result of the of the summary right at first you need to do like gather informations and to gather informations you need to do a lot search queries uh you like gathers

30:07

grab the search results and then from the search results you like uh aggregate and then maybe say okay I'm still missing on this you have to go another route and like another route is another summary so every step of the way you need evaluations right you don't need to

30:21

end to end so maybe the first search query you might first think about like okay now I write five search queries I might look into like how good are the search queries like do they like are they like similar to each other because you need five such queries that are very

30:33

similar like Lenny podcast nanny podcast uh last month landed podcast like two months ago right it's not it's not very very exciting but like if the quer is a podcast like the the keywords are like more um more diverse right and then look at the results of the of the search

30:50

query let's say you enter the search query like Lenny Parscat data label like and then they come up with like 10 pages uh 10 results and then you come up with like oh Lenny podcast on uh I don't know m um I don't know like uh Frontier Labs and have like 10 results. I might look

31:04

I might look at a different web page like how much of them overlapping like are we are we doing both like the breath like getting a lot of page but also like do we have depth and also like have relevance because we come up with a search queries that completely irrelevant to the to to the original prompt.

31:20

So I feel like every aspect of it it would need a way of evaluating right so I don't think it's just like how many evolve should I get but like how many evolve should do I need to get a good coverage a high confidence in my application's performance and also to help me understand like where it is not performing well so that I can fix it. >> Awesome.

31:43

And I'm hearing also just especially for the very core use case like the most common path people take in your product is where you want to focus. >> Yeah. So yeah. >> Okay.

31:54

Let me there's one more term I want to cover and I want to go in a somewhat different direction.

31:58

Rag people see this term a lot. RA what does it mean?

32:04

>> So rag is stand for retrieval augmented generations.

32:07

It also not specific to JD AI.

32:10

So um the idea is just like for a lot of questions we need context to answer.

32:14

So I think it came pretty oh I think it's from the paper 2017.

32:20

So so someone was like um so they realizes like for a bunch of like benchmark when the question answering benchmarks they realize it's like okay if we give the model informations about the questions then the answer can be much much better.

32:33

So what they do was they try to retrieve information from Wikipedia.

32:37

So for for questionable topics just like retrieve that and then put in the context and like answer it does much better.

32:42

So I feel like it sounds like a no-brainer, right? I mean like obviously.

32:45

like obviously. So, so I think that's what racket as a simplest sense is just like providing the model with a relevant context so so that it can answer the questions and and that's where like things get like uh really uh more more interesting because traditionally when

33:00

it started out rack is mostly like text um so so we we talk about like a lot of way like how to prepare data so that the model can retrieve u effectively let's say there like not everything is a wikipia page right like Wikipedia page is pretty contained and like you know everything about it is about a topic. Uh

33:16

Uh but a lot of time you have documents extremely long, right?

33:19

And like they have a weird way of like structures the documents.

33:22

Let's say that um you have documents about Lenny uh podcast, right?

33:27

And in the in the future in the beginning documents like from now on podcast wouldn't refer to Lenny's podcast, right?

33:33

So let's say somebody in the future like tell me about Lenny, right? Lenny's work.

33:37

And because the rest of the document does not have the term leni you just don't know uh you might not read through it and the document is long enough that it chunk into a different part.

33:46

So like the second part doesn't have the the word manic so you cannot reach.

33:50

So I have to find a way to like process data.

33:52

So that makes sure it's like it can retrieve the information just relevant to the query even though it might not immediately like obvious that is related.

34:00

like obvious that is related. So people come up with like only thing I think like contextual visual um like uh giving x chunk of the data uh the relevant like maybe like summary meta data so that it knows um or like some people use like as a hypothetical

34:17

questions it's very interesting like for even the chunk of like documents I generated a bunch of questions that the chunks can help answer so that when they have a a query it was like okay does it match any of the like hypothetical questions so it can it can fetch it so it's very interesting approach Okay. So maybe before I go to the next

34:33

Okay. So maybe before I go to the next thing I just want to say this like data preparations for rack is extremely important and I would say this like in the a lot of the companies that I have seen that's like the biggest performance in their rack solutions coming from like better data preparations not agonizing over like what better databases to use uh because database of course it's very

34:54

important to care about like things like latency or like if you have like very specific access patterns like read heavy or write heavy Of course it's like it matters but in term of like pure quality answers right I think the data preparation is just like hands out >> when you say data preparation what's an example to make that real and concrete for us to understand >> so so like one way is like uh um

35:15

mentioned as in like um you have like chunks of data so we think about like how big of each chunk should be right because um if it's like so let's think about like if a context you want to maximize maybe you can it's very simple example right you want to retrieve like a thousand words right so If HM's data is too long um then so if if a data cham is long then it's more likely to contain more relevant metadata so you can

35:40

retrieve more but if it's too long like then you have a thousand word and so chunk is like a thousand words you can reach one chunk so it's not very useful but it's too short then you can retrieve more relevant information like oh so it can retrieve a wider range of like documents and chunk but at the same time a chunk is too small to contain relevant information. So you have like very nice

36:01

So you have like very nice like chunk design like how big a chunk should be.

36:05

Uh you add like contextual informations like summary, metadata, hypothetical questions.

36:10

Uh somebody was telling me like um a very big performance they got is that from um rewriting their data in the question answering format.

36:18

So like instead of having like so they have a podcast right instead of just chunking the podcast you just like reframe rewrite it into like here's a question here's answers um like and and produce a lot of them.

36:28

you can use AI for that as well.

36:29

So that's one example of data processing.

36:31

A lot of example we I see is like for people helping like using AI to have like specific uh tool use and documentations right and a lot and we write documentation usually a lot of document documentation today is written for human um reading and AI reading is different because it's different because humans we have like common sense and we kind know what it is.

36:53

Um so so one one things or like uh human for human experts they have the context that AI doesn't quite have.

36:59

have. So somebody told me that like what's the big change they have is like let's say that um you have a you have a function a document uh documentation for this maybe this library and this library say okay the output of this one is like maybe talking for like I know some crazy

37:14

term maybe some uh temperature something under graph should be like one zero or minus one and as a human expert maybe understand the scale like what one in the scale mean but like for AI just really doesn't understand what that means so so actually have like another

37:28

allot annotation layer for AI it's like okay got temperatures equal one mean like that it's not like it's the abstra it's like associated with the scale over there so just saving all this data processing to make it easier for AI to retrieve the relevant information to answer the

37:44

questions >> this episode is brought to you by persona the verified identity platform helping organizations onboard users fight fraud and build trust we talk a lot on this podcast about the amazing advances in AI but this can be a double-edged sword report. For every wow

37:58

For every wow moment, there are fraudsters using the same tech to wreak havoc, laundering money, taking over employee identities, and impersonating businesses.

38:06

Persona helps combat these threats with automated user, business, and employee verification.

38:12

Whether you're looking to catch candidate fraud, meet age restrictions, or keep your platform safe, Persona helps you verify users in a way that's tailored to your specific needs.

38:22

Best of all, Persona makes it easy to know who you're dealing with without adding friction for good users.

38:28

This is why leading platforms like Etsy, LinkedIn, Square, and Lyft trust Persona to secure their platform.

38:33

Persona is also offering my listeners 500 free services per month for one full year.

38:40

Just head to withpersona.

38:40

com/lenny to get started. That's withpersona. com/lenny.

38:48

Thanks again to Persona for sponsoring this episode. Awesome. Okay.

38:49

So, you've talked a bit about how you work with companies on these sorts of things, on their AI strategies, on their AI products, how they build, which tools they build, all these things.

39:01

I want to spend a little time here because a lot of companies are building AI products.

39:04

A lot of companies are not having a good time building AI products.

39:07

Let me ask a few questions along these lines of what you've learned working with companies that are doing this.

39:12

Well, one is just, I guess, in terms of AI tool adoption and adoption in general within companies.

39:18

There's all this talk recently of just like all this AI hype.

39:21

The data is actually showing most companies try it doesn't do a lot they stop and so there's all this just like maybe this isn't going anywhere.

39:26

maybe this isn't going anywhere. So in terms of just adoption of tools and AI within companies what are you seeing there >> for gen AI in company I think there are two type of genai toolings that have been uh I have seen like ones is to like um internal productivity right like have coding tool slack chatbot um uh internal knowledge like a lot of big enterprises have some kind like a wrapper around

39:51

like um model so but like with access like maybe some different kind of racket I think we talk about that Um okay like text based rack I haven't talked about like Asian tech rack or like multim motor rack yet but it's like yes there a

40:04

whole very exciting area around that um yeah so like b b b b b b b b b b b b b b b b b b b b b b b b b b b b b b b b b b b b b b b basically to allows the employee to like access internal document. Uh somebody somebody asked

40:12

document. Uh somebody somebody asked like okay um I'm I'm having a baby what could be the maternal or paternal policy right or like am I having this operations with the health benefit like cover that or like I want to like interview or I want to like refer my friend what could be the process for

40:29

that so a lot of this like having chatbot internal chatbot to help with internal operations um and um another things um another category is more like customerf facing um so or like partner facing um so so customer support chatbot is a big one you have a hotel chain you

40:47

might have like a booking chatbot which is like somehow massive like a lot of booking chatbot because I guess it's it's it's I do have this theory of like a lot of applications uh companies pursue because they can't measure the concrete outcome and I feel like booking or sale chatbot it's very clear right

41:03

like what's a conversion rate right now with a chatbot with human operators and what could be conversion rate with a chatbot and then some somehow I think it's like very clear outcome comes and companies are easier to buy into this um this solutions. So a lot of companies

41:16

So a lot of companies have that like customer uh facing chatbot.

41:20

Uh so yeah so so that is um another category of tool um and I think that um I think for customers or external facing tools um because people are driven to um people are driven to choose applications with clear outcomes.

41:39

So, so the questions of uh adopting them is really based on like whether they see the outcome or not.

41:44

Of course, it's not perfect because sometimes uh the outcome can be bad not because the idea or like the applications idea itself is bad.

41:52

It's just because the the I know the the process of building it is like not that great. Um yeah.

41:57

So, so it's tricky for the internal adoptions of like tooling.

42:02

So, internal productivity that's where it gets tricky.

42:05

tricky. I would say like a lot of companies uh what they think of as a strategy like I think of as have like usually have very um have like two key aspect right it's like use cases and the second is talent you might have like great data for great use cases but you don't have talents and you cannot do it

42:22

so a lot of time in the beginning with geni and it's still and sometimes I'm really admire a lot of companies for that it's just like exactly like okay we need our employees to be very geni aware like very AI literate right so what to do is they start like maybe like adopting a bunch of tools for for the team to use. They have a lot of upskill

42:40

They have a lot of upskill upscaling uh workshops like they encourage learning.

42:44

I think it's like a really really good thing and it's also like willing to spend a lot of money into like adopting like um giving people like touchd subscriptions uh cloud code subscriptions like to get the employees to like to to be more AI literate.

42:59

to like to to be more AI literate. Um and the other thing is like a lot of the secretary come say okay we spend a ton of money as this tool but then we don't see because you can see the usage it's like but people don't seem to use them as much and what is the issue so so yeah so I think is that that is um that is

43:18

tricky yeah >> what do you think is the issue is it just they're not they're like they don't know how to use them like what do you think is the gap here do you think we'll get to a place of just like wow work is completely different because of AI for a lot of companies The main thing is like it's really hard to measure productivity uh gain. So I

43:34

to measure productivity uh gain. So I talked to a lot of people and they was like first of all on on sample is coding right a lot of companies not using coding agents uh or coding aided coding uh and um I was asking I was like I was like okay do do you think that like it helps with your productivity and a lot of time the question is very handwavy just like okay like okay I feel like

43:58

it's been better right and okay because we have more PRs uh we see more code and then immediate correct me okay but of course code number of life code is not a good metric for that right so so it's it's really really tricky and it's something funny uh so so so I do ask um people to ask their managers because I work with like usually the VP level so they have like multiple teams under them

44:21

so I ask them like okay do you ask managers um like okay would you rather have access uh would you rather have give everyone on the team like very expensive coding agent um subscriptions or you get an extra headcount right let's say it's like maybe like um and and almost everyone could say the managers could say headcount but if you ask VP level or like someone who manage a lot of teams they could say just like

44:47

they would want AI assistant assistance tools and the reason is that people say like okay because as managers right because you are still growing like you're not as a level when you you manage hundreds of thousands of people so for you like having one HR hash count is like is big so you want that not for productivity reasons but because you just want you have more people working for you. Whereas for executive, you care

45:06

Whereas for executive, you care more about like the um the maybe you have more like business metrics that you care about.

45:14

So so you actually think about like what actually drive drive productivity uh metrics for you.

45:19

Uh so so yeah, so it's tricky.

45:22

Um and I think that's like the question of like productivity.

45:27

um it's not I'm not sure it's like fundamentally is the some people are more productive but it's just like we don't have a good way of measuring productivity improvement.

45:35

productivity improvement. Uh another thing is also varies wily um and I think that people do tell me that they notice different buckets of of employees like different reactions to AI assisted tools like first of all I keep going back to

45:50

coding because it's big and it's like easier to like reason about um so it says like um I have different reports like one team would tell me that like um one of the people tell me okay amongst all his engineers he think it's like senior engineers would get the most

46:07

output like would be more productive because it's like okay so that person very interesting so so he actually divided his team to like three bucket but he didn't tell them obviously he was like okay here's more like currently like u best performing average performing and lowest performing and

46:23

then there's a randomized trial so like they give like half of each of each group like access to like to like cursor and then was noticed like over time was like okay the something funny like the the group that get the biggest performance boost like in his opinion. So it goes very close to his team as the

46:38

So it goes very close to his team as the biggest boom Bruce like is the senior engineer the highest performing.

46:41

So the highest performing engineer get the biggest boost out of it and then the second group is just like the um the average performing.

46:49

So so he so his opinion is like okay the highest performing engineers they also more proactive they say know how to solve problem.

46:58

So AI helps them to solve problem better.

46:59

Whereas the people who are lowest performing they only don't care much about work right?

47:04

So like it's easier to just like go on autopilot get it to like generate like bad code and just like do it and I just don't know what to do with it.

47:12

As another company however they told me just like actually senior engineers are the one most resistant to like using AI as tooling because they said it's like okay but AI because they are more opinionated and they have very high standard was like okay but AI code jet code just sucks.

47:28

So just like very very resistant in using this.

47:32

So I don't know I I haven't quite be able to reconcile like very different reports on that yet.

47:39

>> This is so interesting.

47:39

So just to make sure I'm hearing what you're the story.

47:43

So there's a company work with that did a three bucket test with their engineering team where they created three sorts of groups.

47:50

The highest performing engineers, mid-performing engineers, lowest performing engineers uh and gave some of them so they gave some of them access to say cursor.

47:59

Was it cursor or what did they give them access to? It was cursor.

48:03

>> I think by then it was cursor. >> Okay, cool.

48:04

And so >> I didn't work with them.

48:05

This more like a friend company. >> Okay. It's a friends company.

48:08

So do they give like half of the higher performing engineers cursor and half not or how did they do the split? >> Yeah.

48:14

So like they give like half of the entire company but like half for each bucket. Yeah.

48:17

And then they observe the difference in like productivity. >> I see. Yeah.

48:22

>> So how do they even do that?

48:22

They're just like okay you get cursor, you don't get cursors.

48:25

That how did they do that? That's so interesting. >> Yeah.

48:27

I I didn't again just the mechanics of it.

48:28

Uh but but I was like at Raspia for doing a randomized trial. >> That is so cool. Yeah. >> Okay. Wow.

48:34

How large was this engineering team?

48:36

Was it like hundreds of people?

48:38

>> Um it's it's not that large.

48:38

It's about like maybe um 30 40. Yeah. >> 30 to 40. Okay. Yeah. >> Wow. Okay.

48:44

So they found that the highest performing engineers had the most benefit from using AI tools and then behind them was the middle tier engineers and the worst performers >> were the lowest performers. Okay.

48:59

>> But also not the same everywhere. Um like companies Yeah. Yeah. different. >> Right.

49:03

This other example you shared of just senior engineers in this one example are most resistant to changing the way they work which I get.

49:09

I I do feel like the most valuable people right now other than ML researchers uh and AI researchers like yourself are senior engineers because it feels like junior engineers are just like so much of this is now done by AI but an engineer that knows what they're doing that understands how things work at a large scale with AI tools just basically like infinite junior engineers doing their bidding feels like an extremely valuable and powerful asset. >> Yeah.

49:39

>> Yeah. Uh I definitely like really appreciate as you see companies like we appreciate engineers who are um have a good understanding of the whole systems and be able to have good problem solving skill like thinking holistically instead of like local uh locally um when our

49:55

company have seen as the way they work as they told me they work completely different now I like so they actually restructured engineering arc so that like they get more senior engineer should be more in the PR review because they like to get like sort of writing guidelines on like what is a good engineering practices. Um what is a

50:08

Um what is a process would be like they be like okay so they write like a lot of like processes uh on how to work well and then they um and then they have more more junior engineers just like produce code and and like submit PR but senior engineer more in the reviewing case.

50:25

So I think is it might be prepare for the future.

50:29

So another company actually told me something very similar.

50:31

me something very similar. So that kind preparing the future when they only need a very small group of like very very strong engineers to like create like processes and like reviewing code to get into production but I get like AI or like junior

50:47

engineers should like produce code but then the question becomes like how does one become a very strong >> right that's right that's right I feel like >> yeah so so I don't know what's the process was thinking about like yeah um >> no one's thinking It's just it's a problem. We won't have any more in 10 20

51:03

We won't have any more in 10 20 years.

51:05

There'll be no more engineers because no one's hiring junior engineers.

51:08

Although I could make the case junior engineers, people just getting into computer science right now are just native AI native.

51:12

And in theory, you could argue they will become really good really fast if they're curious, aren't just, you know, delegating learning and thinking to AI, but learning how to actually using it to learn how to code well and architect correctly.

51:28

like you could argue they will be the most successful engineers in the future.

51:33

>> I do think that what I mentioned is like loading to architect um I think I I group that in like system thinking I do think it's very important skill because I think AI can help automate a lot of like um destroy the skills but like knowing how to like utilize the skills together to solve a problems is is very uh it's is it's hard.

51:53

So there's a a webinar between um Mer Sami who was my one of my favorite professors.

52:00

He was a chair of the curriculum of the CS department at Stford.

52:04

department at Stford. We spend a lot of time thinking about CIS educations right like what what what should students learn nowaday like AI coding and then the other person is like Andrew which is of course it's like a legend in the AI space and NAMI person like Sami said something very interesting is like he said like a lot of people think that CS is about coding but it's not like coding is just a means to an end like CS is

52:27

about system thinking like using like coding to solve actual problem and problem solving will never go way because like what like AI can automate more stuff the problem just get bigger but like the process of understanding

52:40

what cause the issue and like how to like design step-by-step solution to it will always be there um so I think an example of um of like I actually have a lot of issues with like AI for like um in the way of like is debugging so I'm

52:56

not sure you use a lot of air for coding but like something I've noticed and also seen from my friends it's like it is pretty good when you have very clear welldefy task maybe write do re documentations fixes specific features or like build an

53:09

app from scratch right like doesn't have to interact with a large existing code base but it added something like a little bit more complicated um maybe require interacting with a lot of components and stuff is usually like not

53:20

that good um and and for example like I was using AI to like use um to deploy an applications um and I was testing out a new uh hosting service I was not familiar with I was like okay like usually they for me so what they think AI does give me is like confidence to try new tool like before with AI is like trying new tools has a lot of documentations for the beginning but I was like okay just try it out and and learn. So I was testing out this new

53:42

So I was testing out this new hosting service and it kept getting a bug.

53:46

bug. It was like very very annoying and I was like okay I asked uh car codes like fix it and it keep g keeping like it keep changing the way like maybe change environment variable fix the code maybe I change from the function to this function maybe change the language maybe

54:01

it doesn't process JavaScript well I don't know whatever and it didn't work and I was like okay that's it I'm just going to read the document document uh documentation myself and see what's wrong and it turns out it's like I'm on another tier like the fish I want did not is not available in this tier. Right? So I feel like okay so the issue Right?

54:18

So I feel like okay so the issue with clon is trying to focus on fixing things from a very a different component where the issue is from a different component.

54:26

component. So I think I think of like okay be understanding like how different components um work together and where the source of issue might come from you need to you need to have a holistic view of it and this made me think is like okay how do we teach AI like system thinking like like that right I think I

54:41

have all the human experts like having like right like very much built scaffold uh just like okay for this kind of problem look into this look into that look into that and then stuff so so I think is that could be one way but also made me think is like how do we teach humans like system thinking. Um yeah, so Um yeah, so so yeah.

54:59

So I think it's very interesting um skill.

55:00

I I do think it's very important.

55:04

>> That's exactly the same insight Brett Taylor shared on the podcast.

55:06

He's the co-founder of Sierra. He created Google Maps.

55:09

He was CEO of Salesforce, Quip, a few other things.

55:12

And I asked him just like should people learn to code?

55:13

And his point is exactly what you said, which is learning taking computer science classes is not about learning Java and Python.

55:20

It's learning how systems work and how code operates and how software works broadly, not just here's like a function to do a thing.

55:32

One thing that I wanted to help people understand, you you wrote this book called AI engineering, which is essentially helping people understand this new genre of engineer.

55:38

And you have this really simple way of thinking about the difference between an ML engineer and an AI engineer which has a really good correlary to product managers now of just like an AI product manager versus a non-AI product manager.

55:52

The way you describe it and fill in what I'm missing is just ML engineers built models themselves.

55:58

AI engineers use existing models to build products.

56:04

Anything you want to add there?

56:04

One thing I really dislike about writing books is that you have to defy like this and and I think it's like no definition to be perfect because they always be like edge cases.

56:14

Um but yeah in general I think it's like just like gen like AI as a service like more as a service like when somebody build the models for you and the base model performance is a pretty shock.

56:25

pretty shock. So, so it's like it's enable people to just like okay now I want she in integrate AI into my product I don't need to learn what green is even though knowing that could really help uh but but yeah it's like it makes the entry barrier really low for people who

56:39

want to use AI to build product and at the same time AI capabilities are like so strong like it's also like increase like the possibilities like the type applications that AI can be used for so I think like yeah so like both entry barriers like super low and like the demand And for like a applications like a lot bigger. So it feels very very

56:57

So it feels very very exciting.

56:59

It opens up like a whole new world of possibilities. >> Oh yeah.

57:04

It's like now you don't have to time you don't even have to spend time building this AI brain.

57:07

Now you can just use it to do stuff.

57:08

Uh such a such an unlock. Okay.

57:11

Maybe just a vital question.

57:14

You get to see a lot of where what's working, what's not working, where things are heading.

57:19

I'm curious just if you had to think about in the next two or three years just where things are heading.

57:24

What do you think what do you think how do you think building products will be different?

57:28

How do you think companies working will be different?

57:32

If you had to think of maybe the biggest change we expect to see in the next few years in terms of how companies work.

57:40

>> I think in a lot of organizations they don't move that fast, right?

57:42

Um but at the same time they also move faster than I expected.

57:48

I expected. uh because again I think it's like biased like and don't work with dinosaur companies don't care like a lot of executive who comes to me are like very forwardlooking so maybe for me I'm very biased uh towards towards like organization just like move fast um so

58:05

so yeah so I think one big change I see is just like in organizational structure um I think it's like a lot of value place um in like um so before like we have like a lot of disjointed team like we have very clear like engineering ing team, product team. But then there's a

58:19

But then there's a question of like who should write Eva, right?

58:23

Like who should own the matrix?

58:25

And it turns out it's like Eva is not a um it's not a it's not a separate problem.

58:29

It's a system problem, right?

58:31

Because you you you need to look into different components how interest each other.

58:34

You need user behaviors because you need to know what users care about so that you can so that you can like write write about like reflect what users care about.

58:42

So, so all of that like you can sort it from like you know look into different component architectures uh place guardrails and stuff.

58:49

So it's just engineering but understanding users is like what product right so so because of like a lot of things and they are extremely important.

58:56

So like that guy bring product team and like engineering team even like marketing team like user acquisition like very close to each other so so yeah since in a way people are structuring so that's more communications between like previously very distinct functions.

59:11

Another thing is I also see as teams um of course like think about like what can be automated in the next few years and what what cannot be automated and I see that people already like shedding like actually is it's a little bit like scary to think about it but I also think it's

59:27

like the team told me it's like okay this is between you and me but we have we like got rid of these functions right like for a lot of thing like uh previously outsource for example like traditionally is a business outsourcing this core to them and like can done with

59:41

like not um can be a more um system uh um systematized um so so with that you can actually like use AI like automate a lot of that and also like a separation think more like what is the value of like junior engineers or like senior engineers how you restructure engineering for that um so so yeah so I

1:00:00

do definitely think that um is one thing to success organization people are just moving pieces around and like thinking about like use cases um whether you need to like spin out new use cases and who would lead a new effort and like yeah um that is one big uh change. Another

1:00:15

Another things in ter of like AI, I think this is um I'm not sure how true this is.

1:00:21

Um I guess I'm I'm also like on the camp of like thinking that is has merit is it's a camp of like okay uh base models we have probably like not quite max out but we want we unlikely to see like really really strong like craziness strong models.

1:00:43

models. So like you remember like when we have like GBT right then GB2 which is a big step up like an order of magnitude like like better than like GBD and then GB3 which like much much bigger GB4 much much bigger and then of course I have GBD 5 but like is GB 5 like that scale of like much bigger like a step jump

1:01:03

compared to like the previous I think it's a debatable right so so I think that it's like we had reached a point where like the base model um performance improvement is not going to be like mind-blowing it was in the last three years uh so so I think it's like a lot of like improvements we're going to see

1:01:21

in the post training phase in the application building phase um and um and yes also I think that's where I feel I was see a lot of improvement there so very like interesting like multimodality um so we've seen a lot of uh text based uh but I think there a lot of um audio videos use cases uh that is very very

1:01:44

exciting and I think audio is not quite as soft as thinking because I do work with like with with like a couple of like voice startups and when I talk to think about voice it's a entirely different beast uh so let's say have chatbot right we go from a text chatbot to voice chatbot it's like the consoles

1:02:03

are completely different because now with voice chatbot right we need to think about like latency because like multiple steps uh first like like text like like voice to text text to text and text question into text answer and then like and then text to voice answer right so it's like manable hops and like

1:02:18

latency become very important and there's a question like what does make you sound natural so for example like people think like um in in in AI and humans so like when humans talk to each other like if I say if I say you try to interrupt me it's like um chip that right I would like pause and I try to hear you out right but sometime I may

1:02:39

just say say some word not like acknowledge when I mhm that I shouldn't stop I just continue so the question of like force interruption like whether it's like I should should I stop or not like it's is a big and like what perceived as like natural conversations and that's also regulations right because like because like a lot of time

1:02:59

people want to build AI chatbot voice chat bots that sound like humans try to like trick users into thinking they're talking to humans but also of like maybe potential regulations saying like okay you have to disclose to users when to talk if the if the bot is talking to is human or um or AI. So, so I think just

1:03:15

human or um or AI. So, so I think just like um there's a whole space I think it's not quite as so as as you think is it but it's al it's not quite like an AI foundation model problem right because like a human interruption detection is actually a classical machineing problem

1:03:32

like you you it's is a different uh framing but like you can view classifier for that or or like the question of like let's see actually have a massive engineuric challenge not an AI challenge of course it can be an AI challenge because people are trying to build like voicetovoice model. So instead of having

1:03:47

So instead of having like having to first like transcribe the voice from me into text and then get a model generous text answer and get another model to like turn from text to speech, you just like voice your voice directly.

1:03:59

So that is something who are working on but it's like very hard. Um yeah. So so yeah.

1:04:03

So like even audio I think of it is like the easier than video right because video have like both image and voice.

1:04:10

Uh it's already like pretty hard.

1:04:12

So I think it's a lot of challenges in that space.

1:04:16

That was an awesome list of things.

1:04:16

Let me mirror them back real quick.

1:04:18

So what you're predicting in the next few years, things that will change in the way we work and these actually resonate with so many conversations I've had on this podcast.

1:04:27

So this is just kind of doubling doubling down on where things are heading.

1:04:30

One is the blurring of lines between different functions instead of just like design engineering.

1:04:36

Everyone's going to be doing a lot of different things now.

1:04:38

Uh, two is just more of work being automated with agents and all these AI tools and just in theory productivity going up.

1:04:45

Third is a shifting from pre-training models to post-training fine-tuning and things like that because to your point model models maybe are slowing down and how smart they're getting.

1:04:56

Although I'll point folks to the ed chat with the co-founder of Anthropic.

1:04:59

He made a really good point here.

1:05:01

He's like we're really bad at understanding what exponentials feel like when we're in the middle of that.

1:05:06

And also models are being released more often.

1:05:07

So the difference between them we may not notice because they're just happening more often versus GPT3 came out like a year I don't know a before after JPT2.

1:05:18

So uh maybe true maybe not.

1:05:18

And then the fourth point you made is this idea of multimodal investing in multimodal experiences.

1:05:24

I cannot wait for JPT voice mode to get better at interruption.

1:05:26

Like exactly what you're saying.

1:05:28

I'm just like talking to it and then someone makes a little sound.

1:05:31

It's like okay and then you have to and then it's like and then it stops talking. It's so annoying.

1:05:36

I'm shocked that we don't have better voice assistant at home yet.

1:05:39

I think I have been testing out a bunch.

1:05:41

Like I keep hoping, oh my god, Zach could be the one and then I know how many of them I just like had to get away because they're not that good. >> I think it's coming. I hear it's coming.

1:05:50

Anthropic is working with someone uh that I don't know if it's launched or not yet.

1:05:54

>> Yeah, I want to bring back to what you mentioned about like the your guest like from Antropic mentioned about the performance uh improvement.

1:06:00

I think there's a big change.

1:06:02

Um I think like um this difference between um a model based capability so I'm not talking about like the pre-trained model right versus a perceived performance.

1:06:11

So, so let's say just like um a machine thought about like are you familiar with the term test time compute? >> Uh I don't think so.

1:06:21

>> Yeah, help us understand.

1:06:21

So um so so the idea is like okay like um you have some a fixed amount of compute, right?

1:06:27

So you're going to spend a lot of compute on pre-shooting or training the model pre-tuning and then have spend a lot of uh some computing and the ratio like pre-tuning to the post training compute is like crazy varies different between different lab um um and also like since then has to spend comput uh on like Jerry inference when I have a train and 500 model now it want to like serve to users so I might type a

1:06:51

questions or prompt and like Jerry like do inference like and that requires a compute and I guess I feel about discussion of like uh should I spend more compute on like pre-tuning or fetuning or inference right because like inference and people found I was just like test time compute so like spending more compute on inference is like call like test time like u compute uh like the strategy of like just allocating

1:07:14

more resources compute resource to generate uh inference when I bring better performance and how does that do it like let's say um let's say you have a math questions right and maybe instead of just generic one answer I can gen

1:07:26

four different answers and say okay whichever is uh the best according to some standard uh or like okay have four answers and then maybe like three of them say 42 and one of them said like 20 okay three of them in the in in

1:07:40

agreement so the answer should be 42 right so like just people shouldn't generate a bunch of it or another thing is like a lot of time like reasoning uh thinking just like people should like generate more thickening tokens I spend more time thinking before showing the final answers uh it's like require more compute but also like give it more uh more and more more better performance. So so yes. So so so I think it's like So so yes.

1:07:58

So so so I think it's like from the user perspective right like when the model spend more time exploring different potential answers thinking longer it can give you much better final answers but the base model itself does not change. >> Awesome. >> Does it make sense? Yes. >> Yes. Absolutely.

1:08:17

Uh that is a good correlary to uh to Ben man's point. >> Yeah.

1:08:23

Chip, we covered a lot of ground.

1:08:25

I've gone through everything I was hoping to learn and more.

1:08:27

Before we get to our very exciting lightning round, is there anything else that you wanted to share?

1:08:32

Anything else you want to leave listeners with?

1:08:34

>> So, I do work at a few companies that does this things of like they want employees to like come up with ideas.

1:08:40

So, there's a big debate on like what is a better way for a strategy, right?

1:08:44

Should it be top down or like bottom up, right?

1:08:46

Should like executive come up with like one or two like killer use case and like everyone like allocate resource to that or like should you give engineers and PMs and smart people like come up with ideas and I think it's a mixture of both.

1:08:58

mixture of both. So, so some companies it was like okay we hire a bunch of smart people like let's see like what they come up with and they they organize like hackathons or like internal challenge to get people to to build product and one things that um I noticed

1:09:13

is like a lot of people just like don't know what to build uh and it shocked me like why I feel like we are in some kind like an idea crisis right now we have all this really cool tools to have you like do everything from scratch I can have you like design it can have you like write code it can have build website. So in theory we should see a

1:09:28

website. So in theory we should see a lot more but at the same time people are like somehow stuck like they don't know what to build and and I think it's like maybe a lot of had to do with like maybe like um society expectations because

1:09:41

like we have gone through uh we have gone into this phase of like specializations like people like uh very highly uh specialized and people are supposed to do like focus on one thing really well instead of like a big picture and we don't have a big picture of you. it's hard to come up with like

1:09:55

it's hard to come up with like ideas of what to build.

1:09:56

ideas of what to build. So, so I know what like uh when when I work with this company on this hackathon like we do work out like a how come up with a guideline like how to come up with ideas and usually what we think of is like okay like one tip is like go look from the last week right like for a week just like pay attention to what you do and

1:10:14

what frustrate you and when something frustrate you think about like is there anything we can do is there like can you be done a different way so it's not frustrating and you can talk like people can swap accept notebooks or teams and if you see common frustrations maybe just something you can think about like just to build something around that. So yeah so I feel

1:10:31

So yeah so I feel like um just like notice like how we work uh thinking of like ways to like constantly ask questions like how can be better and then I just build something to like address the frustrations.

1:10:41

I think it's a good way to just like learn and adopt AI.

1:10:46

>> I think people have felt exactly what you're describing every time they open up one of these vibe coding tools where they could just describe anything you want.

1:10:52

I'm like I don't know what do I want?

1:10:54

And and I love this very tactical piece of advice, just like what frustrates you, just pay attention to where you're frustrated.

1:10:59

For example, I just built a very cool little VIP coded app.

1:11:03

I was working on a newsletter post inside Google Docs and I I pasted all these images into the Google Doc from screenshots and stuff and and then I forgot, oh yeah, you can't take images out of Google Docs.

1:11:14

It's like this Hotel California experience where you can paste stuff into it.

1:11:17

Very hard to get images back out.

1:11:19

So, I just went to all the VIP coded tools and just build an app that I can give you a Google doc URL and it let me download all the images automatically and it worked amazingly well and it made it really cute and I'll I'll link to it in the show notes.

1:11:32

>> Oh, I would love to see that.

1:11:32

I do I'm very bullish on like using AI just create like micro tools like just something just like make your life a bit easier >> and 100%.

1:11:40

I feel like that's one of the main ways people are using these tools just like a little niche problem they have.

1:11:46

With that, Chip, we've reached our very exciting lightning round.

1:11:48

I've got five questions for you. Are you ready? >> Yeah. Always. No. No.

1:11:54

Uh, depends on how hard the questions are.

1:11:58

>> They're very consistent across every guest.

1:12:00

So, uh, I imagine you've heard them before.

1:12:02

First question, what are two or three books that you find yourself recommending most to other people?

1:12:09

O I'm really terrified of like book recommendations because I feel like what books a person should read really depends on what they want and where they in life and where they want to get to.

1:12:18

Uh but there's several books that I do think is really change the way I think and see the world.

1:12:23

So one thing is a selfish gene that's like understand uh it actually changed uh it actually helped me with the question like whether I want to have kids or not.

1:12:31

I want to have kids or not. uh because it's like uh understanding more of like um yeah a lot of our functions of way we operate is the functions of our genes uh and genes want to do one thing was like to procreate uh so so yes in a little way but I like the book also proposed

1:12:49

another thing it's like um so everyone wants to live forever right now and maybe it's not like consciously but subconsciously we do we do want that and and I said two ways like one is like by genes like genans wants just like want to continue forever But also there are two ideas. Um I think there's something

1:13:03

Um I think there's something going meme.

1:13:05

Uh it's like being a boy if you have some ideas out there and then it's like last for a long time.

1:13:08

That's the one you like live on.

1:13:09

I know it's like it's a little bit like um abstract but I thought it's very interesting.

1:13:14

The other books I really really like.

1:13:16

It's like from like u the book from um Singaporean um previous um I think he's known as the father of Singapore.

1:13:24

known as the father of Singapore. I know like Limi Guango I'm not so sure what's the title it but like he did so he was the one who led Singapore from uh he changed uh Singapore from a third country to a forceful country within 25 years and I have never seen any country leaders spend so much effort into like

1:13:41

pushing down his thought on like how to build a country uh like like that um yeah I talk a lot about like public policy like how to like create policies that encourage people to do the right thing that is good for the nation And also talking about like uh foreign affairs, foreign policies like the relation of like the country with other. So it's a really good book to think

1:14:01

So it's a really good book to think about.

1:14:03

For me it's like system thinking but like it's a different kind system which is country which a lot of us don't get a chance to like ever experiment in our life.

1:14:10

So it's good to learn about that.

1:14:12

>> What was the name of that second book?

1:14:14

>> Uh it's called like from third to first world fashion.

1:14:17

I think I have it somewhere here. Yeah, >> there it is. Show and tell. That's that's awesome.

1:14:23

I definitely want to read that.

1:14:24

That's a really good tip.

1:14:25

I've heard a lot about just the impact he's had and I've seen all these videos on Twitter of just his really wise insights into how to build a thriving society and clearly >> believe like how does he have time to write is such a thick book. It's like insane. >> That is Claude. Please summarize. I'm just joking. Uh by the way, selfish.

1:14:41

I also absolutely love that book.

1:14:44

That is such a good choice.

1:14:45

It's such an under the radar kind of book that really changed the way I see the world as well. Uh so really good pick. Okay, next question.

1:14:52

Do you have a favorite recent movie or TV show you really enjoyed?

1:14:56

>> So I watch a lot of movie and TV shows as a research uh because I I working on my first novel and I recently uh sold it.

1:15:03

So I'm interesting like what makes So it's a drama.

1:15:05

It's not a science fictions or uh anything that like tech people usually read.

1:15:10

So it's it's very like I know it's a very um out of the left field out of left field and like very um so like reading watching TV to see like what kind of stories become popular trying to understand the trope and and stuff like that.

1:15:24

So I'm not sure if the audience like >> well what's one what's one that taught you something about writing?

1:15:31

>> I think it like uh Yami Palace is a Chinese TV show. >> Cool. Okay.

1:15:37

I haven't that one on the podcast before. Okay. Yeah. >> Next question.

1:15:42

Do you have a life motto that you often think about come back to when you're dealing with something hard whether it's in work or in life?

1:15:51

>> This sounds very nihilist. I think so.

1:15:53

Say it's like in the end nothing really matters.

1:15:55

Uh usually think of like in the grand scheme of thing like in a billion years nothing will like no one would ever be there.

1:16:01

I think okay some people argue with me about that.

1:16:03

argue with me about that. So I go to like I so my theory is like in a billion years like none of us would ever exist like so like whatever like messy things like crazy things we do or like how bad we do it I mean no one would be remember wouldn't be there to remember it and I

1:16:19

think in a way it's like it sounds scary but it's very liberating because it just allows me okay let's just try things out right like why does it matter and there a story of like recently um so we have some family member who passed away recently and I was talking to my that because I couldn't be home for that. I

1:16:35

because I couldn't be home for that. I was asking my dad like okay is there anything I can do to make the person like oh something like comfort anything you can get that person and my dad was just like what can he possibly want at this moment like and this made me real

1:16:49

like at the end of life like there's nothing that can bring you like like material can bring you joy there's no like money no product nothing and in a way feel like okay what really do I really care about at the end of the day um so I guess it's like I think about it if it's like okay maybe I fail it maybe

1:17:06

don't get the contract maybe things like in but in the end at the end of life like I don't think that actually really matters so in a way it's like it's quite liberating >> uh I know you said it might be nihilistic this is what Steve Jobs shared too in one of his most famous speeches just we will all die someday so

1:17:22

don't take things so seriously and it is freeing absolutely it just makes you appreciate every moment every day you have just like yeah let's just do something hard and scary okay final question you talked about how you're writing a novel most people in tech uh have never written something creative and fiction. What's just like one thing

1:17:39

What's just like one thing you learned in the process about how to write better stories, better fiction?

1:17:46

>> A lot of time when we read uh we get trip up by some small things.

1:17:48

So I think like I I want to do writing because I just want to go a better writer and I thought like maybe try my like a different audience could help me like become better like anticipating what this different type of audience would want to hear and like like what they care about.

1:18:05

So this is a way for me to get a so I think about writing or like even like any kind of like content creations is about like predicting the users's reactions right >> just kidding. >> Yeah.

1:18:17

So, so like you do a podcast, it's like okay, what kind of things that the users could find engaging, right?

1:18:21

And and I find this like a little bit like uh in a lot of companies like you have like launch a product, you have a narrative coming out, okay, what kind do we position this product in a way that like users would want, right?

1:18:32

So, I feel like I have done technical writing for a while and I felt like I have had some experience like trying to predict what engineers would want to hear all care about.

1:18:44

But then I don't have an experience like this completely different type of audience.

1:18:47

So that's what I want you to like career writing creating a story and that's why I was doing a lot of research on like watching I mean going research enjoy a lot like watching a lot of dramas.

1:18:56

I just see like what what people like.

1:18:58

Um so so one things that I care about is just like I think I learned is like what like emotional journey was from my editor, right?

1:19:05

So like when we write something we we care about like how users would feel like across the the story like we want something in the beginning, right?

1:19:13

We want something just like we need to have a hook so that people continue reading.

1:19:17

But we also don't want too much of like drama because we'll get like too tired, right?

1:19:21

like uh because like the emotionally exhausted like because it's like you're being like emotionally manipulated like a lot of time.

1:19:28

manipulated like a lot of time. So it give like emotional emotional journey maybe have like some some climax or like some something more chill or like maybe like I so care about another things I I didn't realize like for me for for technical writing you entirely focus on the content like the argument is very impersonal right like it it like for example like people like ML compilers

1:19:49

like doesn't matter if they like the person telling them about compiler or not right because it's just like objective like like but like for for novel people care about like character likability So, so like in in the first version is my story and makes the characters like a little bit more like uh very uh very logical, very rational and just does everything just like very rationally. And then the feedback I got is I have a

1:20:11

And then the feedback I got is I have a very good friend read it and he was he's a amazing person.

1:20:14

He's a great person and he was like chip I be honest you I hate that person.

1:20:18

So it doesn't matter as a story.

1:20:19

It's just like the person is so unlikable.

1:20:21

So let's say he doesn't find so is a second version and makes a person the character more likable like what how she makes that character more likable is that you put in some vulnerability like sometime that okay maybe a person like has setback because some people can relate to it.

1:20:34

See in a lot of ways it's like it's very interesting.

1:20:38

It's like a lot of it is like um yeah, a lot of it is it's about like understand the emotional bits uh like how the users feel not just about the story but also about the characters.

1:20:50

>> That is so interesting.

1:20:50

Wow, I learned a lot more there than I thought. That was awesome. Really good example.

1:20:54

Chip, two final questions.

1:20:56

Where can folks find you online if they want to reach out and maybe work with you or uh maybe even just share the stuff that you offer if folks want to reach out?

1:21:04

And then how can listeners be useful to you?

1:21:06

I'm like I'm on social media, LinkedIn, Twitter.

1:21:11

I don't post a lot, but I keep telling myself that I should do more because I kind like the conversation with uh with um with with readers.

1:21:17

Uh so I'm actually about to start a a SL a subspect.

1:21:20

Um so I have like a placeholder for subspect right now and I'm thinking of doing it for more system thinking because I think it's a very interesting skill.

1:21:29

Um and so like thinking of doing a YouTube channel on book reviews and basically books that help you think better.

1:21:35

So I think it's the first book I'm going to review is probably like this book because it's like my favorite book growing up.

1:21:40

Uh and I have been like keep on reading it. Uh so yes.

1:21:44

So how can it be helpful like send me books that you like uh books that have you have changed the way you think or change you the way you do anything.

1:21:57

So, I would appreciate it. >> Amazing.

1:22:00

I'm I'm excited to read that book.

1:22:03

>> Uh, Chip, thank you so much for being here.

1:22:05

>> Thank you so much, Lenny, for having me. >> Bye, everyone.

1:22:11

Thank you so much for listening.

1:22:11

If you found this valuable, you can subscribe to the show on Apple Podcasts, Spotify, or your favorite podcast app.

1:22:16

Also, please consider giving us a rating or leaving a review as that really helps other listeners find the podcast.

1:22:23

You can find all past episodes or learn more about the show at lennispodcast. com.

1:22:30

See you in the next episode.