searchlore

Back to Resource

All Segments

Why AI evals are the hottest new skill for product builders | Hamel Husain & Shreya Shankar

Why AI evals are the hottest new skill for product builders | Hamel Husain & Shreya Shankar

65 segments available

Hamel Husain and Shreya Shankar teach the world’s most popular course on AI evals and have trained over 2,000 PMs and engineers (including many teams at OpenAI and Anthropic). In this conversation, they demystify the process of developing effective evals, walk through real examples, and share practical techniques that’ll help you improve your AI product. *What you’ll learn:* 1. WTF evals are 2. Why they’ve become the most important new skill for AI product builders 3. A step-by-step walkthrough of how to create an effective eval 4. A deep dive into error analysis, open coding, and axial coding 5. Code-based evals vs. LLM-as-judge 6. The most common pitfalls and how to avoid them 7. Practical tips for implementing evals with minimal time investment (30 minutes per week after initial setup) 8. Insight into the debate between “vibes” and systematic evals *Brought to you by:* Fin—The #1 AI agent for customer service: https://fin.ai/lenny Dscout—The UX platform to capture insights at every stage: from ideation to production: https://www.dscout.com/ Mercury—The art of simplified finances: https://mercury.com/ *Transcript:* https://www.lennysnewsletter.com/p/why-ai-evals-are-the-hottest-new-skill *My biggest takeaways (for paid newsletter subscribers):* https://www.lennysnewsletter.com/i/173871171/my-biggest-takeaways-from-this-conversation *Where to find Shreya Shankar* • X: https://x.com/sh_reya • LinkedIn: https://www.linkedin.com/in/shrshnk/ • Website: https://www.sh-reya.com/ • Maven course: https://bit.ly/4myp27m *Where to find Hamel Husain* • X: https://x.com/HamelHusain • LinkedIn: https://www.linkedin.com/in/hamelhusain/ • Website: https://hamel.dev/ • Maven course: https://bit.ly/4myp27m *Where to find Lenny:* • Newsletter: https://www.lennysnewsletter.com • X: https://twitter.com/lennysan • LinkedIn: https://www.linkedin.com/in/lennyrachitsky/ *In this episode, we cover:* (00:00) Introduction to Hamel and Shreya (04:57) What are evals? (09:56) Demo: Examining real traces from a property management AI assistant (16:51) Writing notes on errors (23:54) Why LLMs can’t replace humans in the initial error analysis (25:16) The concept of a “benevolent dictator” in the eval process (28:07) Theoretical saturation: when to stop (31:39) Using axial codes to help categorize and synthesize error notes (44:39) The results (46:06) Building an LLM-as-judge to evaluate specific failure modes (48:31) The difference between code-based evals and LLM-as-judge (52:10) Example: LLM-as-judge (54:45) Testing your LLM judge against human judgment (01:00:51) Why evals are the new PRDs for AI products (01:05:09) How many evals you actually need (01:07:41) What comes after evals (01:09:57) The great evals debate (1:15:15) Why dogfooding isn’t enough for most AI products (01:18:23) OpenAI’s Statsig acquisition (1:23:02) The Claude Code controversy and the importance of context (01:24:13) Common misconceptions around evals (1:22:28) Tips and tricks for implementing evals effectively (1:30:37) The time investment (1:33:38) Overview of their comprehensive evals course (1:37:57) Lightning round and final thoughts *LLM Log Open Codes Analysis Prompt:* _Please analyze the following CSV file. There is a metadata field which has an nested field called z_note that contains open codes for analysis of LLM logs that we are conducting. Please extract all of the different open codes. From the _note field, propose 5-6 categories that we can create axial codes from._ *Referenced:* • Building eval systems that improve your AI product: https://www.lennysnewsletter.com/p/building-eval-systems-that-improve • Mercor: https://mercor.com/ • Brendan Foody on LinkedIn: https://www.linkedin.com/in/brendan-foody-2995ab10b • Nurture Boss: https://nurtureboss.io/ • Braintrust: https://www.braintrust.dev/ • Andrew Ng on X: https://x.com/andrewyng • Carrying Out Error Analysis: https://www.youtube.com/watch?v=JoAxZsdw_3w • Julius AI: https://julius.ai/ • Brendan Foody on X—“evals are the new PRDs”: https://x.com/BrendanFoody/status/1939764763485171948 ...References continued at: https://www.lennysnewsletter.com/p/why-ai-evals-are-the-hottest-new-skill *Recommended books:* • Pachinko: https://www.amazon.com/Pachinko-National-Book-Award-Finalist/dp/1455563935 • Apple in China: The Capture of the World’s Greatest Company: https://www.amazon.com/Apple-China-Capture-Greatest-Company/dp/1668053373/ • Machine Learning: https://www.amazon.com/Machine-Learning-Tom-M-Mitchell/dp/1259096955 • Artificial Intelligence: A Modern Approach: https://www.amazon.com/Artificial-Intelligence-Modern-Approach-Global/dp/1292401133/ _Production and marketing by https://penname.co/._ _For inquiries about sponsoring the podcast, email podcast@lennyrachitsky.com._ Lenny may be an investor in the companies discussed.

Segments Timeline

1
0:00 - 0:20
0:20 duration93 words

The Power of Evals in AI Development

Hamel Husain introduces the concept of evals as a crucial skill for building AI products. He emphasizes that evals are the highest ROI activity for product builders, highlighting their addictive nature and the potential for significant learning. The goal is to improve products through actionable insights rather than achieving perfection.

"To build great AI products, you need to be really good at building evals. It's the highest ROI activity you can engage in. This process is a lot of fun. Everyone that does this immediately gets addict..."

2
0:20 - 0:39
0:18 duration82 words

Common Misconceptions About Evals

Hamel discusses the controversies surrounding evals, noting that many people have strong opinions due to past negative experiences. He addresses the misconception that AI can effectively conduct evals, emphasizing the need for human oversight in the process.

"It's to actionably improve your product. >> I did not realize how much controversy and drama there is around eval. There's a lot of people with very strong opinions. People have been burned by evals i..."

3
0:39 - 1:05
0:26 duration104 words

The Benevolent Dictator Concept

Shreya Shankar introduces the idea of a 'benevolent dictator' in the eval process, suggesting that appointing a single trusted individual with domain expertise can streamline the process. This approach avoids the inefficiencies of committee-based decision-making.

"The top one is we live in the age of AI. Can't the AI just eval it? But it doesn't work. A term that you used in your post that I love is this idea of a benevolent dictator. When you're doing this ope..."

4
1:05 - 1:44
0:38 duration128 words

The Rise of Evals in AI Product Building

The hosts discuss the growing importance of evals in AI product development, referencing insights from chief product officers at leading AI companies. They highlight how evals have transitioned from an obscure topic to a vital skill for product managers and engineers.

"Oftentimes it is the product manager. Today my guests are Hamill Hussein and Shrea Shankar. One of the most trending topics on this podcast over the past year has been the rise of evals. Both the chie..."

5
1:44 - 2:28
0:44 duration146 words

Understanding Evals: A Primer

Hamel and Shreya explain what evals are, describing them as systematic methods for measuring and improving AI applications. They emphasize that evals should not be intimidating and can be approached through data analytics to create actionable metrics.

"from being an obscure mysterious subject to one of the most necessary skills for AI product builders. They teach the definitive online course on evals, which happens to be the number one course on Mav..."

6
2:28 - 3:00
0:31 duration108 words

Real Estate Assistant Example

Hamel provides a concrete example of evals in action using a real estate assistant application. He explains how evals help identify errors and improve the assistant's performance, moving beyond guesswork to a structured approach for enhancement.

"This episode is the deepest yet most understandable primer you will find on the world of evals and honestly got me excited to write evals. Even though I have nothing to write eels for, I think you'll ..."

7
3:00 - 3:39
0:39 duration121 words

The Importance of Error Analysis

The conversation shifts to the significance of error analysis in the eval process. Hamel discusses how analyzing data from AI applications can reveal issues and guide improvements, making the case for a systematic approach to understanding performance.

"number one AI agent for customer service. If your customer support tickets are piling up, then you need Finn. Finn is the highest performing AI agent on the market with a 65% average resolution rate. ..."

8
3:39 - 4:02
0:23 duration75 words

Introducing Nurture Boss

Hamel introduces Nurture Boss, an AI assistant for property managers, as a case study for evals. He outlines the various tasks the assistant handles and sets the stage for a deeper exploration of how evals can be applied to improve its functionality.

"by the Finn AI engine, which is a continuously improving system that allows you to analyze, train, test, and deploy with ease. Finn can continuously improve your results, too. So, if you're ready to t..."

9
14:03 - 15:01
0:57 duration182 words

Understanding AI Traces

In this segment, Hamel Husain explains the concept of 'traces' in AI applications, detailing how they log sequences of events crucial for understanding AI behavior. He emphasizes the importance of system prompts and how they guide AI assistants in responding to user inquiries, using a real example from a property management AI assistant.

"of events. It's been a the concept of a trace has been around for a really long time but it's especially really important when it comes to AI applications. So we have all the different components and ..."

10
15:01 - 16:03
1:02 duration166 words

Analyzing User Interactions

Hamel discusses the significance of analyzing user interactions with AI systems. He illustrates how an AI assistant responds to a user's inquiry about apartment availability, highlighting the importance of accurate and helpful responses in enhancing user experience and product effectiveness.

">> That's amazing because that's really it's rare you see actual company products system prompt. That's like their crown jewels a lot of times. So this is actually very cool on its own. >> Yeah. Yeah...."

11
16:03 - 17:10
1:06 duration183 words

The Challenge of Data Analysis

This segment addresses the challenges of analyzing messy data logs from AI applications. Hamel introduces the concept of error analysis as a systematic approach to managing and interpreting data, emphasizing the need for product managers to engage in this process to improve AI performance.

"AI responds, hey, we we have several one-bedroom apartments available, but none specifically listed with a study. Here are a few options. Uh, and then it says, "Can you let me know when one with a stu..."

12
17:10 - 18:01
0:51 duration160 words

The Role of Product Managers in Error Analysis

Hamel highlights the critical role of product managers in the error analysis process. He explains that product managers must be involved in reviewing AI interactions to ensure the system meets user expectations and provides valuable insights for improvement.

"is completely manageable and it's not something that we invented. It's been around in machine learning and data science for a really long time and it's called error analysis. And what you do is the fi..."

13
18:01 - 19:10
1:09 duration189 words

Writing Effective Notes

In this segment, Hamel discusses the importance of writing effective notes during error analysis. He emphasizes that capturing the first observed error is crucial for understanding AI performance and improving the overall user experience.

"conversation. Okay. A user asked about availability. The AI said, "Oh, we don't really have that. Have a nice day." Now, for a product that is helping you with lead management, is that good? Like, do ..."

14
19:10 - 20:02
0:51 duration176 words

Sampling Data for Insights

Hamel explains the strategy of sampling data to gain insights without being overwhelmed. He encourages product teams to focus on a manageable number of interactions to identify patterns and areas for improvement in AI applications.

"quick note um should you know should have handed off to a human >> and as we watch this happening it's like you mentioned this and you'll explain more you're doing this this feels very ma manual and u..."

15
20:02 - 21:19
1:17 duration235 words

Identifying Upstream Errors

This segment focuses on the process of identifying upstream errors in AI interactions. Hamel advises that teams should document the most significant issues first, allowing for a more efficient and effective error analysis process.

"So, this is the next trace. I just pushed a hot key on my keyboard. Let me go back to uh looking at it. >> And these tools make it easy to go through a bunch and add these notes quickly. >> Yes. And s..."

16
21:19 - 22:30
1:10 duration227 words

Handling Text Message Interactions

Hamel discusses the challenges of handling text message interactions in AI applications. He illustrates how garbled messages can lead to misunderstandings and emphasizes the need for clear communication in AI responses.

">> This is more of hey we're not handling this interaction correctly. This is more of a technical problem. um rather than hey the AI is not doing exactly what we want. So we would write down too like ..."

17
22:30 - 23:45
1:14 duration210 words

Documenting AI Errors

In this segment, Hamel explains how to document AI errors effectively. He stresses the importance of clarity in notes to ensure that future analyses can accurately reflect the issues encountered during user interactions.

"first two or three can be very painful, but you know, it doesn't we can, you know, do a bunch of them really fast. So, here's another one. And um let's skip the system prompt again. And the user asks,..."

18
23:45 - 24:52
1:07 duration223 words

The Limitations of LLMs in Error Analysis

Hamel addresses the limitations of using large language models (LLMs) for error analysis. He explains that LLMs often lack the contextual understanding necessary to identify product-specific issues, making manual analysis essential.

"different kinds of errors that we're seeing and we're actually learning a lot about your application um in a very short amount of time. >> One common question that we get from people at this stage is ..."

19
24:52 - 26:39
1:47 duration338 words

The Benevolent Dictator Concept

Hamel introduces the concept of a 'benevolent dictator' in the error analysis process. He explains that appointing a knowledgeable individual to guide the analysis can streamline the process and enhance decision-making efficiency.

"are like, let me automate this with an LLM. >> Do you think they'll we'll get to a place where where an agent can do this? >> Oh, no, no, no. Sorry. There are parts of error analysis that an LLM is su..."

20
26:39 - 28:04
1:25 duration286 words

Choosing the Right Expert for Analysis

In this segment, Hamel discusses the importance of selecting the right expert for conducting error analysis. He emphasizes that the chosen individual should have domain expertise to ensure accurate assessments and effective improvements.

"hey, you need to simplify this across as many dimensions as you can. Another thing that we'll talk about later is when you goes to building an LLM as a judge, you need a binary score. You don't want t..."

21
28:04 - 30:26
2:21 duration404 words

Setting Goals for Error Analysis

Hamel concludes this segment by discussing the importance of setting clear goals for error analysis. He suggests that teams should aim to analyze a specific number of interactions to gain valuable insights and improve AI performance.

"become infinitely expensive if you're not careful. >> Yeah. Okay, cool. Let's go back to your examples. Yeah, no problem. So this is another example where we have someone saying, "Okay, do you have an..."

22
30:29 - 31:36
1:07 duration224 words

Leveraging AI for Error Categorization

This segment covers how to utilize AI tools to categorize notes from error analysis. Hamel emphasizes the importance of human oversight in this process, ensuring that the AI's categorization aligns with the actual issues encountered. The discussion highlights the balance between automation and human judgment in improving AI products.

"should talk about >> yeah so there's actually a term in data analysis and quant qualitative analysis called theoretical saturation So what this means is when you do all of these processes of looking a..."

23
31:36 - 32:43
1:07 duration203 words

Basic Counting as a Powerful Tool

Hamel explains the significance of basic counting in data analysis, describing it as a simple yet powerful technique. He discusses how to categorize notes using AI and the importance of clear, actionable open codes to facilitate effective categorization and analysis.

"you're like not going to discover new types of problems. >> Yeah. Awesome. So let's say you did a 100 of these. What's the next step? >> Yeah. Okay. So you did 100 of these. Now you have all these not..."

24
32:43 - 34:30
1:46 duration321 words

Creating Axial Codes from Open Codes

In this segment, Hamel and Shreya delve into the process of creating axial codes from open codes. They explain how to synthesize error notes into actionable categories, emphasizing the need for specificity in labeling failure modes to enhance the effectiveness of the analysis.

"different uh coding agents or you know uh AI tools and had it categorize these notes. So one is okay, I uploaded into a cloud project. I uploaded a CSV of these notes and I just exported them directly..."

25
34:30 - 36:01
1:30 duration308 words

Synthesizing Categories for Improvement

Hamel discusses the importance of refining axial codes to make them more actionable. He shares examples of how to improve generic categories into specific issues that can be addressed, highlighting the iterative nature of this process and the role of AI in facilitating categorization.

"what is the most prevalent. So then you can go and run and attack that problem. >> That is really helpful. Basically, you're just synthesizing all these categories and themes. >> Super cool. And we'll..."

26
36:01 - 37:10
1:08 duration255 words

The Role of AI in Error Analysis

This segment emphasizes the role of AI in synthesizing information from error analysis. Hamel and Shreya discuss how AI can help categorize problems and the importance of detailed open codes for effective AI categorization, reinforcing the need for clarity in documentation.

">> But, you know, now we can wade through this mess of open codes a lot easier. Another thing that's interesting here in this prompt to generate the axial codes is you can be very detailed if you want..."

27
37:10 - 38:02
0:52 duration171 words

Grounding Techniques in Established Theory

Hamel and Shreya clarify that their approach to error analysis is grounded in established social science theories, rather than being entirely new. They emphasize the importance of using proven techniques while adapting them to the context of AI and LLMs.

"demystify some of this magic. >> And just to double click on this point, like this is not a thing everyone does or knows. This is something you two developed based on your experience doing data analys..."

28
38:02 - 39:10
1:08 duration248 words

The Value of Iteration in AI Analysis

In this segment, Hamel and Shreya discuss the iterative process of refining error analysis techniques. They stress the importance of continuously evaluating and improving open codes and axial codes to ensure they remain relevant and actionable in the context of AI product development.

"social science. >> Amazing. Okay. Like what's funny about you guys doing this is I just want to go do this somewhere. I don't have I don't have an AI product to do this on, but it's just like oh this ..."

29
41:31 - 42:43
1:12 duration230 words

Categorizing Open Codes with AI

In this segment, Hamel and Shreya discuss the process of categorizing open codes using AI tools like Gemini. They emphasize the importance of detailed open codes for effective categorization and the need for human oversight in the AI's suggestions. This segment highlights the significance of clarity in coding to enhance the AI's ability to categorize effectively.

"collect these codes into a list. So now we have a commaepparated list of these codes. And then what you can simply do is you could take your notes that you have those open codes and you can tell an AI..."

30
42:43 - 44:00
1:17 duration286 words

Iterating on Open Codes

Hamel and Shreya explain the iterative process of refining open codes and axial codes. They introduce the concept of a 'none of the above' category to identify gaps in categorization. This segment underscores the importance of continuous improvement in coding practices to ensure comprehensive error analysis.

"somewhat detailed in your open code. >> Okay. So avoid the word janky is a good rule of thumb >> or other words. >> Okay. >> I was being funny. >> Yeah. Okay. What are some of those other words just t..."

31
44:00 - 45:36
1:36 duration335 words

From Chaos to Clarity: Analyzing Errors

This segment focuses on the transition from chaotic error data to structured insights using pivot tables. Hamel and Shreya illustrate how categorizing errors helps identify major issues in AI interactions, emphasizing the need for targeted evaluations based on the categorized data. They discuss the balance between obvious errors and those requiring deeper analysis.

">> Awesome. And what's cool about this is you don't need to do this many many times. Like no, >> for most products, you do this process once and then you build on it, I imagine, and you just tweak it ..."

32
45:36 - 46:56
1:20 duration246 words

Evaluating Human Handoff Issues

Hamel and Shreya delve into the complexities of human handoff issues in AI systems. They discuss the subjective nature of these issues and the role of LLMs as judges in evaluating such scenarios. This segment highlights the importance of understanding when to implement evaluations and the potential pitfalls of rushing into testing without proper analysis.

"gone from chaos to some kind of thinking around oh you know what these are my biggest problems I need to fix conversational issues you know maybe these human handoff issues it's not necessarily the co..."

33
46:56 - 48:03
1:06 duration188 words

Code-Based vs. LLM-as-Judge Evaluations

In this segment, the discussion shifts to the differences between code-based evaluations and LLM-as-judge evaluations. Hamel and Shreya explain the advantages of automated evaluations for straightforward failure modes and the necessity of LLMs for more complex scenarios. They emphasize the importance of defining clear evaluation criteria to ensure effective testing.

"want to get carried away with it. Um, but you want to start, you want to usually ground yourself in your actual errors. You don't want to skip this step. And so, the reason I'm kind of spending so muc..."

34
48:03 - 49:56
1:53 duration387 words

Building an LLM as a Judge

Hamel and Shreya provide insights into constructing an LLM as a judge for specific failure modes. They discuss the importance of creating binary evaluations to simplify decision-making and the need for precise prompts. This segment illustrates the practical steps involved in developing an effective LLM evaluation system.

"judge, for example. So, there's different kinds of evals. One is codebased which you should try to do if you can because they're cheaper. You don't have to, you know, LM as a judge is something it's l..."

35
49:56 - 51:28
1:31 duration304 words

Testing and Monitoring with LLM Judges

This segment covers the application of LLM judges in both pre-production testing and real-time monitoring of AI systems. Hamel and Shreya highlight the benefits of using LLMs to assess application quality and the importance of continuous evaluation in production environments. They stress the need for reliable metrics to maintain trust in the evaluation process.

"describing, you want to test what your say agent or AI product is doing. You ask it a question, it gets back with something. One way to test if it's giving you the right answer is if it's consistently..."

36
51:28 - 52:44
1:16 duration266 words

The Importance of Clear Evaluation Metrics

Hamel and Shreya discuss the significance of establishing clear and actionable evaluation metrics for AI products. They caution against vague scoring systems and advocate for binary evaluations to enhance clarity and trust in the evaluation results. This segment emphasizes the need for transparency in the evaluation process to foster confidence among stakeholders.

"thing about LLM judge is you can use them in unit test or CI sure but you could also use it online for monitoring right like I can sample like thousand traces every day run my LLM judge real productio..."

37
52:44 - 54:03
1:18 duration241 words

Creating Effective Judge Prompts

In this segment, Hamel shares practical tips for crafting effective judge prompts for LLM evaluations. He emphasizes the importance of iterating on prompts and ensuring they align with the evaluation goals. This segment provides valuable insights into the nuances of prompt design to enhance the accuracy of LLM evaluations.

"decision. Like no, you need to make a decision. Is this good enough or not? Yes or no? can be painful to think about what that is, but you should absolutely do it. Otherwise, this thing becomes very u..."

38
54:52 - 56:17
1:25 duration288 words

Iterating on LLM Prompts

In this segment, Hamel and Shreya discuss the importance of iterating on prompts when using LLMs as judges in AI evals. They emphasize that while LLMs can assist in creating prompts, it's crucial to review and edit the outputs to ensure accuracy and relevance. The conversation highlights the need for human oversight to maintain trust in the evaluation process.

"create it, but again, put yourself in the loop. Don't just blindly accept what the LLM does. And in all of these cases, that's what we did. Like with the axial codes, we kind of iterated on this. You ..."

39
56:17 - 57:00
0:43 duration148 words

Aligning LLM Judges with Human Evaluation

Hamel and Shreya explain how to ensure that LLM judges align with human evaluations. They discuss the use of axial codes to measure agreement between the LLM's judgments and human assessments, stressing the importance of this alignment to maintain trust in the eval process. The segment underscores the potential pitfalls of relying solely on LLM outputs without proper validation.

"so we can take this prompt and I'm going to use a spreadsheet again. So the first step is okay when I'm doing this judge I wrote the prompt. Now a lot of people stop there and they say okay I have my ..."

40
57:00 - 1:00:21
3:20 duration597 words

Understanding Agreement Metrics

This segment delves into the complexities of using agreement metrics in AI evals. Hamel warns against the misleading nature of high agreement percentages, especially when errors are infrequent. They discuss the importance of analyzing specific types of errors to improve the evaluation process and ensure that the LLM judge is effectively capturing the nuances of human judgment.

"as a judge, you want to make sure it's aligned to the human. So how do you do that is you actually you have those axial codes and you want to like measure your judge against the axial code and say lik..."

41
1:00:21 - 1:02:58
2:37 duration542 words

Evals as the New PRDs

Hamel and Shreya explore the concept of evals as the new Product Requirements Documents (PRDs) for AI products. They argue that evals provide a continuous feedback loop for product performance, allowing teams to refine their expectations based on real data. The discussion emphasizes the evolving nature of product development and the necessity of adapting PRDs to incorporate insights gained from evals.

"have some misalignment that's okay we talk about in our course also how to code correct that misalignment but in this stage if you're a product manager and the person who's building the LLM judge eval..."

42
1:02:58 - 1:06:07
3:08 duration623 words

Prioritizing Evals in Development

In this segment, the speakers discuss the importance of prioritizing which evals to implement based on the complexity of the product and the potential risks involved. They highlight that not every issue requires an eval, and teams should focus on the most critical failure modes that could impact the business. This approach ensures efficient use of resources while maintaining product quality.

">> I love that. And Haml's pulling up some cool research report. What's this about? >> Oh, this is one of the coolest research reports you can possibly read if you want to know about evals. So, it was..."

43
1:06:07 - 1:09:46
3:39 duration694 words

Leveraging LLM Judges for Continuous Improvement

Hamel and Shreya explain how to effectively use LLM judges for ongoing product improvement. They discuss integrating LLM judges into unit tests and online monitoring systems to ensure consistent performance. The segment emphasizes the value of using these tools to drive product enhancements and maintain a competitive edge in the market.

">> and probably the ones that are most risky to your business if they say something like Mecca Hitler Grock and >> cool okay so that's that's very uh relieving that this because this is this prompt is..."

44
1:09:46 - 1:10:25
0:38 duration130 words

The Evals Debate

The conversation shifts to the ongoing debate surrounding the value and importance of evals in AI development. Hamel and Shreya provide insights into the differing opinions within the community, highlighting the necessity of understanding both sides of the argument. They stress that despite the controversy, the ultimate goal remains the same: improving AI products through effective evaluation.

"Um and I think what's very empowering now is that product managers are doing this and can do this and can really build very very profitable products with this skill set. >> Okay, great segue to a deba..."

45
1:11:36 - 1:12:53
1:17 duration251 words

The Role of Dogfooding in AI Products

The discussion highlights the concept of dogfooding in AI development, particularly for coding agents. Shreya explains how developers' close interaction with their products can lead to effective error analysis and improvement, contrasting this with other domains where such practices may not be feasible.

"then unfortunately X or Twitter is like a medium where you know people are misinterpreting what everybody is saying all the time and you just get all these strong opinions of like don't do eval it's b..."

46
1:12:53 - 1:14:31
1:37 duration281 words

Evaluating the Evals vs. A/B Testing Debate

Hamel and Shreya delve into the debate between evals and A/B testing, discussing how both are essential for understanding product performance. They argue that A/B tests should be informed by thorough error analysis to be effective, emphasizing the need for a systematic approach.

"probably very systematic about the error analysis to some extent. I bet you that they are monitoring who is using Claude, how many people are using Claude, how many chats are being created, how long t..."

47
1:14:31 - 1:16:02
1:30 duration280 words

Data Science Thinking in AI Products

The segment focuses on the necessity of data science thinking in AI product development. Shreya argues that understanding data and conducting thorough analysis is crucial for improving applications, and that the term 'eval' should not be seen as a new concept but rather an extension of existing data science practices.

">> because you're seeing the code like you see the code it's generating you can tell this is great this is terrible >> yeah yeah and so and so I think a lot of people had generalized coding agents bec..."

48
1:16:02 - 1:17:50
1:48 duration367 words

The Future of Evals and OpenAI's Acquisition

Shreya shares her thoughts on OpenAI's acquisition of Statsig, discussing its implications for the future of evals and A/B testing. She expresses hope that this acquisition will lead to a greater focus on product-specific evaluations and improved methodologies in AI development.

"now. >> Do fooding is is a dangerous one only because a lot of people will say they're dog fooding. They're like, "Yeah, we dog fooded." But are they really? And a lot of people aren't really dog food..."

49
1:17:50 - 1:20:00
2:10 duration414 words

The Need for Structured Evaluation Processes

Hamel and Shreya conclude by stressing the importance of structured evaluation processes in AI development. They express their desire for more practitioners to adopt systematic approaches to evals, highlighting the demand for their course and the growing interest in effective evaluation methodologies.

"thought the errors might be. They were these like weird handoff issues or like I don't know like the text message thing was strange. Um, so I would say that like if you're going to do AB tests and the..."

50
1:22:13 - 1:23:03
0:50 duration161 words

The Need for Structured Evals

Hamel and Shreya discuss the limitations of existing AI eval tools that lack proper error analysis. They emphasize the importance of structured thinking in developing application-specific evals and express their hope for broader adoption of these methods in the industry.

"up until recently that some of the big labs have they don't have error analysis. They have gener a suite of generic tools cosign similarity hallucination score whatever and that doesn't work. It's a g..."

51
1:23:03 - 1:24:07
1:04 duration211 words

The Claude vs. Codex Debate

The conversation shifts to the ongoing debate about AI coding agents, particularly Claude and Codex. Hamel and Shreya highlight the inconsistency in public opinion regarding these tools and share anecdotes from their discussions about which AI performs better.

">> Well, the fact that your course on Maven is the number one highest grossing course on Maven, clearly there's demand and interest and there's more people I think on your side. Interestingly, uh, jus..."

52
1:24:07 - 1:25:00
0:52 duration172 words

Common Misconceptions About Evals

Hamel outlines common misconceptions surrounding AI evals, particularly the belief that AI can fully automate the eval process. He stresses the necessity of human involvement and the importance of data analysis in creating effective evals.

"I was like, "Oh my god, >> so true. >> This is the world we live in." Oh my god. >> Okay. So, I want to ask about just top misconceptions people have with Evals and top tips and tricks for being succe..."

53
1:25:00 - 1:26:28
1:28 duration264 words

The Importance of Data Analysis

Shreya shares insights from her consulting experience, emphasizing the power of data analysis in identifying problems. She encourages listeners to engage with their data actively and highlights the learning opportunities that come from examining individual traces.

">> The second one that you know I see a lot is hey um just not looking at the data you know. So in my consulting people come to me with problems all the time and the first thing I'll say is let's go l..."

54
1:26:28 - 1:27:14
0:46 duration157 words

Tips for Successful Evals

Hamel and Shreya provide practical tips for conducting successful evals, including the importance of not being intimidated by data analysis and leveraging AI tools to enhance the eval process. They stress that the goal is actionable improvement rather than perfection.

">> Amazing. Okay. What are a couple just tips and tricks you want to leave people with as they start on their eval journey or just try to get better at something they're already doing? >> So tip numbe..."

55
1:27:14 - 1:28:49
1:34 duration273 words

Creating Custom Tools for Evals

Hamel discusses the benefits of creating custom tools for data analysis in the eval process. He explains how AI can facilitate the development of these tools, making it easier for product builders to analyze data and improve their applications.

"would say is we're very pro- AI. Use LLM to help you organize any thoughts that you have throughout this entire process. So this could be everything ranging from like initial product requirements, rig..."

56
1:28:49 - 1:30:10
1:21 duration236 words

Maximizing ROI Through Evals

The duo emphasizes that the ultimate goal of conducting evals is to enhance product quality and user experience. They discuss how effective evals can lead to significant improvements in AI products, ultimately benefiting businesses.

"the nurture boss use case, we wanted to remove all the friction of looking at data. And so what you see here is just some screenshots of uh what the application that they created looks like. It's just..."

57
1:30:10 - 1:31:47
1:36 duration337 words

Time Investment in Evals

Hamel shares his experience regarding the time investment required for initial error analysis and ongoing evals. He reassures listeners that while the upfront time commitment may seem daunting, the ongoing maintenance is manageable and yields high returns.

"AI products better because the experience is how users interact with your AI. Absolutely. If any, you know, we teach our students, hey, when you're doing these evals, if you see something that's wrong..."

58
1:31:47 - 1:34:05
2:18 duration481 words

Course Insights and Perks

Hamel and Shreya provide an overview of their comprehensive evals course, detailing the syllabus and unique perks for students. They highlight the importance of structured learning and the resources available to help students excel in AI evals.

">> Yeah, it's really not that much time. I think people just get overwhelmed by how much time they spend up front and then thinking that they have to keep doing this all the time. >> Amazing. Is there..."

59
1:36:03 - 1:36:50
0:47 duration127 words

AI-Powered Learning: The Future of Education

Hamel and Shreya discuss the integration of AI into education, emphasizing that learning shouldn't be limited to lectures and assignments. They introduce their AI tool, which compiles all course materials and resources, providing students with 10 months of free access to enhance their learning experience.

"Education shouldn't be this thing where you you're only watching lectures and doing homework assignments. So students should have access to an AI that also helps them. So what we've done is we've uh y..."

60
1:36:50 - 1:37:43
0:52 duration222 words

The Lightning Round Begins!

The conversation shifts to a fun lightning round where Hamel and Shreya answer rapid-fire questions. They share their favorite books, with Shreya recommending 'Pachinko' and 'Apple in China,' while Hamel highlights classic textbooks on machine learning and algorithms.

"students 10 months free unlimited access to that alongside the course. >> Amazing. And then you'll charge for that later down the road is the idea. I just take one month at a time. I don't know what w..."

61
1:37:43 - 1:38:31
0:47 duration145 words

Favorite Movies and TV Shows

Hamel and Shreya share their favorite recent movies and TV shows. Hamel humorously mentions watching 'Frozen' multiple times with his kids, while Shreya discusses her newfound appreciation for 'The Wire,' reflecting on its depth and storytelling.

">> Amazing. Okay. We also have a discord >> of all the students who have ever taken the class and that discord is so active >> I I can't go on vacation without getting notified on the plane or >> bitt..."

62
1:38:31 - 1:39:22
0:51 duration146 words

Discovering New Products

In this segment, Shreya talks about her enthusiasm for AI-assisted coding tools like Cursor and Cloud Code, explaining how they enhance her productivity as a researcher. Hamel agrees, praising the user experience of Cloud Code and its impressive design.

"reading Apple in China which the name of the author is slipping my mind but this is kind of more of an exposition written by a journalist on how Apple did a lot of manufacturing processes in Asia over..."

63
1:39:22 - 1:40:01
0:38 duration124 words

Life Mottos That Inspire

Hamel shares his life motto of 'Keep learning and think like a beginner,' while Shreya emphasizes the importance of understanding opposing viewpoints in discussions, particularly in the context of eval debates. They both reflect on the collaborative spirit of their work.

"simplest and also engineering. So like a lot of times the simpler approach generalizes better. And so that's the thing I kind of internalized deeply from that book. And um also really like this one. S..."

64
1:40:01 - 1:41:03
1:02 duration222 words

Compliments and Reflections

Hamel and Shreya take a moment to compliment each other. Hamel admires Shreya's wisdom and grounded perspective, while Shreya appreciates Hamel's energy and enthusiasm for evals, highlighting the importance of their partnership in the AI space.

"Okay, next question. Favorite recent movie or TV show? I'll jump to Haml first. Okay. So, I'm a dad of two parents. I have two parents. >> I don't get to Oh, sorry. Uh, two kids. So, yeah, I'm a dad o..."

65
1:41:03 - 1:46:08
5:04 duration886 words

Connecting with the Audience

As the episode wraps up, Hamel and Shreya share how listeners can connect with them and access their AI eval course. They encourage audience engagement through questions and success stories, emphasizing their desire to foster a community around AI product development.

">> Worth it. Okay, next question. Do you have a favorite product you've recently discovered that you really love? And we'll start with Shrea. >> Yeah, I really like using cursor. Honestly, now Cloud C..."