searchlore

Back to Resource

All Segments

The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon)

The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon)

53 segments available

Ronny Kohavi, PhD, is a consultant, teacher, and leading expert on the art and science of A/B testing. Previously, Ronny was Vice President and Technical Fellow at Airbnb, Technical Fellow and corporate VP at Microsoft (where he led the Experimentation Platform team), and Director of Data Mining and Personalization at Amazon. He was also honored with a lifetime achievement award by the Experimentation Culture Awards in September 2020 and teaches a popular course on experimentation on Maven. In today’s podcast, we discuss: • How to foster a culture of experimentation • How to avoid common pitfalls and misconceptions when running experiments • His most surprising experiment results • The critical role of trust in running successful experiments • When not to A/B test something • Best practices for helping your tests run faster • The future of experimentation Enroll in Ronny’s Maven class, Accelerating Innovation with A/B Testing, at https://bit.ly/ABClassLenny. Promo code “LENNYAB” will give $500 off the class for the first 10 people to use it. — Brought to you by Mixpanel—Event analytics that everyone can trust, use, and afford: https://mixpanel.com/startups | Round—The private network built by tech leaders for tech leaders: https://www.round.tech/apply?utm_campaign=lennys-letter&utm_medium=email-ad&utm_source=email-marketing&utm_content=send-2-2023-07-27 | Eppo—Run reliable, impactful experiments: https://www.geteppo.com/ Find the full transcript at: https://www.lennysnewsletter.com/p/the-ultimate-guide-to-ab-testing Where to find Ronny Kohavi: • Twitter: https://twitter.com/ronnyk • LinkedIn: https://www.linkedin.com/in/ronnyk/ • Website: http://ai.stanford.edu/~ronnyk/ Where to find Lenny: • Newsletter: https://www.lennysnewsletter.com • Twitter: https://twitter.com/lennysan • LinkedIn: https://www.linkedin.com/in/lennyrachitsky/ In this episode, we cover: (00:00) Ronny’s background (04:29) How one A/B test helped Bing increase revenue by 12% (09:00) What data says about opening new tabs (10:34) Small effort, huge gains vs. incremental improvements  (13:16) Typical fail rates (15:28) UI resources (16:53) Institutional learning and the importance of documentation and sharing results (20:44) Testing incrementally and acting on high-risk, high-reward ideas (22:38) A failed experiment at Bing on integration with social apps (24:47) When not to A/B test something (27:59) Overall evaluation criterion (OEC) (32:41) Long-term experimentation vs. models (36:29) The problem with redesigns (39:31) How Ronny implemented testing at Microsoft (42:54) The stats on redesigns  (45:38) Testing at Airbnb (48:06) Covid’s impact and why testing is more important during times of upheaval  (50:06) Ronny’s book, Trustworthy Online Controlled Experiments: A Practical Guide to A/B Testing (51:45) The importance of trust (55:25) Sample ratio mismatch and other signs your experiment is flawed (1:00:44) Twyman’s law (1:02:14) P-value (1:06:27) Getting started running experiments (1:07:43) How to shift the culture in an org to push for more testing (1:10:18) Building platforms (1:12:25) How to improve speed when running experiments (1:14:09) Lightning round Referenced: • Trustworthy Online Controlled Experiments: A Practical Guide to A/B Testing: https://experimentguide.com/ • Seven rules of thumb for website experimenters: https://exp-platform.com/rules-of-thumb/ • GoodUI: https://goodui.org • Defaults for A/B testing: http://bit.ly/CH2022Kohavi • Ronny’s LinkedIn post about A/B testing for startups: https://www.linkedin.com/posts/ronnyk_abtesting-experimentguide-statisticalpower-activity-6982142843297423360-Bc2U • Sanchan Saxena on Lenny’s Podcast: https://www.lennyspodcast.com/sanchan-saxena-vp-of-product-at-coinbase-on-the-inside-story-of-how-airbnb-made-it-through-covid-what-he8217s-learned-from-brian-chesky-brian-armstrong-and-kevin-systrom-much-more/ • Optimizely: https://www.optimizely.com/ • Optimizely was statistically naive: https://analythical.com/blog/optimizely-got-me-fired • SRM: https://www.linkedin.com/posts/ronnyk_seat-belt-wikipedia-activity-6917959519310401536-jV97 • SRM checker: http://bit.ly/srmCheck • Twyman’s law: http://bit.ly/twymanLaw • “What’s a p-value” question: http://bit.ly/ABTestingIntuitionBusters • Fisher’s method: https://en.wikipedia.org/wiki/Fisher%27s_method • Evolving experimentation: https://exp-platform.com/Documents/2017-05%20ICSE2017_EvolutionOfExP.pdf • CUPED for variance reduction/increased sensitivity: http://bit.ly/expCUPED • Ronny’s recommended books: https://bit.ly/BestBooksRonnyk • Chernobyl on HBO: https://www.hbo.com/chernobyl • Blink cameras: https://blinkforhome.com/ • Narrative not PowerPoint: https://exp-platform.com/narrative-not-powerpoint/ Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email podcast@lennyrachitsky.com. Lenny may be an investor in the companies discussed.

Segments Timeline

1
0:00 - 1:04
1:04 duration168 words

The Power of Testing Everything

Ronny Kohavi emphasizes the importance of testing every code change and feature introduction through A/B testing. He shares insights on how even small bug fixes can lead to unexpected impacts, advocating for a culture that embraces high-risk, high-reward experimentation. Kohavi cautions that while many ideas may fail, the potential for significant wins makes testing essential.

"I'm very clear that I'm a big fan of test everything which is any code change that you make any feature that you introduce has to be in some experiment because again I've observed this sort of surpris..."

2
1:04 - 2:36
1:31 duration270 words

Ronny Kohavi: A/B Testing Expert

In this segment, Ronny Kohavi introduces himself as a leading expert in A/B testing and experimentation, detailing his impressive background at companies like Airbnb, Microsoft, and Amazon. He discusses his book, 'Trustworthy Online Controlled Experiments,' and sets the stage for a tactical conversation about implementing effective A/B testing strategies in organizations.

"welcome to Lenny's podcast where I interview world-class product leaders and growth experts to learn from their hard-winning experiences building and growing today's most successful products today my ..."

3
2:36 - 5:01
2:25 duration407 words

Surprising A/B Test Results

Kohavi shares a surprising A/B test conducted at Bing that involved changing the display of ads. By moving the second line of an ad to the first line, Bing experienced a 12% increase in revenue. This segment highlights the unexpected outcomes of seemingly trivial changes and the importance of validating results through repeated experiments.

"after a short word from our sponsors this episode is brought to you by mixpanel get deep insights into what your users are doing at every stage of the funnel at a fair price that scales as you grow mi..."

4
5:01 - 6:27
1:25 duration245 words

The Impact of Opening New Tabs

Ronny discusses a significant experiment at Airbnb where opening a new tab for search results led to substantial improvements in user engagement. He reflects on the initial skepticism surrounding the idea and emphasizes the importance of institutional memory in retaining successful strategies within a company.

"test you've run or maybe the most surprising result from an A B test that you run yeah so I think the the opening example that I use in my book and in my class is the most surprising public example we..."

5
6:27 - 8:41
2:13 duration399 words

Finding Gold Nuggets in Experiments

Kohavi addresses the rarity of discovering high-impact changes from A/B tests, explaining that most successful outcomes come from incremental improvements rather than major breakthroughs. He shares examples from Bing and Airbnb, illustrating how small, consistent efforts can lead to significant overall gains in revenue.

"that he did he spent a couple of days implementing it and as is you know the common uh thing at Bing he launched the experiment uh and a funny thing happened we had an alarm big escalation something i..."

6
8:41 - 10:34
1:53 duration355 words

Understanding Experiment Failure Rates

In this segment, Ronny reveals the typical failure rates of A/B experiments, citing that around 66% of ideas fail at Microsoft, with even higher rates at Bing and Airbnb. He discusses the humbling nature of these statistics and the importance of learning from failures to improve future experimentation efforts.

"idea in Bing's history and and rated properly right nobody gave this the the priority that in hindsight it deserves and that that's something that happens often I mean we are often humbled by how bad ..."

7
10:34 - 12:59
2:24 duration427 words

Patterns of Success in A/B Testing

Kohavi shares insights on common patterns that lead to successful A/B testing outcomes. He references a paper on rules of thumb and introduces GoodUI.org, a resource that compiles successful experiment results and patterns, helping product teams identify strategies that have historically yielded positive results.

"shout out to Ricardo or mutual friend who helped make this conversation happen there's this like holy grail of experiments that I think people are always looking for of like one you know hour of work ..."

8
14:35 - 15:17
0:42 duration124 words

The Humbling Reality of Experimentation

Ronny Kohavi discusses the high failure rates of A/B experiments, noting that 80-92% of experiments can fail due to implementation issues rather than bad ideas. He emphasizes the importance of iterating and pivoting to achieve successful launches, highlighting the common misconception that new teams will have higher success rates.

"many of them launched not every experiment maps to an idea so it's possible that when you have an idea your first implementation you start an experiment boom it's egregiously bad because you have a bu..."

9
15:17 - 16:08
0:50 duration176 words

Resources for Successful Experiments

Kohavi shares valuable resources for A/B testing, including a paper on rules of thumb from Microsoft and the website GoodUI.org, which compiles patterns from successful experiments. He explains how these resources can help teams identify effective strategies for experimentation.

"very humbling I know that every group that starts to run experiments they always start off by thinking that somehow they're different and their successor is going to be much much higher and they're al..."

10
16:08 - 17:04
0:56 duration167 words

Institutional Learning in Experimentation

Ronny emphasizes the importance of documenting and sharing results from experiments to foster institutional learning. He suggests holding quarterly meetings to review surprising experiments, both successful and unsuccessful, to ensure that teams learn from their experiences.

"um but there's another more more accurate I would say uh resource that's useful that I recommend to people and it's a site called goodui.org and good ui.org is exactly the site that tries to do what y..."

11
17:04 - 18:15
1:10 duration195 words

Understanding Surprising Experiment Results

Kohavi defines a surprising experiment as one where the expected outcome significantly differs from the actual result. He illustrates this with examples, stressing the value of learning from both unexpected successes and failures to inform future experimentation.

"is institutional memory right which is can you document things well enough so that the organization remembers the successes and failures and learns from them I think one of the mistakes that some comp..."

12
18:15 - 19:21
1:06 duration188 words

The Unexpected Consequences of Changes

Ronny recounts an experiment at Microsoft aimed at improving the Windows indexer, which unexpectedly harmed battery life despite improving indexing relevance. He highlights the importance of documenting such outcomes to inform future design iterations.

"negative and that's interesting so we focus not just on the winners but also surprising losers things that people thought would be a no-brainer to run and then for some reason it was very negative and..."

13
19:21 - 20:45
1:23 duration249 words

Documenting Experimentation for Future Success

Kohavi advises on the necessity of documenting experiments and their outcomes to maintain institutional memory. He suggests creating a searchable history of experiments and conducting regular reviews to ensure that valuable insights are not lost.

"next iteration what advice do you have for people to actually remember these surprises you said that a lot of it is institutional what do you recommend people do so that they can actually remember thi..."

14
20:45 - 21:57
1:12 duration195 words

Balancing Experimentation and Innovation

Ronny addresses concerns that excessive experimentation may lead to micro-optimizations at the expense of innovation. He advocates for a balanced approach, encouraging teams to pursue both incremental improvements and high-risk, high-reward ideas.

"this will get segue to something I wanted to touch on which is there's often a I guess a weariness of running too many experiments and being too data driven and the sense that experimentation just lea..."

15
21:57 - 23:40
1:43 duration283 words

Learning from Failed Experiments

Kohavi shares a cautionary tale about a failed integration of social features at Bing, emphasizing the importance of recognizing when to abort an experiment. He discusses the challenges of determining when to pivot or continue based on experimental data.

"be successful over time if you just try enough but some experiments have you have to allocate sometimes to these high risk High reward ideas we're going to try something that's most likely to fail but..."

16
23:40 - 25:01
1:20 duration242 words

When Not to A/B Test

Ronny explains that not all scenarios are suitable for A/B testing, particularly in cases like mergers and acquisitions. He outlines the necessary conditions for effective A/B testing, including having a sufficient user base for statistical validity.

"them were a breakthrough and I remember sort of mailing uh chilu with some statistics showing that you know it's time to abort it's time to fail on this uh and you know he decided to continue more and..."

17
25:01 - 26:19
1:17 duration234 words

The Right Time to Start A/B Testing

Kohavi provides guidance on when startups should begin A/B testing, recommending that they wait until they have tens of thousands of users to ensure statistical reliability. He emphasizes the importance of building a culture of experimentation early on.

"ingredients to a b testing and I'll just say I'll write not every domain is a minimal thing to be testing right you can't a B test mergers and Acquisitions right it's something that happens once you e..."

18
26:19 - 27:43
1:24 duration271 words

Overall Evaluation Criterion (OEC)

Ronny introduces the Overall Evaluation Criterion (OEC), a framework for determining what to optimize in A/B testing. He stresses the need to balance revenue goals with user experience metrics to ensure sustainable growth.

"zero but when you're not there it's still expensive and there are there may be reasons why not to run AB thefts you talked about how you may be too small to run a b tests and this is a constant questi..."

19
27:43 - 29:31
1:48 duration300 words

Optimizing Revenue Without Sacrificing User Experience

Kohavi discusses how to optimize for revenue while maintaining a positive user experience. He provides examples from search and the hotel industry, illustrating the importance of considering long-term user satisfaction in revenue strategies.

"you're not degrading uh and getting value out of response so you ask for rule of thumb 200 000 users you're magical below that start building the culture start building the platform start integrating ..."

20
29:14 - 30:31
1:17 duration236 words

Optimizing Revenue Without Sacrificing User Experience

Ronny Kohavi discusses the balance between increasing revenue through ads and maintaining a positive user experience. He introduces the concept of Overall Evaluation Criterion (OEC) to measure the trade-offs between ad placements and user satisfaction, emphasizing the importance of long-term growth over short-term gains.

"those experiments and we were able to map out you know this number of ads causes this much increase to churn this number of ads causes this much increase to the time that users take to find a successf..."

21
30:31 - 31:56
1:24 duration291 words

The Importance of Long-Term User Satisfaction

In this segment, Kohavi highlights the significance of considering long-term user satisfaction when optimizing conversion rates. He explains how predicting user happiness with a listing on platforms like Airbnb can influence A/B testing strategies, stressing the need for metrics that reflect lifetime value.

"then you need to insert these other criteria and what am I doing through the user experience one way around it is to put this constraint another one is just to have these other metrics again something..."

22
31:56 - 33:00
1:04 duration213 words

Learning from Long-Term Experiments

Kohavi shares insights on running long-term experiments to gather valuable data. He illustrates this with examples from Bing and Amazon, explaining how understanding user behavior and preferences can lead to better decision-making and improved metrics.

"basically have a kind of a drag metric that makes sure you're not hurting something that's really important to the business and then being very clear on what's the long-term metric we care most about ..."

23
33:00 - 34:20
1:19 duration246 words

The Cost of Spamming Users

This segment focuses on the challenges of measuring the effectiveness of email campaigns at Amazon. Kohavi discusses how the initial approach led to spam and how they adjusted their strategy by incorporating unsubscribe metrics to better understand the trade-offs between email frequency and user retention.

"about it one is you can run long term experiments for the goal of learning something so I mentioned that at Bing we did run these experiments where we increased the odds and decreased the odds so that..."

24
34:20 - 36:01
1:40 duration277 words

Incremental Changes vs. Major Redesigns

Kohavi reflects on the pitfalls of large redesigns in product development. He advocates for incremental changes and testing to avoid the common failures associated with sweeping redesigns, emphasizing the importance of learning from smaller adjustments.

"then so then we backed up and then we said okay we can either phrase this as a constraint satisfaction problem you're allowed to send user and email every X days or which is what we ended up doing is ..."

25
36:01 - 37:43
1:42 duration314 words

The Reality of Idea Failures

In this candid discussion, Kohavi shares his experiences with failed ideas in tech companies. He stresses the importance of acknowledging that many ideas will fail and the need for organizations to embrace experimentation to learn and improve continuously.

"understand based on what users were unsubscribing from which ones are really beneficial I love the surprising results we all love them I mean this is this is The Humbling reality and you know people t..."

26
37:43 - 39:10
1:26 duration237 words

The Case for Data-Driven Decisions

Kohavi discusses the necessity of data-driven decision-making in organizations. He recounts his experience at Microsoft, where he built an experimentation platform to foster a culture of testing, highlighting the resistance to acknowledging failure and the importance of using data to guide product development.

"launch this even though it's bad for the user no that's terrible yeah yeah so uh this is this is the other advantage of of recognizing this humble reality that most ideas fail right if if you believe ..."

27
39:10 - 40:57
1:47 duration301 words

Building a Culture of Experimentation

This segment explores how Kohavi helped Microsoft adopt a culture of experimentation. He shares insights on overcoming initial resistance and the importance of demonstrating the value of A/B testing to encourage teams to embrace a more experimental approach.

"just as you said yourself they be ready to fail right I mean do you do you really want to work on something for six months or a year and then run the A B test and realize that you've hurt revenues or ..."

28
40:57 - 42:55
1:57 duration383 words

Allocating Resources for Big Bets

Kohavi discusses the balance between optimizing existing products and making big bets on new ideas. He emphasizes the importance of being prepared for failure when pursuing large redesigns while also recognizing the potential for breakthroughs.

"was and this was you know credited to chilu and uh Sacha Nadella they were ones that says Ronnie you know you try to get office to run experiments we'll give you the air support uh and it was hard but..."

29
42:55 - 45:01
2:05 duration393 words

Avoiding Flat Outcomes in Product Launches

In this concluding segment, Kohavi warns against launching products that yield flat results. He stresses the importance of ensuring that every launch adds value and encourages teams to think critically about the implications of their decisions on user experience and business outcomes.

"is it ever worth just going let's just rethink this whole thing and just give it a shot to break out of a local Minima or local Maxima essentially so I think what you said is fair I mean I I do want t..."

30
44:48 - 45:37
0:48 duration168 words

Legal Requirements vs. Experimentation

In this segment, Kohavi addresses the misconception that legal requirements justify shipping projects with negative or flat results. He encourages teams to explore multiple options and choose the least harmful approach, emphasizing the need for data-driven decision-making even under legal constraints.

"you know let's make sure that we understand that shipping this project has no value is complicating the code Base maintenance costs will go up you don't ship on flat unless it's a sort of a legal requ..."

31
45:37 - 46:11
0:34 duration121 words

Airbnb's Experimentation Culture

Kohavi shares insights from his time at Airbnb, where every aspect of search relevance was A/B tested. He contrasts this with the current trend at Airbnb, where fewer experiments are being run. This segment reflects on the importance of maintaining a strong experimentation culture to drive data-driven decisions.

"in the end from what I remember speaking of Airbnb I want to chat about Airbnb briefly I know there's and you're limited in what you can share but uh it's interesting that Airbnb seems to be moving in..."

32
46:11 - 47:11
1:00 duration178 words

The Impact of COVID on Experimentation

Kohavi discusses the challenges Airbnb faced during the COVID-19 pandemic and the importance of running A/B tests in rapidly changing environments. He argues that data-driven experimentation is crucial for making informed decisions, especially during crises.

"then roughly what's your sense of how things are going where it's going so as you as you know I'm restricted from talking about Airbnb I will say a few things that I am allowed to say one is in my tea..."

33
47:11 - 48:02
0:50 duration170 words

Lessons from Airbnb's Online Experiences

In this segment, Kohavi reflects on Airbnb's investment in online experiences during the pandemic. He notes that initial data was not promising, yet it has become a significant part of Airbnb's strategy. This highlights the unpredictable nature of experimentation and the potential for long-term success.

"counter factual we don't know it's a really interesting perspective yeah there may be such an interesting natural experiment of a way of doing things differently there's like de-emphasizing experiment..."

34
48:02 - 49:01
0:59 duration175 words

The Importance of Trust in Experiments

Kohavi introduces his book, 'Trustworthy Online Controlled Experiments,' discussing the significance of trust in experimentation. He emphasizes that a reliable experimentation platform acts as a safety net, allowing teams to abort bad launches quickly and maintain organizational trust in data-driven results.

"could be better all right mysterious one more question Airbnb you were there during kovid which was quite a wild time for Airbnb we had sunshine on the podcast talking about all the craziness that wen..."

35
49:01 - 50:01
0:59 duration157 words

Building Trust Through Accurate Results

Kohavi elaborates on how to build trust in experimentation results. He stresses the importance of presenting trustworthy data and ensuring that experiments are designed correctly to avoid misleading conclusions. This segment underscores the need for rigorous checks in the experimentation process.

"experiment so sometimes it means that you might have to replicate them six months down when covid say uh is not as impactful as it is saying that you have to make decisions quickly to me I'll point yo..."

36
50:01 - 51:16
1:15 duration217 words

Surprising Success of the A/B Testing Book

Kohavi shares his experience writing 'Trustworthy Online Controlled Experiments,' revealing that it exceeded sales expectations. He discusses the book's focus on practical aspects of A/B testing and the decision to donate all proceeds to charity, highlighting the impact of the book on promoting data-driven experimentation.

"yeah whatever another case study for the history books Airbnb experiences I want to shift a little bit and talk about your book which you mentioned a couple of times it's called trustworthy online con..."

37
51:16 - 52:01
0:44 duration149 words

Why Trust is Key in Experimentation

In this segment, Kohavi explains why trust is essential in the context of experimentation. He discusses how a reliable experimentation platform can prevent mistakes and build confidence in results, emphasizing that trust is easily lost but hard to regain.

"was translated to Chinese Korean Japanese and Russian and so it's it's great to see that we helped the world become more data driven with experimentation and I'm happy because of that and I was pleasa..."

38
52:01 - 53:04
1:03 duration175 words

Avoiding Common Experimentation Pitfalls

Kohavi identifies common pitfalls in running experiments, such as sample ratio mismatch. He explains how this issue can invalidate results and shares strategies for diagnosing and preventing it, emphasizing the importance of maintaining rigorous standards in experimentation.

"important in writing experiments so to me the experimentation platform is the safety net and it's an oracle so it serves really two purposes the safety net means that if you launch something bad you s..."

39
53:04 - 54:05
1:00 duration179 words

The Dangers of Real-Time P-Value Monitoring

Kohavi warns against the dangers of real-time p-value monitoring in A/B testing. He explains how it can inflate the false positive rate and lead to misleading conclusions, stressing the importance of proper statistical methods in experimentation.

"and the nice thing is when we we built all these checks to make sure that the experiment is correct if there was something wrong with it we would stop and say hey something is wrong with the experimen..."

40
54:05 - 55:20
1:15 duration237 words

Trust Issues with Experimentation Platforms

Kohavi shares a cautionary tale about early issues with the Optimizely platform, which led to a loss of trust among users. He discusses how inflated error rates can mislead teams into believing they are achieving success when they are not, highlighting the need for accurate experimentation tools.

"error rate so what this led is that people that started using optimizely thought that the platform was telling them they're very successful but when they actually started to see while it told us this ..."

41
55:20 - 56:58
1:37 duration298 words

Identifying Signs of Invalid Experiments

Kohavi discusses how to identify signs that an experiment may not be valid, focusing on sample ratio mismatch as a key indicator. He explains how to diagnose this issue and the importance of understanding the underlying causes to maintain the integrity of experimentation.

"trust because they built something that had very much inflated erroring that is uh pretty scary to think about you've been running all these experiments and they weren't actually telling you accurate ..."

42
56:58 - 1:00:00
3:01 duration516 words

Common Causes of Sample Ratio Mismatch

In this segment, Kohavi outlines common causes of sample ratio mismatch in experiments, such as bot traffic and data pipeline issues. He emphasizes the need for vigilance in monitoring experiments to ensure accurate results and maintain trust in the experimentation process.

"versus 49.8 split and therefore something is wrong with the experiment right now people I remember when we first implemented this check we were surprised to see how many experiments suffered from this..."

43
1:00:00 - 1:01:38
1:38 duration292 words

Toyman's Law: Investigating Surprising Results

Kohavi introduces Toyman's Law, which states that any figure that looks interesting or different is usually wrong. He stresses the importance of skepticism when results appear too good to be true, advocating for thorough investigation to uncover potential flaws in experiments. This segment underscores the necessity of critical thinking in data analysis.

"mismatch so we blanked out the scorecard we have this button and then we started to see that people press the button is still presented the results of experiments with sample ratio of this method so w..."

44
1:01:38 - 1:04:50
3:11 duration570 words

Understanding P-Values and False Positives

In this segment, Ronny Kohavi clarifies the common misconceptions surrounding p-values in A/B testing. He explains that a p-value does not represent the probability that a treatment is better than control, and discusses the implications of false positive rates, particularly in high-failure environments like Airbnb. Kohavi emphasizes the need for accurate interpretation of statistical results.

"and I will say that nine out of ten when we call out time is law it is the case that we find some flaw in the experiment now there are obviously outliers right that first experiment that I shared wher..."

45
1:04:50 - 1:06:28
1:37 duration307 words

The Importance of Historical Success Rates

Kohavi highlights the significance of understanding historical success rates in A/B testing. He shares insights on how to assess the likelihood of false positives based on past performance, advocating for a more nuanced approach to interpreting statistical significance. This segment provides valuable context for organizations looking to improve their experimentation practices.

"right it's not five percent it's 26 so that's the number that you should have in your mind and that's why when I worked at Airbnb one of the things we did is we said okay if you're less than 0.05 but ..."

46
1:06:28 - 1:09:01
2:32 duration458 words

Getting Started with A/B Testing

Ronny Kohavi offers practical advice for organizations looking to implement A/B testing. He discusses the importance of having experienced personnel, the decision to build or buy experimentation platforms, and the need for a clear optimization criterion (OEC). This segment serves as a guide for teams eager to start running effective experiments.

"say someone listening wants to start running experiments they say they have tens of thousands of users at this point what would be the first couple steps you'd recommend well so if they have somebody ..."

47
1:09:01 - 1:12:30
3:29 duration651 words

Building a Culture of Experimentation

Kohavi shares his experiences at Microsoft and Bing, detailing how to foster a culture of experimentation within organizations. He emphasizes the importance of frequent launches and clear optimization goals, providing insights into how to overcome resistance to A/B testing. This segment is essential for leaders aiming to drive innovation through experimentation.

"understand the question of the oec is it clear what they're optimizing for right there are some groups where you can come up with a good oec some groups are harder you know I remember one funny exampl..."

48
1:12:30 - 1:13:40
1:09 duration209 words

Speeding Up Experimentation Processes

In this final segment, Ronny Kohavi discusses strategies for accelerating the experimentation process. He highlights the role of effective platforms in providing quick feedback and introduces variance reduction techniques to achieve faster results. This segment is crucial for teams looking to enhance their experimentation efficiency.

"ends up being most important still I want to ask you about speed is there anything you recommend for helping people run experiments faster and get results more quickly that they can Implement yeah so ..."

49
1:14:09 - 1:15:00
0:50 duration156 words

The Importance of Structured Narratives

Ronny reflects on a significant change he implemented at Amazon: replacing PowerPoint presentations with structured narratives. He explains how this approach fosters better feedback and understanding among teams, leading to improved decision-making and execution in product development.

"fewer users Ronnie is there anything else you want to share before we get to our very exciting lightning round uh no I think we've asked a lot of good questions uh hope people enjoy this I know they w..."

50
1:15:00 - 1:16:39
1:39 duration277 words

Lessons from Chernobyl

In a light-hearted discussion, Ronny shares his thoughts on the acclaimed series 'Chernobyl.' He reflects on its portrayal of true events and connects it to his personal history, revealing how his family was indirectly affected by the disaster. This segment highlights the intersection of storytelling and real-world implications in understanding complex events.

"dangerous half-truths and total nonsense uh by the Stanford professors from The Graduate School of Business very interesting to see many of the things that we grew up with as sort of well understood t..."

51
1:16:39 - 1:18:43
2:04 duration359 words

Interview Insights and Favorite Products

Ronny shares his favorite interview question that often stumps candidates, revealing insights into technical knowledge gaps. He also discusses his recent discovery of Blink cameras, explaining how they have transformed his ability to monitor his home and wildlife, showcasing the practical applications of technology in everyday life.

"but uh but yeah we were in the vicinity that's pretty scary my wife thinks I've yeah every every time something's wrong with me she's like oh that must be a Chernobyl Chernobyl thing okay next questio..."

52
1:18:43 - 1:21:30
2:46 duration477 words

Hierarchy of Evidence in Decision Making

Ronny emphasizes the importance of understanding the hierarchy of evidence when making decisions. He discusses how different types of evidence, from anecdotal to controlled experiments, should influence trust levels in information. This segment serves as a reminder of the critical thinking necessary in both personal and professional contexts.

"point we had a false alarm and the cops came in and had this amazing video of how they're entering the house uh and pulling the guns out you gotta share that on Tick Tock that's good content wow okay ..."

53
1:21:30 - 1:23:02
1:32 duration292 words

Final Thoughts and Resources

In the closing segment, Ronny shares where listeners can find him online and encourages them to engage with the concepts of controlled experiments. He highlights his book and upcoming class, reinforcing the value of data-driven decision-making and the importance of experimentation in innovation.

"another I think there's a book that's based on this like how to read a book well Ronnie the experiment of us recording a podcast I think is a hundred percent positive p-value 0.0 thank you so much for..."