searchlore

Back to Resource

All Segments

Jeff Dean & Noam Shazeer — 25 years at Google: from PageRank to AGI

Jeff Dean & Noam Shazeer — 25 years at Google: from PageRank to AGI

27 segments available

This week I welcome two of the most important technologists in any field. Jeff Dean is Google's Chief Scientist, and through 25 years at the company, has worked on basically the most transformative systems in modern computing: from MapReduce, BigTable, Tensorflow, AlphaChip, to Gemini. Noam Shazeer invented or co-invented all the main architectures and techniques that are used for modern LLMs: from the Transformer itself, to Mixture of Experts, to Mesh Tensorflow, to Gemini and many other things. We talk about their 25 years at Google, going from PageRank to MapReduce to the Transformer to MoEs to AlphaChip – and soon to ASI. 𝐄𝐏𝐈𝐒𝐎𝐃𝐄 𝐋𝐈𝐍𝐊𝐒 * Transcript: https://www.dwarkesh.com/p/jeff-dean-and-noam-shazeer * Apple Podcasts: https://podcasts.apple.com/us/podcast/jeff-dean-noam-shazeer-25-years-at-google-from-pagerank/id1516093381?i=1000691556147 * Spotify: https://open.spotify.com/episode/4atx1POpKIL8WGvdVfdnbb?si=DLn5uQYMQMWKPTTkj5pt_A 𝐒𝐏𝐎𝐍𝐒𝐎𝐑𝐒 * Meter wants to radically improve the digital world we take for granted. They’re developing a foundation model that automates network management end-to-end. To do this, they just announced a long-term partnership with Microsoft for tens of thousands of GPUs, and they’re recruiting a world class AI research team. To learn more, go to https://meter.com/dwarkesh * Scale partners with major AI labs like Meta, Google Deepmind, and OpenAI. Through Scale’s Data Foundry, labs get access to high-quality data to fuel post-training, including advanced reasoning capabilities. If you’re an AI researcher or engineer, learn about how Scale’s Data Foundry and research lab, SEAL, can help you go beyond the current frontier at https://scale.com/dwarkesh * Curious how Jane Street teaches their new traders? They use Figgie, a rapid-fire card game that simulates the most exciting parts of markets and trading. It’s become so popular that Jane Street hosts an inter-office Figgie championship every year. Download from the app store or play on your desktop at https://www.figgie.com/ To sponsor a future episode, visit https://www.dwarkesh.com/p/advertise 𝐓𝐈𝐌𝐄𝐒𝐓𝐀𝐌𝐏𝐒 00:00:00 - Intro 00:03:29 - Joining Google in 1999 00:06:20 - Future of Moore's Law 00:11:04 - Future TPUs 00:13:56 - Jeff’s undergrad thesis: parallel backprop 00:15:54 - LLMs in 2007 00:25:09 - “Holy shit” moments 00:27:28 - AI fulfills Google’s original mission 00:32:00 - Doing Search in-context 00:36:12 - The internal coding model 00:37:29 - What will 2027 models do? 00:43:20 - A new architecture every day? 00:49:10 - Automated chips and intelligence explosion 00:53:07 - Future of inference scaling 01:02:38 - Already doing multi-datacenter runs 01:08:15 - Debugging at scale 01:12:41 - Fast takeoff and superalignment 01:20:51 - A million evil Jeff Deans 01:24:22 - Fun times at Google 01:27:51 - World compute demand in 2030 01:34:37 - Getting back to modularity 01:44:48 - Keeping a giga-MoE in-memory 01:49:35 - All of Google in one model 01:57:59 - What’s missing from distillation 02:03:10 - Open research, pros and cons 02:09:58 - Going the distance

Segments Timeline

1
0:46 - 3:29
2:42 duration477 words

Meet the Innovators

In this segment, we are introduced to Jeff Dean and Noam Shazeer, two pivotal figures in the tech world. Jeff Dean, Google's Chief Scientist, has been instrumental in developing transformative systems like MapReduce and TensorFlow. Noam Shazeer is recognized for his groundbreaking work on architectures for modern LLMs, including the Transformer. Their collaboration on Gemini at Google DeepMind highlights their significant contributions to AI and computing.

"Today I have the honor of chatting with Jeff  Dean and Noam Shazeer. Jeff is Google's Chief   Scientist, and through his 25 years at the  company, he has worked on basically the most   transformative ..."

2
3:29 - 6:20
2:51 duration495 words

The Evolution of Google

Jeff and Noam reflect on their long tenures at Google, discussing how the company has evolved from a small team to a massive organization. They share personal anecdotes about their early days, including mentorship experiences and the challenges of keeping up with the rapid growth of the company. This segment captures the essence of adapting to change in a fast-paced tech environment.

"How did Google recruit you, by the way? I kind of reached out to them, actually.   And Noam, how did you get recruited? I actually saw Google at a job fair in 1999,   and I assumed that it was already..."

3
6:20 - 11:04
4:44 duration704 words

Moore's Law and Its Impact

The conversation shifts to the implications of Moore's Law on technology development. Jeff explains how advancements in hardware have influenced system design and project feasibility over the decades. He discusses the transition from general-purpose CPUs to specialized hardware like TPUs, emphasizing the importance of adapting algorithms to leverage these advancements for better performance.

"decades changed the kinds of considerations you  have to take on board when you design new systems,   when you figure out what projects  are feasible? What are still the   limitations? What are things..."

4
11:04 - 13:56
2:52 duration429 words

The Future of TPUs

In this segment, Jeff discusses the future of Tensor Processing Units (TPUs) and how they are evolving to meet the demands of modern AI workloads. He highlights the trend towards reduced precision models and the importance of co-designing algorithms and hardware to maximize efficiency. This insight into TPU development showcases the ongoing innovation at Google.

"considering changing for future versions of TPU to  integrate how you're thinking about algorithms?   I think one general trend is we're getting  better at quantizing or having much more   reduced pre..."

5
13:56 - 15:54
1:58 duration320 words

Backpropagation Breakthrough

Jeff shares the story of his undergraduate thesis on parallel backpropagation training for neural networks. He reflects on the early days of neural networks and the challenges faced in training them effectively. This segment highlights the foundational work that laid the groundwork for future advancements in AI.

"Let me start with the undergrad thesis. I got  introduced to neural nets in one section of one   class on parallel computing that I was taking  in my senior year. I needed to do a thesis to   graduate..."

6
15:54 - 25:09
9:15 duration1192 words

LLMs in 2007

The discussion turns to the early exploration of large language models (LLMs) in 2007. Jeff and Noam recount their experiences training a two trillion token N-gram model for language modeling. They reflect on the significance of this work and its implications for the future of AI, showcasing their foresight into the potential of language models.

"me about how the 2007 paper came together. Oh yeah, so that, we had a machine translation   research team at Google led by Franz Och,  who had joined Google maybe a year before,   and a bunch of other..."

7
25:09 - 27:28
2:19 duration344 words

Moments of Realization

Jeff and Noam share their 'holy shit' moments in AI development, discussing breakthroughs that changed their understanding of what was possible. This segment captures the excitement and challenges of working at the forefront of technology, illustrating the passion that drives innovation.

"you're looking at a research area, you come up  with this idea, and you have this feeling of,   "Holy shit, I can't believe that worked?" One thing I remember was in the early days of   the Brain tea..."

8
27:28 - 32:00
4:32 duration661 words

AI and Google's Mission

The conversation shifts to how AI fulfills Google's original mission of organizing the world's information. Jeff emphasizes the broad mandate of the company and how AI technologies are integral to achieving this goal. This segment highlights the alignment between Google's ambitions and the advancements in AI.

"These examples illustrate how these AI systems  fit into what you were just mentioning:   that Google is fundamentally a company that  organizes information. AI, in this context,   is finding relation..."

9
32:00 - 36:12
4:12 duration631 words

In-Context Search

Jeff discusses the concept of doing search in-context, exploring how AI can enhance search capabilities. He explains the potential for AI to provide more relevant and contextualized search results, showcasing the innovative approaches being developed at Google.

"I know one thing you're working on right now is  longer context. If you think of Google Search,   it's got the entire index of the internet in  its context, but it's a very shallow search.   And then ..."

10
36:12 - 37:29
1:17 duration214 words

The Internal Coding Model

In this brief segment, Jeff touches on the internal coding model used at Google. He highlights the importance of having a robust coding framework to support the development of advanced AI systems, emphasizing the technical foundations that enable innovation.

"I want to talk more about the thing you mentioned  about: look, Google is a company with lots of   code and lots of examples. If you just think  about that one use case and what that implies,   so you..."

11
37:29 - 43:20
5:51 duration869 words

Future Models in 2027

Looking ahead, Jeff speculates on what AI models might look like in 2027. He discusses the anticipated advancements in AI capabilities and the potential for new architectures to emerge. This forward-looking segment captures the excitement of future possibilities in AI.

"your own personal work? What will it be like to  be a researcher at Google? You have a new idea   or something. With the way in which you're  interacting with these models in a year,   what does that ..."

12
43:20 - 49:10
5:50 duration729 words

Daily Architectural Innovations

The conversation explores the rapid pace of architectural innovations in AI. Jeff and Noam discuss the frequency of new architectures being developed and the implications for the field. This segment highlights the dynamic nature of AI research and development.

"idea. Suppose in the world today there are  on the order of 10,000 AI researchers in this   community coming up with a breakthrough- Probably more than that. There were   15,000 at NeurIPS last week. ..."

13
49:10 - 53:07
3:57 duration678 words

Automated Chips and Intelligence Explosion

Jeff discusses the concept of automated chips and the potential for an intelligence explosion. He explores how advancements in hardware and AI could lead to unprecedented capabilities, raising important questions about the future of technology.

"we could dramatically speed up the chip design  process. As we were talking earlier, the current   way in which you design a chip takes you roughly  18 months to go from "we should build a chip" to   ..."

14
53:07 - 1:02:38
9:31 duration1517 words

Scaling Inference

In this segment, Jeff addresses the future of inference scaling in AI systems. He discusses the challenges and opportunities associated with scaling AI models to meet growing demands, emphasizing the importance of efficient design and implementation.

"applying more compute at inference time. I  guess the way I like to describe it is that   even a giant language model, even if you’re doing  a trillion operations per token, which is more   than most ..."

15
1:02:38 - 1:08:15
5:37 duration867 words

Multi-Datacenter Operations

The conversation shifts to the logistics of running AI models across multiple datacenters. Jeff shares insights into the current capabilities and the strategies being employed to optimize performance and reliability in distributed environments.

"already tapping out nuclear power plants in  terms of delivering power into one single   campus. Do we have to have just two gigawatts  in one place, five gigawatts in one place,   or can it be more d..."

16
1:08:15 - 1:12:41
4:26 duration561 words

Debugging at Scale

Jeff discusses the complexities of debugging AI systems at scale. He highlights the challenges faced by engineers and the innovative approaches being developed to ensure reliability and performance in large-scale AI deployments.

"decode? You've got these things, some of which are  making the model better, some of which are making   it worse. When you go into work tomorrow, how do  you figure out what the most salient inputs ar..."

17
1:12:41 - 1:20:51
8:10 duration1259 words

Fast Takeoff and Superalignment

In this thought-provoking segment, Jeff and Noam explore the concepts of fast takeoff and superalignment in AI. They discuss the implications of rapid advancements in AI capabilities and the importance of aligning AI systems with human values.

"and the models get better and better over time”,  even if you take the hardware part out of it.   Should the world be thinking more about, and  should you guys be thinking more about this?   There's o..."

18
1:20:51 - 1:24:22
3:31 duration582 words

A Million Evil Jeff Deans

The conversation takes a humorous turn as Jeff and Noam joke about the idea of a million 'evil Jeff Deans.' This light-hearted moment contrasts with the serious discussions about AI, showcasing the camaraderie between the two innovators.

"of nuclear war or something. Just think about  it, like a million evil Jeff Deans or something.   Where do we get the training data? But, to the extent that you think that's   a plausible output of so..."

19
1:24:22 - 1:27:51
3:29 duration582 words

Fun Times at Google

Jeff and Noam reminisce about their experiences at Google, sharing anecdotes that highlight the fun and collaborative culture within the company. This segment captures the human side of working in tech and the friendships formed over the years.

"All right, let's talk about a few more fun topics.  Make it a little lighter. Over the last 25 years,   what was the most fun time? What period of  time do you have the most nostalgia over?   I think ..."

20
1:27:51 - 1:34:37
6:46 duration1020 words

World Compute Demand in 2030

Looking towards the future, Jeff discusses the anticipated demand for computing power by 2030. He explores the implications of this demand for AI development and the need for innovative solutions to meet growing computational needs.

"What I find remarkable about some  of the calls you guys have made   is you're anticipating a level of demand for  compute, which at the time wasn't obvious or   evident. TPUs being a famous example o..."

21
1:34:37 - 1:44:48
10:11 duration1675 words

Getting Back to Modularity

In this segment, Jeff emphasizes the importance of modularity in AI systems. He discusses the benefits of modular design for scalability and flexibility, highlighting the need for a structured approach to AI development.

"Yeah, I've been thinking about this more and  more. I've been a big fan of models that are   sparse because I think you want different parts  of the model to be good at different things. We   have our..."

22
1:44:48 - 1:49:35
4:47 duration744 words

Keeping a Giga-MoE In-Memory

Jeff discusses the challenges of maintaining a giga-MoE (Mixture of Experts) model in-memory. He explores the technical considerations and strategies for optimizing performance in large-scale AI systems.

"Man, there are so many interesting implications  of this that I could just keep asking you about   this- I would regret not asking you more about  this, so I'll keep going. One implication is,   curre..."

23
1:49:35 - 1:57:59
8:24 duration1326 words

All of Google in One Model

The conversation shifts to the ambitious goal of integrating all of Google's knowledge into a single AI model. Jeff and Noam discuss the challenges and potential benefits of such an endeavor, highlighting the vision for the future of AI at Google.

"Right now, language models,  obviously, you put in language,   you get language out. Obviously, it's multimodal. But the Pathways blog post talks about so many   different use cases that are not obvio..."

24
1:57:59 - 2:03:10
5:11 duration840 words

What's Missing from Distillation

In this segment, Jeff addresses the limitations of current distillation techniques in AI. He discusses the challenges of transferring knowledge from large models to smaller ones and the implications for AI development.

"How would you characterize what's missing  from current distillation techniques?   Well, I just want it to work faster. A related thing is I feel like we   need interesting learning techniques during ..."

25
2:03:10 - 2:09:58
6:48 duration988 words

Open Research: Pros and Cons

Jeff and Noam discuss the pros and cons of open research in AI. They explore the benefits of collaboration and transparency, as well as the challenges associated with sharing knowledge in a competitive field.

"So here's the question I have. What you've just  laid out over the last hour is potentially just   like the big next paradigm shift in AI.  That's a tremendously valuable insight,   potentially. Noam,..."

26
2:09:58 - 2:10:00
0:02 duration19 words

Going the Distance

The episode concludes with reflections on the journey of innovation at Google. Jeff and Noam share their thoughts on the future of AI and the importance of perseverance in the face of challenges.

"So we've discussed some of the things you guys  have worked on over the last 25 years, and there  "

27
2:12:02 - 2:15:16
3:13 duration420 words

Balancing Innovation and Collaboration

Dean elaborates on the dynamics of resource allocation within Google Brain, contrasting top-down and bottom-up approaches. He discusses the need for a balance that fosters both collaboration and flexibility, which are crucial for driving innovation. He also mentions his internal slide deck of 'wacky ideas' as a way to inspire new projects and directions.

"I'd say probably a big thing is humility, like  I’d say I’m the most humble. But seriously,   to say what I just did is nothing compared to what  I can do or what can be done. And to be able to   drop..."