
27 segments available
This week I welcome two of the most important technologists in any field. Jeff Dean is Google's Chief Scientist, and through 25 years at the company, has worked on basically the most transformative systems in modern computing: from MapReduce, BigTable, Tensorflow, AlphaChip, to Gemini. Noam Shazeer invented or co-invented all the main architectures and techniques that are used for modern LLMs: from the Transformer itself, to Mixture of Experts, to Mesh Tensorflow, to Gemini and many other things. We talk about their 25 years at Google, going from PageRank to MapReduce to the Transformer to MoEs to AlphaChip – and soon to ASI. 𝐄𝐏𝐈𝐒𝐎𝐃𝐄 𝐋𝐈𝐍𝐊𝐒 * Transcript: https://www.dwarkesh.com/p/jeff-dean-and-noam-shazeer * Apple Podcasts: https://podcasts.apple.com/us/podcast/jeff-dean-noam-shazeer-25-years-at-google-from-pagerank/id1516093381?i=1000691556147 * Spotify: https://open.spotify.com/episode/4atx1POpKIL8WGvdVfdnbb?si=DLn5uQYMQMWKPTTkj5pt_A 𝐒𝐏𝐎𝐍𝐒𝐎𝐑𝐒 * Meter wants to radically improve the digital world we take for granted. They’re developing a foundation model that automates network management end-to-end. To do this, they just announced a long-term partnership with Microsoft for tens of thousands of GPUs, and they’re recruiting a world class AI research team. To learn more, go to https://meter.com/dwarkesh * Scale partners with major AI labs like Meta, Google Deepmind, and OpenAI. Through Scale’s Data Foundry, labs get access to high-quality data to fuel post-training, including advanced reasoning capabilities. If you’re an AI researcher or engineer, learn about how Scale’s Data Foundry and research lab, SEAL, can help you go beyond the current frontier at https://scale.com/dwarkesh * Curious how Jane Street teaches their new traders? They use Figgie, a rapid-fire card game that simulates the most exciting parts of markets and trading. It’s become so popular that Jane Street hosts an inter-office Figgie championship every year. Download from the app store or play on your desktop at https://www.figgie.com/ To sponsor a future episode, visit https://www.dwarkesh.com/p/advertise 𝐓𝐈𝐌𝐄𝐒𝐓𝐀𝐌𝐏𝐒 00:00:00 - Intro 00:03:29 - Joining Google in 1999 00:06:20 - Future of Moore's Law 00:11:04 - Future TPUs 00:13:56 - Jeff’s undergrad thesis: parallel backprop 00:15:54 - LLMs in 2007 00:25:09 - “Holy shit” moments 00:27:28 - AI fulfills Google’s original mission 00:32:00 - Doing Search in-context 00:36:12 - The internal coding model 00:37:29 - What will 2027 models do? 00:43:20 - A new architecture every day? 00:49:10 - Automated chips and intelligence explosion 00:53:07 - Future of inference scaling 01:02:38 - Already doing multi-datacenter runs 01:08:15 - Debugging at scale 01:12:41 - Fast takeoff and superalignment 01:20:51 - A million evil Jeff Deans 01:24:22 - Fun times at Google 01:27:51 - World compute demand in 2030 01:34:37 - Getting back to modularity 01:44:48 - Keeping a giga-MoE in-memory 01:49:35 - All of Google in one model 01:57:59 - What’s missing from distillation 02:03:10 - Open research, pros and cons 02:09:58 - Going the distance
In this segment, we are introduced to Jeff Dean and Noam Shazeer, two pivotal figures in the tech world. Jeff Dean, Google's Chief Scientist, has been instrumental in developing transformative systems like MapReduce and TensorFlow. Noam Shazeer is recognized for his groundbreaking work on architectures for modern LLMs, including the Transformer. Their collaboration on Gemini at Google DeepMind highlights their significant contributions to AI and computing.
"Today I have the honor of chatting with Jeff Dean and Noam Shazeer. Jeff is Google's Chief Scientist, and through his 25 years at the company, he has worked on basically the most transformative ..."
Jeff and Noam reflect on their long tenures at Google, discussing how the company has evolved from a small team to a massive organization. They share personal anecdotes about their early days, including mentorship experiences and the challenges of keeping up with the rapid growth of the company. This segment captures the essence of adapting to change in a fast-paced tech environment.
"How did Google recruit you, by the way? I kind of reached out to them, actually. And Noam, how did you get recruited? I actually saw Google at a job fair in 1999, and I assumed that it was already..."
The conversation shifts to the implications of Moore's Law on technology development. Jeff explains how advancements in hardware have influenced system design and project feasibility over the decades. He discusses the transition from general-purpose CPUs to specialized hardware like TPUs, emphasizing the importance of adapting algorithms to leverage these advancements for better performance.
"decades changed the kinds of considerations you have to take on board when you design new systems, when you figure out what projects are feasible? What are still the limitations? What are things..."
In this segment, Jeff discusses the future of Tensor Processing Units (TPUs) and how they are evolving to meet the demands of modern AI workloads. He highlights the trend towards reduced precision models and the importance of co-designing algorithms and hardware to maximize efficiency. This insight into TPU development showcases the ongoing innovation at Google.
"considering changing for future versions of TPU to integrate how you're thinking about algorithms? I think one general trend is we're getting better at quantizing or having much more reduced pre..."
Jeff shares the story of his undergraduate thesis on parallel backpropagation training for neural networks. He reflects on the early days of neural networks and the challenges faced in training them effectively. This segment highlights the foundational work that laid the groundwork for future advancements in AI.
"Let me start with the undergrad thesis. I got introduced to neural nets in one section of one class on parallel computing that I was taking in my senior year. I needed to do a thesis to graduate..."
The discussion turns to the early exploration of large language models (LLMs) in 2007. Jeff and Noam recount their experiences training a two trillion token N-gram model for language modeling. They reflect on the significance of this work and its implications for the future of AI, showcasing their foresight into the potential of language models.
"me about how the 2007 paper came together. Oh yeah, so that, we had a machine translation research team at Google led by Franz Och, who had joined Google maybe a year before, and a bunch of other..."
Jeff and Noam share their 'holy shit' moments in AI development, discussing breakthroughs that changed their understanding of what was possible. This segment captures the excitement and challenges of working at the forefront of technology, illustrating the passion that drives innovation.
"you're looking at a research area, you come up with this idea, and you have this feeling of, "Holy shit, I can't believe that worked?" One thing I remember was in the early days of the Brain tea..."
The conversation shifts to how AI fulfills Google's original mission of organizing the world's information. Jeff emphasizes the broad mandate of the company and how AI technologies are integral to achieving this goal. This segment highlights the alignment between Google's ambitions and the advancements in AI.
"These examples illustrate how these AI systems fit into what you were just mentioning: that Google is fundamentally a company that organizes information. AI, in this context, is finding relation..."
Jeff discusses the concept of doing search in-context, exploring how AI can enhance search capabilities. He explains the potential for AI to provide more relevant and contextualized search results, showcasing the innovative approaches being developed at Google.
"I know one thing you're working on right now is longer context. If you think of Google Search, it's got the entire index of the internet in its context, but it's a very shallow search. And then ..."
In this brief segment, Jeff touches on the internal coding model used at Google. He highlights the importance of having a robust coding framework to support the development of advanced AI systems, emphasizing the technical foundations that enable innovation.
"I want to talk more about the thing you mentioned about: look, Google is a company with lots of code and lots of examples. If you just think about that one use case and what that implies, so you..."
Looking ahead, Jeff speculates on what AI models might look like in 2027. He discusses the anticipated advancements in AI capabilities and the potential for new architectures to emerge. This forward-looking segment captures the excitement of future possibilities in AI.
"your own personal work? What will it be like to be a researcher at Google? You have a new idea or something. With the way in which you're interacting with these models in a year, what does that ..."
The conversation explores the rapid pace of architectural innovations in AI. Jeff and Noam discuss the frequency of new architectures being developed and the implications for the field. This segment highlights the dynamic nature of AI research and development.
"idea. Suppose in the world today there are on the order of 10,000 AI researchers in this community coming up with a breakthrough- Probably more than that. There were 15,000 at NeurIPS last week. ..."
Jeff discusses the concept of automated chips and the potential for an intelligence explosion. He explores how advancements in hardware and AI could lead to unprecedented capabilities, raising important questions about the future of technology.
"we could dramatically speed up the chip design process. As we were talking earlier, the current way in which you design a chip takes you roughly 18 months to go from "we should build a chip" to ..."
In this segment, Jeff addresses the future of inference scaling in AI systems. He discusses the challenges and opportunities associated with scaling AI models to meet growing demands, emphasizing the importance of efficient design and implementation.
"applying more compute at inference time. I guess the way I like to describe it is that even a giant language model, even if you’re doing a trillion operations per token, which is more than most ..."
The conversation shifts to the logistics of running AI models across multiple datacenters. Jeff shares insights into the current capabilities and the strategies being employed to optimize performance and reliability in distributed environments.
"already tapping out nuclear power plants in terms of delivering power into one single campus. Do we have to have just two gigawatts in one place, five gigawatts in one place, or can it be more d..."
Jeff discusses the complexities of debugging AI systems at scale. He highlights the challenges faced by engineers and the innovative approaches being developed to ensure reliability and performance in large-scale AI deployments.
"decode? You've got these things, some of which are making the model better, some of which are making it worse. When you go into work tomorrow, how do you figure out what the most salient inputs ar..."
In this thought-provoking segment, Jeff and Noam explore the concepts of fast takeoff and superalignment in AI. They discuss the implications of rapid advancements in AI capabilities and the importance of aligning AI systems with human values.
"and the models get better and better over time”, even if you take the hardware part out of it. Should the world be thinking more about, and should you guys be thinking more about this? There's o..."
The conversation takes a humorous turn as Jeff and Noam joke about the idea of a million 'evil Jeff Deans.' This light-hearted moment contrasts with the serious discussions about AI, showcasing the camaraderie between the two innovators.
"of nuclear war or something. Just think about it, like a million evil Jeff Deans or something. Where do we get the training data? But, to the extent that you think that's a plausible output of so..."
Jeff and Noam reminisce about their experiences at Google, sharing anecdotes that highlight the fun and collaborative culture within the company. This segment captures the human side of working in tech and the friendships formed over the years.
"All right, let's talk about a few more fun topics. Make it a little lighter. Over the last 25 years, what was the most fun time? What period of time do you have the most nostalgia over? I think ..."
Looking towards the future, Jeff discusses the anticipated demand for computing power by 2030. He explores the implications of this demand for AI development and the need for innovative solutions to meet growing computational needs.
"What I find remarkable about some of the calls you guys have made is you're anticipating a level of demand for compute, which at the time wasn't obvious or evident. TPUs being a famous example o..."
In this segment, Jeff emphasizes the importance of modularity in AI systems. He discusses the benefits of modular design for scalability and flexibility, highlighting the need for a structured approach to AI development.
"Yeah, I've been thinking about this more and more. I've been a big fan of models that are sparse because I think you want different parts of the model to be good at different things. We have our..."
Jeff discusses the challenges of maintaining a giga-MoE (Mixture of Experts) model in-memory. He explores the technical considerations and strategies for optimizing performance in large-scale AI systems.
"Man, there are so many interesting implications of this that I could just keep asking you about this- I would regret not asking you more about this, so I'll keep going. One implication is, curre..."
The conversation shifts to the ambitious goal of integrating all of Google's knowledge into a single AI model. Jeff and Noam discuss the challenges and potential benefits of such an endeavor, highlighting the vision for the future of AI at Google.
"Right now, language models, obviously, you put in language, you get language out. Obviously, it's multimodal. But the Pathways blog post talks about so many different use cases that are not obvio..."
In this segment, Jeff addresses the limitations of current distillation techniques in AI. He discusses the challenges of transferring knowledge from large models to smaller ones and the implications for AI development.
"How would you characterize what's missing from current distillation techniques? Well, I just want it to work faster. A related thing is I feel like we need interesting learning techniques during ..."
Jeff and Noam discuss the pros and cons of open research in AI. They explore the benefits of collaboration and transparency, as well as the challenges associated with sharing knowledge in a competitive field.
"So here's the question I have. What you've just laid out over the last hour is potentially just like the big next paradigm shift in AI. That's a tremendously valuable insight, potentially. Noam,..."
The episode concludes with reflections on the journey of innovation at Google. Jeff and Noam share their thoughts on the future of AI and the importance of perseverance in the face of challenges.
"So we've discussed some of the things you guys have worked on over the last 25 years, and there "
Dean elaborates on the dynamics of resource allocation within Google Brain, contrasting top-down and bottom-up approaches. He discusses the need for a balance that fosters both collaboration and flexibility, which are crucial for driving innovation. He also mentions his internal slide deck of 'wacky ideas' as a way to inspire new projects and directions.
"I'd say probably a big thing is humility, like I’d say I’m the most humble. But seriously, to say what I just did is nothing compared to what I can do or what can be done. And to be able to drop..."