Mark Zuckerberg — Llama 3, $10B models, Caesar Augustus, & 1 GW datacenters

0:00

That's not even a question for me - whether  we're going to go take a swing at building the next thing.

0:03

I'm just incapable of not doing  that.

0:03

There's a bunch of times when we wanted to launch features and then Apple's just like  nope you're not launching that I was like that sucks.

0:12

Are we set up for that with AI where  you're going to get a handful of companies that run these closed models that are going to be in  control of the apis and therefore are going to be able to tell you what you can build?

0:22

Then when  you start getting into building a data center that's like 300 Megawatts or 500 Megawatts or a  Gigawatt - just no one has built a single Gigawatt data center yet.

0:33

From wherever you sit there's  going to be some actor who you don't trust - if they're the ones who have the super strong AI I  think that that's potentially a much bigger risk Mark, welcome to the podcast. Thanks for having me. Big fan of your podcast.

0:47

Thank you, that's very nice of you to say.

0:47

Let's start by talking about the releases that will go out when this interview  goes out.

0:52

Tell me about the models and Meta AI.

0:57

What’s new and exciting about them?

0:57

I think the main thing that most people in the world are going to see is the new version of  Meta AI.

1:02

The most important thing that we're doing is the upgrade to the model.

1:08

We're  rolling out Llama-3.

1:08

We're doing it both as open source for the dev community and it is  now going to be powering Meta AI.

1:12

There's a lot that I'm sure we'll get into around Llama-3,  but I think the bottom line on this is that we think now that Meta AI is the most intelligent,  freely-available AI assistant that people can use.

1:30

We're also integrating Google  and Bing for real-time knowledge.

1:34

We're going to make it a lot more prominent across  our apps.

1:34

At the top of Facebook and Messenger, you'll be able to just use the search box right  there to ask any question.

1:42

There's a bunch of new creation features that we added that I think are  pretty cool and that I think people will enjoy.

1:54

I think animations is a good one.

1:54

You can  basically take any image and just animate it.

2:00

One that people are going to find pretty wild  is that it now generates high quality images so quickly that it actually generates it as  you're typing and updates it in real time.

2:12

So you're typing your query and it's honing  in.

2:12

It’s like “show me a picture of a cow in a field with mountains in the background, eating  macadamia nuts, drinking beer” and it's updating the image in real time. It's pretty wild.

2:29

I  think people are going to enjoy that.

2:29

So I think that's what most people are going to see in  the world.

2:35

We're rolling that out, not everywhere, but we're starting in a handful of countries and  we'll do more over the coming weeks and months.

2:46

I think that’s going to be a pretty big deal  and I'm really excited to get that in people's hands.

2:50

It's a big step forward for Meta AI.

2:50

But I think if you want to get under the hood a bit, the Llama-3 stuff is obviously the most  technically interesting.

2:57

We're training three versions: an 8 billion parameter model and a 70  billion, which we're releasing today, and a 405 billion dense model, which is still training.

3:11

So  we're not releasing that today, but I'm pretty excited about how the 8B and the 70B turned out.

3:20

They're leading for their scale.

3:20

We'll release a blog post with all the benchmarks so people can  check it out themselves.

3:31

Obviously it's open source so people get a chance to play with it.

3:34

We have a roadmap of new releases coming that are going to bring multimodality, more  multi-linguality, and bigger context windows as well.

3:46

Hopefully, sometime later in the  year we'll get to roll out the 405B.

3:46

For where it is right now in training, it is already  at around 85 MMLU and we expect that it's going to have leading benchmarks on a bunch of the  benchmarks.

4:09

I'm pretty excited about all of that.

4:14

The 70 billion is great too.

4:14

We're releasing that  today.

4:14

It's around 82 MMLU and has leading scores on math and reasoning.

4:22

I think just getting this  in people's hands is going to be pretty wild. Oh, interesting.

4:26

That's the first I’m hearing  of it as a benchmark. That's super impressive.

4:30

The 8 billion is nearly as powerful as the  biggest version of Llama-2 that we released.

4:38

So the smallest Llama-3 is basically  as powerful as the biggest Llama-2.

4:43

Before we dig into these models, I want to go  back in time.

4:43

I'm assuming 2022 is when you started acquiring these H100s, or you can tell me  when.

4:49

The stock price is getting hammered.

4:49

People are asking what's happening with all this  capex.

4:56

People aren't buying the metaverse.

5:00

Presumably you're spending that capex to get  these H100s.

5:00

How did you know back then to get the H100s?

5:04

How did you know that you’d need the GPUs?

5:04

I think it was because we were working on Reels.

5:14

We always want to have enough capacity to build  something that we can't quite see on the horizon yet.

5:23

We got into this position with Reels where we  needed more GPUs to train the models.

5:23

It was this big evolution for our services.

5:31

Instead of just  ranking content from people or pages you follow, we made this big push to start recommending what  we call unconnected content, content from people or pages that you're not following.

5:49

The corpus of content candidates that we could potentially show you expanded from  on the order of thousands to on the order of hundreds of millions.

6:01

It needed a completely  different infrastructure.

6:01

We started working on doing that and we were constrained on  the infrastructure in catching up to what TikTok was doing as quickly as we wanted to.

6:14

I  basically looked at that and I was like “hey, we have to make sure that we're never in this  situation again.

6:19

So let's order enough GPUs to do what we need to do on Reels and ranking content  and feed.

6:25

But let's also double that.

6:25

” Again, our normal principle is that there's going to be  something on the horizon that we can't see yet.

6:35

Did you know it would be AI?

6:35

We thought it was going to be something that had to do with training large models.

6:40

At the time  I thought it was probably going to be something that had to do with content.

6:44

It’s just the pattern  matching of running the company, there's always another thing.

6:52

At that time I was so deep into  trying to get the recommendations working for Reels and other content.

7:00

That’s just such a big  unlock for Instagram and Facebook now, being able to show people content that's interesting to  them from people that they're not even following.

7:09

But that ended up being a very good decision  in retrospect.

7:09

And it came from being behind.

7:18

It wasn't like “oh, I was so far ahead.

7:18

”  Actually, most of the times where we make some decision that ends up seeming good  is because we messed something up before and just didn't want to repeat the mistake.

7:29

This is a total detour, but I want to ask about this while we're on this.

7:32

We'll get back  to AI in a second.

7:32

In 2006 you didn't sell for $1 billion but presumably there's some amount you  would have sold for, right?

7:37

Did you write down in your head like “I think the actual valuation  of Facebook at the time is this and they're not actually getting the valuation right”?

7:45

If they’d  offered you $5 trillion, of course you would have sold.

7:48

So how did you think about that choice?

7:48

I think some of these things are just personal.

7:58

I don't know that at the time I was sophisticated  enough to do that analysis.

7:58

I had all these people around me who were making all these arguments for  a billion dollars like “here's the revenue that we need to make and here's how big we need to be.

8:10

It's clearly so many years in the future.

8:10

” It was very far ahead of where we were at the time.

8:16

I  didn't really have the financial sophistication to really engage with that kind of debate.

8:23

Deep down I believed in what we were doing.

8:30

I did some analysis like “what would I do if I  weren’t doing this?

8:30

Well, I really like building things and I like helping people communicate.

8:40

I  like understanding what's going on with people and the dynamics between people.

8:46

So I think if I sold  this company, I'd just go build another company like this and I kind of like the one I have. So why?

8:51

” I think a lot of the biggest bets that people make are often just based on conviction and  values.

9:03

It's actually usually very hard to do the analyses trying to connect the dots forward.

9:12

You've had Facebook AI Research for a long time.

9:18

Now it's become seemingly central to  your company.

9:18

At what point did making AGI, or however you consider that mission,  become a key priority of what Meta is doing?

9:33

It's been a big deal for a while.

9:33

We started  FAIR about 10 years ago.

9:33

The idea was that, along the way to general intelligence or whatever  you wanna call it, there are going to be all these different innovations and that's going to  just improve everything that we do.

9:48

So we didn't conceive of it as a product.

9:52

It was  more of a research group.

9:52

Over the last 10 years it has created a lot of different things  that have improved all of our products.

10:00

It’s advanced the field and allowed other people in  the field to create things that have improved our products too.

10:11

I think that that's been great.

10:11

There's obviously a big change in the last few years with ChatGPT and the diffusion  models around image creation coming out.

10:24

This is some pretty wild stuff that is  pretty clearly going to affect how people interact with every app that's out there.

10:29

At that  point we started a second group, the gen AI group, with the goal of bringing that stuff into our  products and building leading foundation models that would power all these different products.

10:46

When we started doing that the theory initially was that a lot of the stuff we're doing is  pretty social.

10:54

It's helping people interact with creators, helping people interact with  businesses, helping businesses sell things or do customer support.

11:07

There’s also basic assistant  functionality, whether it's for our apps or the smart glasses or VR.

11:13

So it wasn't completely  clear at first that you were going to need full AGI to be able to support those use cases.

11:24

But in  all these subtle ways, through working on them, I think it's actually become clear that you do.

11:29

For example, when we were working on Llama-2, we didn't prioritize coding because people  aren't going to ask Meta AI a lot of coding questions in WhatsApp. Now they will, right? I don't know.

11:44

I'm not sure that WhatsApp, or  Facebook or Instagram, is the UI where people are going to be doing a lot of coding questions.

11:47

Maybe  the website, meta.

11:47

ai, that we’re launching.

11:47

But the thing that has been a somewhat surprising  result over the last 18 months is that it turns out that coding is important for a lot of domains,  not just coding.

12:08

Even if people aren't asking coding questions, training the models on coding  helps them become more rigorous in answering the question and helps them reason across a lot of  different types of domains.

12:21

That's one example where for Llama-3, we really focused on training  it with a lot of coding because that's going to make it better on all these things even if  people aren't asking primarily coding questions.

12:36

Reasoning is another example.

12:36

Maybe you want  to chat with a creator or you're a business and you're trying to interact with a customer.

12:43

That interaction is not just like “okay, the person sends you a message and you  just reply.

12:47

” It's a multi-step interaction where you're trying to think through “how do I  accomplish the person's goals?

12:53

” A lot of times when a customer comes, they don't necessarily  know exactly what they're looking for or how to ask their questions.

13:01

So it's not really the  job of the AI to just respond to the question.

13:06

You need to kind of think about it  more holistically.

13:06

It really becomes a reasoning problem.

13:09

So if someone else solves  reasoning, or makes good advances on reasoning, and we're sitting here with a basic chat bot,  then our product is lame compared to what other people are building.

13:19

At the end of the day, we  basically realized we've got to solve general intelligence and we just upped the ante and the  investment to make sure that we could do that.

13:32

So the version of
Llama that's going to solve  all these use cases for users, is that the version that will be powerful enough to replace  a programmer you might have in this building?

13:46

I just think that all this stuff is  going to be progressive over time.

13:49

But in the end case: Llama-10.

13:49

I think that there's a lot baked into that question.

13:55

I'm not sure that we're  replacing people as much as we’re giving people tools to do more stuff.

14:00

Is the programmer in this building 10x more productive after Llama-10? I would hope more.

14:03

I don't believe that there's a single threshold of intelligence for  humanity because people have different skills.

14:14

I think that at some point AI is probably going to  surpass people at most of those things, depending on how powerful the models are.

14:21

But I think it's  progressive and I don't think AGI is one thing.

14:29

You're basically adding different capabilities.

14:29

Multimodality is a key one that we're focused on now, initially with photos and images and text but  eventually with videos.

14:34

Because we're so focused on the metaverse, 3D type stuff is important  too.

14:40

One modality that I'm pretty focused on, that I haven't seen as many other people in the  industry focus on, is emotional understanding.

14:46

So much of the human brain is just dedicated  to understanding people and understanding expressions and emotions.

15:00

I think that's  its own whole modality, right?

15:00

You could say that maybe it's just video or image, but it's  clearly a very specialized version of those two.

15:10

So there are all these different capabilities  that you want to train the models to focus on, in addition to getting a lot better at  reasoning and memory, which is its own whole thing.

15:22

I don't think in the future we're going to  be primarily shoving things into a query context window to ask more complicated questions.

15:29

There  will be different stores of memory or different custom models that are more personalized to  people.

15:35

These are all just different capabilities.

15:42

Obviously then there’s making them big and small. We care about both.

15:42

If you're running something like Meta AI, that's pretty server-based.

15:47

We also  want it running on smart glasses and there's not a lot of space in smart glasses.

15:55

So you want to  have something that's very efficient for that.

16:01

If you're doing $10Bs worth of  inference or even eventually $100Bs, if you're using intelligence in an industrial  scale what is the use case? Is it simulations?

16:11

Is it the AIs that will be in the metaverse?

16:11

What will we be using the data centers for?

16:19

Our bet is that it's going to basically change  all of the products.

16:19

I think that there's going to be a kind of Meta AI general assistant  product.

16:24

I think that that will shift from something that feels more like a chatbot, where  you ask a question and it formulates an answer, to things where you're giving it more complicated  tasks and then it goes away and does them.

16:37

That's going to take a lot of inference and it's going  to take a lot of compute in other ways too.

16:48

Then I think interacting with other agents for  other people is going to be a big part of what we do, whether it's for businesses or creators.

16:56

A  big part of my theory on this is that there's not going to be just one singular AI that you interact  with.

17:02

Every business is going to want an AI that represents their interests.

17:09

They're not going to  want to primarily interact with you through an AI that is going to sell their competitors’ products.

17:13

I think creators is going to be a big one.

17:13

There are about 200 million creators on our platforms.

17:25

They basically all have the pattern where they want to engage their community but they're limited  by the hours in the day.

17:31

Their community generally wants to engage them, but they don't know that  they're limited by the hours in the day.

17:35

If you could create something where that creator  can basically own the AI, train it in the way they want, and engage their community, I think  that's going to be super powerful.

17:47

There's going to be a ton of engagement across all these things.

17:55

These are just the consumer use cases.

17:55

My wife and I run our foundation, Chan Zuckerberg Initiative.

18:04

We're doing a bunch of stuff on science and there's obviously a lot of AI work that is going  to advance science and healthcare and all these things.

18:17

So it will end up affecting basically  every area of the products and the economy.

18:25

You mentioned AI that can just go out and do  something for you that's multi-step. Is that a bigger model?

18:30

With Llama-4 for example, will  there still be a version that's 70B but you'll just train it on the right data and that will  be super powerful?

18:36

What does the progression look like? Is it scaling?

18:40

Is it just the same size  but different banks like you were talking about?

18:49

I don't know that we know the answer to that.

18:49

I  think one thing that seems to be a pattern is that you have the Llama model and then you build some  kind of other application specific code around it.

19:06

Some of it is the fine-tuning for the use case,  but some of it is, for example, logic for how Meta AI should work with tools like Google or Bing  to bring in real-time knowledge.

19:14

That's not part of the base Llama model.

19:21

For Llama-2, we had some  of that and it was a little more hand-engineered.

19:30

Part of our goal for Llama-3 was to bring more  of that into the model itself.

19:30

For Llama-3, as we start getting into more of these agent-like  behaviors, I think some of that is going to be more hand-engineered.

19:41

Our goal for Llama-4  will be to bring more of that into the model.

19:48

At each step along the way you have a sense of  what's going to be possible on the horizon.

19:48

You start messing with it and hacking around it.

19:54

I  think that helps you then hone your intuition for what you want to try to train into the next  version of the model itself.

19:59

That makes it more general because obviously for anything that you're  hand-coding you can unlock some use cases, but it's just inherently brittle and non-general.

20:10

When you say “into the model itself,” you train it on the thing that you want in the model itself?

21:21

What do you mean by “into the model itself”?

21:33

For Llama- 2, the tool use was very specific,  whereas Llama-3 has much better tool use.

21:33

We don't have to hand code all the stuff to have  it use Google and go do a search. It can just do that.

21:49

Similarly for coding and running code and  a bunch of stuff like that.

21:49

Once you kind of get that capability, then you get a peek at what we  can start doing next.

22:00

We don't necessarily want to wait until Llama-4 is around to start building  those capabilities, so we can start hacking around it.

22:10

You do a bunch of hand coding and that  makes the products better, if only for the interim.

22:16

That helps show the way then of what we  want to build into the next version of the model.

22:21

What is the community fine tune of Llama-3  that you're most excited for?

22:21

Maybe not the one that will be most useful to you, but the  one you'll just enjoy playing with the most.

22:29

They fine-tune it on antiquity and  you'll just be talking to Virgil or something.

22:32

What are you excited about?

22:32

I think the nature of the stuff is that you get surprised.

22:39

Any specific thing that I thought  would be valuable, we'd probably be building.

22:39

I think you'll get distilled versions.

22:53

I  think you'll get smaller versions.

22:53

One thing is that I think 8B isn’t quite small  enough for a bunch of use cases.

22:58

Over time I'd love to get a 1-2B parameter model, or even a 500M  parameter model and see what you can do with that.

23:18

If with 8B parameters we’re nearly as  powerful as the largest Llama-2 model, then with a billion parameters you should be able  to do something that's interesting, and faster.

23:28

It’d be good for classification, or a lot of  basic things that people do before understanding the intent of a user query and feeding it  to the most powerful model to hone in on what the prompt should be.

23:41

I think that's one  thing that maybe the community can help fill in.

23:46

We're also thinking about getting around to  distilling some of these ourselves but right now the GPUs are pegged training the 405B.

23:52

So you have all these GPUs.

23:52

I think you said 350,000 by the end of the year. That's the whole fleet.

24:00

We built two, I think 22,000 or 24,000 clusters that are the  single clusters that we have for training the big models, obviously across a lot of the stuff that  we do.

24:13

A lot of our stuff goes towards training Reels models and Facebook News Feed and Instagram  Feed.

24:18

Inference is a huge thing for us because we serve a ton of people.

24:24

Our ratio of inference  compute required to training is probably much higher than most other companies that are doing  this stuff just because of the sheer volume of the community that we're serving.

24:37

In the material they shared with me before, it was really interesting that you  trained it on more data than is compute optimal just for training.

24:45

The inference is such a big  deal for you guys, and also for the community, that it makes sense to just have this thing  and have trillions of tokens in there.

24:53

Although one of the interesting  things about it, even with the 70B, is that we thought it would get more saturated.

24:57

We  trained it on around 15 trillion tokens.

24:57

I guess our prediction going in was that it was going  to asymptote more, but even by the end it was still learning.

25:12

We probably could have fed it more  tokens and it would have gotten somewhat better.

25:19

At some point you're running a company and you  need to do these meta reasoning questions.

25:19

Do I want to spend our GPUs on training the 70B model  further?

25:24

Do we want to get on with it so we can start testing hypotheses for Llama-4?

25:31

We needed  to make that call and I think we got a reasonable balance for this version of the 70B.

25:39

There'll  be others in the future, the 70B multimodal one, that'll come over the next period.

25:45

But that  was fascinating that the architectures at this point can just take so much data.

25:53

That's really interesting.

25:53

What does this imply about future models?

25:57

You mentioned that  the Llama-3 8B is better than the Llama-2 70B.

26:03

No, no, it's nearly as good.

26:03

I don’t want to overstate it.

26:06

It’s in a similar order of magnitude.

26:06

Does that mean the Llama-4 70B will be as good as the Llama-3 405B?

26:10

What  does the future of this look like?

26:14

This is one of the great questions, right? I think  no one knows.

26:14

One of the trickiest things in the world to plan around is an exponential  curve.

26:22

How long does it keep going for?

26:29

I think it's likely enough that we'll keep going.

26:29

I think it’s worth investing the $10Bs or $100B+ in building the infrastructure and assuming that  if it keeps going you're going to get some really amazing things that are going to make amazing  products.

26:43

I don't think anyone in the industry can really tell you that it will continue scaling  at that rate for sure.

26:49

In general in history, you hit bottlenecks at certain points.

26:56

Now there's so much energy on this that maybe those bottlenecks get knocked over pretty  quickly.

27:01

I think that’s an interesting question.

27:08

What does the world look like where there aren't  these bottlenecks?

27:08

Suppose progress just continues at this pace, which seems plausible.

27:13

Zooming out and forgetting about Llamas… Well, there are going to be different bottlenecks.

27:18

Over the last few years, I think there was this issue of GPU production.

27:28

Even companies that had  the money to pay for the GPUs couldn't necessarily get as many as they wanted because there were all  these supply constraints.

27:39

Now I think that's sort of getting less.

27:44

So you're seeing a bunch of  companies thinking now about investing a lot of money in building out these things.

27:52

I think  that that will go on for some period of time.

28:00

There is a capital question.

28:00

At what point does  it stop being worth it to put the capital in?

28:06

I actually think before we hit that, you're  going to run into energy constraints.

28:06

I don't think anyone's built a gigawatt single training  cluster yet.

28:14

You run into these things that just end up being slower in the world.

28:21

Getting energy  permitted is a very heavily regulated government function.

28:30

You're going from software, which  is somewhat regulated and I'd argue it’s more regulated than a lot of people in the tech  community feel.

28:37

Obviously it’s different if you're starting a small company, maybe you  feel that less.

28:42

We interact with different governments and regulators and we have lots  of rules that we need to follow and make sure we do a good job with around the world.

28:53

But  I think that there's no doubt about energy.

28:59

If you're talking about building large new  power plants or large build-outs and then building transmission lines that cross other  private or public land, that’s just a heavily regulated thing.

29:11

You're talking about many  years of lead time.

29:11

If we wanted to stand up some massive facility, powering that is a very  long-term project.

29:17

I think people do it but I don't think this is something that can be quite  as magical as just getting to a level of AI, getting a bunch of capital and putting it in, and  then all of a sudden the models are just going to… You do hit different bottlenecks along the way.

29:42

Is there something, maybe an AI-related project or maybe not, that even a company like Meta doesn't  have the resources for?

29:47

Something where if your R&D budget or capex budget were 10x what it is  now, then you could pursue it?

29:51

Something that’s in the back of your mind but with Meta today,  you can't even issue stock or bonds for it?

30:01

It's just like 10x bigger than your budget?

30:01

I think energy is one piece.

30:01

I think we would probably build out bigger clusters than we  currently can if we could get the energy to do it.

30:18

That's fundamentally money-bottlenecked  in the limit?

30:18

If you had $1 trillion… I think it’s time.

30:23

It depends on how far the  exponential curves go.

30:23

Right now a lot of data centers are on the order of 50 megawatts or  100MW, or a big one might be 150MW.

30:36

Take a whole data center and fill it up with all the stuff  that you need to do for training and you build the biggest cluster you can.

30:46

I think a bunch  of companies are running at stuff like that.

30:53

But when you start getting into building a  data center that's like 300MW or 500MW or 1 GW, no one has built a 1GW data center yet. I think  it will happen.

31:04

This is only a matter of time but it's not going to be next year.

31:09

Some of these  things will take some number of years to build out.

31:18

Just to put this in perspective, I think a  gigawatt would be the size of a meaningful nuclear power plant only going towards training a model. Didn't Amazon do this?

31:31

They have a 950MW– I'm not exactly sure what they  did. You'd have to ask them.

31:44

But it doesn’t have to be in the  same place, right?

31:44

If distributed training works, it can be distributed.

31:45

Well, I think that is a big question, how that's going to work.

31:49

It seems quite possible that  in the future, more of what we call training for these big models is actually more along the lines  of inference generating synthetic data to then go feed into the model.

32:05

I don't know what that ratio  is going to be but I consider the generation of synthetic data to be more inference than training  today.

32:11

Obviously if you're doing it in order to train a model, it's part of the broader  training process.

32:16

So that's an open question, the balance of that and how that plays out.

32:24

Would that potentially also be the case with Llama-3, and maybe Llama-4 onwards?

32:30

As in, you  put this out and if somebody has a ton of compute, then they can just keep making these things  arbitrarily smarter using the models that you've put out.

32:37

Let’s say there’s some  random country, like Kuwait or the UAE, that has a ton of compute and they can actually  just use Llama-4 to make something much smarter.

32:52

I do think there are going to be  dynamics like that, but I also think there is a fundamental limitation on the model  architecture.

32:59

I think like a 70B model that we trained with a Llama-3 architecture can get  better, it can keep going.

33:13

As I was saying, we felt that if we kept on feeding it more data  or rotated the high value tokens through again, then it would continue getting better.

33:24

We've  seen a bunch of different companies around the world basically take the Llama-2 70B model  architecture and then build a new model.

33:31

But it's still the case that when you make a generational  improvement to something like the Llama-3 70B or the Llama-3 405B, there isn’t anything like  that open source today.

33:46

I think that's a big step function.

33:54

What people are going to be able to  build on top of that I think can’t go infinitely from there.

33:59

There can be some optimization in  that until you get to the next step function.

34:05

Let's zoom out a little bit from specific  models and even the multi-year lead times you would need to get energy approvals and so  on.

34:11

Big picture, what's happening with AI these next couple of decades?

34:15

Does it feel like  another technology like the metaverse or social, or does it feel like a fundamentally  different thing in the course of human history?

34:29

I think it's going to be pretty fundamental.

34:29

I  think it's going to be more like the creation of computing in the first place.

34:34

You'll get all  these new apps in the same way as when you got the web or you got mobile phones.

34:44

People basically  rethought all these experiences as a lot of things that weren't possible before became possible.

34:50

So I think that will happen, but I think it's a much lower-level innovation.

34:56

My sense is  that it's going to be more like people going from not having computers to having computers.

35:01

It’s very hard to reason about exactly how this goes.

35:16

In the cosmic scale obviously it'll happen  quickly, over a couple of decades or something.

35:27

There is some set of people who are afraid of it  really spinning out and going from being somewhat intelligent to extremely intelligent overnight.

35:33

I just think that there's all these physical constraints that make that unlikely to happen.

35:37

I  just don't really see that playing out.

35:37

I think we'll have time to acclimate a bit.

35:45

But it will  really change the way that we work and give people all these creative tools to do different things.

35:51

I think it's going to really enable people to do the things that they want a lot more.

36:00

So maybe not overnight, but is it your view that on a cosmic scale we can think of  these milestones in this way?

36:05

Humans evolved, and then AI happened, and then they went out  into the galaxy.

36:09

Maybe it takes many decades, maybe it takes a century, but is that the grand  scheme of what's happening right now in history? Sorry, in what sense?

36:22

In the sense that there were other technologies, like computers and even  fire, but the development of AI itself is as significant as humans evolving in the first place. I think that's tricky.

36:29

The history of humanity has been people basically thinking that certain  aspects of humanity are really unique in different ways and then coming to grips with the fact that  that's not true, but that humanity is actually still super special.

36:57

We thought that the earth  was the center of the universe and it's not, but humans are still pretty  awesome and pretty unique, right?

37:12

I think another bias that people tend  to have is thinking that intelligence is somehow fundamentally connected to life.

37:17

It's not actually clear that it is.

37:17

I don't know that we have a clear enough definition of  consciousness or life to fully interrogate this.

37:42

There's all this science fiction about creating  intelligence where it starts to take on all these human-like behaviors and things like that.

37:47

The  current incarnation of all this stuff feels like it's going in a direction where intelligence  can be pretty separated from consciousness, agency, and things like that, which I  think just makes it a super valuable tool.

38:06

Obviously it's very difficult to predict  what direction this stuff goes in over time, which is why I don't think anyone should be  dogmatic about how they plan to develop it or what they plan to do.

38:16

You want to look  at it with each release.

38:16

We're obviously very pro open source, but I haven't committed  to releasing every single thing that we do.

38:27

I’m basically very inclined to think that  open sourcing is going to be good for the community and also good for us because we'll  benefit from the innovations.

38:32

If at some point however there's some qualitative change in what  the thing is capable of, and we feel like it's not responsible to open source it, then we  won't.

38:43

It's all very difficult to predict.

38:52

What is a kind of specific qualitative change  where you'd be training Llama-5 or Llama-4, and if you see it, it’d make you think “you know  what, I'm not sure about open sourcing it”?

39:05

It's a little hard to answer that in  the abstract because there are negative behaviors that any product can exhibit  where as long as you can mitigate it, it's okay.

39:15

There’s bad things about social media  that we work to mitigate.

39:15

There's bad things about Llama-2 where we spend a lot of time trying  to make sure that it's not like helping people commit violent acts or things like that.

39:28

That  doesn't mean that it's a kind of autonomous or intelligent agent.

39:34

It just means that it's learned  a lot about the world and it can answer a set of questions that we think would be unhelpful for it  to answer.

39:38

I think the question isn't really what behaviors would it show, it's what things would  we not be able to mitigate after it shows that.

39:59

I think that there's so many ways in which  something can be good or bad that it's hard to actually enumerate them all up front.

40:03

Look at  what we've had to deal with in social media and the different types of harms.

40:10

We've basically  gotten to like 18 or 19 categories of harmful things that people do and we've basically built  AI systems to identify what those things are and to make sure that doesn't happen on our network  as much as possible.

40:23

Over time I think you'll be able to break this down into more of a  taxonomy too.

40:29

I think this is a thing that we spend time researching as well, because we  want to make sure that we understand that.

41:46

It seems to me that it would be a good idea.

41:46

I would be disappointed in a future where AI systems aren't broadly deployed and everybody  doesn't have access to them.

41:50

At the same time, I want to better understand the mitigations.

41:55

If the mitigation is the fine-tuning, the whole thing about open weights is that you  can then remove the fine-tuning, which is often superficial on top of these capabilities.

42:06

If it's  like talking on Slack with a biology researcher… I think models are very far from this.

42:12

Right  now, they’re like Google search.

42:12

But if I can show them my Petri dish and they can explain why  my smallpox sample didn’t grow and what to change, how do you mitigate that?

42:23

Because somebody  can just fine-tune that in there, right? That's true.

42:29

I think a lot of people will  basically use the off-the-shelf model and some people who have basically bad faith are going to  try to strip out all the bad stuff.

42:35

So I do think that's an issue.

42:41

On the flip side, one of the  reasons why I'm philosophically so pro open source is that I do think that a concentration of AI in  the future has the potential to be as dangerous as it being widespread.

43:02

I think a lot of people think  about the questions of “if we can do this stuff, is it bad for it to be out in the wild and just  widely available?

43:08

” I think another version of this is that it's probably also pretty bad  for one institution to have an AI that is way more powerful than everyone else's AI.

43:25

There’s one security analogy that I think of.

43:31

There are so many security holes in so many  different things.

43:31

If you could travel back in time a year or two years, let's say you just have  one or two years more knowledge of the security holes.

43:50

You can pretty much hack into any system. That’s not AI.

43:50

So it's not that far-fetched to believe that a very intelligent AI probably would  be able to identify some holes and basically be like a human who could go back in time a  year or two and compromise all these systems.

44:07

So how have we dealt with that as a society?

44:07

One big part is open source software that makes it so that when improvements are made to  the software, it doesn't just get stuck in one company's products but can be broadly deployed to  a lot of different systems, whether they’re banks or hospitals or government stuff.

44:24

As the software  gets hardened, which happens because more people can see it and more people can bang on it, there  are standards on how this stuff works.

44:31

The world can get upgraded together pretty quickly.

44:37

I think that a world where AI is very widely deployed, in a way where it's gotten hardened  progressively over time, is one where all the different systems will be in check in a way.

44:52

That  seems fundamentally more healthy to me than one where this is more concentrated.

44:58

So there are  risks on all sides, but I think that's a risk that I don't hear people talking about quite as  much.

45:05

There's the risk of the AI system doing something bad.

45:13

But I stay up at night worrying  more about an untrustworthy actor having the super strong AI, whether it's an adversarial government  or an untrustworthy company or whatever.

45:27

I think that that's potentially a much bigger risk.

45:39

As in, they could overthrow our government because they have a weapon that nobody else has?

45:47

Or just cause a lot of mayhem.

45:47

I think the intuition is that this stuff ends up being  pretty important and valuable for both economic and security reasons and other things.

46:01

If someone whom you don't trust or an adversary gets something more powerful, then I think that  that could be an issue.

46:11

Probably the best way to mitigate that is to have good open source  AI that becomes the standard and in a lot of ways can become the leader.

46:24

It just ensures that  it's a much more even and balanced playing field.

46:33

That seems plausible to me.

46:33

If that works out,  that would be the future I prefer.

46:33

I want to understand mechanistically how the fact that  there are open source AI systems in the world prevents somebody causing mayhem with their AI  system?

46:47

With the specific example of somebody coming with a bioweapon, is it just that we'll do  a bunch of R&D in the rest of the world to figure out vaccines really fast? What's happening?

46:55

If you take the security one that I was talking about, I think someone with  a weaker AI trying to hack into a system that is protected by a stronger AI will  succeed less.

47:03

In terms of software security– How do we know everything in the world is like  that?

47:12

What if bioweapons aren't like that?

47:16

I mean, I don't know that everything in the  world is like that.

47:16

Bioweapons are one of the areas where the people who are most worried about  this stuff are focused and I think it makes a lot of sense.

47:33

There are certain mitigations.

47:33

You  can try to not train certain knowledge into the model.

47:42

There are different things but at  some level if you get a sufficiently bad actor, and you don't have other AI that can balance  them and understand what the threats are, then that could be a risk.

48:00

That's one of  the things that we need to watch out for.

48:05

Is there something you could see in the deployment  of these systems where you're training Llama-4 and it lied to you because it thought you weren't  noticing or something and you're like “whoa what's going on here?

48:17

” This is probably not  likely with a Llama-4 type system, but is there something you can imagine like that where  you'd be really concerned about deceptiveness and billions of copies of this being out in the wild?

48:27

I mean right now we see a lot of hallucinations. It's more so that.

48:37

I think it's an interesting  question, how you would tell the difference between hallucination and deception.

48:43

There are  a lot of risks and things to think about.

48:43

I try, in running our company at least, to balance  these longer-term theoretical risks with what I actually think are quite real risks that  exist today.

49:07

So when you talk about deception, the form of that that I worry about most is  people using this to generate misinformation and then pump that through our networks or  others.

49:18

The way that we've combated this type of harmful content is by building AI systems  that are smarter than the adversarial ones.

49:33

This informs part of my theory on this.

49:33

If you  look at the different types of harm that people do or try to do through social networks, there are  ones that are not very adversarial.

49:38

For example, hate speech is not super adversarial in the sense  that people aren't getting better at being racist.

50:03

That's one where I think the AIs are generally  getting way more sophisticated faster than people are at those issues.

50:08

And we have issues both  ways.

50:08

People do bad things, whether they're trying to incite violence or something, but  we also have a lot of false positives where we basically censor stuff that we shouldn't.

50:20

I think  that understandably makes a lot of people annoyed.

50:25

So I think having an AI that gets increasingly  precise on that is going to be good over time.

50:30

But let me give you another example: nation  states trying to interfere in elections.

50:30

That's an example where they absolutely have cutting edge  technology and absolutely get better each year.

50:35

So we block some technique, they learn what we did  and come at us with a different technique.

50:41

It's not like a person trying to say mean things, They  have a goal. They're sophisticated.

50:46

They have a lot of technology.

50:56

In those cases, I still think  about the ability to have our AI systems grow in sophistication at a faster rate than theirs do.

51:04

It's an arms race but I think we're at least winning that arms race currently.

51:09

This is a lot  of the stuff that I spend time thinking about.

51:18

Yes, whether it's Llama-4 or Llama-6, we need to  think about what behaviors we're observing and it's not just us.

51:26

Part of the reason why you make  this open source is that there are a lot of other people who study this too.

51:29

So we want to see what  other people are observing, what we’re observing, what we can mitigate, and then we'll make  our assessment on whether we can make it open source.

51:40

For the foreseeable future I'm  optimistic we will be able to.

51:40

In the near term, I don't want to take our eye off the ball  in terms of what are actual bad things that people are trying to use the models for today.

51:53

Even if they're not existential, there are pretty bad day-to-day harms that we're familiar  with in running our services.

51:58

That's actually a lot of what we have to spend our time on as well.

52:05

I found the synthetic data thing really curious.

52:14

With current models it makes sense why there might  be an asymptote with just doing the synthetic data again and again.

52:19

But let’s say they get smarter  and you use the kinds of techniques—you talk about in the paper or the blog posts that are coming out  on the day this will be released—where it goes to the thought chain that is the most correct.

52:29

Why do you think this wouldn't lead to a loop where it gets smarter, makes better output, gets  smarter and so forth.

52:36

Of course it wouldn't be overnight, but over many months or years of  training potentially with a smarter model.

52:45

I think it could, within the parameters of  whatever the model architecture is.

52:45

It's just that with today's 8B parameter models, I don't  think you're going to get to be as good as the state-of-the-art multi-hundred billion  parameter models that are incorporating new research into the architecture itself.

53:08

But those will be open source as well, right?

53:15

Well yeah, subject to all the questions that we  just talked about but yes.

53:15

We would hope that that'll be the case.

53:23

But I think that at each  point, when you're building software there's a ton of stuff that you can do with software but  then at some level you're constrained by the chips that it's running on.

53:34

So there are always  going to be different physical constraints.

53:34

How big the models are is going to be constrained  by how much energy you can get and use for inference.

53:49

I'm simultaneously very optimistic  that this stuff will continue to improve quickly and also a little more measured than I think  some people are about it.

53:59

I don’t think the runaway case is a particularly likely one.

54:11

I think it makes sense to keep your options open.

54:17

There's so much we don't know.

54:17

There's a  case in which it's really important to keep the balance of power so nobody becomes a totalitarian  dictator.

54:22

There's a case in which you don't want to open source the architecture because China can  use it to catch up to America's AIs and there is an intelligence explosion and they win that.

54:32

A lot  of things seem possible.

54:32

Keeping your options open considering all of them seems reasonable. Yeah.

54:42

Let's talk about some other things. Metaverse.

54:42

What time period in human history would you be most interested in going into?

54:48

100,000 BCE to  now, you just want to see what it was like? It has to be the past?

54:53

Oh yeah, it has to be the past.

55:04

I'm really interested in American history and  classical history.

55:04

I'm really interested in the history of science too.

55:10

I actually think seeing  and trying to understand more about how some of the big advances came about would be interesting.

55:19

All we have are somewhat limited writings about some of that stuff.

55:24

I'm not sure the metaverse  is going to let you do that because it's going to be hard to go back in time for things that  we don't have records of.

55:29

I'm actually not sure that going back in time is going to be that  important of a thing.

55:38

I think it's going to be cool for like history classes and stuff,  but that's probably not the use case that I'm most excited about for the metaverse overall.

55:47

The main thing is just the ability to feel present with people, no matter where you are.

55:53

I think that's going to be killer.

55:53

In the AI conversation that we were having, so much of it  is about physical constraints that underlie all of this.

56:08

I think one lesson of technology is  that you want to move things from the physical constraint realm into software as much as possible  because software is so much easier to build and evolve.

56:20

You can democratize it more because  not everyone is going to have a data center but a lot of people can write code and take open  source code and modify it.

56:26

Τhe metaverse version of this is enabling realistic digital  presence.

56:33

That’s going to be an absolutely huge difference so people don't feel like they have  to be physically together for as many things.

56:51

Now I think that there can be things that are  better about being physically together.

56:51

These things aren't binary.

56:57

It's not going to be like  “okay, now you don't need to do that anymore.

56:57

” But overall, I think it's just going to be  really powerful for socializing, for feeling connected with people, for working, for parts  of industry, for medicine, for so many things.

57:20

I want to go back to something you said at the  beginning of the conversation.

57:20

You didn't sell the company for a billion dollars.

57:23

And with  the metaverse, you knew you were going to do this even though the market was hammering  you for it. I'm curious.

57:26

What is the source of that edge?

57:31

You said “oh, values, I have  this intuition,” but everybody says that.

57:31

If you had to say something that's specific to  you, how would you express what that is?

57:37

Why were you so convinced about the metaverse?

57:41

I think that those are different questions.

57:52

What are the things that power me?

57:52

We've  talked about a bunch of the themes.

57:52

I just really like building things.

58:02

I specifically like  building things around how people communicate and understanding how people express themselves  and how people work.

58:10

When I was in college I studied computer science and psychology.

58:13

I  think a lot of other people in the industry studied computer science.

58:18

So, it's always been  the intersection of those two things for me.

58:27

It’s also sort of this really deep drive.

58:27

I  don't know how to explain it but I just feel constitutionally that I'm doing something wrong if  I'm not building something new.

58:36

Even when we were putting together the business case for investing  a $100 billion in AI or some huge amount in the metaverse, we have plans that I think made  it pretty clear that if our stuff works, it'll be a good investment.

59:03

But you can't know  for certain from the outset.

59:03

There are all these arguments that people have, with advisors  or different folks.

59:10

It's like, “how are you confident enough to do this?

59:19

” Well the day I stop  trying to build new things, I'm just done.

59:19

I'm going to go build new things somewhere else.

59:26

I'm  fundamentally incapable of running something, or in my own life, and not trying to build new  things that I think are interesting.

59:37

That's not even a question for me, whether we're going to  take a swing at building the next thing.

59:43

I'm just incapable of not doing that. I don't know.

59:51

I'm kind of like this in all the different aspects of my life.

1:00:01

Our family built this ranch in Kauai  and I worked on designing all these buildings.

1:00:01

We started raising cattle and I'm like “alright, I  want to make the best cattle in the world so how do we architect this so that way we can figure  this out and build all the stuff up that we need to try to do that.

1:00:24

” I don't know, that's  me.

1:00:24

What was the other part of the question?

1:01:37

I'm not sure but I'm actually curious  about something else.

1:01:37

So a 19-year-old Mark reads a bunch of antiquity and  classics in high school and college.

1:01:48

What important lesson did you learn from  it?

1:01:48

Not just interesting things you found, but there aren't that many tokens you consume by  the time you're 19.

1:01:50

A bunch of them were about the classics.

1:01:55

Clearly that was important in some way.

1:01:55

There aren't that many tokens you consume... That's a good question.

1:02:06

Here’s one of the things  I thought was really fascinating.

1:02:06

Augustus became emperor and he was trying to establish peace.

1:02:19

There was no real conception of peace at the time.

1:02:30

The people's understanding of peace was  peace as the temporary time between when your enemies inevitably attack you.

1:02:36

So you get a  short rest.

1:02:36

He had this view of changing the economy from being something mercenary and  militaristic to this actually positive-sum thing.

1:02:53

It was a very novel idea at the time.

1:02:53

That’s something that's really fundamental: the bounds on what people can conceive  of at the time as rational ways to work.

1:03:17

This applies to both the metaverse and the AI  stuff.

1:03:17

A lot of investors, and other people, can't wrap their head around why we would open  source this.

1:03:22

It’s like “I don't understand, it’s open source.

1:03:29

That must just be the temporary time  between which you're making things proprietary, right?

1:03:34

” I think it's this very profound thing in  tech that it actually creates a lot of winners.

1:03:49

I don't want to strain the analogy too  much but I do think that a lot of the time, there are models for building things that  people often can't even wrap their head around.

1:04:06

They can’t understand how that would be a  valuable thing for people to do or how it would be a reasonable state of the world.

1:04:11

I think there  are more reasonable things than people think.

1:04:20

That's super fascinating.

1:04:20

Can I give you what  I was thinking in terms of what you might have gotten from it?

1:04:24

This is probably totally off,  but I think it’s just how young some of these people are, who have very important roles  in the empire.

1:04:29

For example, Caesar Augustus, by the time he’s 19, is already one of the most  important people in Roman politics.

1:04:33

He's leading battles and forming the Second Triumvirate.

1:04:39

I  wonder if the 19-year-old you was thinking “I can do this because Caesar Augustus did this.

1:04:42

” That's an interesting example, both from a lot of history and American history too.

1:04:48

One of my  favorite quotes is this Picasso quote that all children are artists and the challenge is to  remain an artist as you grow up.

1:04:56

When you’re younger, it’s just easier to have wild ideas.

1:05:02

There are all these analogies to the innovator’s dilemma that exist in your life as well as for  your company or whatever you’ve built.

1:05:14

You’re earlier on in your trajectory so it's easier to  pivot and take in new ideas without disrupting other commitments to different things.

1:05:26

I think that's an interesting part of running a company. How do you stay dynamic?

1:05:33

Let’s go back to the investors and open source.

1:05:41

The $10B model, suppose it's totally safe.

1:05:41

You've  done these evaluations and unlike in this case the evaluators can also fine-tune the model, which  hopefully will be the case in future models.

1:05:47

Would you open source the $10 billion model?

1:05:52

As long as it's helping us then yeah. But would it?

1:05:57

$10 billion of  R&D and now it's open source.

1:06:01

That’s a question which we’ll have to evaluate  as time goes on too.

1:06:01

We have a long history of open sourcing software.

1:06:11

We don’t tend to open  source our product.

1:06:11

We don't take the code for Instagram and make it open source.

1:06:18

We take  a lot of the low-level infrastructure and we make that open source.

1:06:24

Probably the biggest  one in our history was our Open Compute Project where we took the designs for all of our servers,  network switches, and data centers, and made it open source and it ended up being super helpful.

1:06:36

Although a lot of people can design servers the industry now standardized on our design, which  meant that the supply chains basically all got built out around our design.

1:06:46

So volumes went  up, it got cheaper for everyone, and it saved us billions of dollars which was awesome.

1:06:50

So there's multiple ways where open source could be helpful for us.

1:06:56

One is if people figure  out how to run the models more cheaply.

1:06:56

We're going to be spending tens, or a hundred billion  dollars or more over time on all this stuff.

1:07:01

So if we can do that 10% more efficiently, we're  saving billions or tens of billions of dollars.

1:07:12

That's probably worth a lot by itself.

1:07:12

Especially  if there are other competitive models out there, it's not like our thing is giving  away some kind of crazy advantage.

1:07:22

So is your view that the  training will be commodified?

1:07:29

I think there's a bunch of ways that this could  play out and that's one.

1:07:29

So “commodity” implies that it's going to get very cheap because there  are lots of options.

1:07:39

The other direction that this could go in is qualitative improvements.

1:07:44

You  mentioned fine-tuning.

1:07:44

Right now it's pretty limited what you can do with fine-tuning major  other models out there.

1:07:51

There are some options but generally not for the biggest models.

1:07:56

There’s  being able to do that, different app specific things or use case specific things or building  them into specific tool chains.

1:08:05

I think that will not only enable more efficient development, but  it could enable qualitatively different things.

1:08:18

Here's one analogy on this.

1:08:18

One thing that I think  generally sucks about the mobile ecosystem is that you have these two gatekeeper companies, Apple and  Google, that can tell you what you're allowed to build.

1:08:32

There's the economic version of that which  is like when we build something and they just take a bunch of your money.

1:08:38

But then there's the  qualitative version, which is actually what upsets me more.

1:08:45

There's a bunch of times when we've  launched or wanted to launch features and Apple's just like “nope, you're not launching that. ” That  sucks, right?

1:08:51

So the question is, are we set up for a world like that with AI?

1:09:01

You're going to  get a handful of companies that run these closed models that are going to be in control of the APIs  and therefore able to tell you what you can build?

1:09:13

For us I can say it is worth it to go build  a model ourselves to make sure that we're not in that position.

1:09:19

I don't want any of those  other companies telling us what we can build.

1:09:26

From an open source perspective, I think a lot of  developers don't want those companies telling them what they can build either.

1:09:30

So the question is,  what is the ecosystem that gets built out around that?

1:09:36

What are interesting new things?

1:09:36

How much  does that improve our products?

1:09:36

I think there are lots of cases where if this ends up being like  our databases or caching systems or architecture, we'll get valuable contributions from the  community that will make our stuff better.

1:09:54

Our app specific work that we do will then still  be so differentiated that it won't really matter.

1:10:00

We'll be able to do what we do.

1:10:00

We'll benefit  and all the systems, ours and the communities’, will be better because it's open source.

1:10:03

There is one world where maybe that’s not the case.

1:10:10

Maybe the model ends up  being more of the product itself.

1:10:10

I think it's a trickier economic calculation then, whether  you open source that.

1:10:16

You are commoditizing yourself then a lot.

1:10:22

But from what I can see so  far, it doesn't seem like we're in that zone.

1:10:26

Do you expect to earn significant revenue  from licensing your model to the cloud providers?

1:10:30

So they have to pay you  a fee to actually serve the model.

1:10:36

We want to have an arrangement like that but  I don't know how significant it'll be.

1:10:36

This is basically our license for Llama.

1:10:42

In a lot of ways  it's a very permissive open source license, except that we have a limit for the largest companies  using it.

1:10:51

This is why we put that limit in.

1:10:51

We're not trying to prevent them from using it.

1:10:56

We just  want them to come talk to us if they're going to just basically take what we built and resell it  and make money off of it.

1:11:00

If you're like Microsoft Azure or Amazon, if you're going to be reselling  the model then we should have some revenue share on that.

1:11:12

So just come talk to us before you  go do that.

1:11:12

That's how that's played out.

1:11:15

So for Llama-2, we just have deals with basically  all these major cloud companies and Llama-2 is available as a hosted service on all those  clouds.

1:11:23

I assume that as we release bigger and bigger models, that will become a bigger  thing.

1:11:30

It's not the main thing that we're doing, but I think if those companies are going to be  selling our models it just makes sense that we should share the upside of that somehow.

1:11:37

Regarding other open source dangers, I think you have genuine legitimate points about  the balance of power stuff and potentially the harms you can get rid of because we have better  alignment techniques or something.

1:11:48

I wish there were some sort of framework that Meta had.

1:11:52

Other  labs have this where they say “if we see this concrete thing, then that's a no go on the open  source or even potentially on deployment.

1:11:57

” Just writing it down so the company is ready for it and  people have expectations around it and so forth.

1:12:09

That's a fair point on the existential risk  side.

1:12:09

Right now we focus more on the types of risks that we see today, which are more of these  content risks.

1:12:14

We don't want the model to be doing things that are helping people commit violence  or fraud or just harming people in different ways.

1:12:30

While it is maybe more intellectually  interesting to talk about the existential risks, I actually think the real harms that need more  energy in being mitigated are things where someone takes a model and does something to hurt a  person.

1:12:31

In practice for the current models, and I would guess the next generation  and maybe even the generation after that, those are the types of more mundane harms that we  see today, people committing fraud against each other or things like that.

1:13:07

I just don't want to  shortchange that.

1:13:07

I think we have a responsibility to make sure we do a good job on that. Meta's a big company. You can handle both.

1:13:22

As far as open source goes, I'm actually  curious if you think the impact of open source, from PyTorch, React, Open Compute and other  things, has been bigger for the world than even the social media aspects of Meta.

1:13:30

I've  talked to people who use these services and they think that it's plausible because a  big part of the internet runs on these things.

1:13:39

It's an interesting question.

1:13:39

I mean almost  half the world uses our consumer products so it's hard to beat that.

1:13:48

But I think open  source is really powerful as a new way of building things. I mean, it's possible.

1:13:56

It  may be one of these things like Bell Labs, where they were working on the transistor because  they wanted to enable long-distance calling.

1:14:08

They did and it ended up being really profitable for  them that they were able to enable long-distance calling.

1:14:20

5 to 10 years out from that, if you  asked them what was the most useful thing that they invented it's like “okay, we enabled  long distance calling and now all these people are long-distance calling.

1:14:32

” But if you asked a  hundred years later maybe it's a different answer.

1:14:38

I think that's true of a lot of the things that  we're building: Reality Labs, some of the AI stuff, some of the open source stuff.

1:14:44

The specific  products evolve, and to some degree come and go, but the advances for humanity persist and  that's a cool part of what we all get to do.

1:14:58

By when will the Llama models be  trained on your own custom silicon? Soon, not Llama-4.

1:15:06

The approach that we took is  we first built custom silicon that could handle inference for our ranking and recommendation  type stuff, so Reels, News Feed ads, etc.

1:15:16

That was consuming a lot of GPUs.

1:15:24

When we were able  to move that to our own silicon, we're now able to use the more expensive NVIDIA GPUs only for  training.

1:15:31

At some point we will hopefully have silicon ourselves that we can be using for at  first training some of the simpler things, then eventually training these really large models.

1:15:48

In  the meantime, I'd say the program is going quite well and we're just rolling it out methodically  and we have a long-term roadmap for it. Final question.

1:16:02

This is totally out of  left field.

1:16:02

If you were made CEO of Google+ could you have made it work? Google+? Oof. I don't know.

1:16:14

That's a very difficult counterfactual.

1:16:14

Okay, then the real final question will be: when Gemini was launched, was  there any chance that somebody in the office uttered: “Carthago delenda est”.

1:16:24

No, I think we're tamer now. It's a good question.

1:16:38

The problem is there was no CEO of Google+.

1:16:38

It  was just a division within a company.

1:16:38

You asked before about what are the scarcest commodities  but you asked about it in terms of dollars.

1:16:45

I actually think for most companies, of this scale  at least, it's focus.

1:16:51

When you're a startup maybe you're more constrained on capital.

1:16:58

You’re just  working on one idea and you might not have all the resources.

1:17:04

You cross some threshold at some  point with the nature of what you're doing.

1:17:04

You're building multiple things.

1:17:10

You're creating  more value across them but you become more constrained on what you can direct to go well.

1:17:14

There are always the cases where something random awesome happens in the organization and I  don't even know about it. Those are great.

1:17:22

But I think in general, the organization's capacity  is largely limited by what the CEO and the management team are able to oversee and manage.

1:17:37

That's been a big focus for us.

1:17:37

As Ben Horowitz says “keep the main thing, the main thing” and  try to stay focused on your key priorities.

1:17:59

Awesome,
that was excellent, Mark. Thanks so much. That was a lot of fun. Yeah, really fun. Thanks for having me. Absolutely.