@Asianometry & Dylan Patel — How the semiconductor industry actually works

0:46

Today, I'm chatting with Dylan Patel, who  runs SemiAnalysis, and Jon, who runs the Asianometry YouTube channel.

0:51

Does he have a last name? No, I do not. No, just kidding. Jon Y.

0:54

Why is it only one letter?

1:01

Because Y is the best letter.

1:01

Why is your face covered? Why not?

1:07

No, seriously why is it covered?

1:11

Because I'm afraid of looking at myself  getting older and fatter over the years.

1:16

But seriously, it's for anonymity, right? Anonymity, yeah.

1:20

By the way, do you know  what Dylan's middle name is? Actually, no. I don't know. What's my father's name?

1:26

I'm not going to say it, but I remember. You could say it. It's fine. Sanjay? Yes. What's his middle name? Sanjay? That's right.

1:32

So I'm Dwarkesh Sanjay Patel.

1:34

He's Dylan Sanjay  Patel.

1:34

It's like literally my white name.

1:41

It's unfortunate my parents decided between my  older brother and me to give me a white name.

1:45

I could have been Dwarkesh Sanjay.

1:45

You know how amazing it would have been if we had the same name?

1:48

Butterfly effect and all, it probably would’ve turned out the same way, but...

1:52

Maybe it would have been even closer.

1:56

We would have met each other sooner, you know?

1:56

Who else would be named Dwarkesh Sanjay Patel in the world?

1:59

Alright here’s my first question.

2:01

If you're Xi Jinping and you're  scaling-pilled, what is it that you do?

2:05

Don't answer that question,  Jon, that's bad for AI safety.

2:08

I would basically be contacting every Chinese  national with family back home and saying, "I want information.

2:14

I want to know your recipes.

2:16

I want to know suppliers."

2:16

Lab foreigners or hardware foreigners? Everyone. Honey potting OpenAI?

2:24

This is totally off-cycle, off the reservation,  but I was doing a video about Yugoslavia's nuclear weapons program.

2:33

It started with absolutely  nothing.

2:33

One guy from Paris showed up.

2:42

He knew a little bit about  making atomic nuclear weapons.

2:45

He was like, "Okay, well, do I need help?"

2:45

Then the state's secret police is like, "I will get you everything."

2:49

For a span of like four years, they basically drew up a list. "What do you  need? What do you want?

2:57

What are you going to do?

3:03

What is it going to be for?"

3:03

And the state police just got everything.

3:07

If I were running a country and  I needed to catch up on that, that's the sort of thing that I would be doing.

3:09

Okay, let's talk about espionage.

3:09

What is the most valuable piece, if you could have this  blueprint, this one megabyte of information?

3:23

Do you want it from TSMC?

3:23

Do you want it from NVIDIA?

3:25

Do you want it from OpenAI?

3:25

What is the first thing you would try to steal?

3:30

You have to stack every layer, right?

3:30

The beautiful thing about AI is that because it's growing so fast, every layer is  being stressed to an incredible degree.

3:39

Of course, China has been hacking ASML for  over five years and ASML is kind of like, "Oh, it's fine."

3:45

The Dutch government's really pissed off, but it's fine.

3:46

They already have those files in my view.

3:51

It's just a very difficult thing to build.

3:51

The same applies for fab recipes.

3:51

They can poach Taiwanese nationals.

3:58

It’s not  that difficult because TSMC employees do not make absurd amounts of money.

4:03

You can just poach them and give them a much better life and they have.

4:07

A lot of SMIC's employees are TSMC, Taiwanese nationals, especially a  lot of the really good ones high up.

4:17

You go up the next layers of the stack.

4:17

Of course, there are tons of model secrets.

4:23

But how many of those model secrets do  you not already have and just haven't deployed or implemented or organized?

4:28

That's the one thing I would say.

4:33

China just clearly is still  not scale-pilled in my view.

4:40

If you could hire these people, it would probably  be worth a lot to you because you're building a fab that's worth tens of billions of dollars.

4:44

This  talent knows a lot.

4:44

How often do they get poached?

4:51

Do they get poached by foreign adversaries or do  they just get poached by other companies within the same industry but in the same country?

4:56

Why doesn't that drive up their wages?

5:02

It's because it's very compartmentalized.

5:02

Back in the 2000s, before SMIC got big, it was actually much more open and more flat.

5:11

After that, after Liang Mong Song and after all the Samsung issues and after  SMIC's rise, you literally saw— You should tell that story,  actually, about the TSMC guy that went to Samsung and SMIC and all that.

5:26

I think you should tell that story. There are two stories.

5:28

There's a  guy who ran a semiconductor company in Taiwan called Worldwide Semiconductor.

5:31

This guy, Richard Chang, was very religious.

5:36

All the TSMC people are pretty religious.

5:36

He particularly was very fervent and wanted to bring religion to China.

5:40

So after he sold his company to TSMC—which was a huge coup for TSMC—he worked there for about  eight or nine months and then went back to China.

5:50

Back then, the relations between China  and Taiwan were much more different.

5:55

So he goes over to Shanghai and they  say, "We'll give you a bunch of money."

5:59

Richard Chang basically recruits  a whole conga line of Taiwanese who just get on the plane and fly over.

6:04

Generally that’s actually true of a lot of the acceleration points within  China’s semiconductor industry.

6:13

It’s from talent flowing from Taiwan.

6:13

The second story is about Liang Mong Song.

6:18

Liang Mong Song is a nut. I’ve not met him.

6:18

I’ve met people who worked with him and they say he is a nut.

6:23

He's probably on the spectrum.

6:23

He  doesn't care about people, business, or anything.

6:31

He wants to take it to the limit.

6:31

That’s the only thing he cares about.

6:34

He worked at TSMC, a literal genius  with 300 patents or whatever, 285.

6:39

He works his way all the way to the top  tier and then one day he loses out on some power game within TSMC and gets demoted.

6:47

He was like the head of R&D or something right?

6:52

He was like one of the top R&D people.

6:52

He was in like second or third place.

6:55

It was for the head of R&D position, basically.

6:55

Correct, it was for the head of R&D position.

6:58

He’s like, "I can’t deal with this."

6:58

He goes to Samsung and steals a bunch of talent from TSMC.

7:03

Literally, again, it’s a conga line.

7:09

At some point, some of these people  were getting paid more than the Samsung chairman, which is not really comparable… Isn't the Samsung chairman usually like part of the family that owns Samsung? Correctamundo.

7:20

Okay, yeah so it’s kind of irrelevant.

7:20

So he goes over there and says, "We will make Samsung into this monster. Forget  everything.

7:25

Forget all the stuff you’ve been trying to do incrementally. Toss that out.

7:31

We  are going to the leading edge and that is it."

7:36

They go to the leading edge.

7:36

They win a big portion of Apple's business back from TSMC.

7:47

And then at TSMC, Morris Chang is like, "I'm not letting this happen."

7:52

That guy is toxic to work for as well but also goddamned brilliant.

7:57

He’s also very good at motivating people.

8:02

He sets up what is called the Nightingale Army.

8:02

They split a bunch of people and they say, "You are working R&D night shift.

8:11

There is no rest at the TSMC fab.

8:19

As you go in, there will  be a day shift going out."

8:21

They called it "burning your liver."

8:21

In Taiwan, they say as you get old and as you work, you're sacrificing your liver.

8:27

They called it the liver buster.

8:31

They basically did this Nightingale Army for  a year or two years. They finished FinFET.

8:38

They basically just blow away Samsung.

8:38

At the same time, they sue Liang Mong Song directly for stealing trade secrets.

8:44

Samsung basically separates from Liang Mong Song and Liang Mong Song went to SMIC.

8:50

So Samsung at one point was better than TSMC.

8:56

Then he goes to SMIC and  SMIC caught up rapidly after. Very rapid. That guy's a genius. That guy's a  genius.

9:00

I don't even know what to say about him.

9:05

He's like 78 and he's beyond brilliant.

9:05

He does not care about people.

9:11

What does research to make the  next process node look like?

9:15

Is it just a matter of a hundred  researchers going in, they do the next n+1?

9:20

The next morning, the next  hundred researchers go in? It’s experiments.

9:24

They have a  recipe and that's what they do.

9:27

Every TSMC recipe is the culmination of long years  of research. It's highly secret.

9:27

The idea is that you're going to look at one particular part of it  and say, "Run an experiment. Is it better? Is it not? Is it better or not?" It’s a thing like that.

9:43

It's basically a multivariable problem.

9:43

Every single tool sequentially you're  processing the whole thing.

9:51

You turn knobs up and down on every single tool.

9:51

You can increase the pressure on this one specific deposition tool.

9:56

What are you trying to measure? Does it increase yield?

9:58

It's yield, it's performance, it's power.

10:05

It's not just better or worse.

10:05

It's a multivariable search space.

10:08

What do these people know  such that they can do this?

10:09

Is it that they understand  the chemistry and physics?

10:12

It's a lot of intuition, but yeah.

10:12

It's PhDs in chemistry, PhDs in physics, PhDs in electrical engineering… Brilliant geniuses.

10:20

They don't even know about  the end chip a lot of times.

10:23

It's like, "Oh, I am an etch engineer and  all I focus on is how hydrogen fluoride etches this and that's all I know.

10:29

If I do it at different pressures… If I do it at different temperatures… If I do it  with a slightly different recipe of chemicals… It changes everything.

10:37

I remember someone told me this when I was speaking.

10:39

How did America lose the ability to do this sort of thing?

10:43

I’m talking about etch and hydrofluoric acid and all of that.

10:44

He told me basically it's very master-apprentice.

10:52

You know like in Star Wars with the Sith,  there's only one, right?

10:52

Master apprentice, master apprentice.

10:55

It used to be that  there is a master, there's an apprentice, and they pass on this secret knowledge.

11:00

This guy knows nothing but etch, nothing but etch.

11:04

Over time, the apprentices stopped coming.

11:04

In the end, the apprentices moved to Taiwan.

11:10

That's the same way it's still run.

11:10

Like you have NTHU, National Tsing Hua University.

11:15

There's a bunch of masters.

11:15

They teach apprentices, and they just pass this secret, sacred knowledge down.

11:19

Who are the most AGI-pilled people in the supply chain?

11:24

I got to have my phone call with Colette right now. Okay, go for it.

11:27

Could we mention to the podcast that NVIDIA is calling  Dylan to update him on the earnings call?

11:36

Well, it's not exactly that, but… Go for it, go for it… Dylan is back from his call with Jensen Huang.

11:40

It was not with Jensen, Jesus.

11:45

What did they tell you, huh?

11:45

What did they tell you about next year’s earnings?

11:47

No, it was just color around like  Hopper, Blackwell, and margins.

11:51

It's quite boring stuff for most  people, I think it's interesting though.

11:56

I guess we could start talking about NVIDIA.

11:56

You know what, before we do… I think we should go back to China.

11:59

There's a lot of points there.

12:01

Alright, we covered the chips themselves.

12:01

How do they get the 10 gigawatts data center up? What else do they need?

12:05

There is a true question of how decentralized do you go versus centralized.

12:12

In the US, as far as labs and such, you have OpenAI, xAI, Anthropic.

12:20

Microsoft has their own effort, Anthropic has their own efforts, despite having  their partner. Then you have Meta.

12:25

You also have all the interesting startups doing stuff.

12:29

You go down the list and there's quite a decentralization of efforts.

12:37

Today in China, it is still quite decentralized.

12:41

It's not like, "Alibaba,  Baidu, you are the champions."

12:45

You have DeepSeek doing amazing stuff  and it’s like, "Who the hell are you?

12:47

Does the government even support you?"

12:51

If you are Xi Jinping and scale-pilled, you  must now centralize the compute resources.

12:57

Because you have sanctions on how  many NVIDIA GPUs you can get in now.

13:01

They're still north of a million a year,  even post-October last year sanctions.

13:06

We still have more than a million  H20s, and other Hopper GPUs getting in through other means but legally the H20s.

13:11

On top of that, you have your domestic chips, but that's less than a million chips.

13:18

When you look at it, it's like, "Oh, well, we're still talking about a million chips."

13:21

The scale of data centers people are training on today slash over the next  six months is 100,000 GPUs.

13:30

OpenAI, xAI, these are quite  well documented and others.

13:34

But in China, they have no  individual system of that scale yet.

13:40

Then the question is, "How do we get there?"

13:40

No company has had the centralization push to have a cluster that large and train on  it yet, at least publicly and well-known.

13:51

The best models seem to be from a company  that has got like 10,000 GPUs or 16,000 GPUs.

13:58

It's not quite as centralized as the US companies  are and the US companies are quite decentralized.

14:02

If you're Xi Jinping and you're scale-pilled,  do you just say, "XYZ company is now in charge and every GPU goes to one place"?

14:09

Then you don't have the same issues as in the US.

14:15

In the US, we have a big problem with being  able to build big enough data centers, being able to build substations and transformers and  all this that are large enough in a dense area.

14:23

China has no issue with that at all because their  supply chain adds as much power as like half of Europe every year.

14:30

It’s some absurd statistic.

14:30

They're building transformer substations or building new power plants constantly.

14:36

They have no problem with getting power density.

14:40

You go look at Bitcoin mining.

14:40

Around the Three Gorges Dam, at one point at least, there was like 10  gigawatts of Bitcoin mining estimated.

14:51

We're talking about gigawatt data centers  coming over in 2026 or 2027 in the US.

15:01

This is an absurd scale relatively.

15:01

We don't have gigawatt data centers ready.

15:06

China could just build it in six months, I think,  around the Three Gorges Dam or many other places.

15:12

They have the ability to do the substations.

15:12

They have the power generation capabilities.

15:15

Everything can be done like a flip of  a switch, but they haven't done it yet.

15:18

Then they can centralize the chips like crazy.

15:18

Right now they can be like "Oh, a million chips that NVIDIA's shipping in Q3 and Q4, the H20,  let's just put them all in this one data center."

15:28

They just haven't had that  centralization effort yet.

15:31

You can argue that the more you  centralize it, the more you start building this monstrous thing within the  industry, you start getting attention to it.

15:39

Suddenly, lo and behold, you have a  little bit of a little worm in there.

15:44

Suddenly while you're doing your big  training run, "Oh, this GPU is off. Oh, this GPU... Oh no, oh no, oh no..."

15:48

I don't know if it's like that.

15:53

Is that a Chinese accent by the way?

15:53

Just to be clear, Jon is East Asian. He's Chinese.

15:58

I'm of East Asian descent.

15:58

Half Taiwanese, half Chinese? That is right.

16:02

But I don't know if that's as simple as that because training systems  are like… Is it water gated? Firewalled? What is it called? Not firewalled. I don't know. There's a word for that.

16:13

Where they're… Air-gapped.

16:16

I think they’re Chinese-walled.

16:16

You’re going through all the four elements. Earth, water, fire!

16:20

If you’re Xi Jinping and you’re scale-pilled… You got to unite the four forces. Fuck the airbenders. Fuck the firebenders. We got the Avatar.

16:34

You have to build the Avatar. Okay. I think that's possible.

16:34

The question is, "Does that slow down your research?"

16:41

Do you crush people like DeepSeek who are clearly not being influenced by  the government?

16:46

and put some idiot… You put an idiot bureaucrat at the top.

16:51

Suddenly, he's all thinking about these politics.

16:58

He's trying to deal with  all these different things.

17:00

Suddenly, you have a single point of failure, and that's bad.

17:03

On the flip side, there are obviously immense gains from being  centralized because of the scaling laws.

17:12

The flip side is compute efficiency which is  obviously going to be hurt because you can't experiment and have different people lead and try  their efforts as much if you're more centralized.

17:24

There is a balancing act there.

17:24

That is actually really interesting, the fact that they can centralize.

17:26

I didn't think  about this.

17:26

Even if America as a whole is getting millions of GPUs a year, the fact that any one  company is only getting hundreds of thousands or fewer means that there's no one person who can  do a single training run as big in America as if China as a whole decides to do one together.

17:44

The 10 gigawatts you mentioned near the Three Gorges Dam, how widespread is it? Is it a state? Is it like one wire?

17:55

Would you do a sort of distributed training run?

17:55

It’s not just the dam itself, but also all of the coal.

17:59

There's some nuclear reactors there as well I believe.

18:00

Between all of that and renewables like solar and wind, in that region there is an absurd  amount of concentrated power that could be built.

18:13

I'm not saying it's like one button, but  it's more like, "hey within X mile radius."

18:18

That's more of the correct way to frame it.

18:18

That's how the labs are also framing it in the US.

18:26

If they started right now, how long would it take  to build the biggest AI data center in the world?

18:33

Actually the other thing is, could we  notice it? I don't think so.

18:33

With the amount of factories that are being spun up—the  amount of other construction, manufacturing, etc.

18:44

that's being built—a gigawatt is  actually like a drop in the bucket.

18:48

A gigawatt is not a lot of power.

18:48

10 gigawatts  is not an absurd amount of power, it's okay.

18:54

Yes, it's like hundreds of thousands  of homes, millions of people. But you’ve got 1. 4 billion people.

18:57

You’ve got most of the world's extremely energy intensive refining and rare earth refining  and all these manufacturing industries here.

19:09

It would be very easy to hide it.

19:09

It would be very easy to just shut down like… I think the largest aluminum mill in the world  is there and it's north of 5 gigawatts alone.

19:19

Could we tell if they stopped making aluminum  there and instead started making AI there?

19:25

I don't know if we could tell because they could  also just easily spawn like 10 other aluminum mills to make up for the production and be fine.

19:30

There's many ways for them to hide compute as well.

19:33

To the extent that you could just take out a five gigawatt aluminum refining center  and build a giant data center there, then I guess the way to control Chinese AI has to be the chips?

19:42

Just walk me through how many chips they have now.

19:52

How many will they have in the future?

19:52

What will that be in comparison to the US and the rest of the world?

19:55

In the world we live in, they are not restricted at all in the physical infrastructure side of  things in terms of power, data centers, etc.

20:05

Their supply chain is built for that.

20:05

It's pretty easy to pivot that.

20:09

Whereas the US adds so little power each  year and Europe loses power every year.

20:13

The Western industry for power  is non-existent in comparison.

20:20

On the flip side, "Western" manufacturing when  you include Taiwan is way, way, way larger than China's, especially on leading-edge where China  theoretically has—depending on the way you look at it—either zero or a very small percentage share.

20:32

There you have equipment, wafer manufacturing, and then you have advanced packaging capacity.

20:40

Where can the US control China?

20:47

Advanced packaging capacity is kind  of shot because the vast majority… The largest advanced packaging company  in the world was Hong Kong-headquartered.

20:53

They just moved to Singapore, but that's  effectively in a realm where the U. S. can't sanction it.

20:58

A majority of these other companies are in similar places.

21:01

Advanced packaging capacity is very hard.

21:06

Advanced packaging is useful for stacking memory,  stacking chips on co-ops, things like that.

21:11

Then the step down is wafer fabrication.

21:11

There is immense capability to restrict China there.

21:15

Despite the US making some sanctions, China in the most recent quarters was like 48% of  ASML's revenue and like 45% of Applied Materials’.

21:29

You just go down the list.

21:29

Obviously it's not being controlled that effectively.

21:32

But it could be on the equipment side of things.

21:35

The chip side of things is actually being  controlled quite effectively, I think.

21:40

Yes, there is shipping of GPUs through Singapore  and Malaysia and other countries in Asia to China.

21:46

But the amount you can smuggle is quite small.

21:46

Then the sanctions have limited the chip performance to a point where it's like,  "You know, this is actually kind of fair."

21:55

But there is a problem with  how everything is restricted.

22:00

You want to be able to restrict China from  building their own domestic chip manufacturing industry that is better than what we ship them.

22:04

You want to prevent them from having chips that are better than what we have.

22:09

And then you want to prevent them from having AIs that are better.

22:12

That’s the  ultimate goal.

22:12

If you read the restrictions, it’s very clear that it's about AI.

22:16

Even in 2022, which is amazing, at least the Commerce Department was kind of AI-pilled.

22:21

It was like, "You want to restrict them from having AIs better than us."

22:25

So starting on the right end, it's like, "Okay, well, if you want to restrict  them from having better AIs than us.

22:27

you have to restrict chips. Okay.

22:30

If you want to  restrict them from having chips, you have to let them have at least some level of chip that  is better than what they can build internally."

22:41

But currently the restrictions  are flipped the other way.

22:44

They can build better chips in China than what  we restrict them from in terms of chips that NVIDIA or AMD or Intel can sell to China.

22:49

So there's sort of a problem there in that the equipment that is shipped can be used  to build chips that are better than what the Western companies can actually ship them.

22:59

Jon, Dylan seems to think the export controls are kind of a failure. Do you agree with him?

23:03

That is a very interesting question because I think it's like… Why thank you.

23:13

Dwarkesh, you're so good.

23:13

Yeah, Dwarkesh, you're the best.

23:16

Failure is a tough word to say  because what are we trying to achieve?

23:23

Let’s just take lithography.

23:23

If your goal is  to restrict China from building chips and you just boil it down to like, "Hey, lithography is  25-30% of making a chip.

23:37

Cool, let's sanction lithography.

23:42

Okay, where do we draw the line?

23:42

Let me figure out where the line is."

23:49

If I'm a bureaucrat or lawyer at the Commerce  Department or what have you, obviously I'm going to go talk to ASML and ASML is going to  tell me, "This is the line," because they know this and this… there's some blending over.

23:59

They're looking at it like, "What's going to cost us the most money?"

24:02

Then they all constantly say, "If you restrict us, then China will have their own industry."

24:07

The way I like to look at it is that chip manufacturing is like 3D chess  or like a massive jigsaw puzzle.

24:18

If you take away one piece, China can be like,  "Oh, that's the piece. Let's put it in."

24:18

Currently year by year by year, they keep updating export  restrictions ever since like 2018 or 2019 when Trump started and now Biden's accelerated them.

24:30

They haven't just taken a bat to the table and broken it.

24:36

It's like, "Let's take one jigsaw puzzle piece out. Walk away. Oh shit. Let's take two more out. Oh shit."

24:39

You either have to go full "bat to the fricking table/wall"  or chill out and let them do whatever they want.

24:56

Because the alternative is everything is  focused on this thing and they make that.

25:01

Then now when you take out another two pieces  they can be like, "Well, I have my domestic industry for this.

25:04

I can also now make a domestic industry for these."

25:06

You go deeper into the tech tree or what have you.

25:09

It's art in the sense that there are  technologies out there that can compensate.

25:16

The belief that lithography is a linchpin  within the system, it's not exactly true.

25:23

At some point, if you keep pulling a thread, other  things will start developing to close that loop.

25:32

That's why I say it's an art.

25:32

I don't think you can stop the Chinese semiconductor industry from progressing.

25:35

That's  basically impossible.

25:35

The Chinese nation, the Chinese government, believes in the  primacy of semiconductor manufacturing.

25:47

They believed it for a long time,  but now they really believe it.

25:51

To some extent, the sanctions have made China  believe in the importance of the semiconductor industry more than anything else.

25:56

So from an AI perspective, what's the point of export controls then?

25:59

If they're going to be able to get these...

26:07

Well they’re not centralized though.

26:07

That's  the big question: are they centralized?

26:11

Also, I'm not sure if I really believe it but  on prior podcasts, there have been people who talked about nationalization.

26:16

Why are you referring to this ambiguously?

26:22

"My opponent…" No, I love Leopold, but there have been a couple where people have talked about nationalization.

26:29

If you have nationalization, then all of the sudden you aggregate all the flops. There's no  fucking way.

26:35

China can be centralized enough to compete with each individual US lab.

26:39

They could have just as many flops in 2025 and 2026 if they decided they were scale pilled.

26:44

Just from foreign chips, for an individual model.

26:50

Like they can release a 1e27 model by 2026?

26:50

And a 1e28 model in the works.

27:00

They totally could do this  just with foreign chip supply.

27:02

It’s just a question of centralization.

27:02

Then the question is, do you have as much innovation and compute  efficiency wins when you centralize?

27:09

Or do Anthropic, OpenAI, xAI, and Google develop  things, and secrets shift a bit between each other, resulting in a better long-term  outcome versus nationalization in the US?

27:29

China could absolutely have it in 2026-27  if they desire to, just from foreign chips.

27:36

Domestic chips are the other question.

27:36

You have 600,000 of the Ascend 910B, which is roughly 400 teraflops or so.

27:43

If they put them all in one cluster, they could have a bigger model  than any of the labs next year.

27:54

I have no clue where all the Ascend  910Bs are going, but there are rumors about them being divvied up between the  majors, Alibaba, ByteDance, Baidu, etc.

28:04

Next year, you have more than a million.

28:04

It's possible they actually have 1e30 before the US because data center isn't as big an issue.

28:10

A 10 gigawatt data center… I don't think anyone is even trying to build that today  in the US, even out to 2027-28.

28:21

They're focusing on linking  many data centers together.

28:23

There's a possibility that come 2028-2029, China  can have more flops delivered to a single model, even once the centralization question is solved.

28:32

That's clearly not happening today for either party.

28:37

I'd bet if AI is as important as you and I believe, they will centralize sooner  than the West does.

28:42

So there is a possibility.

28:58

How many more wafers could they make and how many  of those wafers could be dedicated to the 910B?

29:01

I assume there's other things they  want to do with these semiconductors. There's two parts there.

29:04

The way the US  has sanctioned SMIC is really stupid.

29:11

They've sanctioned a specific spot  rather than the entire company.

29:16

SMIC is still buying a ton of tools that  can be used for their 7nm and their 5.

29:16

5nm, or 6nm process, for the 910C  which releases later this year.

29:27

They can build as much of that  as long as it's not in Shanghai.

29:31

Shanghai has anywhere from 45 to 50  high-end immersion lithography tools.

29:39

That’s what is believed by  intelligence and many other folks.

29:44

That roughly gives them as much as 60,000  wafers a month of 7 nanometer, but they also make their 14 nanometer in that fab.

29:50

The belief is that they actually only have about 25,000-35,000 capacity  of 7 nanometer wafers a month.

30:00

Doing the math of the chip die size and all  these things—Huawei also uses chiplets so they can get away with using less leading-edge  wafers but then their yields are bad—you can roughly say something like 50 to 80  good chips per wafer with their bad yield.

30:20

Why do they have bad yield? Because it's hard.

30:24

Even if everyone knows the number,  like say there’s a 1000 steps.

30:28

Even if you're 98-99% for each, in  the end you'll still get a 40% yield.

30:37

If it's six sigma of perfection and you  have your 10,000 plus steps, you end up with yield that's still dog shit by the end.

30:44

That is a scientific measure, dog shit percent.

30:52

Yeah, as a multiplicative effect.

30:52

Yields are bad  because they have hands tied behind their back.

31:01

They are not getting to use EUV.

31:01

On 7 nanometer Intel never used EUV, but TSMC eventually started using  EUV.

31:06

Initially, they used DUV.

31:10

Doesn't that mean the export controls succeeded?

31:10

They have bad yield because they have to use...

31:15

It’s a brand new process.

31:15

Again, they're still determined. Success means they stop. They're not stopping.

31:20

Let’s go back to the yield question.

31:23

It’s theoretically 60,000 wafers a month  times 50-100 dies per wafer with yielded dies. Holy shit. That's millions of GPUs.

31:30

Now,  what are they doing with most of their wafers?

31:35

They still have not become scale pilled,  so they're still throwing them out.

31:38

Let's make 200 million Huawei phones, right? Okay, cool. I don't care.

31:38

As the West, you don't care as much, even though  Western companies will get screwed, like Qualcomm and MediaTek Taiwanese companies. Obviously there's that.

31:46

The same applies to the US, but when you flip to like… Sorry, I  don't fucking know what I was going to say. Nailed it!

32:02

We're keeping this in That's fine, that's fine.

32:05

In 2026 if they're centralized, they can have as big training  runs as any one US company— Oh, the reason why I was bringing up Shanghai.

33:17

They're building 7nm capacity in Beijing.

33:22

They're building 5nm capacity in Beijing,  but the US government doesn't care.

33:25

They're importing dozens of tools into  Beijing and saying to the US government and ASML, "This is for 28nm, obviously."

33:29

In the background, they’re making 5nm there.

33:37

Are they doing it because they believe in AI,  or because they want to make Huawei phones?

33:41

Huawei was the largest TSMC customer for  a few quarters before they got sanctioned.

33:46

Huawei makes most of the  telecom equipment in the world.

33:50

Phones, modems, accelerators, networking  equipment, video surveillance chips, you go through the whole gamut.

33:57

A lot of that could use 7 and 5 nanometer.

34:01

Do you think the dominance of Huawei is actually  bad for the rest of the Chinese tech industry?

34:08

Huawei is so cracked that it's hard to say that.

34:08

Huawei out-competes Western firms regularly with two hands tied behind their back.

34:15

What the hell are Nokia and Sony Ericsson?

34:23

They’re trash compared to Huawei.

34:23

Huawei isn't  allowed to sell to European or American companies and they don't have TSMC.

34:30

Yet they still  destroy them.

34:30

The new phone is as good as a year-old Qualcomm phone on a process node that’s  equivalent to something three or four years old.

34:45

They actually out-engineered us with the  worst process node. Huawei is crazy cracked.

34:52

Where do you think that culture comes from?

34:52

The military, because it's the PLA.

34:56

It's generally seen as an arm of the PLA.

34:56

How do you square that with the fact that sometimes the PLA seems to mess stuff up?

35:04

Oh, like filling water in rockets?

35:10

I don't know if that was true. I'm not denying it.

35:10

There is that crazy conspiracy… You don't know what the hell to believe in China,  especially as a non-Chinese person.

35:20

Nobody knows, even Chinese people  don't know what's going on in China.

35:23

There's all sorts of stuff like, "Oh,  they're filling water in their rockets, clearly they're incompetent."

35:26

If I'm the Chinese military, I want the Western world to believe I'm  completely incompetent because one day, I can just destroy everything with hypersonic  missiles and drones.

35:34

"No, no, we're filling water in our missiles. These are all fake.

35:42

We  don't actually have a hundred thousand missiles that we manufacture in a super advanced facility  and Raytheon is stupid as shit because they can't make missiles nearly as fast."

35:50

That’s also the  flip side.

35:50

How much false propaganda is there?

35:58

There's a lot of, "SMIC could never,  they don't have the best tools."

36:02

Then it's like, "Motherfucker, they  just shipped 60 million phones last year with this chip that performs only  one year worse than what Qualcomm has."

36:10

The proof is in the pudding. There's a lot of cope.

36:16

I just wonder where that culture comes from.

36:16

There's something crazy about them.

36:16

Everything they touch, they seem to succeed in. I wonder why. They're making cars.

36:23

I wonder what’s going on there.

36:27

If we imagine historically… Do you think they're getting something from somewhere? Espionage, you mean? Obviously.

36:40

East Germany and the Soviet industry was  basically a conveyor belt of secrets coming in, and they used that to run everything.

36:44

But the Soviets were never good at it.

36:48

They could never mass produce  it. But now you have China.

36:49

How would espionage explain how they can  make things with different processes? It’s not just espionage.

36:54

They’re just literally cracked. That’s why.

36:57

It has to be something else.

36:57

They have the espionage without a doubt.

36:59

ASML is known to have been  hacked at least a few times.

37:06

People have been sued who made it  to China with a bunch of documents.

37:08

It’s not just ASML, but every  company in the supply chain.

37:11

Cisco code was literally in early Huawei routers.

37:11

You go down the list… But architecturally, the Ascend 910B  looks nothing like a GPU or TPU.

37:22

It is its own independent thing.

37:22

Sure, they probably learned some things from some places, but they're good at engineering. It's 9-9-6.

37:26

Wherever that culture comes from, they do good. They do very good.

37:32

Another thing I'm curious about is where that culture comes from, but also how it stays there.

37:37

With American firms or any other firm, you can have a company that's very good, but over  time it gets worse, like Intel or many others.

37:47

I guess Huawei just isn't that old, but  it's hard to be a big company and stay good. That is true.

37:53

A word I hear a lot  regarding Huawei is "struggle."

38:01

China has a culture where the  Communist Party is big on struggle.

38:06

Huawei brought that culture into their  way of doing things.

38:06

You said this before, right?

38:12

They go crazy because they think in five  years they're going to fight the United States.

38:20

Everything they do, every second it’s  like their country depends on it.

38:24

It's the Andy Grove-ian mindset.

38:24

Shout  out, the based Intel of Andy Grove.

38:24

Only the paranoid survive.

38:28

Paranoid Western companies  do well.

38:28

Why did Google really screw the pooch on a lot of stuff and then resurge now?

38:34

It’s because they got paranoid as hell.

38:41

If Huawei is just constantly paranoid about the  external world and thinking, "Oh fuck, we're gonna die. They're gonna beat us.

38:45

Our country depends  on it.

38:45

We’re going to get the best people from the entire country at whatever they do…" And tell them, "If you do not succeed, our country will die.

39:00

Your family will  be enslaved. It will be terrible."

39:03

"By the evil western pigs."

39:03

"Capitalist…" Or not capitalist, they don't say that anymore.

39:08

It's more like, "Everyone is against China. China is being defiled.

39:11

That  is all on you, bro.

39:11

If you can’t do that…" "If you can't get that radio to be slightly less  noisy and transmit 5% more data, we are fucked."

39:28

"It's like the great palace fire all over again.

39:28

The British are coming and they will steal all the trinkets. That's on you."

39:31

Why isn't there more vertical integration in the semiconductor industry?

39:37

Why is it like, "This subcomponent requires this subcomponent from this other company,  which requires this subcomponent from this other company…" Why isn’t more of it done in-house?

39:44

The way to look at it today is that it’s super stratified.

39:48

Every industry has anywhere from one to three competitors.

39:50

The most competitive it gets is like 70% share, 25% share, 5% share, in any layer  of manufacturing chips, anything, chemicals, different types of chips, etc.

40:02

It used to be vertically integrated.

40:07

At the very beginning it was integrated. Why did that stop?

40:14

You had companies that used to do it all in one.

40:14

Then suddenly a guy would be like, "I hate this.

40:20

I know how to do it better."

40:20

He’d spin off, do his own thing, start his own company, then go back to his old company  and say, "I can sell you a product that’s better."

40:27

That's the beginning of what we call  the semiconductor equipment industry.

40:31

In the seventies, everyone  made their own equipment. Sixties and seventies.

40:34

All these people spin off.

40:34

What happened was that the companies that accepted outside products and equipment  got better stuff and did better.

40:46

There were companies that were totally  vertically integrated for decades.

40:50

They are still good, but nowhere near competitive.

40:50

One thing I'm confused about is the actual foundries themselves, there's  fewer and fewer of them every year.

40:59

There's maybe more companies overall,  but fewer final wafer makers.

41:09

It's similar to AI foundation models where  you need revenues from a previous model.

41:18

You need market share to fund the next  round of ever more expensive development.

41:23

When TSMC launched the foundry  industry, there was a wave of Asian companies that funded semiconductor foundries.

41:29

Malaysia with Silterra, Singapore with Chartered, one from Hong Kong… A bunch in Japan. A bunch in Japan. They all did this thing.

41:41

When going to leading-edge, it got harder, which means you had to aggregate more demand  from all the customers to fund the next node.

41:56

Technically you’re aggregating all  this profit to fund the next node to the point where now there's no  room in the market for an N2 or N3.

42:08

You could argue that economically, N2 is  a monstrosity that doesn't make sense.

42:17

It should not exist without  the immense concentrated spend of like five players in the market.

42:23

I'm sorry to completely derail you, but there's this video where it's like, "This  is an unholy concoction of meat slurry." Yes! What?

42:35

There’s this video that’s  like, "Ham is disgusting.

42:37

It’s an unholy concoction of meat  with no bones or collagen." I don’t know.

42:42

The way he was describing 2nm is like that.

42:42

It’s like the guy who pumps his right arm so much. He's super muscular.

42:49

"The human body  was not meant to be so muscular!" What’s the point?

42:55

Why is  2 nanometer not justified?

42:57

I'm not saying it for N2 specifically, but N2 as  a concept.

42:57

The next node should technically...

43:06

There will come a point where economically,  the next node will not be possible at all.

43:11

Unless more technology spawned, like AI now  makes one nanometer or whatever, A16, viable.

43:18

Makes it viable in what  sense? It makes it worth it? Money. Money.

43:23

Every two years you get a shrink, like clockwork, Moore's law.

43:26

Then five nanometer  happened. It took three years. Holy shit.

43:26

Then three nanometer happened. It took three  years. Is Moore's law dead? Because TSMC didn't...

43:40

and then what did Apple do?

43:40

When three nanometer finally launched, Apple only moved half of the  iPhone volume to three nanometer.

43:52

Now they did a fourth year of five  nanometer for a big chunk of iPhones.

43:58

Is the mobile industry petering out?

43:58

Then you look at two nanometer and it's going to be similarly difficult  for the industry to pay for this.

44:06

Apple because they get to make the phone,  they have so much profit they can funnel into more and more expensive chips.

44:10

But finally that was really running out.

44:16

How economically viable is 2nm just  for one player, TSMC?

44:16

Ignore Intel, ignore Samsung.

44:21

Samsung is paying for it  with memory, not with their actual profit.

44:27

Intel is paying for it from  their former CPU monopoly… Private equity money, chips money, debt,  and the salaries of laid-off people.

44:39

There's a strong argument that funding  the next node wouldn't be economically viable anymore if it weren’t for AI taking off  and generating humongous demand for the most leading-edge chip.

44:51

Dwarkes Patel How big is the difference  between 7nm to 5nm to 3nm?

44:57

Is it a huge deal in terms of who  can build the biggest cluster?

45:00

There's this simplistic argument that moving  a process node only saves X percent in power.

45:06

That has been petering out.

45:06

When you moved  from 90 nanometer to 80 or 70 something, you got 2x.

45:14

Dennard scaling was still intact.

45:14

But  now when you move from 5 nanometer to 3 nanometer, you don't double density.

45:19

SRAM doesn't scale  at all.

45:19

Logic does scale, but it's like 30%.

45:25

All in all, you only save about  20% in power per transistor.

45:29

But because of data locality and movement  of data, you actually get a much larger improvement in power efficiency by moving  to the next node than just the individual transistors' power efficiency benefit.

45:38

For example, if you're multiplying a matrix that's 8,000 by 8,000 by 8,000,  you can't fit that all on one chip.

45:47

But if you could fit more and more, you have  to move off chip less, go to memory less, etc.

45:52

The data locality helps a lot too.

45:52

AI really, really wants new process nodes because power usage is a lot less, you  get higher density and higher performance.

46:05

The big deal is if I have a gigawatt data  center, how much more flops can I get?

46:09

If I have a two gigawatt data center,  how much more flops can I get?

46:12

If I have a ten gigawatt data center,  how much more flops can I get?

46:14

You look at the scaling and everyone needs to go to the most recent process  node as soon as possible.

46:20

I want to ask a normie question…  I won’t phrase it that way. Not for you nerds.

46:28

I think Jon and I could communicate to the point where you even  wouldn't know what we're talking about.

46:39

Suppose Taiwan is invaded  or Taiwan has an earthquake.

46:43

Nothing is shipped out of Taiwan from now  on. What happens next?

46:43

How would the rest of the world feel its impact a day  in, a week, a month in, a year in?

46:54

It's a terrible thing to talk about.

46:54

Can you just say it’s all just terrible?

47:00

It's not just leading-edge.

47:00

People will  focus on leading-edge, but there's a lot of trailing-edge stuff that people depend on  every day. We all worry about AI.

47:04

The reality is you're not going to get your fridge.

47:09

You're not going to get your cars.

47:12

You're not going to get everything. It's terrible.

47:12

Then there's the human part of it. It's all terrible. It's depressing. And I live there.

47:15

Day one, the market crashes a lot.

47:25

The six or seven biggest companies, the  Magnificent Seven, are like 60-75% of the S&P 500 and their entire business relies on chips.

47:32

Google, Microsoft, Apple, Nvidia, Meta, they all entirely rely on AI.

47:39

You would have an extremely insane tech reset.

47:48

So the market would crash.

47:48

A  couple weeks in, people are preparing now.

47:54

People are like, "Oh shit, let's start building  fabs.

47:54

Fuck all the environmental stuff."

47:54

War's probably happening.

47:58

The supply chain is trying  to figure out what the hell to do to refix it.

48:04

Six months in, the supply of chips for making new  cars is gone or sequestered to make military shit.

48:12

You can no longer make cars.

48:12

We don't even know how to make non-semiconductor induced cars, this  unholy concoction with all these chips.

48:22

Cars are like 40% chips now.

48:22

There are chips in the tires.

48:26

There's like 2,000+ chips in every car.

48:26

Every Tesla door handle has like four chips in it.

48:30

It’s like, "What the fuck?" Why?

48:30

It’s like shitty  microcontrollers and stuff but there’s 2000+ chips even in an internal combustion engine vehicle.

48:37

Every engine has dozens and dozens of chips.

48:43

Anyways, this all shuts down.

48:43

Not all of  the production, there's some in Europe, some in the US, some in Japan, some in Singapore.

48:48

In Europe they're going to bring in a guy to work on Saturday, until four. Yeah.

48:52

TSMC always builds new fabs.

48:59

They tweak production up in old fabs.

48:59

New designs move to the next nodes and old stuff fills in the old nodes.

49:05

Ever since TSMC has been the most important player… It’s not just TSMC,  there's UMC there, PSMC, and other companies.

49:15

Taiwan's share of total manufacturing  has grown every single process node.

49:19

In 130 nanometers, there's a lot.

49:19

That’s including many chips from Texas Instruments, Analog Devices,  or NXP.

49:24

100% of it is manufactured in Taiwan by PSMC, TSMC, UMC or whatever.

49:29

But then you step forward to 28 nanometer, 80% of the world's production is in Taiwan. Oh  fuck, right?

49:37

What’s made on 28 nanometer today?

49:45

It’s tons of microcontrollers and stuff  but also every display driver I see.

49:48

Cool, even if I could make my Mac chip, I  can’t make the chip that drives the display.

49:53

You just go down the list.

49:53

That means no fridges, no automobiles, no weed whackers, because that stuff has chips.

49:57

My toothbrush has Bluetooth in it. Why? I don’t know.

50:02

There are so many things that would  just go poof. We'd have a tech reset.

50:06

We were supposed to do this  interview many months ago.

50:09

I kept delaying because I was like, "Ah,  I don't understand any of this shit."

50:13

But it is a very difficult thing to  understand.

50:13

Whereas with AI, it’s like… You’ve just spent the time to— Sure but it feels like the kind of thing where you can pick up what's going  on in the field in an amateur kind of way.

50:29

In this field, I'm curious about how one  learns the layers of the stack.

50:29

It's not just papers online.

50:36

You can't just look up  a tutorial on how the transformer works.

50:40

It's many layers of really difficult shit.

50:40

There are 18-year-olds who are cracked at AI already.

50:45

There are high school dropouts who get jobs at OpenAI.

50:49

This existed in the  past.

50:49

Pat Gelsinger, the current CEO of Intel, grew up in the Amish area of Pennsylvania.

50:57

He went straight to work at Intel because he's just cracked.

50:59

That's not possible in semiconductors today.

51:04

You can't get a job at a tool company without at  least a master's in chemistry, probably a PhD.

51:12

Of the 75,000 TSMC workers, like  50,000 have a PhD or something insane.

51:20

There’s a next-level amount of how  specialized everything's gotten.

51:25

Whereas today, you can take someone like Sholto,  who started working on AI not that long ago.

51:31

Not to say anything bad about Sholto. No, he’s cracked.

51:31

He’s omega-cracked at what he does.

51:35

You could drop him into another part of the AI stack.

51:37

He understands it already and could probably become cracked at that too.

51:43

That's not the case in semiconductors.

51:48

You specialize like crazy  and can't just pick it up.

51:56

Sholto, what did he say… He just started— He was a consultant at McKinsey, and at night he would read papers  about robotics and run experiments.

52:04

And then people noticed him and  were like, "Who is this guy?

52:10

I thought everyone who knew about this was at  Google already. Come to Google."

52:10

That can't happen in semiconductors. It's just not possible. ArXiv is a free thing.

52:15

The paper publishing industry is abhorrent everywhere else.

52:24

You can't just download IEEE papers or SPIE papers or from other organizations.

52:29

At least up until late 2022 or early 2023 with Google's PaLM inference paper, all the  best stuff was just posted on the internet.

52:46

After that, there was a little bit of clamping  down by the labs, but there are also still all these companies making innovations in the public.

52:50

What is state-of-the-art is public.

52:50

That is not the case in semiconductors.

52:56

Semiconductors have been shut down since the 1970s basically.

52:58

It's crazy how little information has been formally transmitted from one country to another.

53:04

The last time you could really think of this was maybe the Samsung era.

53:08

So then how do you guys keep up with it? We don't know it.

53:13

I don't  personally think I know it.

53:16

If you don’t know it, what  are you making videos about? I spoke to one guy.

53:21

He’s a PhD in  etch, one of the top people in etch.

53:26

He’s like, "Man, you really know lithography."

53:26

I don't feel like I know lithography.

53:31

But then you talk to the people who know  lithography and they’re like, "You’ve done pretty good work in packaging." Nobody knows anything.

53:35

They all have Gell-Mann amnesia?

53:39

They’re all in this single well.

53:39

They’re digging deep for what they’re getting at.

53:46

They don’t know the other stuff well enough.

53:46

In some ways, nobody knows the whole stack.

53:51

The stratification of just  manufacturing is absurd.

53:56

The tool people don't even know exactly  what Intel and TSMC do in production, and vice versa, they don’t know exactly  how the tool is optimized like this.

54:04

How many different types of tools are there?

54:04

Each of those has an entire tree of all the things we've built, invented, and continue to  iterate upon, and then there’s the breakthrough innovation that happens every few years in it too.

54:15

If that’s the case and nobody knows the whole stack, how does the industry coordinate?

54:19

"In two years we want to go to the next process which has gate all-around and for that  we need X tools and X technologies developed…" It's a fascinating social phenomenon. You  can feel it.

54:33

I went to Europe earlier this year. Dylan had allergies.

54:40

I was talking to  those people. It's like gossip.

54:40

You start feeling people coalescing around something.

54:49

Early on, we used to have SEMATECH where American companies came together  and talked and hammered it out.

55:00

But in reality it was dominated by a single  company.

55:00

Nowadays it's more dispersed.

55:00

It's a blue moon arising kind of thing.

55:10

They are  going towards something. They know it.

55:10

Suddenly, the whole industry suddenly is  like, "This is it. Let's do it."

55:20

It's like God came and proclaimed it: "We  will shrink density 2x every two years."

55:25

Gordon Morris made an observation. It didn’t  go nowhere.

55:25

It went way further than he ever expected because it’s like, "There’s a  line of sight to get to here and here."

55:33

He predicted 7-8 years out,  multiple orders of magnitude of increases in transistors and it came true.

55:37

But by then, the entire industry was like, "This is obviously true.

55:42

This is the word of God."

55:44

Every engineer in the entire industry, tens  of millions of people, were driven to do this.

55:51

Now not every single engineer believed it but  people were like, "Yes, to hit the next shrink, we must do this, this, this and  these are the optimizations we make."

55:58

You have this stratification, every single  layer and abstraction layers through the entire stack.

56:03

It's an unholy concoction.

56:03

No  one knows what's going on because there's an abstraction layer between every single layer.

56:12

On this layer, the people below you and above you know what's going on.

56:17

Beyond that, you can try to understand, but not really...

56:22

I watched a video about IRDS or whatever, 10 or 20 years ago, where they're like, "We're  going to do EUV instead of the other thing.

56:29

This is the path forward."

56:36

How do they do that if  they don't have the whole picture of different constraints, trade-offs, and so on?

56:42

They kind of argue it out.

56:46

They get together and they talk and argue.

56:46

Basically, at some point, a guy somewhere says, "I think we can move forward with this."

56:53

Semiconductors are so siloed.

56:53

The data and knowledge within each layer is: A)  Not documented online at all because it's all siloed within companies.

57:04

B) There's a lot of human element to it because a lot of the knowledge, as Jon was  saying, is apprentice-master type of knowledge.

57:14

Or it's "I've been doing this for 30 years,"  and there's an amazing amount of intuition on what to do just when you see something.

57:19

AI can't just learn semiconductors like that.

57:26

But at the same time, there's a massive talent  shortage and ability to move forward on things.

57:37

Most of the equipment in  semiconductor fabs runs on Windows XP.

57:43

Each tool has a Windows XP server on it.

57:43

All the chip design tools have CentOS version 6, which is old as hell.

57:49

There are so many areas where it's so far behind.

57:57

At the same time, it's so hyper-optimized  that the tech stack is broken in that sense.

58:03

They're afraid to touch it.

58:03

Yeah, because it's an unholy amalgamation.

58:09

This thing should not work.

58:09

It's literally a miracle.

58:11

So you have all the abstraction layers.

58:11

One, there's a lot of breakthrough innovation that can happen now  stretching across abstraction layers.

58:19

Two, because there's so much inherent knowledge in  each individual one, what if I can just experiment and test at a 1000x or 100,000x velocity?

58:25

Some examples of where this is already shown true are some of NVIDIA's AI layout tools, and  Google as well, laying out the circuits within a small blob of the chip with AI.

58:39

Some of these RL design things, various simulation things...

58:45

But is that design or is that manufacturing?

58:49

It's all design, most of it is design.

58:49

Manufacturing has not really seen much of this yet, although it's starting to come in.

58:53

Inverse lithography, maybe. ILT and… maybe.

58:57

I don't know if that's AI.

58:57

Anyway, there's a tremendous opportunity to bring breakthrough innovation simply because there  are so many layers where things are unoptimized.

59:11

You see all these single-digit to low  double-digit advantages just from RL techniques from AlphaGo-type stuff, or not AlphaGo  but 5-8 year-old RL techniques being brought in.

59:27

Generative AI being brought in could  really revolutionize the industry, although there's a massive data problem.

59:31

Can you give the possibilities here in numbers in terms of maybe like a FLOP per  dollar or whatever the relevant thing is?

59:42

How much do you expect in the future  to come from process node improvements?

59:46

How much from just how the hardware is  designed because of AI?

59:46

We're talking specifically for GPUs.

59:53

If you had  to disaggregate future improvements.

1:00:00

It's important to state that semiconductor  manufacturing and design is the largest search space of any problem that humans do because it  is the most complicated industry that humans do.

1:00:12

When you think about it, there are 1e10, 1e11,  100 billion transistors, on leading-edge chips.

1:00:22

Blackwell has 220 billion transistors or something  like that.

1:00:22

Those are just on-off switches.

1:00:22

Think about every permutation of putting those together,  contact, ground, drain source, with wires.

1:00:35

There are 15 metal layers connecting every  single transistor in every possible arrangement.

1:00:40

This is a search space that  is literally almost infinite.

1:00:43

The search space is much larger than any  other search space that humans know of.

1:00:46

And what is the nature of the search?

1:00:46

What are you trying to optimize over? Useful compute.

1:00:50

If the goal is to optimize  intelligence per picojoule—and intelligence is some nebulous nature of what the  model architecture is, and picojoule is a unit of energy—how do you optimize that?

1:01:04

There are humongous innovations possible in architecture because the vast majority of  the power on an H100 does not go to compute.

1:01:16

There are more efficient ALU  (Arithmetic Logic Unit) designs.

1:01:25

But even then, the vast majority  of the power doesn't go there.

1:01:28

The vast majority of the power  goes to moving data around.

1:01:32

When you look at what the movement of  data is, it’s either networking or memory.

1:01:36

You have a humongous amount of movement relative  to compute and a humongous amount of power consumption relative to compute.

1:01:44

So how can you minimize that data movement and maximize the compute?

1:01:48

There are 100x gains possible from architecture.

1:01:53

Even if we literally stopped shrinking, we could  have 100x gains from architectural advancements. Over what time period?

1:01:58

The question is how much can we advance the architecture.

1:02:01

The other challenge is that the number of people designing chips has  not necessarily grown in a long time.

1:02:12

Company to company, it shifts, but within the  semiconductor industry in the US—the US designs the vast majority of leading-edge chips—the number  of people designing chips has not grown much.

1:02:23

What has happened is the output per  individual has soared because of EDA (Electronic Design Assistance) tooling.

1:02:27

Now this is all still classical tooling.

1:02:32

There's just a little bit of AI in there yet.

1:02:32

The question is what happens when we bring this in and how you can solve this search space somehow,  with humans and AI working together to optimize this so most of the power doesn’t just go to data  movement and the compute is actually very small.

1:02:51

The compute can get like 100x more  efficient just with design changes, and then you could minimize  that data movement massively.

1:03:01

You can get a humongous gain in  efficiency just from architecture itself.

1:03:04

Process node helps you innovate that there.

1:03:04

Power delivery helps you innovate that.

1:03:09

System design, chip-to-chip  networking helps you innovate that.

1:03:12

Memory technologies, there's  so much innovation there.

1:03:14

There are so many different vectors of innovation  that people are pursuing simultaneously.

1:03:21

NVIDIA gen to gen to gen will do more than 2x  performance per dollar.

1:03:21

I think that's very clear.

1:03:28

Hyperscalers are probably going to try and shoot  above that, but we'll see if they can execute.

1:03:32

There are two narratives you can  tell here of how this happens.

1:03:36

One is that these AI companies training the  foundation models understand the trade-offs of.

1:03:45

How much is the marginal increase in  compute versus memory worth to them?

1:03:49

What trade-offs do they want  between different kinds of memory?

1:03:52

They understand this, so the accelerators  they build can make these trade-offs in a way that's most optimal.

1:03:57

They can also design the architecture of the model itself in a way  that reflects the hardware trade-offs. Another is NVIDIA.

1:04:07

I don’t know how this works.

1:04:07

Presumably, they have some sort of know-how.

1:04:14

They're accumulating all this knowledge  about how to better design this architecture and also better search tools.

1:04:18

Who has the better moat here?

1:04:27

Will NVIDIA keep getting better at  design, getting this 100x improvement?

1:04:30

Or will it be OpenAI and Microsoft and  Amazon and Anthropic who are designing their accelerators and will keep getting  better at designing the accelerator?

1:04:39

There are a few vectors to go here.

1:04:39

One you mention is important to note.

1:04:45

Hardware has a huge influence on the  model architecture that's optimal.

1:04:48

It's not a one-way street  that better chip equals...

1:04:52

The optimal model for Google to run on TPUs,  given a certain amount of dollars and compute, is different architecturally than what it is  for OpenAI with NVIDIA stuff.

1:04:58

It's absolutely different.

1:05:04

Even down to networking  decisions and data center designs that different companies make, the optimal  solution—X amount of compute of TPU vs.

1:05:14

GPU compute optimally—will diverge in what  the architecture is.

1:05:14

That's important to note.

1:05:21

Can I ask about that real quick?

1:05:21

Earlier we were talking about how China has the H20s or B20s, and there's much less compute  per memory bandwidth than the amount of memory.

1:05:36

Does that mean that Chinese models  will actually have very different architecture and characteristics  than American models in the future?

1:05:42

You can take this to a very large leap and  say, "Oh, neuromorphic computing or whatever is the optimal path and that looks very  different than what a transformer does."

1:05:51

Or you could take it to a simple thing, the  level of sparsity and coarse-grain sparsity, experts, and all this sort of stuff.

1:05:57

You have the arrangement of what exactly the attention mechanism is,  because there are a lot of tweaks.

1:06:04

It's not just pure transformer attention.

1:06:04

Or how wide versus tall the model is, that's very important.

1:06:11

D-mod versus number of layers.

1:06:11

These are all things that would be different.

1:06:19

I know they're different between, say,  Google and OpenAI and what is optimal.

1:06:23

But it really starts to get  like, "Hey, if you were limited on a number of different things like..."

1:06:26

China invests hugely in compute and memory.

1:06:34

The memory cell is directly  coupled or is the compute cell.

1:06:40

These are things that China's investing hugely in.

1:06:40

You go to conferences and there are like 20 papers from Chinese companies/universities  about compute and memory.

1:06:50

Because the FLOP limitation is here, maybe  NVIDIA pumps up the on-chip memory and changes the architecture because they still  stand to benefit tens of billions of dollars by selling chips to China.

1:06:59

Today, it's just neutered American chips that go to China.

1:07:01

But it'll start to diverge more and more architecturally because they'd be stupid not to  make chips for China.

1:07:06

Huawei, obviously, has their constraints.

1:07:12

Where are they limited on memory?

1:07:12

Oh, they have a lot of networking capabilities and they could move to certain optical  networking technologies directly onto the chip much sooner than we could.

1:07:22

Because that is what's optimal for them within their search space of solutions,  because this whole area is blocked off.

1:07:30

It's really interesting to think about  how the development of Chinese AI models will differ from American AI models  because of these changes or constraints. It applies to use cases. It applies to  data.

1:07:40

American models are very focused on learning from you, being able to  use you directly as a random consumer.

1:07:51

That's not the case for Chinese models, I assume.

1:07:51

There are probably very different use cases for them.

1:07:56

China crushes the West at video and image recognition.

1:07:58

At ICML, Albert Gu of Cartesia, who invented state space models, was there.

1:08:04

Every single Chinese person was like, "Can I take a selfie with you?" The man  was harassed.

1:08:07

In the US, you see Albert and it's awesome, he invented state space models.

1:08:10

It's not like state space models are revered here.

1:08:15

But that's because state space models potentially  have a huge advantage in video, image, and audio, which is stuff that China does more of and is  further along in and has better capabilities in.

1:08:29

Because of all the surveillance cameras there. Yeah.

1:08:29

That's the quiet part out loud.

1:08:33

But there's already divergence  in capabilities there.

1:08:36

If you look at image recognition,  China destroys American companies on that because of the surveillance.

1:08:40

You have this divergence in tech tree and people can start to design different architectures  within the constraints they're given.

1:08:50

Everyone has constraints, but the constraints  different companies have are even different.

1:08:56

Google's constraints have shown them that  they built a genuinely different architecture.

1:09:00

But now if you look at Blackwell and what's said  about TPU v6, they're not exactly converging but they are getting a little bit closer in  terms of how big the MatMul unit size is and some of the topology and world size of the  scale-up versus scale-out network.

1:09:14

There is some convergence slightly.

1:09:19

I’m not saying  they're similar yet, but they're starting to.

1:09:24

Then there are different architectures  that people could go down.

1:09:27

You see stuff from all these startups  that are trying to go down different tech trees because maybe that'll work.

1:09:30

There's a self-fulfilling prophecy here too.

1:09:34

All the research is in transformers that are very  high arithmetic intensity because the hardware we have is very high arithmetic intensity and  transformers run really well on GPUs and TPUs.

1:09:44

You sort of have a self-fulfilling prophecy.

1:09:44

If all of a sudden you have an architecture which is theoretically way better, but you  can get only like half of the usable FLOPs out of your chip, it's worthless because  even if it's a 30% compute efficiency win, it's half as fast on the chip.

1:09:58

There are all sorts of trade-offs and self-fulfilling prophecies  of what path people go down.

1:10:57

If you were made head of compute of a new AI  lab, if Ilya Sutskever’s new lab SSI came to you and they're like, "Dylan, we give you  $1 billion.

1:11:02

You’re our head of compute. Help us get on the map.

1:11:10

We're going to compete with the frontier labs." What is your first step? Okay.

1:11:12

The constraints are that you're a US/Israeli firm because that's what SSI is.

1:11:17

Your researchers are in the US and Israel.

1:11:24

You probably can't build data centers  in Israel because power is expensive as hell and it's probably risky.

1:11:28

So it’s still in the US most likely.

1:11:33

Most of the researchers are here in  the US, in Palo Alto or wherever.

1:11:40

You need a significant chunk of compute.

1:11:40

Obviously, the whole pitch is you're going to make some research breakthrough, compute  efficiency, data efficiency, or whatever it is.

1:11:50

You’re going to make some research  breakthroughs but you need compute to get there.

1:11:53

Your GPUs per researcher  is your research velocity.

1:11:58

Obviously, data centers are very tapped out.

1:11:58

Maybe not tapped out but in terms of every new data center coming up,  most of them have been sold.

1:12:05

That’s led people like Elon to go  through this insane thing in Memphis.

1:12:10

I'm just trying to square the circle.

1:12:10

On that question, I kid you not, in my group house group chat, there have  been two separate people who have been like, "I have a cluster of H100s and I have a long  lease on them, but I'm trying to sell them off."

1:12:27

Is it like a buyer's market right now?

1:12:27

Because it does seem like people are trying to get rid of them.

1:12:30

For the Ilya question, a cluster of like 256 GPUs or even 4K GPUs is kind of cope. It's not enough.

1:12:36

Yes, you're going to make compute efficiency wins, but with a billion dollars you probably just  want the biggest cluster in one individual spot.

1:12:47

Small amounts of GPUs are probably  not possible to use for them, and that's what most of the sales are.

1:12:54

You go and look at GPU List or Vast or Foundry or a hundred different GPU  resellers, the cluster sizes are small.

1:13:05

Now, is it a buyer's market? Yeah.

1:13:05

Last year  you would buy H100s for like $4 or $3 an hour for shorter-term or mid-term deals.

1:13:12

Right now, if you want a six-month deal, you could get it for like $2. 15 or less.

1:13:17

The natural cost… If I have a data center and I'm paying standard data center pricing to purchase  the GPUs and deploy them, it’s like $1. 40.

1:13:29

Add on the debt, because I probably took debt to  buy the GPUs or you have cost of equity, the cost of capital, it gets up to like $1. 70 or something.

1:13:33

You see deals that are… The good deals are like Microsoft renting from  CoreWeave at like $1. 90 to $2.

1:13:44

People are getting closer and closer.

1:13:44

But there's still a lot of profit.

1:13:47

The natural rate even after  debt and all this is $1. 70.

1:13:50

There’s still a lot of profit when  people are selling in the low twos.

1:13:53

GPU companies are deploying them,  but it is a buyer's market in the sense that it's gotten a lot cheaper.

1:13:57

Cost of compute is going to continue to tank.

1:14:02

I don’t remember the exact name of the law.

1:14:02

It's  effectively Moore's Law.

1:14:02

Every two years, the cost of transistors halved, and yet the industry grew.

1:14:08

Every six months or three months, the cost of intelligence… OpenAI and GPT-4 in February  2023, was roughly $120 per million tokens. Now it's like $10.

1:14:26

The cost of intelligence  is tanking, partially because of compute, partially because of the model's compute  efficiency wins.

1:14:33

That's a trend we'll see.

1:14:37

That's gonna drive adoption as you scale up and  make it cheaper and scale up and make it cheaper. Right.

1:14:40

Anyway, if you were head of compute at SSI… Okay, I’m head of compute at SSI.

1:14:46

There's obviously no free data center  lunch, in terms of what we see in the data.

1:14:55

There’s no free lunch if you need compute for  a large cluster size, even six months out.

1:15:02

There's some availability, but not  a huge amount because of what X did.

1:15:05

xAI is like, "Oh shit, we're going to  buy a Memphis factory, put a bunch of mobile generators usually reserved for natural  disasters outside, add a Tesla battery pack, drive as much power as we can from the grid, tap  the natural gas line that's going to the natural gas plant two miles away, the gigawatt  natural gas plant, and just send it.

1:15:28

Get a cluster built as fast as possible."

1:15:28

Now you're running 100K GPUs.

1:15:28

That costs about $4-5 billion, not $1 billion.

1:15:33

The scale that SSC has is much smaller.

1:15:42

Their cluster size will be  maybe 1/3 or 1/4 of that size.

1:15:48

So now you're talking about a 25K to 32K  cluster.

1:15:48

You still don't have that.

1:15:48

No one is willing to rent you a 32K cluster  today, no matter how much money you have.

1:15:57

Even if you had more than a billion dollars.

1:15:57

Now it makes the most sense to build your own cluster instead of renting, or get a very close  relationship like OpenAI/Microsoft with CoreWeave or Oracle/Crusoe The next step is Bitcoin.

1:16:08

OpenAI has a data center in Texas, or it's going to be their data center.

1:16:20

It’s kind of contracted and all that from CoreWeave.

1:16:22

There is a 300 megawatt natural gas plant on site, powering these crypto mining data centers  from a company called Core Scientific.

1:16:34

They're just converting that.

1:16:34

There's a lot  of conversion, but the power's already there.

1:16:38

The power infrastructure is already there.

1:16:38

It's really about converting it, getting it ready to be water-cooled, all that sort of stuff,  and converting it to a 100,000 GB200 cluster.

1:16:45

They have a number of those going up across  the country, but that's also tapped out to some extent because NVIDIA is doing the same  thing in Plano, Texas for a 32,000 GPU cluster that they're building. Is NVIDIA doing that?

1:16:57

Well, they're going through partners.

1:16:57

Because  this is the other interesting thing: the big tech companies can't do crazy shit like Elon did. Why? ESG.

1:17:05

They can't just do crazy shit like… Oh that’s interesting.

1:17:05

Do you expect Microsoft and Google and whoever to like drop their net zero  commitments as the scaling picture intensifies? Yeah.

1:17:17

What xAI is doing isn't that  polluting in the scheme of things, but it's you have 14 mobile generators and you're  just burning natural gas on site on these mobile generators that sit on trucks.

1:17:31

Then you have power directly two miles down the road.

1:17:35

There's no way to say any of the power is green because up to two miles  down the road is a natural gas plant as well.

1:17:43

You go to the CoreWeave thing.

1:17:43

There's a natural gas plant literally on site from Core Scientific and all that.

1:17:47

Then the data centers around it are horrendously inefficient.

1:17:50

There's this metric called PUE, which is basically how much power is brought  in versus how much gets delivered to the chips.

1:17:58

The hyperscalers, because they're so  efficient, their PUE is like 1. 1 or lower. I. e.

1:18:05

, if you get a gigawatt in, 900  megawatts or more gets delivered to chips.

1:18:11

It’s not wasted on cooling  and all these other things.

1:18:14

This Core Scientific one is  going to be like 1. 5 or 1. 6. I. e.

1:18:17

, even though I have 300 megawatts  of generation on site, I only deliver like 180-200 megawatts to the chips.

1:18:21

Given how quickly solar is getting cheaper… There’s also the fact that the  reason solar is difficult elsewhere is you've got to power the homes at night.

1:18:31

Here I guess it's theoretically possible to figure out only running the  clusters in the day or something… Absolutely not. That's not possible.

1:18:42

Because it's so expensive to have these GPUs? Yes.

1:18:47

So when you look at the power cost of a  large cluster, it's trivial to some extent.

1:18:55

The meme that you can't build a data center  in Europe or East Asia because the power is expensive, that's not really relevant.

1:19:00

Or there’s the meme that power is so cheap in China and the US that those are the  only places you can build data centers.

1:19:06

That's not really the real reason.

1:19:06

It's the ability to generate new power for these activities.

1:19:08

That’s why it's really difficult, the economic regulation around that.

1:19:13

But the real thing is if you look at the cost of ownership of an H100.

1:19:17

Let's just say you gave me a billion dollars and I already have a data  center, I already have all this stuff.

1:19:25

I'm paying regular rates for the data centers,  not paying through the nose or anything.

1:19:28

Paying regular rates for power,  not paying through the nose.

1:19:30

Power is sub 15% of the cost.

1:19:30

It's sub 10% of the cost actually.

1:19:35

The biggest, like 75-80% of  the cost, is just the servers.

1:19:40

And this is on a multi-year basis, including debt  financing, including cost of operation, all that.

1:19:45

When you do a TCO (total cost of ownership),  like 80% is the GPUs, 10% is the data center, 10% is the power, rough numbers.

1:19:51

So it's kind of irrelevant how expensive the power is.

1:19:57

You'd rather do what Taiwan does.

1:20:01

What did they do when there were droughts?

1:20:01

They forced people to not shower.

1:20:06

They basically reroute the power… When there  was a power shortage in Taiwan, they basically rerouted power from the residential areas.

1:20:11

And this will happen in a capitalistic society as well, most likely because, "fuck you, you  aren’t going to pay X dollars per kilowatt hour.

1:20:20

To me, the marginal cost of power is irrelevant.

1:20:20

Really it's all about the GPU cost and the ability to get the power.

1:20:24

I don't want to turn it off eight hours a day. Let’s zoom out a bit.

1:20:28

Let's discuss  what would maybe happen if the training regime changes and if it doesn't change.

1:20:30

You could imagine that the training regime becomes much more parallelizable where it's  about coming up with some sort of search and most of the compute for training is used to come  up with synthetic data or do some kind of search.

1:20:45

That can happen across a wide area.

1:20:45

In that world, how fast could we scale?

1:20:53

Let's go through the numbers year after year.

1:20:53

You would know more than me, but then suppose it has to be the current regime.

1:20:59

Just explain what that would mean in terms of how distributed that would  have to be and how plausible it is to get clusters of certain sizes over the next few years.

1:21:08

It's not too difficult for Ilya's company to get a cluster of like 32K of Blackwell next year.

1:21:14

Forget about Ilya's company, let's talk about the clear players. Like 2025, 2026, 2027.

1:21:19

2025, 2026… Before I talk about the US, it's important to note that there's like  a gigawatt plus of data center capacity in Malaysia next year now. That's mostly ByteDance.

1:21:32

Power-wise there's also the humongous damming of the Nile in Ethiopia and the country uses like  one-third of the power that that dam generates.

1:21:44

There's like a ton of power there to… How much power does that dam generate?

1:21:46

It's like over a gigawatt.

1:21:46

The country consumes  like 400 megawatts or something trivial.

1:21:52

Are people bidding for that power?

1:21:52

I think people just don't think they can build a data center in Ethiopia. Why not?

1:21:57

I don't think the dam is filled yet, is it?

1:21:57

No, the dam could generate that power. They just don't.

1:22:02

There's a little bit more  equipment required, but that's not too hard. Why don't they?

1:22:06

There are true security  risks.

1:22:06

If you're China, or if you're a US lab, to build a data center with all your IP in  Ethiopia… You want AGI to be in Ethiopia?

1:22:23

You want it to be that accessible?

1:22:23

People can't even monitor the technicians in the data center or powering  the data center, all these things.

1:22:31

There are so many things you could do...

1:22:31

You could just destroy every GPU in a data center if you want if you just fuck  with the grid, pretty easily, I think.

1:22:40

People talk a lot about it in the Middle East.

1:22:42

There's a 100k GB200 cluster  going up in the Middle East.

1:22:49

The US is clearly doing stuff too.

1:22:49

G42 is the UAE data center company, cloud company.

1:22:55

Their CEO is a Chinese national, or not a Chinese  national but there’s basically Chinese allegiance.

1:23:03

OpenAI wanted to use a data center from  them but instead… The US forced Microsoft—I feel like this is what happened— to do a deal  with them so that G42 has a 100K GPU cluster, but Microsoft is administering and  operating it for security reasons.

1:23:18

There's Omniva in Kuwait, like  the Kuwait super-rich guy spending five plus billion dollars on data centers.

1:23:23

You just go down the list, all these countries.

1:23:27

Malaysia has $10+ billion of AI data center  build outs over the next couple of years.

1:23:35

Go to every country, this stuff is happening.

1:23:35

But in the grand scheme of things, the vast majority of the compute is being  built in the US, then China, then Malaysia, Middle East, and the rest of the world.

1:23:44

Let’s go back to your point. You have synthetic data.

1:23:49

You have the search stuff.

1:23:49

You have all these post-training techniques.

1:23:56

You have all these ways to soak up flops, or you  just figure out how to train across multiple data centers, which I think they have.

1:24:03

At least, Microsoft and OpenAI have figured it out.

1:24:06

What makes you think they figured it out? Their actions.

1:24:09

Microsoft has signed deals  north of $10 billion with fiber companies to connect their data centers together.

1:24:16

There are some permits already filed to show people are digging between certain data centers.

1:24:20

We think, with fairly high accuracy, that there are five regions that they're connecting  together, which comprises many data centers.

1:24:34

What will be the total power usage of the...

1:24:34

Depends on the time, but easily north of a gigawatt.

1:24:37

Which is like close to a million GPUs.

1:24:42

Well, each GPU is getting  higher power consumption too.

1:24:46

The rule of thumb is that a H100 is  like 700 watts, but then total power per GPU all-in is like 1200-1400 watts.

1:24:50

But next-generation NVIDIA GPUs are like 1200 watts for the GPU.

1:24:57

It actually ends up being like 2000 watts all in.

1:25:02

There's a little bit of scaling of power  per GPU.

1:25:02

You already have 100K clusters.

1:25:08

OpenAI in Arizona, xAI in Memphis.

1:25:08

Many others are already building 100K clusters of H100s.

1:25:13

You have multiple, at least five, I believe GB200 100K clusters being built  by Microsoft/OpenAItheir partners for them.

1:25:25

It’s potentially even more.

1:25:25

500K GB200s is  like a gigawatt and that's online next year.

1:25:33

The year after that, if you aggregate all  the data center sites, and how much power… You only look at net adds since 2022, instead  of the total capacity at each data center, then you're still north of multi-gigawatt.

1:25:42

They're spending north of $10+ billion dollars on these fiber deals with a few fiber companies:  Lumen, Zayo, and a couple other companies.

1:25:53

Then they've got all these data centers  where they're clearly building 100K clusters, like old crypto mining sites with CoreWeave  in Texas or this Oracle/Crusoe thing in Texas.

1:26:03

You have them in Wisconsin and  Arizona and a couple other places.

1:26:07

There's a lot of data centers being built up.

1:26:07

You have providers like QTS and Cooper and many other providers and self-build data  centers, data centers I'm building myself.

1:26:20

Let's just give the number to like, "Okay,  it’s 2025 and Elon's cluster is going to be the biggest…" It doesn’t matter who it is.

1:26:26

There's the definition game.

1:26:26

Elon claims he has the largest cluster at 100K GPUs  because they're all fully connected.

1:26:34

Rather than who it is, I just  want to know how many… I don't know if it's better to denominate in H100s… It’s 100k GPUs this year for the biggest cluster.

1:26:45

Next year, 300-500k depending on  whether it's one site or many.

1:26:51

300-700k I think is the upper bound of that.

1:26:51

But it's about when they tier it on, when they can connect them, when the fiber's connected together.

1:26:57

Let's say 300-500k but those GPUs are 2-3X faster versus the 100K cluster.

1:27:05

So on an H100 equivalent basis, you're at a million chips next year  in one cluster by the end of the year.

1:27:14

Well, one cluster is the wishy-washy definition. It’s multisite, right? Can you do multisite?

1:27:20

What's the efficiency loss when you go multisite? Is it possible at all? I truly believe so.

1:27:20

What's the efficiency loss is the question.

1:27:27

Would it be like 20% loss, 50% loss? Great question.

1:27:32

This is  where you need the secrets.

1:27:36

And Anthropic's got similar plans  with Amazon and you go down the list.

1:27:40

And then the year after that? This is 2026.

1:27:40

In 2026 there is a single gigawatt site.

1:27:47

And that's just part of the  multiple sites for Microsoft.

1:27:51

The Microsoft five gigawatt thing happens in 2026?

1:27:51

One gigawatt, one site in 2026.

1:27:56

But then you have a number of others.

1:27:56

You have five different locations, some with multiple sites, some with single sites.

1:28:02

You're easily north of 2-3 gigawatts.

1:28:08

Then the question is, can you start  using the old chips with the new chips?

1:28:11

The flop scaling is going to continue much faster  than people expect, as long as the money pours in.

1:28:19

There's no way you can pay for the scale of  clusters being planned to be built next year for OpenAI unless they raise like $50-100  billion, which I think they will raise late this year or early next year. $50-100 billion? Are you kidding me? No. Sam has a superpower.

1:28:34

It's recruiting and  raising money.

1:28:34

That's what he's like a god at.

1:28:44

Will chips themselves be a  bottleneck to the scaling? Not in the near term.

1:28:46

It's more about  concentration versus decentralization.

1:28:52

The largest cluster is 100,000 GPUs.

1:28:52

NVIDIA has manufactured close to 6 million Hoppers across last year and  this year. So that's fucking tiny.

1:29:02

But then why is Sam talking about $7  trillion to build foundries and whatever? Draw the line.

1:29:08

It’s like a log-log line. Numbers go up, right?

1:29:08

If you do that, you're going from 100K to 300-500K,  where the equivalent is a million.

1:29:18

You just 10X year on year.

1:29:18

Do that again, do that again, or more.

1:29:21

If you increase the pace... What is "do that again"?

1:29:21

So in 2026, the number of...

1:29:25

If you increase the globally produced flops by like 30x year on year  or 10x year on year—and the cluster size grows by 3-7x and you start getting multi-site  going better and better—you can get to the point where multi-million chip clusters,  even if they're regionally not connected right next to each other, are right there.

1:29:50

And in terms of flops it would be 1e… what?

1:29:56

I think 1e30 is very possible in 2028 or 2029. Wow.

1:29:56

Okay you’re saying 1e30 you said by 2028-29.

1:30:05

That is literally six orders of magnitude.

1:30:05

That's like 100,000x more compute than GPT-4. Yes.

1:30:12

The other thing to say is the way you  count flops on a training run is really stupid.

1:30:18

You can't just do like active  parameters x tokens x six.

1:30:22

That's really dumb because the paradigm—as you  mentioned, and you’ve had many great podcasts on this stuff—it's like synthetic data and  RL stuff, post-training, verifying data, all these things generating and throwing  it away, search, inference time compute.

1:30:36

All these things aren't  counted in the training flops.

1:30:39

So you can't say 1e30 is a really stupid number  to say because by then the actual flops of the pre-training may be X, but the data to generate  for the pre-training may be way bigger, or the search inference time may be way, way bigger. Right.

1:30:51

Also, because you're doing adversarial synthetic data where the thing you're weakest at,  you can make synthetic data for that, it might be way more sample efficient.

1:31:02

So even though… Pre-training flops will be irrelevant.

1:31:02

I actually don't think pre-training flops will be 1e30.

1:31:07

I think more reasonably, it will be like the total summation of the flops that you deliver  to the model across pre-training, post-training, synthetic data for that pre-training and  post-training data, as well as some of the inference time compute efficiencies.

1:31:22

It's more like 1e30 in total.

1:31:27

Suppose you really do get to the  world where it's worth investing….

1:31:32

Actually, if you're doing 1e30, is  that like a trillion dollar cluster, hundred billion dollar cluster?

1:31:36

It'll be like multi-hundred billion dollars, but I truly believe people are going to be able  to use their prior generation clusters alongside their new generation clusters.

1:31:48

Obviously it’ll be smaller batch sizes or whatever, or use that to generate and  verify data, all these sorts of things.

1:31:56

And then for 1e30… Right now, I think 5% of  TSMC's N5 is NVIDIA or whatever percent it is.

1:32:06

By 2028, what percentage will it be?

1:32:06

Again, this is a question of how scale-pilled you are and how much money will  flow into this and how you think progress works.

1:32:16

Will models continue to get better  or does the line slope over?

1:32:20

I believe it'll continue to  skyrocket in terms of capability.

1:32:24

In that world—not of 5 nanometer, but of 2  nanometer, A16, A14, these are the nodes that'll be in that timeframe of 2028—used for AI, I  could see it being like 60-80% of it. No problem.

1:32:38

Given the fabs that are currently planned  and being built, is that enough for the 1e30, or will we need more? I think so, yeah.

1:32:44

Okay, so then the chip goal doesn't  make any sense.

1:32:44

The chip goal stuff about how we don't have enough compute… No, I think the plans of TSMC on two nanometer and such are quite aggressive for a reason.

1:32:55

To be clear, Apple, which has been TSMC's largest customer, does not need  how much 2nm capacity they're building.

1:33:07

They will not need A16, they will not need A14.

1:33:07

Apple doesn't need this shit.

1:33:07

Although they did just hire Google’s head of system design for TPU.

1:33:14

So they are going to make an AI accelerator.

1:33:22

But that’s besides the point.

1:33:22

Apple  doesn’t need this for their business.

1:33:26

They have been 25% or so of  TSMC's business for a long time.

1:33:29

And when you zone in on just the leading-edge,  they've been like more than half of the newest node or 100% of the newest node almost constantly. That paradigm goes away.

1:33:33

Let’s say you believe in scaling and you believe the models get better,  that the new models will generate amazing productivity gains for the world and so on.

1:33:47

If you believe in that world, then TSMC needs to act accordingly and the amount of  silicon that gets delivered needs to be there.

1:33:56

So in 2025 and 2026, TSMC is definitely there.

1:33:56

Then on a longer timescale, the industry can be ready for it, but it's going to be a constant  game of convincing them constantly that they must do this. It's not a simple game.

1:34:09

If people  work silently, it's not going to happen.

1:34:17

They have to see the demonstrated growth over  and over and over again across the industry. Who will need to see? Investors or companies?

1:34:26

More so, TSMC needs to see NVIDIA volumes continue  to grow straight up and Google's volumes continue to grow straight up, and so on down the list.

1:34:32

Chips in the near term, next year for example, are less of a constraint than data centers. It’s likewise for 2026.

1:34:39

The question for 2027-28… Always when you grow super rapidly,  people want to say that's the one bottleneck.

1:34:52

Because that's the convenient thing to say.

1:34:52

In 2023 there was a convenient bottleneck, CoWoS (Chip on Wafer on Substrate).

1:34:57

The picture has gotten much cloudier.

1:35:01

Not clouder but we can see  that HBM is a limiter too.

1:35:05

CoWoS is as well, CoWoS-L especially.

1:35:05

You have data centers, transformers, substations, power generation, batteries,  UPSs, CRHs, water cooling stuff.

1:35:16

All of this stuff is now a limitation  next year and the year after. Fabs will be in 2026-27.

1:35:19

Things will get cloudy  because the moment you unlock one… Only 10% higher, the next one is the thing.

1:35:25

Only 20% higher, the next one is the thing.

1:35:29

Today, data centers are like 4-5%  of total US power consumption.

1:35:34

When you think about it as a percentage  of US power, that's not that much.

1:35:37

But on the flip side you also  consider that all this coal has been curtailed and all these other things.

1:35:41

So power is not that crazy on a national basis.

1:35:48

On a localized basis, it is, because  it's about the delivery of it.

1:35:51

It’s the same with the substation  transformer supply chains.

1:35:54

These companies have operated in an environment  where the US power demand has been flat or even slightly down because of efficiency gains.

1:35:58

There has been humongous weakening of the industry.

1:36:07

Now all of a sudden if you tell that industry, "Your business will  triple next year if you can produce more."

1:36:15

They can only produce 50% more. Okay, fine.

1:36:15

Year after that, now we can produce 3x as much.

1:36:20

You do that to the industry, the US  industrial base as well as the Japanese, all across the world can get revitalized  much faster than people realize.

1:36:32

I truly believe that people can  innovate when given the need to.

1:36:35

It's one thing if it's a shitty industry where  my margins are low and we're not growing really.

1:36:44

All of a sudden it’s like, "Oh, this is  the sexiest time to be alive in power.

1:36:49

We're going to do all these different plans  and projects and people have all this demand.

1:36:53

They're begging me for another percent of  efficiency advantage because that gives them another percent to deliver to the chips."

1:36:56

You see all these things happen and innovation is unlocked.

1:37:01

You also bring in AI tools, you bring in all these things, innovation will be unlocked.

1:37:08

Production capacity can grow, not overnight, but it will in 6 months, 18 months, 3 year  timescales. It will grow rapidly.

1:37:12

You see the revitalization of these industries.

1:37:18

Getting people to understand that, getting people to believe… because you know, if  we pivot to like… Yeah, I'm telling you that Sam's going to raise $50-100 billion dollars because  he's telling people he's going to raise this much.

1:37:31

He’s literally having discussions with sovereigns  and Saudi Arabia and the Canadian pension fund and the biggest investors in the world.

1:37:41

Of course, Microsoft as well, but he's literally having these discussions  because they're going to drop their next model or they're going to show it off to people  and raise that money. This is their plan.

1:37:51

If these sites are already planned and... The money's not there. So how do you plan?

1:37:54

How do  you plan a site without...

1:37:57

Today Microsoft is taking on immense credit risk.

1:37:57

They've signed these deals with all these companies to do this stuff.

1:38:03

But Microsoft  doesn't have...

1:38:03

I mean, they could pay for it.

1:38:07

Microsoft could pay for it  on the current timescale.

1:38:11

Their CapEx is going from  $50-80 billion of direct CapEx, and then another $20 billion across Oracle,  CoreWeave, and then like another $10 billion across their data center partners.

1:38:21

They can afford that for next year.

1:38:28

This is because Microsoft  truly believes in OpenAI.

1:38:30

They may have doubts like, "Holy shit,  we're taking a lot of credit risk."

1:38:33

Obviously, they have to message  Wall Street and all these things, but that's affordable for them because they  believe they're a great partner to OpenAI.

1:38:41

They'll take on all this credit risk.

1:38:41

Now, obviously OpenAI has to deliver.

1:38:44

They have to make the next  model that's way better.

1:38:47

They also have to raise the money. And I think  they will.

1:38:47

I truly believe from how amazing 4o, how small it is relative to GPT-4…  The cost of it is so insanely cheap.

1:38:57

It's much cheaper than the API  prices lead you to believe.

1:39:00

You're like, "Oh, what if  you just make a big one?"

1:39:01

It's very clear what's going to happen to me on  the next jump that they can then raise this money and they can raise this capital from the world. This is intense. It's very intense.

1:39:12

Jon, actually, if he's right,  or not him, but in general.

1:39:16

If the capabilities are  there, the revenue is there… Revenue doesn't matter. Revenue matters.

1:39:23

Is there any part of that picture that still seems  wrong to you in terms of displacing so much of TSMC production, wafers and power and so forth?

1:39:28

Does any part of that seem wrong to you?

1:39:33

I can only speak to the semiconductors  part, even though I'm not an expert. I think TSMC can do it. They'll  do it.

1:39:37

I just wonder though… he's right in that 2024-25 is covered.

1:39:43

But 2026-27 is that critical point where you have to say, can the semiconductor  industry and the rest of the industry be convinced that this is where the money is?

1:39:54

That means, is there money by 2024-25?

1:40:00

How much revenue do you think the AI industry as  a whole needs by 2025 in order to keep scaling? Doesn't matter. Compared to smartphones.

1:40:04

I know he says it doesn't matter. I'll get to why. What are smartphones at?

1:40:11

Like Apple's  revenue is like $200 billion. So like...

1:40:14

Yeah, it needs to be another  smartphone-size opportunity, right?

1:40:17

Even the smartphone industry didn't drive this  sort of growth. It's crazy. Don't you think?

1:40:23

The only thing I can really perceive…  AI Girlfriend. You know what I mean.

1:40:30

No, I want a real one, dammit. There’s a few  things.

1:40:30

The return on invested capital for all of the big tech firms is up since 2022.

1:40:38

Therefore, it's clear as day that investing in AI has been fruitful so far for the big tech  firms just based on return on invested capital.

1:40:52

Financially, you look at  Meta's, you look at Microsoft's, you look at Amazon's, you look at Google's.

1:40:55

The return on invested capital is up since 2022. On AI in particular?

1:41:01

No, just generally as a company.

1:41:03

Now, obviously there's other factors here.

1:41:03

Like what is Meta's ad efficiency?

1:41:06

How much of that is AI, right? That’s super messy.

1:41:09

But here's the other thing,  this is Pascal's wager, right?

1:41:11

This is a matrix of like, do you believe in God? Yes or no.

1:41:11

If you believe in God and God's real and you go to heaven, that's great. That's fine. Whatever.

1:41:21

If you don't believe in God and God is real, then you're going to hell.

1:41:25

This is the deep technical analysis you'll subscribe to SemiAnalysis for.

1:41:28

Can you imagine what happens to the stock if Satya starts talking about Pascal's wager?

1:41:35

But this is psychologically what's happening, right?

1:41:39

Satya said it on his earnings call.

1:41:43

The risk of under-investing is worse  than the risk of over-investing.

1:41:47

He has said this word for word. This  is Pascal's wager.

1:41:47

I must believe I am AGI-pilled because if I'm not and my  competitor does it, I'm absolutely fucked.

1:41:55

Other than Zuck, who seems pretty convinced… No, Sundar said this on the earnings call. So Zuck said it. Sundar said it.

1:42:01

Satya's  actions on credit risk for Microsoft do it.

1:42:05

He's very good at PR and messaging,  so he hasn't said it so openly. Sam believes it. Dario believes it.

1:42:11

You look  across these tech titans, they believe it.

1:42:15

Then you look at the capital holders. The UAE believes it. Saudi believes it.

1:42:20

How do you know the UAE and Saudi believe it?

1:42:20

All these major companies and capital holders also believe it because  they're putting their money here.

1:42:27

But it won't last, it can't last unless  there's money coming in somewhere.

1:42:33

Correct, correct, but then the question is...

1:42:33

The simple truth is that GPT-4 costs like $500 million dollars to train.

1:42:38

It has generated billions in recurring revenue.

1:42:43

In the meantime, OpenAI raised $10 billion or  $13 billion and is building a model that costs that much, effectively.

1:42:49

Obviously they're not  making money.

1:42:49

What happens when they do it again?

1:42:56

They release and show GPT-5 with whatever  capabilities that make everyone in the world go, "Holy fuck."

1:43:01

Obviously the revenue takes time after you release the model to show up.

1:43:03

You still have only a few billion dollars or $5 billion of revenue at run rate.

1:43:08

You just raised $50-100 billion dollars because everyone sees this like, "Holy fuck, this  is going to generate tens of billions of revenue."

1:43:15

But that tens of billions takes time to flow  in.

1:43:15

It's not an immediate click.

1:43:15

But the time where Sam can convince, not just Sam… people's  decisions to spend the money are being made then.

1:43:27

Therefore, you look at the data  centers people are building.

1:43:29

You don't have to spend most of  the money to build the data center.

1:43:31

Most of the money's the chips, but you're  already committed to having so much data center capacity by 2027 or 2026 that you're  never gonna need to build a data center again for like 3-5 years if AI is not real.

1:43:41

Basically, that’s what all their actions are.

1:43:46

Or I can spend over a hundred billion  dollars on chips in 2026 and I can spend over a hundred billion dollars on chips in 2027.

1:43:50

These are the actions people are taking and the lag on revenue versus when you spend the money  or raise the money, there's a lag on this.

1:44:04

You don't necessarily need the  revenue in 2025 to support this.

1:44:08

You don't need the revenue  in 2026 to support this.

1:44:10

You need the revenue in 2025-2026 to support  the $10 billion that OpenAI spent in '23, or Microsoft spent in 2023 and early 2024 to build  the cluster, which model they trained in mid 2024, which they then released at the end of 2024, which  then started generating revenue in 2025-2026.

1:44:28

The only thing I can say is that you  look at a chart with three points on a graph: GPT-1, 2, 3, and you’re like… Even that graph… The investment you have to make in GPT-4 over GPT-3 is 100X.

1:44:36

The investment you had to make in GPT-5 over GPT-4 is 100X.

1:44:40

Currently the ROI could be positive—this very well could be true, I think it will be true—but  the revenue has to increase exponentially. Of course. I agree with you.

1:44:53

But I also  agree with Dylan. It can be achieved. ROI, TSMC does this. It invests $16 billion. It  expects ROI. It does that. I understand that. That's fine. Lag all that.

1:45:08

The thing that  I don't expect is that GPT-5 is not here.

1:45:14

It's all dependent on GPT-5 being good.

1:45:14

If GPT-5 sucks, if GPT-5 looks like it doesn't blow people's socks off, this is all void.

1:45:22

What kind of socks are you wearing, bro? Show them. AWS. GPT-5 is not here. GPT-5 is late. We don't know. I don't think it's late. I think it's late. Okay.

1:45:35

I want to zoom out and go back to the end of the decade picture again. We've already lost Jon.

1:45:45

We've already accepted GPT-5 would be good? Hello? You gotta, you know?

1:45:52

Life is so much more fun when  you just are delusionally… We're just ripping bong hits, are we?

1:45:56

When you feel the AGI, you feel your soul.

1:46:03

This is why I don't live in San Francisco.

1:46:03

I have tremendous belief in the GPT-5 area. Why?

1:46:09

Because of what we've seen already.

1:46:11

The public signs all show that  this is very much the case.

1:46:15

What we see beyond that is more questionable  and I'm not sure because I don't know.

1:46:23

We'll see how much they progress.

1:46:23

If things continue to improve, life continues to radically  get reshaped for many people.

1:46:34

Every time you increment up the intelligence,  the amount of usage of it grows hugely.

1:46:39

Every time you increment the cost  down of that amount of intelligence, the amount of usage increases massively.

1:46:43

As you continue to push that curve out, that's what really matters.

1:46:47

It doesn't need to be today.

1:46:52

It doesn't need to be revenue vs. how much CapEx.

1:46:52

In any time in the next few years, it just needs to be, did that last humongous chunk of CapEx  make sense for OpenAI or whoever the leader was?

1:47:03

How does that then flow through?

1:47:03

Or were they able to convince enough people that they can raise this much money?

1:47:06

You think Elon's tapped out of his network with raising $6 billion? No.

1:47:10

xAI is going to be able to raise $30+ billion easily.

1:47:14

You think Sam's tapped out?

1:47:14

You think Anthropic's tapped out?

1:47:14

Anthropic's barely even diluted the company relatively.

1:47:21

There's a lot of capital to be raised.

1:47:27

Call it FOMO if you want, but during the  dot-com bubble, the private industry flew through like $150 billion a year.

1:47:33

We're nowhere close to that yet.

1:47:38

We're not even close to the dot-com bubble.

1:47:38

Why would this bubble not be bigger?

1:47:42

You go back to the prior bubbles: PC bubble,  semiconductor bubble, mechatronics bubble.

1:47:46

Throughout the US, each bubble was smaller.

1:47:46

I don't know if you call it a bubble or not.

1:47:50

Why wouldn't this one be bigger?

1:47:50

How many billions of dollars a year is this bubble right now? For private capital?

1:47:53

It's like $55-60 billion so far for this year. It can  go much higher.

1:47:57

I think it will next year. Let me think about this.

1:48:07

You need another bong rip.

1:48:12

At least like finishing up and looping into the  next question… Prior bubbles also didn't have the most profitable companies that humanity has ever  created investing and they were debt financed.

1:48:21

This is not debt financed yet.

1:48:21

That's the last little point on that one.

1:48:25

Whereas the 90s bubble was very debt financed.

1:48:25

That was disastrous for those companies. Yeah, sure. But so much was built.

1:48:29

You’ve got  to blow a bubble to get real stuff to be built.

1:48:35

It is an interesting analogy.

1:48:35

Even though  the dot-com bubble obviously burst and a lot of companies went bankrupt, they  in fact did lay out the infrastructure that enabled the web and everything.

1:48:43

You could imagine an AI… A lot of the foundation model companies or whatever,  a bunch of companies will go bankrupt, but they will enable the singularity.

1:48:52

At the turn of the 1990s, there was an immense amount of money invested in things like  MEMS and optical technologies because everyone expected the fiber bubble to continue.

1:49:02

That all ended in 2003 or 2002. It started in 94?

1:49:06

There hasn't been revitalization since.

1:49:10

You could risk the possibility of a… Bro, Lumen, one of the companies that's doing the fiber build out for Microsoft, its stock  like fucking 4x'd last month, or this month.

1:49:20

How'd it do from 2002 to 2024?

1:49:20

Oh no, horrible, horrible, but we're gonna rip, baby. Rip that bong, baby!

1:49:23

You could freeze AI for another two decades.

1:49:30

Sure, sure, it’s possible.

1:49:30

Or people can  see a badass demo from GPT-5, a slight release, they raise a fuckload of money.

1:49:36

It could even be like a Devin-like demo, where it's like complete bullshit, but it's fine. Edit that out! Edit that out! No, it's fine. I don’t really care.

1:49:47

The capital's going to flow in.

1:49:47

Now whether it deflates or not is an irrelevant concern in the near term because  you operate in a world where it is happening.

1:50:03

What is that Warren Buffett quote?

1:50:03

I don't even know if it's Warren Buffett.

1:50:07

You don’t know who's swimming  naked until the tide goes out? No, no, no.

1:50:10

The one about how the market is  delusional for longer than you can remain solvent, or something like that. Oh, that's not Buffett. That's not Buffett?

1:50:16

That’s John Maynard Keynes.

1:50:20

Oh shit, that's that old? Okay. So Keynes said  it.

1:50:20

So this is the world you're operating in.

1:50:28

It doesn't matter what exactly happens.

1:50:28

There'll be ebbs and flows, but that's the world you're operating in.

1:50:32

I reckon that if the AI bubble pops, each one of these CEOs lose their jobs. Sure.

1:50:37

Or if you don't invest and you lose, it's a Pascalian wager. That's  much worse.

1:50:43

Across decades, the largest company at the end of each decade of  the largest companies, that list changes a lot.

1:50:52

And these companies are the  most profitable companies ever.

1:50:54

Are they going to let themselves lose it?

1:50:54

Or are they going to go for it?

1:50:59

They have one shot, one opportunity to make  themselves into… the whole Eminem song.

1:51:05

I want to hear the story of how both of you  started your businesses, the thing you're doing now. Jon, how did it begin?

1:51:12

What were  you doing when you started the YouTube channel?

1:51:19

It’s all about your textile company? Oh my god, no way. Please, please. Wait, is he joking?

1:51:25

If he doesn't want to, we'll talk about it later. The story's famous.

1:51:25

I've told it a million times.

1:51:31

Asianometry started off as a tourist channel.

1:51:31

I moved to Taiwan for work. DoIng what?

1:51:38

I was working in cameras.

1:51:44

What was the other company you started?

1:51:44

It tells too much about me.

1:51:52

I worked in cameras and then  basically I went to Japan with my mom.

1:51:56

My mom was like, "Hey, what  are you doing in Taiwan?

1:51:59

I don't know what you're doing."

1:51:59

I was like, "All right, mom, I will go back to Taiwan and I'll make stuff for you." I made  videos.

1:52:02

I would go to the Chiang Kai-shek Park and be like, "Hi, mom, this park was this, this."

1:52:07

Eventually, you run out of stuff.

1:52:11

But then it's a pretty smooth transition from  that into Chinese history, Taiwanese history.

1:52:19

People started calling me Chinanometry. I didn't  like that.

1:52:19

So I moved to other parts of Asia.

1:52:25

What year did people start watching your videos?

1:52:25

Let's say like a thousand views per video or something?

1:52:32

Oh my gosh, I started the channel in 2017 and it wasn't until  2018 or 2019 that it actually… I labored on for the first three years with no one watching.

1:52:40

I got like 200 views and I'd be like, "Oh, this is great."

1:52:45

Were the videos basically like the ones you have now?

1:52:47

Sorry, backing up for the audience who might not know, I imagine basically everybody  knows Asianometry, but if you don't it's the most popular channel about semiconductors, Asian  business history, business history in general, geopolitics, history and so forth.

1:53:03

Honestly, I've done research for different AI guests and different  things I'm trying to understand. How does hardware work? How does AI work? How does a zipper work?

1:53:14

Did you watch that video?

1:53:19

I think it was a span of three videos.

1:53:19

It's like the Russian oil industry in the 1980s and how it funded everything and then when  it collapsed, they were absolutely fucked.

1:53:28

Then the next video was like,  the zipper monopoly in Japan. Not a monopoly.

1:53:33

The next video was about ASML. Not a monopoly.

1:53:33

Strong holding in a mid-tier size.

1:53:33

They're like the luxury zipper makers.

1:53:40

Asianometry is always just  stuff I'm interested in.

1:53:43

I'm interested in a whole  bunch of different stuff.

1:53:46

Then the channel, for some reason  people started watching the stuff I do.

1:53:50

I still have no idea why.

1:53:50

To be honest, I still feel like a fraud.

1:53:55

I sit in front of Dylan and I feel  like a legit fraud, especially when he starts talking about 60,000 wafers and all that.

1:54:00

I feel like I should know this but in the end, I just try my best to bring interesting stories out.

1:54:08

How do you make a video every single week?

1:54:08

These are like… Two a week?

1:54:16

You know how long he had a full-time job? Five years, six years. While doing this?

1:54:19

Sorry, a textile business and a full-time job.

1:54:23

Wait, no it’s a full-time job, textile  business, and Asianometry for a long long time.

1:54:27

I literally just gave up the  textile business this year.

1:54:30

How are you doing research and making a  video twice a week? I don't know.

1:54:30

I do these, I'm just fucking talking. This is all I  do.

1:54:35

I do this once every like two weeks.

1:54:40

The difference is, Dwarkesh, you go to SF  Bay Area parties constantly and Jon is locked in.

1:54:47

He's like locked in 24/7.

1:54:47

He's got the TSMC work ethic and I've got the Intel work ethic.

1:54:52

If I don't… I got the Huawei ethic.

1:54:56

If I do not finish this video,  my family will be pillaged.

1:55:01

He actually gets really stressed about  it, not doing something on his schedule.

1:55:08

I do two videos per week, I  write them both simultaneously.

1:55:12

How are you scouting out  future topics you want to do?

1:55:16

You just pick up random articles, books, whatever.

1:55:16

If you find it interesting, you make a video about it?

1:55:19

Sometimes what I'll do is Google a country, I'll Google an industry, and I'll Google what a country  is exporting now and what it used to export.

1:55:28

I compare that and I say, "That's my video."

1:55:28

Sometimes it’s also just as simple as, "I should do a video about YKK." This zipper is nice.

1:55:34

I should do a video about it. I do.

1:55:40

It literally is… Do you keep a list? "Here's the next one.

1:55:44

Here's the one after that."

1:55:44

I have a long list of ideas.

1:55:49

Sometimes it's as vague as Japanese whiskey.

1:55:49

I have no idea what Japanese whiskey is about. I heard about it before. I watched  that movie.

1:55:55

So I was just like, "Okay, I should do a video about that."

1:55:59

How many research topics do you have on the back burner, basically?

1:56:04

I’m talking about something where you’re reading about it constantly and then in a month  or so you’re like, "I'll make a video about it."

1:56:10

I just finished a video about how IBM lost the PC.

1:56:10

Right now I'm unstressing about that.

1:56:16

But I'll kind of move right on to… The  videos do kind of lead into others.

1:56:21

Right now this one is about how IBM lost the PC.

1:56:21

Now what’s next is how Compaq collapsed, how the wave destroyed Compaq. So I'll do that.

1:56:27

At the  same time, I'm dual lining a video about qubits.

1:56:35

I'm dual lining a video about directed  self-assembly for semiconductor manufacturing, which I'll read a lot of Dylan's work for.

1:56:42

But then a lot of that is in the back of my head.

1:56:49

I'm producing it as I go. Dylan, how do you work?

1:56:53

How does one go from Reddit shitposter to running  a semiconductor research and consulting firm?

1:57:00

Let's start with the shitposting. It's a long line.

1:57:00

I had immigrant parents and I grew up in rural Georgia.

1:57:04

When I was seven, I begged for an Xbox and when I was eight I got it. It was the 360.

1:57:09

They had a  manufacturing defect called the red ring of death.

1:57:14

There are a variety of fixes that I  tried like putting a wet towel around the Xbox, something called the penny trick.

1:57:17

Those all didn't work, my Xbox still didn't work.

1:57:22

My cousin was coming next weekend and he's  like two years older than me. I look up to him.

1:57:27

He's in between my brother and me  but I'm like, "Oh, no, no, we're friends.

1:57:31

You don't like my brother as much as you like me."

1:57:31

My brother's more of a jock-y type. It didn't matter.

1:57:34

He didn't really care  that the Xbox was broken.

1:57:39

He's like, "You better fix it though.

1:57:39

Otherwise parents will be pissed."

1:57:39

I figured out how to fix it online.

1:57:42

I tried a variety of fixes, ended up shorting the temperature sensor.

1:57:46

That worked for long enough until Microsoft did the recall.

1:57:49

But in that, I learned how to do it out of necessity on the forums.

1:57:54

I was a nerdy kid, so I liked games, but whatever.

1:57:59

There was no other outlet so  once I was like, "Holy shit, this is Pandora's box what just got opened up,"  then I just shitposted on the forums constantly.

1:58:05

I did that for many, many years and  then I ended up moderating all sorts of Reddits when I was a tween and teenager.

1:58:11

Then as soon as I started making money… I grew up in a family business, but I didn't get  paid for working, of course, like yourself.

1:58:22

But as soon as I started making money at my  internships, I was like 18 or 19, I started making money.

1:58:28

I started investing in semiconductors.

1:58:28

I  was like, "Of course, this is the shit I like."

1:58:35

By the way, the whole way through as technology  progressed, especially mobile, it went from very shitty chips and phones to very advanced.

1:58:40

Every generation they'd add something and I'd read every comment, I'd read  every technical post about it.

1:58:49

Also, I was interested in all the history around  that technology and who's in the supply chain and just kept building and building and building.

1:58:52

I went to college and did data science type stuff.

1:58:57

I went to work on hurricane/earthquake/wildfire  simulation and stuff for a financial company.

1:59:04

During college I wasn't shitposting  on the internet as much.

1:59:07

I was still posting some, but I was following  the stocks and all these sorts of things, the supply chain, all the way  from the tool equipment companies.

1:59:14

The reason I liked those is because  all this technology, it’s made by them.

1:59:17

Did you have friends in person who were  into this shit, or was it just online?

1:59:22

I made friends on the internet. Oh, that's dangerous.

1:59:27

I've only ever had like literally  one bad experience and that was just because he was drugged out.

1:59:30

One bad experience online?

1:59:34

Meeting someone from the internet in person.

1:59:34

Everyone else has been genuinely… You have enough filtering before that point.

1:59:38

Even if they're hyper mega autistic, it's cool. I am too. No, I’m just  kidding.

1:59:44

You go through the layers and you look at the economic angle, you look at  the technical angle, you read a bunch of books.

1:59:56

You can just buy engineering textbooks and read  them. What's stopping you?

1:59:56

If you bang your head against the wall, you learn it.

2:00:03

While you were doing this, did you expect to work on this at some  point or was it just pure interest?

2:00:09

No, it was an obsessive hobby of  many years and it pivoted all around.

2:00:14

At some point I really liked gaming.

2:00:14

Then I moved into phones and rooting them and underclocking them, and the chips there,  and screens and cameras, and then back to gaming, and then to data center stuff because that was  where the most advanced stuff was happening.

2:00:30

I liked all sorts of telecom  stuff for a little bit.

2:00:33

It bounced all around, but generally it was in  computing hardware. I did data science.

2:00:33

I said I did AI when I interviewed but it was like  bullshit multivariable regression, whatever.

2:00:47

It was simulations of hurricanes, earthquakes,  and wildfires for financial reasons.

2:00:56

I had a job for three years after  college. I was posting.

2:00:56

I had a blog, an anonymous blog for a long time.

2:01:00

I'd even made some YouTube videos and stuff.

2:01:03

Most of that stuff is scrubbed off the  internet, including Internet Archive, because I asked them to remove it.

2:01:11

In 2020, I quiet quit my job and started  shitposting more seriously on the internet.

2:01:16

I moved out of my apartment and  started traveling through the US.

2:01:21

I went to all the national  parks, in my truck/tent.

2:01:24

I also stayed in hotels and motels  for like three or four days a week.

2:01:28

I started posting more frequently on the internet.

2:01:28

I'd already had some small consulting arrangements in the past.

2:01:33

But it really started to pick up in mid-2020, consulting arrangements  from the internet from my persona. What kinds of people?

2:01:40

Was it  investors, hardware companies?

2:01:44

It was people who weren't in hardware  that wanted to know about hardware.

2:01:44

It would be some investors.

2:01:48

Some VCs  did it, some public market folks.

2:01:54

There were times where companies would  ask about three layers up in the stack.

2:01:58

They saw me write some random posts and like "Hey,  can we…" There's all sorts of random stuff.

2:01:58

It was really small money.

2:02:03

Then in 2020, it really picked  up and I thought, "Why don't I just arbitrarily make the price way higher?" And it worked.

2:02:09

I made  a new newsletter as well. I kept posting.

2:02:09

The quality kept getting better because people would  read it and be like, "this is fucking retarded.

2:02:22

Here’s what’s actually right"  over more than a decade.

2:02:29

Towards the end of 2021,  I made a paid post because someone didn't pay for a report or whatever.

2:02:32

It was about photoresist and the developments in that industry, which is the stuff you  put on top of the wafer before you put in the lithography tool. It did great.

2:02:42

I went to  sleep that night and woke up the next day with 40 paid subscriptions. I thought, "What?

2:02:47

Okay,  let's keep going."

2:02:47

I started posting more paid content, partially free, partially paid.

2:02:52

I did all sorts of stuff covering advanced packaging, chips, data center stuff, and AI chips.

2:02:55

It was all sorts of stuff I was interested in and thought was interesting.

2:03:03

I always bridged economically—because I've read all the company's earnings since I  was 18 and I'm 28 now—all the way through to the technical stuff that I could.

2:03:13

In 2022, I also started going to every conference I could.

2:03:18

I go to about 40 conferences a year, not trade show type conferences, but  technical ones: chip architecture, photoresist, AI and NeurIPS, ICML, and so on.

2:03:28

How many conferences do you go to a year? Like 40.

2:03:34

So you basically live at conferences? Yeah.

2:03:36

I've been a digital nomad since 2020.

2:03:36

I've basically stopped and I moved to SF now, kind of not really.

2:03:41

You can't say that the California government… I don't live in SF, come  on… But I basically do now.

2:03:50

California Internal Revenue Service.

2:03:50

Do not joke about those guys.

2:03:56

Seriously, don't joke about this.

2:03:56

They're going to send you a clip of this podcast and be like, "40% please."

2:03:58

I am in San Francisco sub four months a year contiguously, exactly 100 days or whatever it is,  179 days. Let's go, right?

2:04:05

Over the full course of the year…but no, I go to every conference, make  connections with all these very technical things like international electron device manufacturing,  lithography and advanced patterning, very large scale integration.

2:04:26

You have circuits conferences.

2:04:26

You just go to every single layer of the stack.

2:04:33

It's so siloed, there's tens of millions  of people that work in this industry.

2:04:36

But you go to every single one.

2:04:36

You try and understand the presentations.

2:04:38

You do the required  reading.

2:04:38

You look at the economics of it.

2:04:42

You are just curious and want to learn.

2:04:42

You can start to build up more and more.

2:04:47

The content got better and  what I followed got better.

2:04:50

Then I started hiring people in mid-2022, as well.

2:04:57

I got people in different layers of the stack.

2:04:57

Now today, almost every hyperscaler is a customer, not for the newsletter but for the data we sell.

2:05:06

Many major semiconductor companies, many investors, all these people are  customers of the data and stuff we sell.

2:05:17

The company has people all the  way from ex-CYMER, ex-ASML, all the way to ex-Microsoft and an AI company.

2:05:20

Through the stratification, now there's 14 people here in the company and all across  the US, Japan, Taiwan, Singapore, France.

2:05:34

It’s all over the world and across many  ranges of… You have ex-hedge funds as well.

2:05:41

You kind of have this amalgamation  of tech and finance expertise.

2:05:47

We just do the best work there, I think.

2:05:47

Are you still talking about a monstrosity? An unholy concoction.

2:05:52

We have data  analysis, consulting, etc.

2:05:52

for anyone who really wants to get deeper into this.

2:06:02

We can talk about people building big data centers, but how many chips are being made in  every quarter of what kind for each company?

2:06:14

What are the subcomponents of these chips?

2:06:14

What are the subcomponents of the servers?

2:06:19

We try to track all of that.

2:06:19

We follow every server manufacturer, every component manufacturer, every cable manufacturer,  all the way down the stack, tool manufacturer.

2:06:27

We know how much is being sold, where and  how, and project out all the way out to like, "Hey, where's every single data center?

2:06:32

What is the pace that it's being built out?"

2:06:37

This is the sort of data we want to have and sell.

2:06:37

The validation is that hyperscalers purchase it and they like it a lot.

2:06:43

AI companies do and semiconductor companies do.

2:06:51

How it got there to where it is, is just like,  "Just try and do the best and try to be the best."

2:06:56

If you were an entrepreneur who's like, "I  want to get involved in the hardware chain somewhere…" If you could start a business today,  somewhere in the stack, what would you pick?

2:07:07

Jon, tell him about your textile business.

2:07:07

I'd work in memory, something in memory.

2:07:15

The concept is that they have to  hold immense amounts of memory.

2:07:22

I think memory already is tapped technologically.

2:07:22

HBM exists because of limitations in DRAM.

2:07:30

I think it's fundamentally… We've forgotten it  because it is a commodity, but we shouldn't.

2:07:38

I think breaking memory would  change the world in that scenario.

2:07:45

The context here is that Moore's  Law was predicted in 1965.

2:07:49

Intel was founded in 1968 and released  their first memory chips in 1969 and 1970.

2:07:54

So a lot of Moore's Law was about memory.

2:07:54

The memory industry followed Moore's law up until 2012 when it stopped.

2:07:58

Tt became very incremental gains since then.

2:08:04

Whereas logic has continued and people are  like, "Oh, it's dying. It's slowing down."

2:08:04

At least there's still a little bit of coming.

2:08:07

It’s still more than 10-15% a year CAGR of growth and density and cost improvement.

2:08:14

Memory is literally since 2012, really bad.

2:08:20

When you think about the cost of memory…  It's been considered a commodity.

2:08:24

But memory integration with accelerators,  this is something… I don't know if you can be an entrepreneur here though.

2:08:29

That's  the real challenge.

2:08:29

Because you have to manufacture at some really absurdly large scale  or design something in an industry that does not allow you to make custom memory devices.

2:08:36

Or use materials that don't work that way.

2:08:42

So there's a lot of work there.

2:08:42

I don't necessarily agree with you.

2:08:45

But I do agree it's one of the most  important things for people to invest in.

2:08:49

It's really about where you are good at. Where  can you vibe?

2:08:49

Where can you enjoy your work and be productive in society?

2:08:54

There are a thousand different layers of the abstraction stack.

2:08:59

Where can you make it more efficient?

2:09:02

Where can you utilize AI to build better and make  everything more efficient in the world and produce more bounty and iterate feedback loop?

2:09:08

There's more opportunity today than any other time in human history, in my view.

2:09:14

Just go out there and try. What engages you?

2:09:21

Because if you're interested  in it, you'll work harder.

2:09:23

If you have a passion for copper wires…  I promise to God, if you make the best copper wires, you'll make a shitload of money.

2:09:28

If you have a passion for B2B SaaS, I promise to God, you'll make fuck loads of money.

2:09:34

I don't like B2B SaaS, but whatever you have a passion for.

2:09:41

Just work your ass off, try and innovate, bring AI into it.

2:09:44

Try and use AI to make yourself more efficient and make everything more efficient.

2:09:51

I promise, you will be successful. That's really the view.

2:09:56

It’s not necessarily that there's one  specific spot because every layer of the supply chain has… You go to the conferences.

2:10:01

You go to talk to the experts there.

2:10:04

It's like, "Dude, this is the stuff that's  breaking and we could innovate in this way.

2:10:08

These five extraction layers, we could innovate  this way." Yeah, do it.

2:10:08

There's so many layers where we're not at the Pareto optimal.

2:10:12

There's so much more to go in terms of innovation and inefficiency.

2:10:17

That's a great place to close.

2:10:21

Dylan, Jon, thank you so much  for coming on the podcast.

2:10:24

I'll just give people the reminder,  Dylan Patel, semianalysis. com.

2:10:29

That's where you can find all the technical  breakdowns that we've been discussing today.

2:10:33

Asianometry YouTube channel, everybody will  already be aware of Asianometry, but anyways, thanks so much for doing this. It was a lot of fun. Thank you. Yeah, thank you.