0:00
One of the most interesting sub plots of this whole world is the talent wars.
One of the most interesting sub plots of this whole world is the talent wars.
A cool idea is that as these things get better, maybe we begin to automate some of the research function that people formally would have played.
Do you see a world where like we're squeezing down the fewer and fewer number of people that really matter that will have all the impact on where we go in terms of like net new research and that means that all this crazy spending that's happening at Meta or elsewhere makes a lot of sense that like maybe even those numbers should be higher or something like this.
I think it's like tremendously hilarious that people are like, "Oh my god, this person is getting paid a billion dollars." It is infeasible.
It's like, "How could this person possibly be worth that much?"
Well, they're running the experiment ments on chips that cost you know, a hundred billion dollars.
If every wasted experiment they do, if they just used like a third of the compute and and their ideas and their impact on it wasted the compute was an idea that was already done or like, you know, like there's so much wasted compute.
I'll say I call it wasted.
It's trying stuff and failing.
But like none of us know what to try and what not to try.
And these things are so complicated.
There's like a group of people just trying different stuff on the existing data. How do you mix it?
What order do you feed it into the model?
Um, how do you filter it?
Like what's the architecture?
There's different people working on long context.
There's different people working on every single aspect of the model that like if you just make them a little bit more efficient that they come up with the idea that's 5% more efficient. Well, fantastic.
I just saved not only 5% of my compute time training time.
I also save 5% across my entire inference fleet.
And then I do it again and again and again and again because we're so far away from like these models being anywhere near as efficient as a human brain.
And we know it can at least get as efficient as us.
Maybe the compute substrate isn't the same, but like whatever, right?
Adding more people to the problem doesn't make it faster because there's so many things you're trying.
You run these experiments, you learn something, and then you implement it.
You tweak the knobs in these ways in a hundred different ways, and then you see the trend line and you're like, "Oh, so actually I should tweak it this way.
Let's implement that, right?"
There's so much like gut feel, there's so much like reading data, understanding implement re-implementing it into these things that if you add people, you're going to slow it down.
In a sense, a lot of like problems before they did the super intelligence thing was that they just had too many people that weren't led by leadership that was amazing.
Um, and they had like a lot of failed experiments and wasted time doing things that didn't matter.
There's a tweet from one of my friends uh at OpenAI.
Um, he's pretty famous on Twitter. His name is Rune.
He's like, "I get visibly viscerally angry every time I think about how many H100s Meta is wasting."
It's such a funny tweet because it's like, "Well, yeah, they're wasting a ton of compute."
They were, you know, maybe they still are, but you know, like and everyone's wasting compute, right?
OpenAI is wasting tons of compute cuz, you know, what's the Pareto optimal model architecture? Who knows?
Another thing I saw Rune say recently which was so interesting was uh why don't we just go make even more ridiculous offers around the people that have process knowledge for things that we want here in the US in other countries.
Like why don't we if we're getting pretty good at the Arizona, you know, fab that we've built and we we think that we can sort of extract the process knowledge from the people, why don't we like go acquire like all the best people in Shenzhen or all the best people in like other places in the world?
Do you think it starts to escalate to that level like so much is dependent on the process knowledge of a relatively small group of people and that and the war the talent war should actually be it shouldn't be Meta and OpenAI.
It should be like the US maybe through Meta and OpenAI and like people from all over the world.
Like do you think it starts to get that extreme and and should it?
That's almost a function of why like Intel has has fallen off a lot, right?
It's like um, you have all these geniuses in, you know, you know, nano chemistry and PhDs and all these like random like, you know, like things whether it be chemistry, physics.
All these like incredibly smart people but there's a whole class of incredibly smart people that never went that way cuz they're like, "Oh, those guys are making like 200k."
Like why would I do that?
I'm going to go I'm going to go to Google and make 800k.
And now I'm going to go to OpenAI and make 10 mil or now I'm going to go to Meta and go make 100 mil, right?
Like any like smart 18-year-old is going to be like, "Fuck that. I'm doing this." Right?
Why do the smartest doctors and I don't mean to say the smartest doctors in a general sense, but there's skewed really smart population of doctors that want to be dermatologists and anesthesiologists.
It's like, "Is that the most valuable thing for them to do?"
No, but those are the two professions that give you like good working hours and great pay.
Not to say that, you know, the general doctor is not smart as them, but if you took the population of general like family doctors like just the random doctor and you took the population of dermatologists the newest coming out of school, the ones that are being dermatologists and anesthesiologists are way smarter or at least were scored better, were able to get into the field that was coveted.
And so, yeah, talent war is like it it is truly like, you know, we've sort of been through this process of like capital has, you know, it's it's always been human human capital and and capital goods sort of those two vying off of each other.
And for a long time with mechanization, industrialization, we had the human capital decreasing as the industrial capital increased.
Um, and sort of that got to a point where especially in the 70s it really started to tank as the ability to globalize and and all these things started to really hit the US and and that's why you have all the lot of the like population level dynamics and income inequality that we have today that like is very bad for the psyche of the US and and the stability of it.
But then you have now you now have like we're in such a age of like well, actually like manufacturing things is pretty commodity.
Like most of the value doesn't come from the manufacturing of it.
It comes from the creation of the idea.
One thing Jensen told me which I thought was like amazing, right?
He's like, you know, Dylan, he's like, "The reason America's rich, like people have it all wrong.
The reason we're rich is because we've exported all the labor, but we've kept all the value."
And that's what Nvidia does, right?
They've exported the labor of making their chips. And Apple, right? Everyone. It's it's done in Asia.
Um, and those those companies make money.
Not as much money as Nvidia and Apple. Right.
All the gross profits are going to them.
Um, and then they're either reinvesting it or buying back stock or whatever.
However they allocate the capital is like, you know, different concern.
If like as you said the process knowledge is so valuable, why aren't we why aren't we doing this?
That's that's a that's a that's a that's a great idea. Rune's idea, not mine. >> Yeah, I know.
I mean, I think I think the challenge is how to choose people is really difficult.
So, for some roles someone who can talk the talk, they're great, right?
Like people just automatically assume they're great because they can talk the talk.
But you know how many people suck at talking and are really freaking good at doing, yeah.
Yeah, but then you don't know. You don't know, right?
Cuz it's like, well, they but then there's people who talk about being able to do better than the person who's doing and like and then like, you know, these tests are never as good, right?
Like so work trials is like, "How do you select?"
Um, and this was a big challenge for um, so some of the criticisms are like, "They didn't get all of the best people.
They actually got a lot of like bad people."
It's like the cope from like OpenAI and Anthropic and, you know, these kinds of companies are like, "No, no, no, they didn't get our best people." This is Sam said, right?
He's like, "They didn't get our best people. It's okay."
Meanwhile, he did have to do counter offers internally, right?
So, it's like, you know, um, so so as far as the process knowledge, I think I think the ML researchers are an extreme of how much value one can do.
But my favorite analogy that I came up with recently is that ML research is the exact same as semiconductor manufacturing.
You know, there's a ton of jobs in in in semiconductor manufacturing that don't exist in ML research, but it is a ton of tune a thousand different knobs, right?
Oh, you put the wafer in this tool, you're going to change the pressure of the chamber when you're doing the deposition or you're going to change the mix of the chemicals flowing in, which chemicals you're putting in, right?
Like what what speed you do it at?
Do you do it for 30 minutes?
Do you do it for 31 minutes?
Do you do it for, you know, obviously it splices way down.
There's so many knobs on every single tool and you have a thousand plus.
>> Input and process knobs. Right.
Process knobs on each tool plus it's like the sequence of them all.
And and so like you frankly cannot test everything, right? It's impossible.
It's it's too large of a search space.
Just like just like designing a chip is too large of a search space.
You have a hundred trillion transistors.
How are you going to possibly try every single thing? Impossible, right?
You just have to have enough intuition like pick that point, pick that point, pick that point, see the data. Oh, okay.
I think the answer is here, right?
And then just yellow, right?
Um, and and, you know, obviously once you know you think the answer is here, you test here.
You're like, "Okay, here."
But like a different person might have seen these three and then said, "Okay, the answer is actually here, not here."
And like the data is like fuzzy.
It's like somewhere in the center, but like, you know, it's like this this whole like idea of like ML research is you spend a lot of time on compute training doing what effectively were useless things besides teaching yourself what's the right thing to do and what's the wrong thing to do.
And and semiconductor manufacturing is the same way.
And actually all process manufacturing is the same way if you're iterating super fast and you're trying to get better and better and better or you're optimizing a process on a chemistry or whatever it is.
You you try, you fail, you learn, you do.
In semiconductor manufacturing maybe it's just running tens of thousands of wafers.
And so your R&D cost of, you know, an of an Intel and and or or your your cost of your, you know, main fab that is running the R&D is very very high and it's producing zero economic value besides that it's teaching you how to do the next node which then you can deploy at volume and that is like like what actually makes the money. [Music]