- The following is a conversation all
about the state-of-the-art in artificial intelligence, including some of the
exciting technical breakthroughs and developments in AI that happened
over the past year, and some of the interesting things we
think might happen this upcoming year.
0:15
At times, it does
get super technical, but we do try to make sure that
it remains accessible to folks outside the field without
ever dumbing it down.
0:23
It is a great honor and pleasure
to be able to do this kind of episode with two of my
favorite people in the AI community, Sebastian Raschka and Nathan Lambert.
0:38
They are both
widely respected machine learning researchers and engineers
who also happen to be great communicators, educators,
writers, and X posters.
0:49
Sebastian is the author of two books I highly recommend for beginners
and experts alike.
0:53
First is Build a Large Language Model from Scratch and Build a Reasoning
Model from Scratch.
1:01
I truly believe in the
machine learning world, the best way to learn and understand
something is to build it yourself from scratch.
1:13
Nathan is the post-training lead at
the Allen Institute for AI, author of the definitive book on
Reinforcement Learning from Human Feedback.
1:26
Both of them have great X
accounts, great Substacks.
1:30
Sebastian has courses on
YouTube, Nathan has a podcast.
1:34
And everyone should absolutely
follow all of those. those.
1:37
This is the Lex Fridman
podcast.
1:37
To support it, please check out our sponsors in the
description, where you can also find links to contact me, ask
questions, get feedback, and so on.
1:49
And now, dear friends, here's
Sebastian Raschka and Nathan Lambert.
1:57
So I think one useful lens to
look at all this through is the so-called DeepSeek
moment.
2:01
This happened about a year ago in January 2025,
when the open-weight Chinese company DeepSeek released
DeepSeek R1, that I think it's fair to say surprised
everyone with near-state-of-the-art performance, with allegedly much less
compute for much cheaper.
2:16
And from then to today, the AI competition
has gotten insane, both on the research and product
level.
2:28
It's just been accelerating.
2:32
discuss all of this today, and
maybe let's start with some spicy questions if we can.
2:38
Who's winning at the international
level?
2:38
Would you say it's the set of companies in China or the set
of companies in the United States?
2:46
And Sebastian, Nathan,
it's good to see you guys. guys.
2:50
So Sebastian, who
do you think is winning?
2:53
- Winning is a very broad term.
2:57
I would say you mentioned the DeepSeek
moment, and I think DeepSeek is winning the hearts of the people who work on
open-weight models because they share these as open models.
3:05
Winning,
I think, has multiple timescales to it.
3:09
We have today,
we have next year, we have in 10 years.
3:12
One thing I know
for sure is that I don't think nowadays, in 2026,
that there will be any company that has access to
technology that no other company has access to.
3:24
That
is mainly because researchers are frequently changing jobs and labs. They rotate.
3:32
I don't think there
will be a clear winner in terms of technology access.
3:36
However,
I do think there will be, The differentiating factor will be
budget and hardware constraints.
3:43
I don't think the ideas
will be proprietary, but rather the resources needed
to implement them.
3:46
I don't see currently a winner-take-all scenario. I can't see that. At the moment.
3:59
- Nathan, what do you think?
4:00
- You see the labs put different energy
into what they're trying to do, and I think to demarcate the point in time
when we're recording this, the hype over Anthropic's Claude
Opus 4.
4:08
5 model has been absolutely insane, which is just...
4:12
I mean, I've used it and built stuff in the last few weeks, and it's...
4:16
it's almost
gotten to the point where it feels like a bit of a meme in terms of the hype.
4:20
And it's kind of funny because this is very organic,
and then if we go back a few months ago, we can see the release date and
the notes, as Gemini 3 from Google got released, and it seemed like the marketing and just, like, wow
factor of that release was super high.
4:36
But then at the end of November,
Claude Opus 4.
4:36
5 was released and the hype has been growing, but Gemini 3
was before this.
4:40
And it kind of feels like people don't really talk about it as much, even
though when it came out, everybody was like, this is Gemini's moment to retake Google's structural advantages in AI.
4:52
And Gemini 3
is a fantastic model, and I still use it.
4:56
It's just kind of
differentiation is lower.
4:56
And I agree with Sebastian; what you're saying
with all these, the idea space is very fluid, but culturally
Anthropic is known for betting very hard on code, which is the Claude Code thing,
is working out for them right now.
5:08
So I think that even if the ideas flow
pretty freely, so much of this is bottlenecked by human effort and the
culture of organizations, where Anthropic seems to at least be presenting
as the least chaotic.
5:20
It's a bit of an advantage, if they can keep
doing that for a while.
5:24
But on the other side of things, there's a lot of
ominous technology from China where there's way more labs than
DeepSeek.
5:32
So DeepSeek kicked off a movement within China, I
say kind of similar to how ChatGPT kicked off a movement in the US
where everything had a chatbot.
5:39
There's now tons of tech companies in China that are
releasing very strong frontier open-weight models, to the point where I would say that
DeepSeek is kind of losing its crown as the preeminent open model maker
in China, and the likes of Z.
5:56
ai with their GLM
models, Minimax's models, Kimi Moonshot, especially in the
last few months, has shown more brightly.
6:04
The new DeepSeek models are
still very strong, but that's kind of a...
6:08
it could look back as a big
narrative point where in 2025 DeepSeek came and it provided this
platform for way more Chinese companies that are releasing these
fantastic models to kind of have this new type of operation.
6:19
So these models from these
Chinese companies are open-weights, and depending on this trajectory of business
models that these American companies are doing, they could be at risk.
6:27
But
currently, a lot of people are paying for AI software in the US, and
historically in China and other parts of the world, people
don't pay a lot for software.
6:37
- So some of these models like DeepSeek
have the love of the people because they are open-weight.
6:41
How long do
you think the Chinese companies keep releasing open-weight models?
6:47
- I would say for a few years.
6:47
I think
that, like in the US, there's not a clear business model for it.
6:51
I have been
writing about open models for a while, and these Chinese companies have realized
it.
6:55
So I get inbound from some of them.
6:59
And they're smart and realize the same
constraints: a lot of top US tech companies and other IT companies
won't pay for an API subscription to Chinese companies for security
concerns.
7:07
This has been a long-standing habit in tech, and the people at
these companies then see open weight models as an ability to influence
and take part of a huge growing AI expenditure market in the US.
7:18
And
they're very realistic about this, and it's working for them.
7:23
I think that
the government will see that that is building a lot of influence internationally
in terms of uptake of the technology, so there's going to be a lot of
incentives to keep it going.
7:31
But building these models and doing the research is
very expensive, so at some point, I expect consolidation.
7:39
But I don't expect that to
be a story of 2026, where there will be more open model builders throughout
2026 than there were in 2025.
7:44
And a lot of the notable ones will be in China.
7:50
- You were going to say something? - Yes.
7:51
You mentioned DeepSeek losing its
crown.
7:51
I do think to some extent, yes, but we also have to consider though, they are
still, I would say, slightly ahead.
7:58
And the other ones—it's not that DeepSeek
got worse, it's just that the other ones are using the ideas from DeepSeek.
8:06
For example, you mentioned Kimi—same architecture, they're training it.
8:10
And
then again, we have this leapfrogging where they might be at some point in time a
bit better because they have the more recent model.
8:17
And I think this comes back
to the fact that there won't be a clear winner.
8:21
It will just be
like that: one person releases something, the other one comes in, and the
most recent model is probably always the best model. - Yeah.
8:30
We'll also see the Chinese companies
have different incentives.
8:30
Like, DeepSeek is very secretive,
whereas some of these startups are like the MiniMaxs and Z. ais of the
world.
8:37
Those two literally have filed IPO paperwork, and they're
trying to get Western mindshare and do a lot of outreach there.
8:45
So I
don't know if these incentives will change the model development, because DeepSeek
famously is built by a hedge fund, Highflyer Capital, and we don't
know exactly what they use the models for or if they care about this.
8:59
- They're secretive in terms of communication; they're
not secretive in terms of the technical reports that describe how their models work.
9:03
They're
still open on that front.
9:03
And we should also say, on the Claude Opus 4.
9:06
5 hype,
there's the layer of something being the darling of the
X echo chamber, on the Twitter echo chamber, and the actual
amount of people that are using the model.
9:22
I think it's probably
fair to say that ChatGPT and Gemini are focused on the
broad user base that just want to solve problems in their
daily lives, and that user base is gigantic.
9:33
So the hype
about the coding may not be representative of the actual use.
9:38
- I would say also a lot of
the usage patterns are, like you said, name recognition,
brand and stuff, but also muscle memory almost, where, you
know, ChatGPT has been around for a long time.
9:50
People just got used to
using it, and it's almost like a flywheel: they recommend it to other users and that
stuff.
9:54
One interesting point is also the customization of LLMs.
9:58
For example, ChatGPT has a memory feature, right?
10:02
And so you
may have a subscription and you use it for personal stuff, but I don't know
if you want to use that same thing at work.
10:10
Because it's a boundary between private and work.
10:10
If you're working at a company, they might not allow that or you may not want that.
10:14
And
I think that's also an interesting point where you might have multiple
subscriptions. One is just clean code.
10:22
It has nothing of your
personal images or hobby projects in there.
10:26
It's just like the work thing.
10:26
And then the other one is your personal thing.
10:30
So I think that's also something where there are
two different use cases, and it doesn't mean you only have to have one.
10:34
I think
the future is also multiple ones.
10:38
- What model do you think won 2025, and what
model do you think is going to win '26?
10:43
- I think in the context of consumer chatbots,
it's a question of: are you willing to bet on Gemini over ChatGPT?
10:50
Which I would say, in my gut,
feels like a bit of a risky bet because OpenAI has been the incumbent, and
there are so many benefits to that in tech.
10:58
I think the momentum, if you look at 2025, was on
Gemini's side, but they were starting from such a low point.
11:05
And RIP Bard and these
earlier attempts at getting started.
11:13
Huge credit to them for powering through the
organizational chaos to make that happen.
11:17
But also it's hard to bet against
OpenAI because they always come off as so chaotic, but they're very good
at landing things.
11:22
And I think, personally, I have very mixed
reviews of GPT-5, but it must have saved them so much money with the
high-line feature being a router where most users are no longer charging
their GPU costs as much.
11:38
So I think it's very hard to dissociate the things that I like out of models
versus the things that are going to actually be a general
public differentiator.
11:50
- What do you think about
2026? Who's going to win?
11:52
- I'll say something, even though it's risky.
11:52
I think
Gemini will continue to make progress on ChatGPT.
11:56
I think Google's scale,
when both of these are operating at such extreme
scales—and Google has the ability to separate research and product
a bit better, whereas you hear so much about OpenAI being chaotic operationally
and chasing the high-impact thing, which is a very startup culture.
12:11
And then
on the software and enterprise side, I think Anthropic will have continued success,
as they've again and again been set up for that.
12:19
And obviously Google Cloud
has a lot of offerings, but I think this kind of Gemini name
brand is important for them to build.
12:27
Google Cloud will continue to do well, but that's a more complex
thing to explain in the ecosystem, because that's competing with
the likes of Azure and AWS rather than on the model provider side.
12:40
- So in infrastructure, you think
TPU is giving an advantage?
12:45
- Largely because the margin on
NVIDIA chips is insane, and Google can develop everything from top to
bottom to fit their stack and not have to pay this margin.
12:53
And they've
had a head start in building data centers.
12:57
So all of these things that have
both high lead times and very hard margins on high costs, Google has a just kind
of historical advantage there.
13:05
And if there's going to be a new paradigm,
it's most likely to come from OpenAI where their research division
again and again has shown this ability to land a new research
idea or a product.
13:13
Like Deep Research, Sora, o1
thinking models—all these definitional things have come from
OpenAI, and that's got to be one of their top traits as an organization.
13:25
So
it's kind of hard to bet against that, but I think a lot of this
year will be about scale and optimizing what could be described
as low-hanging fruit in models.
13:37
- And clearly there's a trade-off
between intelligence and speed.
13:40
This is what ChatGPT-5 was trying to solve behind the scenes.
13:44
It's like, do people actually want intelligence, the broad
public, or do they want speed?
13:52
- I think it's a nice variety, or
the option to have a toggle there.
13:56
I mean, for my personal usage,
most of the time when I look something up, I use ChatGPT to ask a quick
question, get the information I wanted fast.
14:03
For most daily
tasks, I use the quick model.
14:07
Nowadays, I think the auto mode is pretty
good where you don't have to specifically say thinking or non-thinking.
14:11
Then again, I also sometimes want the pro mode.
14:14
Very often
what I do is, when I have something written, I put it into
ChatGPT and say, "Hey, do a very thorough check.
14:22
Are all my
references correct?
14:22
Are all my thoughts correct?
14:25
Did I make any formatting
mistakes and are the figure numbers wrong?" Or something like
that.
14:29
And I don't need that right away.
14:33
I finish my stuff, maybe
have dinner, let it run, come back and go through this.
14:37
I think
this is where it's important to have this option.
14:41
I would go crazy if for
each query I would have to wait 30 minutes or 10 minutes even. - That's me.
14:48
I'm sitting over here losing my mind that you
use the router and the non-thinking model.
14:52
I'm like, "How do you live with that?" That's like my reaction.
14:55
I've been
heavily on ChatGPT for a while.
15:01
I never touched ChatGPT-5
non-thinking.
15:01
I find its tone and then its propensity for errors—it has a higher
likelihood of errors.
15:05
Some of this is from back when OpenAI released
o3, which was the first model to do this deep search and find many
sources and integrate them for you.
15:17
I became habituated with that.
15:17
So
I will only use GPT-5.
15:17
2 Thinking or Pro when I'm finding any sort
of information query for work, whether that's a paper or
some code reference that I found.
15:28
And I will regularly
have like five Pro queries going simultaneously, each
looking for one specific paper or feedback on an equation or something.
15:38
- I have a fun example where I needed
the answer as fast as possible for this podcast before
I was going on the trip.
15:46
like a local GPU running at home
and I wanted to run a long RL experiment.
15:49
And usually I also unplug
things because you never know if you're not at home, you don't want things
plugged in.
15:53
And I accidentally unplugged the GPU.
15:57
My wife was already in the
car and it's like, "Oh dang."
16:01
Then basically I wanted as
fast as possible a Bash script that runs my different
experiments and the evaluation.
16:05
And it's something I know, I
learned how to use the Bash interface or Bash terminal, but
in that moment I just needed like 10 seconds, give me the command.
16:18
- This is a hilarious situation
but yeah, so what did you use?
16:21
- So I did the non-thinking fastest
model.
16:21
It gave me the Bash command to chain different
scripts to each other and then the thing is like you have the
tee thing where you want to route this to a log file.
16:33
Top of my head I was just
like in a hurry, I could have thought about it myself.
16:37
- By the way I don't know if there's a
representative case, wife waiting in the car- ...
16:40
you have to run, you know, unplug the GPU.
16:40
You
have to generate a Bash script.
16:40
This sounds like a movie, like- Mission Impossible. - I use Gemini for that.
16:46
So I use thinking for
all the information stuff and then Gemini for fast things or stuff that I could sometimes
Google, which is like it's good at explaining things and I trust that it has
this kind of background of knowledge and it's simple.
16:58
And the Gemini app
has gotten a lot better and- It's good for those sorts of things.
17:01
And
then for code and any sort of philosophical discussion, I use Claude
Opus 4. 5.
17:05
Also always with extended thinking.
17:08
Extended thinking and
inference time scaling is just a way to make the models marginally smarter.
17:12
And I will always err on that side when the progress is very high
because you don't know when that'll unlock a new use case.
17:20
And then sometimes
use Grok for real-time information or finding something on AI
Twitter that I knew I saw and I need to dig up and I just
fixated on.
17:28
Although when Grok 4 came out, the Grok
4 SuperGrok Heavy, which was like their pro variant was actually very good
and I was pretty impressed with it, and then it just kind of like muscle memory lost
track of it with having the ChatGPT app open.
17:43
So I use many different things. - Yeah.
17:45
I actually do use Grok 4 Heavy for debugging.
17:49
For like hardcore
debugging that the other ones can't solve, I find that it's the
best at. And...
17:53
it's interesting 'cause you say ChatGPT is
the best interface.
17:57
For me, for that same reason, but
this could be just momentum- Gemini is the better interface for me.
18:04
I think because I fell in love with their best needle in the
haystack.
18:09
If I ever put something that has a lot of context but I'm
looking for very specific kinds of information to make sure it tracks all of
it, I find at least that Gemini for me has been the best.
18:21
So it's funny with some of these models,
if they win your heart over- for one particular feature
on one particular day, for that particular query, that
prompt, you're like, "This model's better."
18:35
And so you'll
just stick with it for a bit until it does something really dumb.
18:40
There's
like a threshold effect.
18:40
Some smart thing and then you fall in love with it and then it
does some dumb thing and you're like, "You know what?
18:47
I'm gonna switch and try Claude or
ChatGPT."
18:47
And all that kind of stuff.
18:51
- This is exactly it: you use it until it
breaks, until you have a problem, and then you change the LLM.
18:55
And I think it's the same as how we use anything,
like our favorite text editor, operating systems, or the browser.
19:02
I mean, there are many options: Safari, Firefox, Chrome.
19:06
They're
relatively similar, but then there are edge cases, extensions
you want, and then you switch.
19:14
But I don't think anyone types
the same thing into different browsers and compares them.
19:19
You
only do that when something breaks. So that's a good point.
19:23
You use it until
it breaks, then you explore other options.
19:28
- On the long context thing, I was
also a Gemini user, but the GPT-5.
19:28
2 release blog had crazy long context
scores.
19:32
People were like, "Did they just figure out some algorithmic
change?"
19:36
It went from 30% to 70% in this minor model
update.
19:40
It's very hard to keep track of all of these
things, but now I look more favorably at GPT-5. 2's long
context.
19:47
So it's just like, "How do I actually get to testing
this?"
19:51
It's a never-ending battle.
19:57
- Well, it's interesting that none of
us talked about the Chinese models from a usage perspective. What
does that say?
20:01
Does it mean the Chinese models are not as good, or are
we just very biased and US-focused?
20:11
- I think currently there's a discrepancy
between the model and the platform.
20:15
The open models are more known for the
open weights, not the platform yet.
20:19
known for the open weights,
not their platform yet.
20:21
- Many companies will sell you open-model
inference at a very low cost.
20:25
With OpenRouter, it's easy to
look at multi-model things.
20:29
You can run DeepSeek on Perplexity.
20:29
Sitting here, we're like, "We use OpenAI GPT-5 Pro consistently."
20:33
We're all willing to pay for the marginal intelligence gain.
20:39
These
models from the US are better in terms of the outputs.
20:43
I think the question is, will they stay better for this
year and for years to come?
20:51
As long as they're better,
I'm gonna pay for them.
20:55
There's also analysis showing that the way the Chinese models are served—you could
argue this is due to export controls— is that they use fewer GPUs per
replica, which makes them slower and have different errors.
21:06
If speed and
intelligence are in your favor as a user, in the US, a lot of users will go for
this.
21:10
And I think that will spur these Chinese companies to want
to compete in other ways, whether it's free or substantially
lower costs, or it'll breed creativity in terms of offerings,
which is good for the ecosystem.
21:26
But the simple thing is: the US models
are currently better, and we use them.
21:30
I tried these other open models, and
I'm like, "Fun, but I don't go back."
21:34
models, and I'm like, "Fun, but not
gonna... I don't go back to it."
21:38
- We didn't really mention
programming.
21:38
That's another use case that a lot of people deeply
care about.
21:42
I use basically half-and-half Cursor and Claude
Code, because they're...
21:46
I fundamentally different
experiences and both are useful. What do you guys...
21:54
You
program quite a bit, so what do you use? What's the current vibe?
21:59
- So, I use the Codeium plugin for
VS Code.
21:59
You know, it's very convenient.
22:03
It's just like a plugin, and then
it's a chat interface that has access to your repository.
22:06
I know that Claude Code
is, I think, a bit different.
22:06
It is a bit more agentic. It touches more things.
22:10
It does the whole project for you.
22:10
I'm not quite there yet where I'm comfortable
with that because maybe I'm a control freak, but I still would like
to see a bit what's going on.
22:18
And Codeium is kind of, right now, for
me, the sweet spot where it is helping me, but it is not
taking completely over.
22:29
- I should mention, one of the reasons
I do use Claude Code is to build the skill of programming with English.
22:33
I
mean, the experience is fundamentally different. You're...
22:37
As
opposed to micromanaging the details of the process of the
generation of the code, and looking at the diff, which you can
in Cursor if that's the IDE you use, and in changing, altering.
22:52
Looking and reading the code and
understanding the code deeply as you progress, versus just thinking in
this design space and just guiding it at this macro level, which I think is another way of thinking
about the programming process.
23:10
Also, we should say that
Claude Code just seems to be somehow a better utilization
of Claude Opus 4. 5.
23:18
- It's a good side-by-side for people to do.
23:18
You
can have Claude Code open, you can have Cursor open, you can have VS Code open, and you
can select the same models on all of them— ...
23:26
and ask questions, and it's very
interesting.
23:26
Claude Code is way better in that domain. It's remarkable.
23:32
- All right, we should say that both
of you are legit on multiple fronts: researchers, programmers,
educators, Tweeters.
23:43
And on the book front,
too.
23:43
So Nathan, at some point soon, hopefully has
an RLHF book coming out.
23:50
- It's available for preorder, and
there's a full digital preprint.
23:50
I'm just making it pretty and better organized for
the physical thing, which is a lot of why I do it, because it's fun to create things
that you think are excellent in the physical form when so much
of our life is digital.
24:05
- I should say, going to Perplexity here,
Sebastian Raschka is a machine learning researcher and author known for several
influential books.
24:08
A couple of them that I wanted to mention—which is a book
I highly recommend—Build a Large Language Model from Scratch,
and the new one, Build a Reasoning Model from Scratch.
24:20
So, I'm really excited about that.
24:23
Building stuff from scratch is one
of the most powerful ways of learning.
24:27
- Honestly, building an LLM from scratch is
a lot of fun.
24:27
It's also a lot to learn.
24:31
And like you said, it's probably the best
way to learn how something really works, 'cause you can look at figures, but
figures can have mistakes.
24:35
You can look at concepts and explanations, but
you might misunderstand them.
24:43
But if there is code, and the code works, you know it's correct.
24:47
I mean,
there's no misunderstanding. It's precise.
24:50
Otherwise, it wouldn't work.
24:50
And I
think that's the beauty behind coding. It doesn't lie. It's math, basically.
24:54
So, even though with math, I think you can have
mistakes in a book you would never notice.
25:02
Because you are not running the math when you
are reading the book, you can't verify this.
25:06
And with code, what's nice
is you can verify it.
25:09
- Yeah, I agree with you about the Build
an LLM from Scratch book.
25:09
It's nice to tune out everything else, the internet
and so on, and just focus on the book.
25:16
But, you know, I read several
history books.
25:16
It's just less lonely somehow. It's really more
fun.
25:23
Like for example, on the programming front, I think it's genuinely
more fun to program with an LLM.
25:31
And I think it's genuinely
more fun to read with an LLM. But you're right.
25:36
That
distraction should be minimized.
25:40
So you use the LLM to basically enrich
the experience, maybe add more context.
25:48
I just find the rate of aha moments
for me is really high with LLMs. - 100%.
25:54
I also want to correct myself:
I'm not suggesting not to use LLMs.
25:58
I suggest doing it in
multiple passes.
25:58
Like, one pass just offline, focus mode, and
then after that...
26:02
I mean, I also take notes, but I, I try
to resist the urge to immediately look things up. I do
a second pass.
26:10
It's just more structured this way.
26:14
Sometimes
things are answered in the chapter, but sometimes also it just
helps to let it sink in and think about it.
26:21
Other people have different
preferences.
26:21
I highly recommend using LLMs when reading books.
26:25
For me, it's not the
first thing to do; it's the second pass.
26:29
- My recommendation is the opposite.
26:29
I
like to use the LLM at the beginning to lay out the full context of
what is this world that I'm now stepping into?
26:39
But I try
to avoid clicking out of the LLM into the world of Twitter and blogs, because then you're down
this rabbit hole.
26:47
You're reading somebody's opinion.
26:50
There's a flame war
about a particular topic and all of a sudden you're in the realm
of the internet and Reddit and so on.
26:58
But if you're
purely letting the LLM give you the context of why this
matters, what are the big picture ideas...
27:06
sometimes books are good
at doing that, but not always.
27:12
- This is why I like the ChatGPT app,
because it gives the AI a home on your computer where you can focus on it,
rather than just being another tab in my mess of internet options.
27:19
And I think Claude Code does a good job of making
that a joy, where it seems very engaging as a product
design to be an interface that your AI will then go out into
the world.
27:31
It's something that is intangible between it and
Codex; it just feels warm and engaging, where Codex can often
be as good from OpenAI, but it just, feel a little bit rough around
the edges.
27:43
Whereas Claude Code makes it fun to build things
from scratch, where you just trust that it'll make something.
27:51
Obviously this is good for websites and kind of refreshing
tooling and stuff like this, which I use it for,
or data analysis.
27:58
For my On my blog, we scrape Hugging Face so we keep
download numbers for every dataset and model.
28:06
over time, so we have them.
28:06
And Claude was
just like, "Yeah, I've made use of that data, no problem."
28:10
And I was like,
"That would've taken me days."
28:13
And then I have enough situational awareness to be
like, "Okay, these trends obviously make sense." You can check things.
28:17
But that's just a
wonderful interface where you can have an intermediary and not have to do
the kind of awful low-level work that you would have to do to
maintain different web projects. - All right.
28:29
So we just
talked about a bunch of the closed-weight models.
28:32
Let's talk about the open ones.
28:36
Tell me about
the landscape of open LLM models. Which are interesting?
28:39
Which stand out to you and why?
28:39
We already mentioned DeepSeek R1.
28:44
- Do you wanna see how many we can
name off the top of our head?
28:47
- Yeah, without looking at notes.
28:48
- DeepSeek, Kimi, MiniMax, Z. ai,
Moonshot.
28:48
We're just going Chinese.
28:57
- Let's throw in Mistral AI, Gemma... ...
29:01
gpt-oss, the open weight
model by OpenAI.
29:01
Actually, NVIDIA had a really cool
one, Nemotron 3.
29:05
There, there's a lot of stuff especially at the
end of the year.
29:09
Qwen might be the one— - Oh, yeah.
29:12
Qwen was the
obvious name I was gonna say.
29:15
You can get at least 10 Chinese
and at least 10 Western.
29:15
I think that OpenAI released
their first open model— ... since GPT-2.
29:22
When I was writing about OpenAI's open model release, they were like,
"Don't forget about GPT-2," which I thought was really funny 'cause it's just such a
different time.
29:29
But gpt-oss-120b is actually a very strong model and
does some things that other models don't do very well.
29:38
Selfishly, I'll promote
a bunch of Western companies in the US and Europe that have these fully open models.
29:44
I work at the Allen Institute for AI, where we've been building OLMo, which
releases data and code.
29:48
And now we have actual competition for people
that are trying to release everything so that others can train these models.
29:56
There's the Institute for Foundation Models/LM360, which has had their K2 models of various types.
30:03
Apertus is a Swiss research consortium.
30:07
Hugging Face has SmolLM, which is very popular.
30:11
And NVIDIA's Nemotron
3 has started releasing data as well.
30:14
And then Stanford's Martini
Community Project, which is kind of making it so there's a pipeline for people
to open a GitHub issue and implement a new idea and then have it run in a
stable language modeling stack.
30:26
This space, that list
was way smaller in 2024— ...
30:31
so I think it was just AI2.
30:31
So it's
a great thing for more people to get involved and to understand language
models, which doesn't really have a Chinese analog.
30:38
While I'm talking, I'll say that the Chinese open
language models tend to be much bigger, and that gives them higher
peak performance as MoEs, where a lot of these things that we like a
lot, whether it was Gemma and Nemotron, have tended to be smaller
models from the US, which is starting to change from the US and Europe.
30:58
Mistral Large 3 came out, which was a giant MoE model, very similar
to DeepSeek architecture in December.
31:05
And then a startup, RCAI, and
both Nemotron and NVIDIA have teased MoE models way bigger than
100 billion parameters- like this 400 billion parameter
range coming in this Q1 2026 timeline.
31:19
So I think this
kind of balance is set to change this year in terms of what people are
using the Chinese versus US open models for, which I'm personally going
to be very excited to watch.
31:32
- First of all, huge props for
being able to name so many of these.
31:36
Did you actually name LLaMA? - No. - I feel like ... - RIP.
31:41
- This was not on purpose. - RIP LLaMA. All right.
31:45
Can you mention some interesting
models that stand out?
31:45
You mentioned Qwen 3 is obviously a standout.
31:51
- So I would say the year's almost
bookended by both DeepSeek V3 and R1.
31:55
And then on the other
hand, in December, DeepSeek-V3. 2.
31:59
Because what I like about those is
they always have an interesting architecture tweak that others don't have.
32:03
But
otherwise, if you want to go with the familiar but really good
performance, Qwen 3 and, like Nathan said, also gpt-oss-120b.
32:11
And I think what's interesting about it is it's kind
of like the first public or open weight model that was really trained
with tool use in mind, which I do think is kind of a paradigm shift
where the ecosystem was not quite ready for it.
32:27
By tool use, I mean
that the LLM is able to do a web search or to call a Python interpreter.
32:33
And I do think it's a
standout because it's a huge unlock.
32:36
Because one of the
most common complaints about LLMs are, for example,
hallucinations, right?
32:43
And so, in my opinion, one of the best
ways to solve hallucinations is to not try to always remember
information or make things up.
32:51
For math, why not use a
calculator app or Python?
32:54
If I ask the LLM, "Who won
the soccer World Cup in 1998?"
32:58
instead of just trying
to memorize, it could go do a search.
33:02
I think mostly
it's still a Google search.
33:06
So ChatGPT and gpt-oss-120b,
they would do a tool call to Google, maybe find the FIFA
website.
33:09
Find, okay, it was France.
33:12
It would get you that information
reliably instead of just trying to memorize it.
33:16
So I think it's
a huge unlock which right now is not fully utilized yet by
the open-source, open-weight ecosystem.
33:24
A lot of people
don't use tool call modes because I think, first, it's a trust thing.
33:27
You don't want to run this on your computer where it has access to tools, could wipe your
hard drive or whatever.
33:31
So you want to maybe containerize that.
33:35
But
I do think that is like a really important step for the
upcoming years to have this ability. - So a few quick things.
33:44
First of all, thank you for defining what you mean by tool use.
33:47
I
think that's a great thing to do in general for the concepts we're talking
about.
33:51
Even things as sort of well-established as MoEs.
33:57
You have to say that means mixture
of experts, and you kind of have to build up an intuition for people what
that means, how it's actually utilized, what are the different flavors.
34:05
So what
does it mean that there's just such an explosion of open models? What's your intuition?
34:13
- If you're releasing an open model, you want people
to use it, is the first and foremost thing.
34:17
And then after that comes things like
transparency and trust.
34:17
I think when you look at China, the biggest reason
is that they want people around the world to use these models, and I think a
lot of people will not.
34:25
If you look outside of the US, a lot of people will not pay for software, but they
might have computing resources where you can put a model on it and run it.
34:33
I think there can also be
data that you don't want to send to the cloud.
34:37
So the number one thing is
getting people to use models, use AI, or use your AI that might not be
able to do it without having access to the model.
34:46
- I guess we should state explicitly, so we've
been talking about these Chinese models and open weight models.
34:50
Oftentimes, the way they're run is locally.
34:54
So it's not like you're sending
your data to China or to whoever developed Silicon Valley, or whoever
developed the model.
35:04
- A lot of American startups
make money by hosting- ...
35:07
these models from China and
selling them.
35:07
It's called selling tokens, which means somebody will call
the model to do some piece of work.
35:15
I think the other reason is for US
companies like OpenAI.
35:15
They are so GPU deprived.
35:19
They're at the limits of the
GPUs.
35:19
Whenever they make a release, they're always talking about
like, "Our GPUs are hurting."
35:23
And I think during one of these
gpt-oss-120b release sessions, Sam Altman said, "Oh, we're releasing this because
we can use your GPUs.
35:31
We don't have to use our GPUs, and OpenAI can
still get distribution out of this," which is another very real thing,
because it doesn't cost them anything.
35:43
- And for the user, I think also,
there are users who just use the model locally how they would use ChatGPT.
35:47
But also for companies I think it's a huge unlock to have these models because you can
customize them, you can train them, you can add post-training,
add more data.
35:55
Like, specialize them into, let's say,
law, medical models, whatever you have.
36:02
And the appeal, you mentioned
Llama, the appeal of the open-weight models from China is that the
open-weight models' licenses are even friendlier.
36:10
I think they are
just unrestricted open source licenses where if we use something like Llama
or Gemma, there are some strings attached.
36:17
I think it's like an upper limit in
terms of how many users you have.
36:17
And then if you exceed, I don't know, so and so many
million users, you have to report your financial situation to, let's
say, Meta or something like that.
36:28
And I think while it is a
free model, there are strings attached, and people do like things
where strings are not attached.
36:32
So I think that's also one of the
reasons, besides performance, why the open-weight models from China are so
popular, because you can just use them.
36:43
There's no catch in that sense.
36:46
- The ecosystem has gotten better on that
front, but mostly downstream of these new providers providing such open licenses.
36:50
That was
funny when you pulled up Perplexity and said, "Kimi K2 Thinking hosted in the US."
36:53
Which is
just like an exact...
36:53
I've never seen this, but it's an exact example of what we're talking
about where people are sensitive to this.
37:01
But Kimi K2 Thinking and Kimi K2 is
a model that is very popular.
37:01
People say that has very good creative writing
and also in doing some software things.
37:09
So it's just these little quirks that
people pick up on with different models that they like.
37:14
- What are some interesting ideas
that some of these models have explored that you can speak to, that
are particularly interesting to you?
37:21
- Maybe we can go chronologically.
37:21
I mean,
there was, of course, DeepSeek.
37:21
DeepSeek R1 that came out in January of 2025,
if we just focus on 2025.
37:25
However, this was based on DeepSeek-V3,
which came out the year before in December 2024.
37:33
There are
multiple things on the architecture side.
37:37
What is fascinating is...
37:37
I mean, that's what I do with my from-scratch coding projects.
37:41
You can
still start with GPT-2, and you can add things to that model to make it
into this other model.
37:44
So it's all still kind of like the same
lineage.
37:48
It is a very close relationship between those.
37:52
But
top of my head, DeepSeek—what was unique there is the Mixture of Experts.
37:55
Not that they were inventing Mixture of Experts—we can maybe talk a bit
more about what Mixture of Experts means—but just to list these
things first before we dive into detail.
38:06
Mixture of Experts,
but then they also had Multi-head Latent Attention, which is a tweak
to the attention mechanism, where this was, I would say,
the main distinguishing factor between these
open-weight models.
38:19
Different tweaks to make inference or KV
cache size...
38:23
We can also define KV cache in a few moments,
but to kind of make it more economical to have long context, to
shrink the KV cache size.
38:31
So what are tweaks that we can do?
38:35
And most of them
focused on the attention mechanism.
38:35
There is Multi-head Latent Attention
in DeepSeek.
38:38
There is Group Query Attention, which is still very popular.
38:43
It's not invented by any of those models.
38:46
It goes back a few years.
38:46
But
that would be the other option.
38:50
Sliding window attention—I think
OLMo 3 uses it, if I remember correctly.
38:53
So there are these
different tweaks that make the models different.
38:57
Otherwise, I put
them all together in an article once where I just
compared them.
39:01
They are very, surprisingly similar.
39:04
It's just different
numbers in terms of how many repetitions of the transformer block you
have in the center.
39:08
And, like, just little knobs that people tune.
39:12
But what's so nice about it is it works no matter what. You can
tweak things.
39:16
You can move the normalization layers around to get some
performance gains.
39:20
And OLMo is always very good in ablation studies,
showing what it actually does to the model if you move something around.
39:28
Ablation studies: does it make it better or worse?
39:32
But there are so many, let's say, ways you
can implement a transformer and make it still work.
39:36
The big ideas that are
still prevalent is Mixture of Experts, multi-head latent attention,
sliding window attention, group query attention.
39:43
And then at the
end of the year, we saw a focus on making the attention mechanism
scale linearly with inference token prediction.
39:51
So there was
Qwen2-VL, for example, which added a gated delta
net.
39:55
It's kind of inspired by State space models, where you have a fixed
state that you keep updating.
39:59
But it makes essentially this attention cheaper, or it replaces attention
with a cheaper operation.
40:08
- And it may be useful to step
back and talk about transformer architecture in general.
40:13
- Yeah, so maybe we should start
with the GPT-2 architecture.
40:13
The transformer that was derived from the
"Attention Is All You Need" paper.
40:21
The "Attention Is All You Need" paper had
a transformer architecture that had two parts, an encoder and a
decoder.
40:25
And GPT went just focusing in on the decoder
part.
40:30
It is essentially still a neural network and it has
this attention mechanism inside.
40:37
And you predict
one token at a time.
40:37
You pass it through an embedding layer.
40:41
There's
the transformer block.
40:41
The transformer block has attention modules and a fully
connected layer.
40:45
And there are some normalization layers in between.
40:49
But it's essentially
neural network layers with this attention mechanism.
40:52
So coming from
GPT-2 when we move on to gpt-oss-120b, there is, for
example, the Mixture of Experts layer.
41:00
It's not invented by
gpt-oss-120b. It's a few years old.
41:04
But it is essentially a tweak to make
the model larger without consuming more compute in each forward pass.
41:11
So
there is this fully connected layer, and if listeners are
familiar with multi-layer perceptrons, you can think of a
mini multi-layer perceptron, a fully connected neural network layer
inside the transformer.
41:22
And it's very expensive, because it's fully connected.
41:26
If you
have a thousand inputs and a thousand outputs, that's like one million connections.
41:30
And it's a very expensive part in this transformer.
41:34
And the idea
is to kind of expand that into multiple feedforward
networks.
41:37
So instead of having one, let's say you have 256, but it
would make it way more expensive, because now you have 256, but you don't
use all of them at the same time.
41:45
So you now have a router that says, "Okay,
based on this input token, it would be useful to use this fully
connected network."
41:53
And in that context, it's called an expert.
41:57
So a
Mixture of Experts means you have multiple experts.
42:01
And depending on what
your input is, let's say it's more math-heavy, it would use different
experts, compared to, let's say, translating input text from English to
Spanish.
42:09
It would maybe consult different experts.
42:12
It's not quite clear, I mean,
not as clear-cut to say, "Okay, this is only an expert for math and
for Spanish." It's a bit more fuzzy.
42:20
But the idea is
essentially that you pack more knowledge into the network, but not all
the knowledge is used all the time.
42:27
That would be very wasteful.
42:27
So, during the token generation, you are more selective.
42:31
There's a router that selects which tokens should go to which expert. It adds more complexity. It's harder to train.
42:38
There's a lot that
can go wrong, like collapse and everything.
42:42
So I think that's why OLMo
3 still uses dense...
42:42
I mean, you have OLMo models with Mixture of
Experts, but dense models, where dense means... So also, it's
jargon.
42:50
There's a distinction between dense and sparse.
42:54
So Mixture
of Experts is considered sparse, because we have a lot of experts, but only
a few of them are active. So that's called sparse.
43:01
And then dense would be the
opposite, where you only have one fully connected module, and
it's always utilized.
43:08
- So maybe this is a good place
to also talk about KV cache.
43:11
But actually, before that,
even zooming out, like fundamentally, how many new ideas have
been implemented from GPT-2 to today?
43:22
Like, how different really
are these architectures?
43:25
- Take the Mixture of Experts.
43:25
The attention mechanism in gpt-oss-120b, that would be the Group
Query Attention mechanism.
43:28
So it's a slight tweak from Multi-Head Attention to
Group Query Attention. So that we have too...
43:36
I think they replaced LayerNorm by RMSNorm, but it's just like a different
normalization there and not a big change. It's just like a tweak.
43:43
The
nonlinear activation function— people familiar with deep neural
networks, I mean, it's the same as changing sigmoid with ReLU.
43:51
It's not
changing the network fundamentally.
43:55
It's just a little tweak.
43:55
And that's about it, I would say.
43:58
It's not really fundamentally
that different.
43:58
It's still the same architecture.
44:02
So you can go from one into the other by just
adding these changes basically.
44:09
- It fundamentally is still
the same architecture. - Yep.
44:12
For example, you mentioned my book
earlier.
44:12
That's a GPT-2 model in the book because it's simple
and it's very small, so 124 million parameters approximately.
44:19
But
in the bonus materials, I do have OLMo from scratch, Gemini 3 from
scratch, and other types of from-scratch models.
44:28
And I always start it with my
GPT-2 model and just tweak the—well, add different components and you get
from one to the other.
44:31
It's kind of like a lineage in a sense.
44:37
- Can you build up an intuition for
people?
44:37
Because when you zoom out, you look at it, there's so much
rapid advancement in the AI world.
44:46
And at the same time, fundamentally
the architectures have not changed.
44:51
So where is all the turbulence, the
turmoil of the advancement happening?
44:58
Where are the gains to be had?
45:01
- So there are different stages where
you develop the network or train the network.
45:04
You have the pre-training.
45:04
Now back in the day, it was just pre-training with GPT-2.
45:08
Now you
have pre-training, mid-training, and post-training.
45:12
So I think
right now we are in the post-training focus stage.
45:16
Pre-training still gives you advantages if you scale it up with
better, higher quality data.
45:20
But then we have capability unlocks that
were not there with GPT-2, for For example, ChatGPT is basically a GPT-3 model.
45:32
And GPT-3 is the same as
GPT-2 in terms of architecture.
45:36
What was new was adding
supervised fine-tuning and reinforcement learning with human feedback.
45:40
So it's more on the algorithmic side than the architecture.
45:44
- I would say that the systems also change a
lot.
45:44
If you listen to NVIDIA's announcements, they talk about things like, "You
now do FP8, you can now do FP4."
45:53
What's happening is these labs are figuring
out how to utilize more compute to put it into one model, which lets
them train faster and put more data in.
46:00
And then you can find better
configurations faster by doing this.
46:04
So you can look at, essentially,
tokens per second per GPU as a metric that you look at when
you're doing large-scale training.
46:12
You can go from 10k to 13k
by turning on FP8 training, which means you're using less
memory per parameter in the model.
46:20
By saving less information,
you do less communication and train faster.
46:23
So all of these
system things underpin way faster experimentation
on data and algorithms.
46:35
It's a loop that keeps going where it's hard to
describe when you look at architectures and they're exactly the same, but the code base used
to train these models is vastly different- -and you could probably...
46:44
the GPUs
are different but you probably train gpt-oss-20b way faster in wall-clock
time than GPT-2 was trained at the time. - Yeah.
46:54
Like you said, they had, for
example, in Mixture of Experts this FP4 optimization where you
get more throughput.
46:57
But I do think, for speed this is true, but it doesn't give the model new capabilities.
47:05
It's just: how much can we make the computation coarser
without suffering in terms of model performance degradation?
47:13
But I do think- I mean, there are alternatives popping up to
the transformer.
47:17
Text diffusion models, a completely different paradigm. And
there is also...
47:21
I mean, although text diffusion models might use
transformer architectures, it's not an autoregressive transformer. And also Mamba models.
47:32
It's a state space model.
47:32
But
they do have trade-offs, and nothing has yet replaced
the autoregressive transformer as the state-of-the-art model.
47:40
For state-of-the-art, you would still go with that, but there are now
alternatives for the cheaper end—alternatives that are kind of making compromises.
47:51
It's not just one
architecture anymore.
47:51
There are little ones coming up.
47:56
But if we talk about
the state-of-the-art, it's pretty much still the transformer architecture, autoregressive,
derived from GPT-2 essentially.
48:06
- I guess the big question here is, we talked
quite a bit about the architecture behind the pre-training.
48:10
Are the scaling laws
holding strong across pre-training, post-training, inference, context
size, data, and synthetic data?
48:20
- I'd like to start with the technical
definition of a scaling law- -which informs all of this.
48:23
The scaling
law is the power law relationship between...
48:26
You can think of the x-axis,
so kind of what you are scaling as a combination of compute and data,
which are kind of similar, and then the y-axis is like the held-out
prediction accuracy over next tokens.
48:38
We talked about models being autoregressive.
48:38
It's like if you keep a set of text that the model has not seen, how
accurate will it get when you train?
48:44
And the idea of scaling laws came
when people figured out that that was a very predictable
relationship.
48:51
And I think that that technical term is continuing,
and then the question is, what do users get out of it?
48:59
Then there
are more types of scaling where, OpenAI's o1 was famous for introducing
inference time scaling.
49:03
And I think less famously for also showing that you
can scale reinforcement learning training and get kind of this log x-axis and then a linear increase in performance on y-axis.
49:15
So there's kind of these three axes now where the traditional scaling laws are talked
about for pre-training, which is how big your model is and how big your dataset is,
and then scaling reinforcement learning, which is like how long can you do this trial
and error learning that we'll talk about.
49:30
We'll define more of this, and then this
inference time compute, which is just letting the model generate more tokens on a specific
problem.
49:34
So I'm kind of bullish, but they're all really still working, but
the low-hanging fruit has mostly been taken, especially in the last year on reinforcement
learning with verifiable rewards, which is this RLVR, and then
inference time scaling, which is just why these models feel
so different to use, where previously you would get that first token immediately.
49:53
And now they'll go off for seconds, minutes, or even hours,
generating these hidden thoughts before giving you the first word of
your answer.
50:01
And that's all about this inference time scaling, which is such a wonderful kind of step function in
terms of how the models change abilities.
50:11
They kind of enabled this tool use
stuff and enabled this much better software engineering that we
were talking about.
50:14
And this, when we say enabled, is almost entirely
downstream of the fact that this reinforcement learning with verifiable
rewards training just kind of let the models pick up these skills very
easily.
50:25
So let the models learn, so if you look at the reasoning
process when the models are generating a lot of tokens, what it'll often be doing is:
it tries a tool, it looks at what it gets back.
50:37
It tries another API, it sees what it
gets back and if it solves the problem.
50:41
So the models, when you're training
them, very quickly learn to do this.
50:45
And then at the end of the day, that
gives this kind of general foundation where the model can use CLI
commands very nicely in your repo and handle Git for you and move
things around and organize things or search to find more information, which if
we were sitting in these chairs a year ago is something that we didn't really think
of the models doing.
51:00
So this is just kind of something that has happened this
year and has totally transformed how has totally transformed how
we think of using AI which evolution and just unlocks so
much value.
51:11
But it's like, just so- pr- unlocks so much
value.
51:15
But it's- it's like, it's not clear what the next avenue will
be in terms of unlocking stuff like this. I think there's...
51:23
we'll get to continual
learning later, but there's a lot of buzz around certain areas of AI, but no
one knows when the next step function will really come.
51:31
- So you've actually said
quite a lot of things there, and said profound things
quickly.
51:35
It would be nice to unpack them a little bit.
51:39
You say you're
bullish basically on every version of scaling.
51:43
So can we just even start
at the beginning?
51:43
Pre-training, are we kind of implying that the
low- hanging fruit on pre-training scaling has been picked?
51:53
Has pre-training hit a plateau, or is even pre-training
still something you're bullish on?
52:01
- Pre-training has gotten extremely
expensive.
52:01
I think to scale up pre-training, it's also implying
that you're gonna serve a very large model to the users.
52:08
So I think that it's been loosely established the likes of GPT-4
and similar models were around one trillion parameters at the biggest size.
52:16
There's a lot of rumors that they've actually gotten smaller as training has
gotten more efficient.
52:20
You want to make the model smaller because then your costs
of serving go down proportionately.
52:28
These models, the cost of training
them is really low relative to the cost of serving them to hundreds of millions of
users.
52:32
I think DeepSeek had this famous number of about five million dollars for
pre-training at cloud market rates. In OLMo 3, section 2.
52:40
4
in the paper, we just detailed how long we had the GPU clusters
sitting around for training which includes engineering
issues, multiple seeds, and it was like about two million dollars to rent
the cluster to deal with all the headaches of training a model.
52:55
So
these models are pretty— like, a lot of people could get one to
10 million dollars to train a model, but the recurring costs of
serving millions of users is really billions of dollars of compute.
53:06
I think that you can look at a thousand GPU rental you can pay
100 grand a day for.
53:10
And these companies could have millions of GPUs.
53:15
Like you can look at how much these things cost to sit around.
53:18
So
that's kind of a big thing, and then it's like, if scaling
is actually giving you a better model, is it gonna be financially worth it?
53:25
And I think we'll slowly push it out as AI solves more compelling tasks,
so like the likes of Claude Opus 4.
53:29
5, making Claude Code just
work for things.
53:33
I— I launched this project called the
ATOM project, which is American Truly Open Models in July, and
that was like a true vibe coded website, and like,
I have a job to make plots and stuff.
53:49
And then I came back to
refresh it in the last few weeks and it's like Claude Opus 4.
53:52
5 versus whatever model at
the time was like, just crushed all the issues that it had from building in
June and July and like, it might be a bigger model.
54:00
There's a lot of things that go
into this, but there's still progress coming.
54:04
- So what you're speaking to is the
nuance of the y-axis of the scaling laws—the way it's experienced
versus on a benchmark, the actual intelligence might be
different.
54:11
But still, your intuition about pre-training, if you
scale the size of compute, will the models get better?
54:20
Not
whether it's financially viable but just from the law aspect of it, do you
think the models will get smarter? - Yeah.
54:28
And I think that there's...
54:28
And this sometimes comes off as almost like disillusionment
from people, leadership at AI companies saying this, but they're like,
"It's held for 13 orders of magnitude of compute, why would it ever
end?"
54:39
So I think fundamentally it is pretty unlikely to stop, it's just
eventually we're not even gonna be able to test the bigger scales because of all the
problems that come with more compute.
54:50
I think that there's a
lot of talk on how 2026 is a year when very large
Blackwell compute clusters, like gigawatt-scale facilities at
hyperscalers, are coming online.
55:03
These were all contracts for
power and data centers that were signed and sought
out in 2022 and 2023.
55:11
So before or right after ChatGPT.
55:11
It took this two-to-three-year lead time to build these bigger clusters to
train the models.
55:15
While there's obviously immense interest in building even more data
centers than that.
55:19
So that is the crux that people are saying: these new clusters
are coming.
55:23
The labs are gonna have more compute for training.
55:27
They're going to utilize this, but it's not a given.
55:30
I've seen so
much progress that I expect it, and I expect a little bit
bigger models, and I expect...
55:39
I would say it's more like we'll see a
$2,000 subscription this year.
55:39
We've seen $200 subscriptions.
55:43
That could 10X again,
and these are the kind of things that could come, and they're all
downstream of this bigger model that offers just a little
bit more cutting edge.
55:53
- So, you know, it's reported that xAI
is gonna hit that one-gigawatt scale early '26, and a full two gigawatts by year end.
56:01
How
do you think they'll utilize that in the context of scaling laws?
56:09
Is a lot of that inference?
56:09
Is a lot of that training?
56:12
- It ends up being all of
the above.
56:12
So I think that all of your decisions when
you're training a model come back to pre-training.
56:20
So if you're going to
scale RL on a model, you still need to decide on your architecture that enables
this.
56:23
We were talking about other architectures and using different types of
attention, or a mixture of experts models.
56:31
The sparse nature of MoE models makes it much more efficient to do
generation, which becomes a big part of post-training, and you need to
have your architecture ready so that you can actually scale up this
compute.
56:43
I still think most of the compute is going in at pre-training.
56:47
Because you can still make a model better, you still want to go and revisit
this.
56:51
You still want the best base model you can.
56:54
And in a few years that'll
saturate and the RL compute will just go longer.
57:00
- Are there people who disagree with
you and say pre-training is dead?
57:06
It's all about scaling inference,
scaling post-training, scaling context, continual learning,
scaling data, synthetic data?
57:15
- People vibe that way and describe it in
that way, but I think it's not the practice that is happening.
57:19
- It's just the general vibe of
people saying this thing is dead- - The excitement is elsewhere.
57:21
So the low-hanging fruit- ... in RL is elsewhere.
57:24
For example,
we released our model in November...
57:28
Every company has deadlines.
57:28
Our
deadline was November 20th, and for that, our run was five
days, which compared to 2024 is a very long time to just be doing
post- training at a model of 30 billion parameters. It's not a big model.
57:39
And
then in December, we had another release, where we let the RL run for
another three and a half weeks, and the model got notably better,
so we released it.
57:47
And that's a to just allocate to something
that is going to be your peak- ... for the year.
57:55
So it's like- - The reasoning is- - There's these types of decisions when
training a model where they just...
58:01
They can't leave it forever.
58:01
You have to keep pulling in the improvements from
researchers.
58:05
So you redo pre-training, you'll do this
post-training for a month, but then you need to give it to your users.
58:12
You
need to do safety testing. So it's just...
58:16
I think there's a lot in
place that reinforces this cycle of updating the models. Things improve.
58:20
You get a new compute cluster that lets you do something more
stably or faster.
58:24
It's like you hear a lot about Blackwell
having rollout issues, where at AI2, most of the models we're
pre-training are on 1,000 to 2,000 GPUs.
58:35
But when pre-training on
10,000 or 100,000 GPUs, you hit very different failures.
58:39
GPUs break
in weird ways, and on a 100,000 GPU run, you're pretty much
guaranteed to have one GPU that is down.
58:47
Your training code must
handle that redundancy, which is a very different problem.
58:51
Whereas what we're doing, like playing
with post-training on a cluster, or for people learning ML,
what they're battling to train these biggest models is just- ...
59:02
mass distributed scale,
and it's very different.
59:05
But that's somewhat different
than- That's a systems problem- ...
59:11
in order to enable scaling
laws, especially at pre-training.
59:15
You need all these GPUs at once.
59:15
When we shift to RL, it actually lends itself to heterogeneous compute
because you have many copies of the model.
59:23
To do a primer for language model reinforcement learning,
what you're doing is having two sets of GPUs.
59:31
One you can call the actor,
and one you call the learner.
59:34
The learner is where your actual
reinforcement learning updates happen.
59:38
These are traditionally
policy gradient algorithms.
59:42
Proximal Policy Optimization,
PPO, and Group Relative Policy Optimization, GRPO, are
the two popular classes.
59:50
And on the other side you have
actors which are generating completions, and these completions
are what you're going to grade.
59:57
Reinforcement learning is all about
optimizing reward.
59:57
In practice, you can have a lot of different actors
in different parts of the world doing different types of problems,
and then you send it back to this highly networked compute
cluster to do this actual learning where you take the gradients.
1:00:12
You need to have a tightly meshed network to do
different types of parallelism and spread out your model for efficient
training.
1:00:20
Every different type of training and serving has these
considerations to scale.
1:00:28
We talked about pre-training and RL,
and then inference time scaling- how do you serve a model that's thinking
for an hour to 100 million users?
1:00:35
I don't know about that, but I know
that's a hard problem.
1:00:35
In order to give people this intelligence, there's
all these systems problems, and we need more compute and you need more
stable compute to do it."
1:00:46
- But you're bullish on all of these kinds
of scaling is what I'm hearing.
1:00:46
On the inference, on the reasoning,
even on the pre-training?
1:00:54
- Yeah, so that's a big can of worms
here, but there are basically two...
1:00:58
The knobs are the training and the
inference scaling where you can get gains.
1:01:02
In a world where
we had, let's say, infinite compute resources, you want to do all of them.
1:01:05
So you have training, you have inference scaling, and training is like a hierarchy:
it's pre-training, mid-training, post-training.
1:01:13
Changing the model size,
more training data, training a bigger model gives you more knowledge in the
model.
1:01:17
Then the model, let's say, has a better base model.
1:01:22
Back in the day,
or still, we call it a foundation model, and it unlocks...
1:01:26
But
you don't, let's say, have the model be able to solve
your most complex tasks during pre-training or after pre-training.
1:01:34
You still have these other unlock phases where you have mid-training
or, for example, post-training with RL that unlocks capabilities that the model
has in terms of knowledge in the pre-training.
1:01:45
And I think, sure, if you do more pre-training, you get a
better base model that you can unlock later.
1:01:53
But like Nathan said, it just becomes
too expensive.
1:01:53
We don't have infinite compute, so you have to decide, do I want to spend
that compute more on making the model larger? It's like a trade-off.
1:02:01
In
an ideal world, you want to do all of them.
1:02:05
And I think in that sense,
scaling is still pretty much alive.
1:02:05
You would still get a better model, but like we saw
with Claude Opus 4.
1:02:08
5, it's just not worth it.
1:02:12
Because you can unlock more performance with other techniques at
that current moment, especially if you look at inference scaling.
1:02:20
That's one
of the biggest gains this year with o1, where it took a smaller model further than pre-training a larger model like
Claude Opus 4. 5.
1:02:29
So I wouldn't say pre-training scaling is dead, it's just that
there are other more attractive ways to scale right now.
1:02:36
But at some point,
you will still want to make some progress on the pre-training.
1:02:40
The thing also to consider is where you want to spend your
money.
1:02:44
If you spend it more on the pre-training, it's like a fixed cost.
1:02:48
You train the model, and then it has this capability forever. You can always use it.
1:02:56
With inference scaling, you don't spend
money during training, you spend money later per query, and then it's also like
math.
1:03:00
How long is my model gonna be on the market if I replace it in
half a year?
1:03:04
Maybe it's not worth spending $5 million, $10 million,
$100 million on training it longer.
1:03:11
Maybe I will just do more inference scaling and get performance there.
1:03:15
It maybe costs me $2 million in terms of user queries.
1:03:19
It becomes a question
of how many users you have and doing the math, and I think that's also where
it's interesting where ChatGPT is in a position.
1:03:26
I think they have a lot of users where
they need to go a bit cheaper, where they have that GPT-5 model that is a bit
smaller.
1:03:30
Other companies that have...
1:03:33
Let's say, if
your customers have other trade-offs.
1:03:37
For example, there was
also the Math Olympiad or some of these math problems where ChatGPT or they had a proprietary
model, and I'm pretty sure it's just like a model that has been fine-tuned
a little bit more, but most of it was during inference scaling to achieve
peak performance in certain tasks. need that all the time.
1:03:56
But
yeah, long story short, I do think all of these pre-training,
mid-training, post-training, inference scaling, they are all still
things you want to do.
1:04:03
It's just finding—at the moment, in this year, it's
finding the right ratio that gives you the best bang for the buck, basically.
1:04:13
- I think this might be a good place to
define pre-training, mid-training, and post-training.
1:04:18
- So, pre-training is the classic training
one next token prediction at a time.
1:04:21
You have a big corpus of data.
1:04:21
And Nathan
probably also has very interesting insights there because of OLMo 3.
1:04:26
A big
portion of the paper focuses on the right data mix.
1:04:29
So, pre-training is
essentially just, you know, training cross entropy loss, training
on next token prediction on a vast corpus of internet data,
books, papers and so forth.
1:04:37
It has changed a little bit over the years
in the sense people used to throw in everything they can.
1:04:45
Now,
it's not just raw data.
1:04:45
It's also synthetic data
where people, let's say, rephrase certain things.
1:04:52
So synthetic
data doesn't necessarily mean purely AI-made data.
1:04:56
It's also taking something from an article, a Wikipedia
article, and then rephrasing it as a Q&A question or summarizing it, rewording
it, and making better data that way.
1:05:12
Because I think of it also
like with humans.
1:05:12
If someone, let's say, reads a book compared to a
messy—no offense, but like—Reddit post or something like that,
I do think you learn— - There's going to be a post
about this, Sebastian.
1:05:28
- Some Reddit data is very coveted
and excellent for training.
1:05:31
You just have to filter it.
1:05:33
- And I think that's the idea.
1:05:33
I think it's like if someone took that and rephrased it in a, let's
say, more concise and structured way, I think it's higher quality data
that gets the LLM there faster.
1:05:46
You get the same LLM out of it at the
end, but it trains faster because if the grammar and the punctuation are correct, it already learns the
correct way, versus getting information from a messy source and then
learning later how to correct that.
1:06:02
So, I think that is how pre-training
evolved and why scaling still works.
1:06:09
It's not just about the
amount of data, it's also the tricks to make that data
better for you, in a sense.
1:06:13
And then mid-training is...
1:06:17
I mean, it
used to be called pre-training.
1:06:21
I think it's called mid-training because it was awkward
to have pre-training and post-training but nothing in the middle, right? It sounds a bit weird.
1:06:25
You have pre-training
and post-training, but what's the actual training?
1:06:29
So, the mid-training is usually
similar to pre-training, but it's a bit more
specialized.
1:06:33
It's the same algorithm, but what you do is
you focus, for example, on long-context documents.
1:06:40
The reason you don't do that during pre-training is
because you don't have that many long context documents.
1:06:48
We have a specific
phase.
1:06:48
And one problem of LLMs is still that it's a neural network.
1:06:52
It has the problem of catastrophic forgetting.
1:06:55
So, you teach it something,
it forgets other things. And you wanna...
1:06:59
I mean, it's not 100% forgetting,
but it's like "no free lunch."
1:07:03
It's the same with humans.
1:07:03
If you
ask me some math I learned 10 years ago, I would have to look at it again.
1:07:09
- Nathan was actually saying that he's consuming
so much content that there's a catastrophic forgetting issue.
1:07:14
- Yeah, I'm trying to learn so much about
AI, and it's like I was learning about pre-training parallelism.
1:07:18
I'm like, "I lost
something and I don't know what it was."
1:07:22
- I don't want to anthropomorphize
LLMs, but it's the same kind of thing in how humans learn.
1:07:25
I mean,
quantity is not always better because you have to be selective.
1:07:29
And mid-training is being selective in terms of quality content
at the end.
1:07:33
So the last thing the LLM has seen is the quality stuff.
1:07:37
And then post-training is all the fine-tuning, supervised fine-tuning, DPO, Reinforcement Learning
with Verifiable Rewards (RLVR), with human feedback, and so forth.
1:07:49
So
the refinement stages.
1:07:49
And it's also interesting, it's a cost thing.
1:07:53
You spend
a lot of money on pre-training right now. RL a bit less.
1:07:57
With RL, you don't really teach it knowledge.
1:08:01
It's more like unlocking
the knowledge; it's more like a skill learning, like how to solve problems with
the knowledge that it has from pre-training.
1:08:09
There are actually three papers
this year, or last year, 2025, on RL for pre-training.
1:08:12
But I don't
think anyone does that in production. - Toy examples for now. - Toy examples, right?
1:08:18
But to generalize,
RL post-training is more like the skill unlock, where pre-training
is like soaking up the knowledge.
1:08:26
- A few things that could be
helpful.
1:08:26
A lot of people think of synthetic data as being bad
for training the models.
1:08:31
You mentioned how DeepSeek got almost...
1:08:37
OCR, which is Optical Character
Recognition. A lot of labs did it.
1:08:41
Ai2 had one, Meta had multiple.
1:08:44
And the reason each of these labs
has these is because there are vast amounts of PDFs and other digital documents
on the web that aren't in formats that are encoded with text easily.
1:08:52
So you use these Almost-OCR, DeepSeek OCR, or what we called our
Almost-OCR, to extract trillions of tokens of candidate data for pre-training.
1:09:04
Pre-training dataset size is
measured in trillions of tokens.
1:09:08
Smaller models from researchers can be
something like five to 10 trillion.
1:09:11
researchers can be something
like five to 10 trillion.
1:09:11
Um, Qwen is documented going up to 50 trillion,
and there are rumors that these closed labs can go to 100 trillion tokens.
1:09:19
Getting this potential data to put in—they have a very big funnel, and
the data you actually train on is a small percentage of this.
1:09:27
This character recognition data would be described as synthetic data for
pre-training in a lab.
1:09:31
And then there's also the fact that ChatGPT now
gives wonderful answers, and you can train on those best answers, and that's
synthetic data.
1:09:39
It's very different than early ChatGPT with lots
of hallucination data.
1:09:46
when people became grounded
in synthetic data.
1:09:48
- One interesting question is, if I recall
correctly, OLMo 3 was trained with less data than specifically some other
open-weight models, maybe even OLMo 2.
1:09:56
But you still got better performance,
and that might be one example of how the data helped.
1:10:01
- It's mostly down to data quality.
1:10:02
I think if we had more compute, we would
train for longer.
1:10:02
I think we'd ultimately see that as something we would want
to do.
1:10:06
And especially with big models, you need more compute, because we
talked about having more parameters and we talked about knowledge.
1:10:14
Essentially,
there's a ratio where big models can absorb more from data, and then you
get more benefit out of this.
1:10:22
Any logarithmic graph in your mind
is like a small model will level off sooner if you're measuring tons
of tokens, and bigger models need more.
1:10:29
But mostly, we aren't training
that big of models right now at AI2, and getting the highest quality data
we can is the natural starting point.
1:10:38
- Is there something to be said about the
topic of data quality?
1:10:38
Is there some low-hanging fruit there still where
the quality could be improved?
1:10:46
- It's like turning the crank.
1:10:46
Historically, in the open, there's been a canonical best pre-training
dataset that has moved around between who has the most
recent one or the best recent effort.
1:10:53
Like AI2's Dolma was very early with the
first OLMo, and Hugging Face had FineWeb.
1:11:00
And there's a DCLM project,
which has been kind of like a, which stands for Data Comp Language
Model.
1:11:04
There's been Data Comp for other machine learning projects,
and they had a very strong dataset.
1:11:12
And a lot of it is
the internet is becoming fairly closed off, so we have Common Crawl,
which is hundreds of trillions of tokens, and you filter it.
1:11:20
It looks
like scientific work where you're training classifiers and
making decisions based on how you prune down this dataset into the
highest quality stuff and the stuff that suits your tasks.
1:11:31
Previously,
language models were tested a lot more on knowledge and conversational
things, but now they're expected to do math and code.
1:11:39
To train a reasoning
model, you need to remix your whole dataset.
1:11:43
And there's a lot of wonderful
scientific methods here where you can, you can take your gigantic dataset,
sample really tiny things from different sources, such as
GitHub, Stack Exchange, Reddit, Wikipedia.
1:11:55
You can sample small things from
them, and train small models on each of these mixes and measure their performance on your
evaluations.
1:11:59
You can just do basic linear regression, and it's like, "Here's your optimal
dataset."
1:12:03
But if your evaluations change, your dataset changes a lot.
1:12:07
So a lot
of OLMo 3 was new sources for reasoning to be better at math and
code, and then you do this mixing procedure and it gives you the answer.
1:12:15
I think that's happened at labs this year; there's new hot things, whether
it's coding environments or web navigation, and you need to bring in new data,
change your whole pre-training so that your post-training can work better.
1:12:26
And that's like the constant evolution and the redetermining of
what they care about for their models.
1:12:35
- Are there fun anecdotes of
what sources of data are particularly high quality that we
wouldn't expect?
1:12:39
You mentioned Reddit sometimes can be a source.
1:12:45
- Reddit was very useful.
1:12:45
I
think PDFs is definitely one. - Oh, especially arXiv.
1:12:52
- Yeah, so AI2 has run Semantic
Scholar for a long time, which is what you can say is a competitor to
Google Scholar with a lot more features.
1:13:01
And to do this, AI2 has found
and scraped a lot of PDFs for openly accessible papers
that might not be behind the closed walled garden of a certain
publisher.
1:13:09
So, truly open scientific PDFs.
1:13:12
And if you sit
on all of these and you process it, you can get value out
of it.
1:13:16
And I think that like, a lot of that style
of work has been done by the frontier labs did much
earlier.
1:13:22
You just need to have a pretty skilled researcher that
understands how things change models; they bring it in, clean
it, and it's a lot of labor.
1:13:33
When frontier labs scale
researchers, a lot more goes into data.
1:13:37
If you join a
frontier lab and you want to have impact, the best way to do
it is just find new data that's better.
1:13:44
And then, the fancy,
glamorous algorithmic things like figuring out how to make o1 is like the
sexiest thought of a scientist.
1:13:49
It's like, "Oh, I figured out how to scale RL."
1:13:52
There's a group that did that, but most of the contribution is like— - On the dataset - ..."
1:13:58
I'm gonna make the data better," or, "I'm gonna
make the infrastructure better so everyone on my team can run experiments 5% faster."
1:14:04
- At the same time, I think it's also one of the
closest guarded secrets, what your training data is, for legal reasons.
1:14:08
And so there's also, I
think, a lot of work that goes into hiding what your training data was
essentially.
1:14:12
Like training the model to not give away the sources
because you have legal reasons.
1:14:19
- The other thing, to be complete, is that
some people are trying to train on only licensed data, whereas Common Crawl
is a scrape of the whole internet.
1:14:26
So if I host multiple websites, I'm happy to have them train language
models, but I'm not explicitly licensing what governs it.
1:14:33
And
therefore, Common Crawl is largely unlicensed, which means that
your consent really hasn't been provided for how to use the data.
1:14:41
There's another
idea where you can train language models only on data that has been licensed
explicitly, so that the kind of governing contract is provided, and I'm not sure
if Apertus is the copyright thing or the license thing.
1:14:53
I know that the reason that they did it was for
an EU compliance thing, where they wanted to make sure that their model
fit one of those checks.
1:15:05
- On that note, there's also the
distinction in licensing.
1:15:05
Some people just purchase the license.
1:15:11
Let's say they buy an Amazon Kindle book, or a Manning book, and then
use that in training.
1:15:15
That is a gray zone 'cause you paid for the content and you
might want to train on it.
1:15:19
But then there are also restrictions where even that
shouldn't be allowed.
1:15:23
And so that is where it gets a bit fuzzy.
1:15:26
And yeah, I think that is right now still a hot topic.
1:15:30
Big companies like OpenAI approached private companies
for their proprietary data and private companies, they become
more and more, let's say, protective of their data because they
know, "Okay, this is going to be my moat in a few years."
1:15:46
And I do think that's
like the interesting question, where if LLMs become more commoditized, and I think
a lot of people learn about LLMs, there will be a lot more people able to train LLMs.
1:15:56
Of course, there are infrastructure challenges.
1:16:00
But if you think of big
industries like pharmaceutical industries, law, finance industries,
I do think they, at some point, will hire people from other
frontier labs to build their in-house models on their proprietary
data, which will be then, again, another unlock with pre-training that is
currently not there.
1:16:15
Because even if you wanted to, you can't get that data.
1:16:19
You
can't get access to clinical trials most of the time and these types of things.
1:16:23
So, I do think scaling, in that sense, might be still pretty much
alive if you also look in domain-specific applications, because we are
still right now, in this year, just looking at general purpose LLMs on, on ChatGPT,
Anthropic and so forth.
1:16:34
They are just general purpose, they're not even, I think,
scratching the surface of what an LLM can do if it is really specifically
trained and designed for a specific task.
1:16:47
- I think on the data thing—this is one of the
things that happened in 2025, and we totally forget it—is Anthropic lost in
court and owed $1. 5 billion to authors.
1:16:54
Anthropic, I think,
bought thousands of books and scanned them and was cleared legally
for that because they bought the books, and that is kind of going through
the system.
1:17:02
And then the other side, they also torrented some books, and I think this
torrenting was the path where the court said that they were then culpable to pay these
billions of dollars to authors, which is just such a mind-boggling lawsuit that
kind of just came and went.
1:17:14
That is so much money- ... from the VC ecosystem.
1:17:22
- These are court cases that will define the
future of human civilization, because it's clear that data drives a lot of this, and
there's this very complicated human tension of...
1:17:29
I mean, you can
empathize. You're both authors.
1:17:34
And there's some degree to which, I
mean, you put your heart and soul and your sweat and tears into
the writing that you do.
1:17:38
It feels a little bit like theft
for somebody to train your data without giving you credit.
1:17:49
- And there are, like Nathan said, also two
layers to it.
1:17:49
Someone might buy the book and then train on it, which could be argued
fair or not fair, but then there are the straight-up
companies who use pirated books where it's not even compensating the
author.
1:18:00
That is, I think, where people got a bit angry about it
specifically, I would say.
1:18:06
- Yeah, but there has to be some kind of
compensation scheme.
1:18:06
This is like moving towards- ...
1:18:11
towards something like
Spotify streaming did- ... originally for music.
1:18:13
You know, what does that- ... compensation look like?
1:18:15
You have to define those
kinds of models.
1:18:15
You have to think through all of that.
1:18:19
One other thing I think people are generally
curious about, I'd love to get your thoughts, as LLMs are used more and
more.
1:18:23
If you look at even arXiv, but GitHub, more and more of
the data is generated by LLMs.
1:18:32
What do you do in that kind of
world?
1:18:32
How big of a problem is that?
1:18:38
- Largest problem's the infrastructure
and systems, but from an AI point of view, it's kind of inevitable.
1:18:45
- So it's basically LLM-generated data that's
curated by humans essentially, right?
1:18:49
- Yes, and I think that a lot of open source
contributors are legitimately burning out.
1:18:53
If you have a popular open source
repo, somebody's like, "Oh, I want to do open source AI.
1:18:56
It's good for
my career," and they just vibe- -code something and they throw it
in.
1:19:00
You might get more of this- - I have a- - - than I do.
1:19:05
- Yeah, so I have actually
a case study here.
1:19:09
I have a repository called
MLxtend that I developed as a student around 10 years
ago, and it is a reasonably popular library still for certain
algorithms, I think especially like frequent data mining stuff.
1:19:20
And
there were recently two or three people who submitted a lot of
PRs in a very short amount of time.
1:19:28
I do think LLMs have been involved
in submitting these PRs.
1:19:28
Me, as the maintainer, there are two things.
1:19:32
First,
I'm a bit overwhelmed.
1:19:32
I don't have time to read through it because, especially as an
older library, that is not a priority for me.
1:19:40
At the same time, I kind of also
appreciate it because I think something people forget is it's not just using the
LLM.
1:19:44
There's still a human layer that verifies something, and that
is in a sense also how data is labeled, right?
1:19:52
One of
the most expensive things is getting labeled data for
RL from human feedback phases.
1:19:59
And this is kind of like
that, where it goes through phases, and then you actually get higher
quality data out of it.
1:20:03
And so I don't mind it in a sense.
1:20:07
It can feel overwhelming,
but I do think there is also value in it.
1:20:11
- It feels like there's a fundamental
difference between raw LLM-generated data and LLM-generated data
with a human in the loop that does some kind of verification, even
if that verification is a small percent of the lines of code.
1:20:25
- I think this goes with anything where people think also sometimes, "Oh,
yeah.
1:20:29
I can just use an LLM to learn about XYZ," which is true.
1:20:33
You can,
but there might be a person who is an expert who might have used
an LLM to write specific code.
1:20:41
There is kind of like this human work
that went into it to make it nice, throwing out the not-so-nice
parts to kind of pre-digest it for you, and that saves
you time.
1:20:49
I think that's the value-add, where you have
someone filtering things or even using the LLMs correctly.
1:20:57
This is still labor that you get for free.
1:21:01
For example, if
you read a Substack article, I could maybe ask an LLM to
give me opinions on that, but I wouldn't even know what to ask.
1:21:08
And I think there is still value in reading that article compared to me
going to the LLM because you are the expert.
1:21:16
You select what knowledge
is actually spot on, should be included, and you give
me this very... this executive summary.
1:21:23
And this is
a huge value-add because now I don't have to waste three to
five hours to go through this myself, maybe get some incorrect
information and so on.
1:21:31
And so I think that's also where the future still is for writers even though there are
LLMs that...
1:21:37
Can kind of save you time.
1:21:43
- It's kinda fascinating to watch.
1:21:43
I'm sure you guys do this, but for me, I look at the
difference between a summary and the original content.
1:21:52
Even
if it's a page-long summary of a page-long content, it's interesting to see how the LLM-based summary
takes the edge off.
1:22:00
What is the signal it removes from the thing?
1:22:07
- The voice is what I talk about a lot. - Voice? Well, voice...
1:22:09
I would love
to hear what you mean by voice, but sometimes there's literally insights.
1:22:16
By removing an insight, you're
changing the meaning of the thing.
1:22:20
So I'm continuously
disappointed how bad LLMs are at really getting to the core
insights, which is what a great summary does.
1:22:29
Yet even if I have
these extremely elaborate prompts where I'm really trying to dig for the
insights, it's still not quite there, which...
1:22:41
I mean,
that's a whole deep philosophical question about what is human knowledge
and wisdom and what does it mean to be insightful and so on.
1:22:49
But when you talk
about the voice, what do you mean?
1:22:52
- When I write, I think a lot
of what I'm trying to do is take what you think as a
researcher, which is very raw.
1:22:55
A researcher is trying to encapsulate
an idea at the frontier of their understanding and they're trying
to put what is a feeling into words.
1:23:07
I try to do this
in my writing, which makes it come across as raw but
also high-information in a way that some people will get it and some won't.
1:23:15
And that's the nature of research.
1:23:19
And language models don't do
this well.
1:23:19
They're all trained with Reinforcement Learning from Human
Feedback, which takes feedback from many people and averages how the
model behaves from this.
1:23:30
And I think it's going to
be hard for a model to be very incisive when there's
that sort of filter.
1:23:34
This is a wonderful fundamental problem
for researchers in RLHF.
1:23:40
This provides so much utility in making
the models better, but also the problem formulation is
kind of...
1:23:48
there's this knot in it that you can't
get past.
1:23:52
These language models don't have this prior in their
deep expression they're trying to get at.
1:23:59
I don't think it's impossible.
1:23:59
There are stories of models that really shock people. Like, I
think of...
1:24:03
I would love to have tried Bing Sydney.
1:24:07
Did that
have more voice?
1:24:07
Because it would so often go off the
rails, which is historically obviously a scary way—like telling a
reporter to leave his wife—is a crazy model to potentially put in general
adoption.
1:24:18
But that's kind of like a trade-off; is this RLHF process
in some ways adding limitations?
1:24:28
- That's a terrifying place to be as one
of these frontier labs and companies, because millions of people are using them.
1:24:35
- There was a lot of backlash
last year with GPT-4o getting removed, and I've personally never used
the model, but I've talked to people at OpenAI where they get emails from users that might be detecting subtle
differences in the deployments in the middle of the night.
1:24:51
And they email
them like, "My friend is different."
1:24:55
And they find these employees' emails
and send them things because they are so attached to this set of model weights and configuration
that is deployed to the users. We see this with TikTok. You
open it...
1:25:05
I don't use TikTok, but supposedly in five minutes the
algorithm gets you. It's locked in.
1:25:12
And those are language models
doing recommendations.
1:25:12
I think there are ways you can do this.
1:25:16
Within
five minutes of chatting with it, the model just gets you.
1:25:20
And that is something that people aren't really ready for.
1:25:23
Like, don't give that to kids.
1:25:28
At least until we know what's happening.
1:25:30
- But there's also going to be this
mechanism...
1:25:30
What's going to happen with these LLMs as they're
used more and more...
1:25:36
Unfortunately, the nature of the human
condition is such that people commit suicide.
1:25:39
And so what journalists will do
is they will report extensively on the people who commit suicide.
1:25:43
And they would very likely link it to the LLMs because they have
that data about the conversations.
1:25:50
If you're really struggling in
your life, if you're depressed, if you're thinking about suicide, you're
going to probably talk to LLMs about it.
1:25:58
And so what journalists will do is
say, "Well, the suicide was committed because of the LLM."
1:26:02
And that's
going to lead to the companies, because of legal issues and so on, more and more taking the edge off
of the LLM.
1:26:10
So it's going to be as generic as possible.
1:26:14
It's so
difficult to operate in this space because you don't want an LLM to cause
harm to humans at that level, but also this is the nature of
the human experience, is to have a rich conversation,
a fulfilling conversation, one that challenges you from which
you grow. You need that edge.
1:26:35
And that's something extremely difficult
for AI researchers on the RLHF front to actually have to solve, because you're
dealing with the human condition.
1:26:47
- A lot of researchers at these
companies are so well-motivated, and definitely Anthropic and OpenAI
culturally want to do good for the world.
1:26:56
And it's such a—I'm like, "Ooh,
I don't wanna work on this," because on the one hand, a lot of
people see AI as a health ally, as somebody they can talk to about
their health confidentially, but then it bleeds all the way into this, like talking about mental
health, where it's heartbreaking that this will be the
thing where somebody goes over the edge, but other people might be saved.
1:27:19
And I'm like, "I don't..."
1:27:23
As a researcher, it's like, I don't want to
train image generation models and release them openly because I don't want to
enable somebody to have a tool on their laptop that can harm other people.
1:27:34
I don't have the infrastructure in
my company to do that safely.
1:27:34
But there's a lot of areas like this
where it just needs people that will approach it with complexity and
conviction.
1:27:42
It's just such a hard problem.
1:27:47
- But also, we as a society, as users of these
technologies, need to make sure that we're having the complicated conversation
about it versus just fearmongering— that Big Tech is causing harm to
humans or stealing your data.
1:27:59
It's more complicated than
that. And you're right.
1:28:03
There's a very large number of people inside
these companies, many of whom I know, who deeply care about
helping people.
1:28:07
They are considering the full human experience of
people from across the world, not just Silicon Valley.
1:28:14
People across the United States
and the world, what their needs are.
1:28:18
It's really difficult to
design this one system that is able to help all these different
kinds of people across different age groups, cultures, and mental states.
1:28:31
- I wish that the timing of AI was
different relative to the relationship of Big Tech to the average person.
1:28:35
Big Tech's reputation was so low, and with how AI is so expensive, it's
inevitably going to be a Big Tech thing.
1:28:42
It takes so many resources,
and people say the US is, "betting the economy on AI" with this
build-out.
1:28:46
To have these be intertwined at the same time makes for such a
hard communication environment.
1:28:54
It would be good for me to go talk to more
people in the world who hate Big Tech and see AI as a continuation of this.
1:29:02
- And one of the things you recommend,
one of the antidotes that you talk about, is to find agency
in this system.
1:29:06
As opposed to sitting back in a powerless
way and consuming the AI slop as it rapidly takes over the internet.
1:29:17
Find agency by using AI to build stuff, build apps, build...
1:29:24
One, that
actually helps you build intuition, but two, it's empowering because
you can understand how it works, what the weaknesses are.
1:29:32
It gives your voice power to say, "This is a bad use of the
technology, and this is a good use."
1:29:40
And you're more plugged
into the system then, so you can understand it better and
you can steer it better as a consumer.
1:29:48
- I think that's a good point you brought up about
agency.
1:29:48
Instead of ignoring it and saying, "Okay, I'm not going to use it,"
I think it's probably long-term healthier to say, "Okay, it's
out there. I can't put it back." when they came out.
1:30:00
How do
I make best use of it, and how does it help me to up-level myself?
1:30:04
The one thing I worry about here, though, is, if you just fully use it for something
you love to do, the thing you love to do is no longer there.
1:30:12
And that could
potentially, I feel, lead to burnout.
1:30:16
For example, if I use an LLM
to do all my coding for me, now there's no coding.
1:30:19
I'm just managing
something that is coding for me.
1:30:23
Two years later, let's say, if I
just do that eight hours a day, having something code for me,
do I feel fulfilled still?
1:30:36
Is this hurting me in terms of being excited
about my job, excited about what I'm doing?
1:30:40
Am I still proud to build something?
1:30:43
- On that topic of enjoyment, it's quite
interesting.
1:30:43
We should just throw this in there, that there's this recent survey of about 791 professional developers, meaning
10-plus years of experience. - That's a long time. As a junior developer?
1:31:01
- Yeah, in this day and age.
1:31:01
So, there's also many fronts that are surprising.
1:31:05
They break it down by junior and senior developers.
1:31:09
But, I mean, it just shows that both junior and senior developers
use AI-generated code in code they ship.
1:31:22
So this is not just for fun or
intermediate learning things. This is code they ship.
1:31:26
25%—like, most of them use around 50% or more.
1:31:30
And what's interesting is, for the category of over 50%
of your code that you ship is AI-generated, senior
developers are much more likely to do so.
1:31:41
But you don't want AI
to take away the thing you love.
1:31:46
I think this speaks to my experience,
these results I'm about to say.
1:31:49
Together, about 80% of people
find it either somewhat more enjoyable or significantly more
enjoyable to use AI as part of the work.
1:31:59
- I think it depends on the task.
1:31:59
From
my personal usage, for example, I have a website where I
sometimes tweak things.
1:32:07
I personally don't enjoy
this.
1:32:07
So in that sense, if the AI can help me to implement
something on my website, I'm all for it. It's great.
1:32:15
But then, at the same
time, when I solve a complex problem— well, if there's a bug,
and I hunt this bug, and I find it, it's the best feeling
in the world. You feel great.
1:32:27
But now, if you don't
even think about the bug, you just go directly to the LLM, well, you
never have this kind of feeling, right?
1:32:35
But then there could be
the middle ground where, well, you try yourself, you
can't find it, you use the LLM, and then you don't get frustrated because it helps
you and you move on to something that you enjoy.
1:32:46
And so, looking at these statistics,
I think what is not factored in is that it's averaging over
all the different scenarios.
1:32:54
We don't know if it's for the core task or if it's for something mundane that
people would not have enjoyed otherwise.
1:33:02
So, in a sense, AI is really great for doing
mundane things that take a lot of work.
1:33:06
For example, my wife the
other day—she has a podcast for book discussions, a book club, and
she was transferring the show notes from Spotify to YouTube, and then
the links somehow broke.
1:33:21
And she had in some episodes, because
it is so many books, like 100 links, and it would have been really painful to
go in there and fix each link manually.
1:33:29
So I suggested, "Hey, let's try ChatGPT."
1:33:29
We copied the text into ChatGPT, and it fixed them.
1:33:33
Instead of two hours going from link to link fixing
that, it made that type of work much more seamless.
1:33:41
I think everyone has a use case where AI is useful
for something that would be really boring, really mundane.
1:33:51
- For me personally, since we're talking about
coding, and you mentioned debugging...
1:33:58
the source of enjoyment for me, more
on the Cursor side than Claude Code, is that I have a friend, I have
a pair programmer. It's less lonely.
1:34:09
You made debugging
sound like this great joy.
1:34:15
No, I would say debugging
is like a drink of water after you've been
going through a desert for days.
1:34:21
You skip the whole desert part where you're suffering.
1:34:25
Sometimes it's nice to have a friend who can't really find the bug,
but can give you some intuition about the code, and together you're going through the desert and finding
that drink of water.
1:34:37
At least for me, maybe it speaks to the loneliness
of the programming experience. That is a source of joy.
1:34:48
- It's maybe also related to delayed
gratification.
1:34:48
I'm a person who, even as a kid, I liked
the idea of Christmas presents—having them, getting
them—better than actually receiving the presents.
1:35:00
I would
look forward to the day I get the presents, but then it's over and I'm
disappointed.
1:35:04
And maybe it's the same with food.
1:35:07
I think food tastes
better when you're really hungry.
1:35:11
You're right, with debugging,
it is not always great.
1:35:11
It's often frustrating, but then if
you can solve it, then it's great.
1:35:22
But there's a sweet Goldilocks
zone; if it's too hard, then it's just wasting your time.
1:35:26
But I think another challenge is how
will people learn?
1:35:30
We looked at the chart and
saw that more senior developers are shipping more
AI-generated code than the junior ones.
1:35:42
It's very interesting, because intuitively
you would think it's the junior developers because they don't know
how to do the thing yet, and so they use AI to do that
thing.
1:35:49
It could mean the AI is not good enough yet to solve that
task, but it could also mean experts are more effective
at using it.
1:35:57
They know how to use it better, review the code,
and then they trust the code more.
1:36:06
One issue for society in the future
will be: how do you become an expert if you never try to do the thing
yourself?
1:36:09
One way I always learned is by trying things myself.
1:36:17
If you
look at math textbooks and the solutions, you learn something,
but you learn actually better if you try first.
1:36:25
Then you
appreciate the solution differently because you know how to put it
into your mental framework.
1:36:29
If LLMs are here all the time, would
you actually go to the length of struggling?
1:36:37
Would you be willing to
struggle?
1:36:37
Struggle is not nice, right?
1:36:37
But if you use the LLM to do everything, at
some point you will never really take the next step, and then you will maybe not get that unlock that
you would get as an expert using an LLM.
1:36:52
So, I think there's
like a Goldilocks sweet spot where maybe the trick
here is you make dedicated offline time where you study two hours
a day, and the rest of the day use LLMs.
1:37:03
But I think it's important
also for people to still invest in themselves, in my opinion,
to not just LLM everything.
1:37:10
- Yeah, as a civilization, we each individually
have to find that Goldilocks zone.
1:37:16
And in the programming context as
developers.
1:37:16
Now, we've had this fascinating conversation that started with
pre-training and mid-training.
1:37:24
Let's get to post-training.
1:37:24
A lot of fun stuff in post-training.
1:37:27
So, what are some of the
interesting ideas in post-training?
1:37:31
- The biggest one from 2025 is learning this reinforcement learning with verifiable
rewards.
1:37:35
You can scale up the training there, which means doing a lot of
this kind of iterative generate-grade loop, and that lets the
models learn both interesting behaviors on the tool use and software
side.
1:37:47
This could be searching, running commands on their own and seeing the
outputs, and then also that training enables this inference time
scaling very nicely.
1:37:55
And it just turned out that this paradigm was
very nicely linked, where this kind of RL training enables inference time
scaling.
1:38:02
But inference time scaling could have been found in different ways.
1:38:06
So, it was kind of
this perfect storm where the models change a lot, and the way that they're trained is a major
factor in doing so.
1:38:10
And this has changed how people approach
post-training dramatically.
1:38:20
- Can you describe RLVR, popularized by
DeepSeek R1?
1:38:20
Can you describe how it works? - Yeah.
1:38:25
Fun fact: I was on the
team that came up with the term RLVR, which is from our Tulu
3 work before DeepSeek.
1:38:29
We don't take a lot of credit for being
the people to popularize the scaling RL, but as fun as what
academics get, as an aside, is the ability
to name and influence the discourse, because the closed
labs can only say so much.
1:38:43
That one of the things you can do as an academic is
you might not have the compute to train the model, but you can frame things
in a way that ends up being described as a community coming
together around this RLVR term, which is very fun.
1:38:59
And then DeepSeek are the
people that did the training breakthrough, which is, they scaled the
reinforcement learning.
1:39:03
You have the model generate answers and
then grade the completion if it was right, and then that accuracy is your reward for reinforcement learning.
1:39:14
So reinforcement learning is classically an agent that acts in an
environment, and the environment gives it a state and a reward back,
and you try to maximize this reward.
1:39:25
In the case of language
models, the reward is normally accuracy on a set of verifiable
tasks, whether it's math problems or coding tasks.
1:39:32
And it
starts to get blurry with things like factual domains.
1:39:38
That is also,
in some ways, verifiable, or constraints on your instruction,
like respond only with words that start with A."
1:39:45
All of
these things are verifiable in some way, and the core
idea of this is you find a lot more of these problems that are
verifiable and you let the model try it many times while taking
these RL steps, these RL gradient updates.
1:40:01
The
infrastructure evolved from reinforcement learning from human
feedback, where in that era the score they were trying to optimize was
a learned reward model of human preferences.
1:40:12
So you kind of changed
the problem domains and that let the optimization go on to
much bigger scales, which kind of kickstarted a major change in what the
models can do and how people use them.
1:40:24
- What kind of domains is RLVR amenable to?
1:40:28
- Math and code are the famous
ones, and then there's a lot of work on what is called the rubrics,
which is related to a word people might have heard: LLM-as-a-judge.
1:40:36
For each problem, I'll have a set of problems in my
dataset.
1:40:40
I will then have an LLM and ask it, "What would a good
answer to this problem look like?"
1:40:48
And then you could try the
problem over and over again and assign a score based on this rubric.
1:40:52
That's not necessarily verifiable like math and code domains, but
this rubrics idea and other scientific problems that might be
a little bit more vague is where the attention is, where they're trying
to push this set of methods into these kind of more open-ended domains
so the models can learn a lot more.
1:41:11
- I think that's called reinforcement
learning with AI feedback, right?
1:41:14
- That's the older term for it coined
in Anthropic's Constitutional AI paper.
1:41:17
It's like a lot of
these things come in cycles.
1:41:21
- Also, just one step back
for RLVR.
1:41:21
I think the interesting thing here is that
you ask the LLM a, let's say, math question, and then you know the
correct answer, and you let the LLM, as you said, figure it out.
1:41:33
How
it does it—you don't constrain it much.
1:41:37
There are some constraints like
"use the same language, don't switch between Spanish and English."
1:41:41
But let's
say you're pretty much hands-off.
1:41:45
You only give the question and
the answer, and then the LLM has the task to arrive at the
right answer, but the beautiful thing here is what
happens in practice: the LLM will do a step-by-step description,
like as a student or as a mathematician would derive the
solution.
1:41:59
It will use those steps, and that helps the model to
improve its own accuracy.
1:42:07
And then, like you said,
the inference scaling.
1:42:11
Inference scaling loosely
means spending more compute during inference, and here
the inference scaling is that the model would use more tokens.
1:42:20
In the DeepSeek R1 paper, they showed the longer they train the
model, the longer the responses are. They grow over time.
1:42:28
They use more
tokens, so it becomes more expensive.
1:42:32
It becomes expensive for simple tasks,
but these explanations help accuracy.
1:42:36
There are also papers showing
what the model explains does not necessarily have to be
correct or maybe it's unrelated to the answer, but for some reason, it
still helps the model that it is explaining.
1:42:47
And I think it's
also—again, I don't want to anthropomorphize these LLMs, but it's
kind of like how we humans operate.
1:42:55
If there's a complex math
problem in a math class, class, you usually have a note paper and
you do it step by step. You cross out things.
1:43:02
And the model also self-corrects,
and that was, I think, the aha moment in the DeepSeek R1 paper.
1:43:06
They called
it the aha moment because the model itself recognized it made a mistake and then said,
"Ah, I did something wrong, let me try again."
1:43:14
And I think that's just
so cool that this falls out of just giving it the correct answer
and having it figure out how to do it, that it kind of does in a
sense what a human would do.
1:43:26
Although LLMs don't think like humans, it's
kind of like an interesting coincidence and it...
1:43:30
And the other nice
side effect is it's great for us humans often to see these steps.
1:43:34
It builds trust, but also we us humans to see these steps.
1:43:38
It builds trust,
but also we learn and can double check things. - There's a lot in here.
1:43:40
I think some of the debate...
1:43:41
There's been a lot of
debate this year on if the language models like these...
1:43:44
I think the aha moments
are kind of fake because in pre-training you essentially have
seen the whole internet.
1:43:51
so you have definitely seen people
explaining their work, even verbally, like a transcript of a math lecture.
1:43:55
"You try
this, oh, I messed this up."
1:43:55
And what RLVR is very good at doing is amplifying these behaviors because they're very useful
in enabling the model to think longer and to check its work.
1:44:06
And I agree that
it is very beautiful that this training kind of...
1:44:10
The model learns to amplify
this in a way that is so useful for the final answers being better.
1:44:16
- I can give you also a hands-on example.
1:44:16
I
was training the Qwen 3 base model with RLVR on MATH-500.
1:44:20
The base model had an
accuracy of about 15%.
1:44:20
Just 50 steps, like in a few minutes with
RLVR, the model went from 15% to 50% accuracy. And the
model...
1:44:32
You can't tell me it's learning anything
fundamentally about math in - The Qwen example is weird because there've been
two papers this year, one of which I was on, about data contamination in Qwen and specifically that they train on a
lot of this special mid-training phase that we should take a minute
on, because it's weird because they train on problems
that are almost identical to MATH. - Exactly.
1:44:53
And so you can
see that basically the RL, it's not teaching the model any new knowledge
about math.
1:44:57
You can't do that in 50 steps.
1:45:01
So the knowledge is already there, in the
pre-training, you're just unlocking it.
1:45:03
- I still disagree with the premise because
there's a lot of weird complexities that you can't prove because one of the things that points to weirdness is
that if you take the Qwen 3 so-called base model and you...
1:45:14
You
could Google like "math dataset, Hugging Face", and you could take
a problem and what you do if you put it into Qwen 3 base...
1:45:22
All these math problems
have words, so it'd be like "Alice has five apples and takes one...
1:45:26
and gives
three to whoever," and there are these word problems.
1:45:30
With these Qwen-based
models, why people are suspicious of them is if you change the
numbers but keep the words- Qwen will produce, without
tools, will produce a very high accuracy decimal representation of the answer, which means there's
some...
1:45:43
At some time, it was shown problems that were almost identical
to the test set, and it was using tools to get a very high precision answer,
but a language model without tools will never actually have this.
1:45:55
So it's kind of been this big debate in the research community:
how much of these reinforcement learning papers that are training on Qwen
and measuring specifically on this math benchmark, where there's been multiple
papers talking about contamination, is like, how much can you believe them?
1:46:10
And I think
this is what caused the reputation of RLVR being about formatting, because you
can get these gains so quickly, therefore it must already be in the model.
1:46:18
But
there's a lot of complexity here that we...
1:46:22
It's not really like controlled
experimentation, so we don't really know.
1:46:26
- But if it weren't true, I would say
distillation wouldn't work, right?
1:46:30
I mean, distillation can work to
some extent, but the thing is that is, I think, the biggest problem, and I research this
contamination because we don't know what's in the data.
1:46:38
Unless you have a new dataset,
it is really impossible.
1:46:38
And the same, you mentioned the math dataset,
where you have a question and then answer and an explanation is given,
but then also even something simpler like MMLU, which is a multiple-choice
benchmark.
1:46:50
If you just change the format slightly,
like, I don't know, if you use a dot instead of a parenthesis
or something like that, the model accuracy will vastly differ.
1:47:04
- I think that that could be like a model
issue rather than a general issue.
1:47:09
- It's not even malicious by the developers of the LLM,
like, "Hey, we want to cheat at that benchmark."
1:47:13
It has seen something at some point.
1:47:13
I
think the only fair way to evaluate an LLM is to have a new benchmark that is after
the cutoff date when the LLM was deployed.
1:47:22
- Can we lay out what would
be the recipe of all the things that go into post-training?
1:47:26
And you mentioned RLVR was a really exciting, effective thing.
1:47:32
Maybe we should elaborate.
1:47:32
RLHF
still has a really important component to play.
1:47:36
What kind of other
ideas are there on post-training?
1:47:40
- I think you can kind of take
this in order.
1:47:40
I think you could view it as what made o1, which
is this first reasoning model, possible, or what will the
latest model be? And they actually...
1:47:50
You're going to have
similar interventions at these, where you start with mid-training,
and the thing that is rumored to enable o1 and similar
models is really careful data curation, where you're providing
a broad set of what is called reasoning traces, which
is just the model generating words in a forward process
that is reflecting, like breaking down a problem into intermediate steps
and trying to solve them.
1:48:13
So at mid-training, you need to have data that is similar to this to make it
so that when you move into post-training, primarily with these verifiable
rewards, it can learn.
1:48:27
And then what is happening
today is you're figuring out which problems to
give the model and how out which problems to give the model and how long
you can train it for and how much inference you can enable the model to use when solving
these verifiable problems.
1:48:38
So as models get better, certain problems models get better, certain problems are no longer...
1:48:45
The model will solve them 100% of the time, and therefore there's very little signal
in this.
1:48:49
If we look at the GRPO equation, this one is famous for
this because essentially the reward given to the agent
is based on how good a given action—an action is a
completion—is relative to the other answers to that same problem.
1:49:04
So if all the problems
get the same answer, there's no signal in these types of algorithms.
1:49:08
So what they're doing
is they're finding harder problems, which is why you hear about things like
scientific domains, where it's so hard to get anything right.
1:49:16
If you
have a lab or something, it just generates so many tokens or much
harder software problems.
1:49:19
So the frontier models are all pushing into these
harder domains when they can train on more problems and the model will learn
more skills at once.
1:49:27
The RLHF link to this is that RLHF has been
and still is kind of like the finishing touch on the models, where
it makes the models more useful by improving the organization or style
or tone.
1:49:38
There are different things that resonate with different audiences, like some
people like a really quirky model and RLHF could be good at enabling that
personality, and some people hate this markdown bulleted list
thing that the models do, but it's actually really good for quickly
parsing information.
1:49:54
In RLHF, this human feedback stage is
really great for putting this into the model at the end of
the day.
1:50:02
It's what made ChatGPT so magical for people.
1:50:06
And that
use has actually remained fairly stable.
1:50:09
This formatting can also help the models get better at math
problems, for example.
1:50:13
So it's like the border between style
and formatting, and like the method that you use
to answer a problem is actually all very closely
linked in terms of when you're training these models, which is
why RLHF can still make a model better at math, but these verifiable domains
are a much more direct process to doing this because it makes
more sense with the problem formulation, which is why it ends up all
forming together.
1:50:39
But to summarize, it's like mid-training is give the model
the skills it needs to then learn.
1:50:47
RL with verifiable
rewards is letting the model try a lot of times, so put a lot of compute
into trial-and-error learning across hard problems.
1:50:54
And then RLHF would be like
finishing the model, making it easy to use and kind of just
rounding the model out.
1:51:02
- Can you comment on the amount
of compute required for RLVR?
1:51:06
- It's only gone up and up.
1:51:06
I think Ilya Sutskever was famous for saying they use a similar
amount of compute for pre-training and post-training.
1:51:12
Back to the scaling
discussion, they involve very different hardware for scaling.
1:51:16
Pre-training
is very compute-bound, which is like this FLOPs discussion, which is just how many matrix
multiplications can you get through at once.
1:51:24
And because with RL you're generating
these answers, you're trying the model in real-world environments, it ends
up being much more memory-bound because you're generating long sequences
and the attention mechanisms have this behavior where you get a
quadratic increase in memory as you're getting to longer sequences.
1:51:40
So the compute becomes very different.
1:51:43
In pre-training we would talk about a
model—if we go back to like the Biden administration executive order, it's
like 10 to the 25th FLOPs to train a model.
1:51:51
If you're using FLOPs in
post-training, it's a lot weirder because the reality is just like: how many hours
are you allocating? How many GPUs for?
1:51:59
And I think in terms
of time, the RL compute is getting much closer because you just
can't put it all into one system.
1:52:06
Pre-training is so computationally dense where
all the GPUs are talking to each other and it's extremely efficient, whereas RL has all
these moving parts and can take a long time to generate a sequence of
100,000 tokens.
1:52:15
If you think about GPT-5.
1:52:18
2 Pro taking an hour,
it's like, what if your training run has a sample for an hour and you have to
make sure that's handled efficiently?
1:52:25
So I think in GPU hours
or just wall-clock hours, the RL runs are probably approaching the
same number of days as pre-training, but they probably aren't using
as many GPUs at the same time.
1:52:36
There are rules of thumb where in
labs you don't want your pre-training runs to last more than a month because
they fail catastrophically.
1:52:40
And if you are planning a huge cluster to be
held for two months and then it fails on day 50, the opportunity
costs are just so big.
1:52:52
So people don't want to put
all their eggs in one basket.
1:52:56
GPT-4 was the ultimate YOLO
run, and nobody ever wanted to do it before, where it took three months
to train and everybody was shocked that it worked.
1:53:04
I think people are a little
bit more cautious and incremental now.
1:53:07
- So RLVR is more, let's say, unlimited in how much you can train or still get benefit,
where RLHF, because it's a preference tuning, you reach a certain point where it
doesn't really make sense to spend more RL budget on that.
1:53:19
So just a step
back with preference tuning: there are multiple people that
can give multiple, let's say, explanations for the same thing and they can both
be correct, but at some point you learn a certain style and it doesn't make
sense to iterate on it.
1:53:30
My favorite example is: if relatives
ask me what laptop they should buy, I give them an explanation or
ask, "What is your use case?"
1:53:42
They, for example,
prioritize battery life and storage.
1:53:46
Other people like us, for
example, we would prioritize RAM and compute.
1:53:50
Both answers are
correct, but different people require different answers.
1:53:54
With preference
tuning, you're trying to average somehow.
1:53:58
You are asking the data labelers
to give you, not the right, but the preferred answer and then you
train on that.
1:54:02
But at some point you learn that average preferred answer.
1:54:05
And there's no reason to keep training longer on it because it's
just a style, whereas with RLVR, you let the model solve more and more complex, difficult problems.
1:54:18
So I
think it makes more sense to allocate more budget long-term to RLVR.
1:54:22
Also,
right now we are in an RLVR 1.
1:54:22
0 blend where it's still that simple
thing where we have a question and answer, but we don't do anything
with the stuff in between.
1:54:38
There were multiple research
papers, also by Google for example, on process reward models that
also give scores for the explanation—how correct is the
explanation.
1:54:46
And I think that will be the next thing, let's say
RLVR 2.
1:54:50
0 for this year, focusing in between question and
answer, like how to leverage that information, the explanation,
to help it get better accuracy. So that's one angle.
1:55:01
And there was a DeepSeek Math-V2 paper where they also had interesting inference scaling
there where, first, they had developed models that grade
themselves, a separate model.
1:55:16
And I think that will be one aspect.
1:55:16
And
the other, like Nathan mentioned, it will be for RLVR branching into other domains.
1:55:23
- The place where people are
excited are value functions, which is pretty similar.
1:55:26
So process
reward models are kind of like...
1:55:30
Process reward models assign how good something is at each
intermediate step in a reasoning process, where value functions apply
value to every token the language model generates.
1:55:40
Both of these have
been largely unproven in the language modeling and reasoning model era.
1:55:46
People are more optimistic about value functions for whatever
reason now.
1:55:50
I think process reward models were tried a lot more
in this pre-o1, pre-reasoning model era, and a lot of people had a lot of
headaches with them.
1:55:57
So I think a lot of it is human nature...
1:56:03
Value models have a very deep history in
reinforcement learning.
1:56:03
They're one of the first things core to deep
reinforcement learning existing, is training value models.
1:56:11
So
right now people are excited about trying value models, but there's
very little proof.
1:56:15
And there are negative examples in trying to scale
up process reward models.
1:56:22
These things don't always hold
in the future.
1:56:22
We came to this discussion by talking about scaling.
1:56:25
The simple way to summarize what you're saying is you don't want to do too much
RLHF, where the signal doesn't scale.
1:56:33
People have worked on RLHF for language
models for years, especially with intense interest after ChatGPT.
1:56:37
And the first release of a reasoning model trained with
RLVR, OpenAI's o1, had a scaling plot where if you increase training
compute logarithmically, you get a linear increase in evaluations.
1:56:49
This has been
reproduced multiple times.
1:56:49
DeepSeek had a plot like this.
1:56:53
But there's no
scaling law for RLHF where if you log-increase the compute, you get
performance.
1:56:57
In fact, the seminal scaling paper for RLHF is scaling laws for
reward model over-optimization.
1:57:05
So that's a big line to
draw with RLVR and the methods we have now.
1:57:08
In the future,
they will follow this scaling paradigm: where you can let
the best runs run for an extra 10x and you get performance,
but you can't do this with RLHF.
1:57:20
And that is just going to be
field-defining in how people approach them.
1:57:24
While I'm a shill for
people to academically do RLHF, to do the best RLHF you might not need the extra 10 or 100x of
compute, but to do the best RLVR you do.
1:57:37
I think
there's a seminal paper from a Meta internship.
1:57:41
It's
called something like "The Art of Scaling Reinforcement Learning with
Language Models."
1:57:45
What they describe as a framework is Scale-RL.
1:57:48
Their
incremental experiment was like 10,000 V100 hours,
which is like thousands or tens of thousands of dollars per
experiment.
1:57:56
They do a lot of them, and This cost is not accessible to the average academic, which is a
hard equilibrium where it's trying to figure out how
to learn from each community.
1:58:11
- I was wondering if we could take
a bit of a tangent and talk about education and learning.
1:58:15
If you're someone listening to this who's a smart person
interested in programming and AI, I presume building something
from scratch is a good beginning.
1:58:28
So can you take me through what
you would recommend people do?
1:58:32
- I would personally start, like you
said, implementing a simple model from scratch that you can run on
your computer.
1:58:35
The goal is not, when you build a model from scratch,
to have something for every day use.
1:58:43
It's not going to be your
personal assistant replacing an existing open-weight model or
ChatGPT.
1:58:46
It's to see what exactly goes into the LLM, what comes out,
and how the pre-training works on your own computer, preferably.
1:58:59
Then you learn about pre-training,
supervised fine-tuning, and the attention mechanism.
1:59:02
You get a solid
understanding of how things work, but at some point you reach a limit, because
small models can only do so much.
1:59:11
The problem with learning about
LLMs at scale is that it's exponentially more complex
to make a larger model, because the model isn't just
larger—you have to shard your parameters across multiple GPUs.
1:59:22
Even
for the KV cache, there are multiple ways to implement it.
1:59:26
One is just to
understand how it works, just to grow the cache.
1:59:30
You grow it step-by-step
by, let's say, concatenating lists, but then that
wouldn't be optimal on GPUs.
1:59:38
You would pre-allocate
a tensor and then fill it in.
1:59:42
But that adds another 20 or 30
lines of code.
1:59:42
And for each thing, you add so much code.
1:59:46
The goal with
the book is basically to understand how the LLM works.
1:59:50
It's not going to be a
production-level LLM, but once you have that, you can understand
the production-level LLM.
1:59:56
- So you're trying to always build an
LLM that's going to fit on one GPU? - Yes. Most of them do.
2:00:00
I have
some bonus materials on some MoE models.
2:00:04
One or two of them
may require multiple GPUs, but the goal is to have it on one
GPU.
2:00:08
And the beautiful thing is, you can self-verify. It's almost
like RLVR.
2:00:12
When you code these from scratch, you can take an existing
model from the Hugging Face Transformers library.
2:00:20
The
library is great, but if you want to learn about LLMs, it's not
the best place to start because the code is so complex to
fit so many use cases.
2:00:32
Because people use it in production,
it has to be really sophisticated, really intertwined, and hard
to read. It's not linear.
2:00:39
- It started as a fine-tuning
library, and then it grew to be the standard representation of
every model architecture.
2:00:45
Hugging Face is the default place to get a
model, and Transformers is the software.
2:00:49
It enables it so people can easily load
a model and do something basic with it.
2:00:56
- And all frontier labs that have
open-weight models have a Transformers version of it, like from DeepSeek
to gpt-oss-120b.
2:01:00
That's the canonical weight format
you can load.
2:01:03
But even even Transformers, the library, is
not used in production.
2:01:07
People use SGLang or vLLM, and it adds
another layer of complexity.
2:01:15
- We should say that the Transformers
library has like 400 models.
2:01:19
- So it's the one library that tries
to implement a lot of LLMs, and so you have a huge codebase,
basically. It's huge.
2:01:23
It's like, I don't know, maybe millions— - That's crazy.
2:01:30
- hundreds of thousands of lines
of code.
2:01:30
Understanding the part you want to understand is finding the needle
in the haystack.
2:01:34
But what's beautiful is you have a working implementation, so
you can work backwards.
2:01:38
What I would recommend doing, or what I also do, is if I
want to understand, for example, how OLMo is implemented, I would look at the
weights in the model hub, the config file, and then you can see, "Oh,
they used so many layers.
2:01:50
They use, let's say, Group Query Attention
or Multi-Head Attention in that case."
2:01:58
Then you see all the components in a
human-readable, 100-line config file.
2:02:01
And then you start, let's say, with
your GPT-2 model and add these things.
2:02:05
The cool thing here is you can
then load the pretrained weights and see if they work in your
model.
2:02:09
You want to match the same output that you get with a Transformer
model, and then you can use that, basically as a verifiable reward
to make your architecture correct.
2:02:21
Sometimes it takes me a day.
2:02:21
With
OLMo 3, the challenge was RoPE for the position embeddings.
2:02:25
They had a YaRN extension and there was some custom scaling there, and I couldn't quite match
these things.
2:02:32
In this struggle, you kind of understand things.
2:02:36
At the
end, you know you have it correct because you can unit test it.
2:02:40
You
can check against the reference implementation.
2:02:44
I think that's one
of the best ways to learn, really.
2:02:48
To basically reverse-engineer something.
2:02:51
- I think that is something everyone
interested in getting into AI today should do.
2:02:55
That's why
I liked your book.
2:02:55
I came to language models from the RL and
robotics field.
2:02:59
I had never taken the time to just learn all the fundamentals.
2:03:06
This transformer architecture
is so fundamental, just as deep learning was in the past,
and people need to do this.
2:03:14
I think where a lot of
people get overwhelmed is, "How do I apply this to have
impact or find a career path?"
2:03:22
Because language models make this
fundamental stuff so accessible, and people with motivation
will learn it.
2:03:26
Then it's like, "How do I get cycles on goal
to contribute to research?"
2:03:34
I'm actually fairly optimistic
because the field moves so fast that a lot of times the best
people don't fully solve a problem because there's a bigger problem to
solve that's very low-hanging fruit, so they move on.
2:03:45
I think that a lot
of what I was trying to do in this RLHF book is take post-training
techniques and describe how people think about them influencing
the model and what people are doing.
2:03:57
Then it's remarkable how many things I just think people stop
studying or don't pursue.
2:04:05
I think people trying to go narrow
after doing the fundamentals is good, and then reading the relevant papers and being engaged in the ecosystem.
2:04:12
It's like you actually... actually...
2:04:16
The proximity that
random people online have to the leading researchers—no
one knows who all the...
2:04:23
The anonymous accounts on X and ML are
very popular, and no one knows who all these people are.
2:04:27
It could just be random people
that study this stuff deeply, especially with the AI tools.
2:04:31
To just be like, "I don't
understand this, keep digging into it," is a very useful thing.
2:04:35
But there's
a lot of research areas that maybe have three papers that
you need to read, and then one of the authors will probably email
you back.
2:04:43
But you have to put in a lot of effort into these emails to
understand the field.
2:04:47
I think it would be for a newcomer easily weeks
of work to feel like they can truly grasp what is a very
narrow area, but I think going narrow after you have the fundamentals
will be very useful to people because I've become very interested in
character training, which is how you make the model funny or
sarcastic or serious, and what do you do to the
data to do this?
2:05:11
A student at Oxford reached out to me and was like, "Hey,
I'm interested in this," and I advised him.
2:05:18
And that paper now exists.
2:05:18
There's like two or three people in the world that were very
interested in this.
2:05:22
He's a PhD student, which gives him an advantage, but for me, that was a topic I was waiting for
someone to be like, "Hey, I have time to spend cycles on this."
2:05:32
I'm sure there's a lot more
very narrow things where you're just like, "It doesn't make sense that there was
no answer to this."
2:05:36
I think it's just there's so much information coming that
people are like, "I can't grab onto any of these," but if you just stick in an area, I think
there's a lot of interesting things to learn.
2:05:48
- Yeah, I think you can't try to do it all
because it would be very overwhelming and you would burn out.
2:05:52
For me, for
example, I haven't kept up with computer vision in a long time; I just focused
on LLMs.
2:05:56
But coming back to your book, I think this is a really great book and
a really good bang for the buck because if you want to learn about RLHF, I
wouldn't go out there and read RLHF papers because you would
be spending two years— - Some of them contradict.
2:06:11
I've just edited the book, and there's
no chapter where I had to be like, "X papers say one thing and Y papers
say another, and we'll see what comes out to be true."
2:06:21
- Just to go through the table of contents, what
are some ideas we might have missed in the bigger picture of post-training?
2:06:25
First of
all, you did the problem setup, training overview, what are preferences,
preference data, and the optimization tools, reward modeling,
regularization, instruction tuning, rejection sampling, and
reinforcement learning.
2:06:36
Then, Constitutional AI and AI feedback,
reasoning, and inference-time scaling to use in function calling,
synthetic data and distillation, evaluation, and then an
open questions section, over-optimization, style and
information, and then product UX, character and post-training.
2:06:58
What are some ideas worth mentioning that connect both the
educational and the research components?
2:07:06
You mentioned character training,
which is pretty interesting.
2:07:08
- Character training is interesting because there's
so little on it.
2:07:08
We talked about how people engage with these models.
2:07:12
We
feel good using them because they're positive, but that can go too far;
it can be too positive.
2:07:16
And it's like, essentially, it's: How
do you change your data and decision-making to make it
exactly what you want?
2:07:23
And like, OpenAI has this thing called a model
spec, which is essentially their internal guideline for what they want the
model to do, and they publish this to developers.
2:07:34
So, essentially, you
can know what is a failure of OpenAI's training—where they have the
intentions and they haven't met them yet— versus what is something that they actually
wanted to do and that you don't like.
2:07:46
And that transparency is very nice,
but all the methods for curating these documents and how easy it is to follow
them is not very well known.
2:07:49
I think the way the book is designed is that the RL
chapter is obviously what people want because everybody hears about it with RLVR, and
it's the same algorithms and the same math, but you can use it in
very different documents.
2:08:05
So I think the core of
RLHF is like how messy preferences are.
2:08:09
It's essentially
a rehash of a paper I wrote years ago, but this is essentially
the chapter that'll tell you why RLHF is never ever fully solvable because, the way that even RL is set up, it
assumes that preferences can be quantified and that multiple
preferences can be reduced to single values.
2:08:32
And I think it relates
in the economics literature to the Von Neumann-Morgenstern utility theorem,
and that is the chapter where all of that philosophical,
economic, and psychological context tells you what gets compressed
into doing RLHF.
2:08:44
So it's like you have all of this and then later in the book
it's like: You use this RL map to make the number go up.
2:08:52
And I think that's why
it'll be very rewarding for people to do research on, because quantifying
preferences is something that humans have designed
a problem in order to make preferences studyable.
2:09:03
But there's
kind of fundamental debates, like, an example is in a language model response you
have different things you care about, like accuracy or style.
2:09:12
And when you're
collecting the data, they all get compressed into: "I like this more
than another."
2:09:15
And that is happening, and there's a lot of
research in other areas of the world that go into how you
should actually do this.
2:09:23
I think social choice theory is the subfield of economics around how
you should aggregate preferences.
2:09:33
And I went to a workshop
that published a white paper on: "How can you think about
using social choice theory for RLHF?"
2:09:40
So I mostly would want people
that get excited about the math to come and find things where they
could stumble into this broader context.
2:09:48
I think there's a fun thing:
I just keep a list of all the tech reports of reasoning models
I like.
2:09:52
So in Chapter 14, where there's a short summary of RLVR,
there's just a gigantic table where I list every single reasoning
model that I like.
2:10:03
I think in education, a lot of it needs
to be like, at this point, what I like, because the language models are so
good at the math.
2:10:08
For example, the famous paper, Direct Preference Optimization,
which is a much simpler way of solving the problem than RL.
2:10:16
The derivations
in the appendix skip steps of math.
2:10:20
And for this book, I redid
the derivations and I'm like, "What the heck is this log trick
that they use to change the math?"
2:10:27
But doing it with language models, they're
like, "This is the log trick."
2:10:27
And I'm like, "I don't know if I like this,
that the math is so commoditized."
2:10:34
I think some of the struggle
in reading this appendix- ...
2:10:38
and following the math
is good for learning.
2:10:43
- Yeah, we're returning to this
often on the topic of education.
2:10:47
You both have brought up the
word "struggle" quite a bit. So there is value.
2:10:51
If you're not
struggling as part of this process, you're not fully following the
proper process for learning.
2:10:59
proper process for learning, I suppose.
2:11:02
- Some providers are working on models
for education designed to not give- actually, I haven't used them, but I'd
guess they're designed to not give all the information at once.
2:11:12
And make people work for it.
2:11:12
Training models
to do this would be a wonderful contribution.
2:11:16
Where, like all of the stuff in the book,
you had to reevaluate every decision.
2:11:19
decision for it- It's a great example.
2:11:19
There's a chance we work on it at Ai2, which I thought would be so fun. - It makes sense.
2:11:26
I did something like
that the other day for video games.
2:11:30
Sometimes for pastime I play
video games, like I like- Video games with puzzles, like
Zelda and Metroid.
2:11:34
And there's this new game where I really got
stuck and was okay with it.
2:11:41
I don't want to struggle for
two days, so I used an LLM.
2:11:45
But then you say, "Hey, please don't add
spoilers.
2:11:45
Just, you know, I'm here and there.
2:11:49
What do I have to do next?"
2:11:49
You can do
the same thing for math where you say, "Okay, I'm stuck at this point.
2:11:53
Don't give me the full solution, but what is something I could try?"
2:11:57
Where you carefully probe it.
2:12:01
But the problem here is I
think it requires discipline.
2:12:05
Many people enjoy math, but
there are also a lot of people who need to do it for their homework,
and then it's like a shortcut.
2:12:13
We could develop an educational LLM,
but other LLMs are still there, and there's still a temptation
to use the other LLMs.
2:12:20
- I think many people in college understand
the stuff they're passionate about- about- ...
2:12:24
they're self-aware and they
understand it shouldn't be easy.
2:12:27
I think we just have to
develop a good taste- ...
2:12:31
talk about research
taste, school taste about stuff that you should be struggling on- ...
2:12:37
and stuff you shouldn't be.
2:12:37
It's tricky,
because you don't have good long-term vision sometimes you don't have
good long-term vision about what would be actually
useful to you in your career.
2:12:48
But you have to develop that taste, yeah.
2:12:51
- I was talking to my fiancee
or friends about this, there's this brief 10-year window
where all of the homework and all the exams could be digital.
2:12:59
Before that, everybody
had to do all the exams in blue books because there was no other way.
2:13:03
And now after AI,
everyone's going to need to be in blue books and oral exams because everyone could cheat so easily.
2:13:07
It's like this brief generation that had a different education system
where everything could be digital, but you still couldn't
cheat.
2:13:15
And now it's just going back. It's just very funny.
2:13:20
- You mention character training.
2:13:20
Just zooming out on a more general topic, for that project
how much compute was required?
2:13:27
And in general,
to contribute as a researcher, are there places where not too much compute is required where you can
actually contribute as an individual researcher?
2:13:39
- For the character training thing, I think
this research is built on fine-tuning about 7 billion parameter
models with LoRA, which is essentially only fine-tuning a small
subset of the weights of the model.
2:13:51
I don't know exactly how many
GPU hours that would take. - But it's doable.
2:13:55
- Not doable for every academic.
2:13:55
The
situation for some academics is so dire that the only work you can do is doing
inference where you have closed models or open models and you get completions from them
and you can look at them and understand the models.
2:14:07
And that's very well-suited
to evaluation, where you want to be the best at creating representative problems that the models fail on or
show certain abilities, which I think that you can break through
with this.
2:14:18
I think that the top-end goal for a researcher
working on evaluation, if you want to have career momentum, is
that Frontier Labs pick up your evaluation.
2:14:29
You don't need to have
every project do this.
2:14:29
But if you go from a small university with no compute and
find something that Claude struggles with, and then the next Claude model has
it in the blog post, there's your career rocket ship.
2:14:41
I think
that's hard, but if you want to scope the maximum possible impact
with minimum compute, it's something like that, which is just get very narrow
and it takes learning of where the models are going.
2:14:53
So you need
to build a tool that tests where Claude 4. 5 will fail.
2:14:57
If I'm going to start a research project, I need to think where the
models in eight months are going to be struggling.
2:15:05
- But what about developing
totally novel ideas? - This is a trade-off.
2:15:08
I think that if you're
doing a PhD, you could also be like, "It's too risky to work in language
models.
2:15:13
I'm going way longer term," which is like what is— what is the thing that's going to define
language model development in 10 years?
2:15:22
I end up being a person that's
pretty practical.
2:15:22
I mean, I went to my PhD where it was like, "I got into Berkeley.
2:15:26
Worst case, I get a master's, and then I go work in tech."
2:15:30
I'm very
practical about it, so I'm like the life afforded to people to work at
these AI companies, the amount of...
2:15:38
OpenAI's average compensation is over
a million dollars in stock a year per employee.
2:15:42
For any normal
person in the US, to get into this AI lab is transformative for
your life.
2:15:46
So I'm pretty practical about it.
2:15:50
there's still a lot of upward mobility
working in language models if you're focused. And look at these jobs.
2:15:53
But from a research perspective, the transformative impact in these academic awards...
2:16:00
to be the next
Yann LeCun is from not working on— not caring about language
model development very much.
2:16:07
- It's a big financial
sacrifice in that case.
2:16:09
- So I work with some awesome students, and
they're like, "Should I go work at an AI lab?"
2:16:13
And I'm like, "You're getting
a PhD at a top school.
2:16:13
Are you gonna leave to go to a lab?" I don't
know.
2:16:17
If you go work at a top lab, I don't blame you.
2:16:21
Don't go work at some random
startup that might go to zero.
2:16:21
But if you're going to OpenAI, I'm like, "It could
be worth leaving a PhD for."
2:16:30
- Let's more rigorously think
through this.
2:16:30
So where would you give a recommendation for people to do
a research contribution?
2:16:33
So the options are academia: get a PhD.
2:16:37
Spend five years publishing.
2:16:44
Compute resources are
constrained.
2:16:44
There's— there's research labs that are more
focused on open-weight models, and working there.
2:16:54
Or closed
frontier research labs.
2:17:01
So OpenAI, Anthropic, xAI, and so on.
2:17:04
- The two gradients are: the more closed,
the more money you tend to get, but you also get less
credit.
2:17:08
In terms of building a portfolio of things that
you've done, it's very clear what you have done as an academic.
2:17:18
Versus if you are going to trade this fairly reasonable progression for being a cog in the
machine, which could also be very fun.
2:17:30
So I think it's very different
career paths.
2:17:30
But the opportunity cost for being a researcher
is very high because PhD students are paid essentially nothing.
2:17:37
So it ends
up rewarding people that have a fairly stable safety net, and they
realize that they can operate in the long term, wanting to do very interesting
work and get a very interesting job.
2:17:49
So it is a privileged position to be like, "I'm gonna see out my PhD and figure it
out after because I want to do this."
2:17:59
At the same time, the academic ecosystem
is getting bombarded by funding getting cut and stuff.
2:18:02
So there's just
so many different trade-offs where I understand plenty of people that are like,
"I don't enjoy it.
2:18:06
I can't deal with this funding search.
2:18:10
My grant got cut
for no reason by the government," or, "I don't know what's gonna
happen."
2:18:14
So I think there's a lot of uncertainty and trade-offs
that, in my opinion, favor just taking the well-paying job with
meaningful impact.
2:18:21
It's not like you're getting paid to sit around at OpenAI.
2:18:25
You're building the cutting edge of things that are— changing millions of
people's relationship to tech.
2:18:34
- But publication-wise, they're being
more secretive, increasingly so.
2:18:38
So you're publishing less and
less.
2:18:38
And so you are having a positive impact at scale, but
you're a cog in the machine.
2:18:47
- I think it honestly
hasn't changed that much. I have been in academia.
2:18:51
I'm
not in academia anymore.
2:18:55
wouldn't want to miss my time in academia.
2:18:55
But what I wanted to say before I get to that is that I think it hasn't changed
that much.
2:18:59
I was working in computational biology, using
AI or machine learning methods with collaborators, and a lot of
people went from academia directly to Google.
2:19:15
And I think it's the
same.
2:19:15
Back then, professors were sad that their students went into industry because they couldn't carry on their
legacy. I think it's the same.
2:19:26
It hasn't changed that much.
2:19:26
The only thing that has changed is the scale.
2:19:30
Cool
stuff was always developed in industry that was closed.
2:19:34
You couldn't talk about it.
2:19:38
And I think the difference
now is your preference.
2:19:42
Do you like to publish your
work, or are you more in a closed lab? That's one difference.
2:19:46
The compensation, of course, is another, but it's
always been like that.
2:19:54
It depends on where you feel
comfortable. And nothing is forever.
2:19:58
Right now, there's
a third option, which is launching a startup.
2:20:02
A lot
of people are doing that.
2:20:05
It's a very risky move, but it can be a high-risk, high-reward situation,
whereas joining an industry lab is pretty safe.
2:20:13
You also
have upward mobility.
2:20:17
I think once you've been at an
industry lab, it's easier to find future jobs.
2:20:21
But then again,
how much do you enjoy the team and working on proprietary
things versus how much you like publishing work?
2:20:32
I mean,
publishing is stressful.
2:20:35
Acceptance rates at
conferences can be arbitrary and very frustrating, but it's
high reward if you have a paper published.
2:20:43
You feel good because your name
is on there.
2:20:43
It's a high accomplishment.
2:20:48
- I feel like my friends who are professors
seem happier than those who work at a frontier lab, to be honest.
2:20:52
There's a grounding there.
2:20:57
The frontier labs
definitely do this 9-9-6, which is shorthand for
working all the time.
2:21:03
- Can you describe 9-9-6?
2:21:03
It's a
culture invented, I believe, in China and adopted in Silicon Valley. What is 9-9-6?
2:21:07
It's 9:00 AM to 9:00 PM, - Six days a week. - six days a week. What is that, 72 hours? Okay.
2:21:15
So, is this basically the standard in AI companies in Silicon Valley?
2:21:21
This kind of grind mindset.
2:21:26
- Yeah, I mean, maybe not exactly like that,
but I think there is a trend towards it. And it's interesting.
2:21:30
I think it
almost flipped because when I was in in academia, I felt like that.
2:21:34
As
a professor, you write grants, you teach, and you do research.
2:21:38
It's like three jobs in one, and it's more than a full-time
job if you want to be successful. successful.
2:21:45
And I feel like
now, like Nathan just said, the professors, in comparison
to a lab, I think they have less pressure or workload than
at a frontier lab because— - I think they work a lot.
2:21:57
They're just
so fulfilled.
2:21:57
By working with students— and having a constant runway of
mentorship and a mission that is very people-oriented, I think in a
era when things are moving very fast and are very chaotic, it's
very rewarding to people.
2:22:11
- Yeah, and I think at a startup,
it's this pressure.
2:22:11
It's like you have to make it.
2:22:15
And it's really
important that people put in the time, but it is really hard because you
have to deliver constantly, and I've been at a startup.
2:22:23
I had a good
time, but I don't know if I could do it forever.
2:22:27
It's an interesting pace and it's exactly like we talked about in the
beginning.
2:22:31
These models are leapfrogging each other, and they are just
constantly trying to take the next step compared to their competitors.
2:22:38
It's just ruthless right now.
2:22:42
- I think this leapfrogging nature and
having multiple players is actually an underrated driver of language modeling
progress where competition is so deeply ingrained in people, and these companies have intentionally created
very strong cultures.
2:22:53
Like, Anthropic is known to be so culturally, like, deeply committed and organized.
2:23:01
I mean,
we hear so little from them, and everybody at Anthropic seems very
aligned.
2:23:05
And it's like being in a culture that is super tight and having this competitive
dynamic is a thing that's gonna make you work hard and create
things that are better.
2:23:20
But that comes at the cost of
human capital, which is like you can only do this for so long, and
people are definitely burning out.
2:23:28
I wrote a post on burnout
as I've tread in and out of this myself, especially trying to
be a manager, full-mode training.
2:23:36
It's a crazy job doing this.
2:23:36
The
book Apple in China by Patrick McGee, he talked about how hard the Apple
engineers worked to set up the supply chains in China, and he was like, they
had "saving marriage" programs, and he told in a podcast, he
was like, "People died from this level of working
hard."
2:23:51
So I think it's just like it's a perfect environment for
creating progress based on human expense, and there's
gonna be a lot of...
2:23:59
the human expense is the 996 that we
started this with, which is like— ... people do really grind. - I also read this book.
2:24:08
I think they had a
code word for if someone had to go home to spend time with their family to save
the marriage, and it's crazy.
2:24:12
Then the colleagues said, "Okay, this is like red
alert for this situation.
2:24:16
We have to let that person go home this weekend."
2:24:20
But at the same time, I don't think they were forced to work.
2:24:24
They were so
passionate about the product, I guess, that you get into that mindset.
2:24:28
And
I had that sometimes as an academic, but also as an independent
person, I have that sometimes.
2:24:32
I overwork, and it's unhealthy.
2:24:36
I had back
issues, I had neck issues, because I did not take the breaks that I maybe
should have taken.
2:24:40
But no one forced me to; it's because I wanted to
work, because it's exciting stuff.
2:24:46
- That's what OpenAI and Anthropic are
like.
2:24:46
They want to do this work.
2:24:49
- Yeah, but there's also a
feeling of fervor that's building, especially in Silicon Valley,
aligned with the scaling laws idea, where there's this hype where the world
will be transformed in a scale of weeks and you want to be at the
center of it.
2:25:00
And then, you know, I have this great fortune of having conversations with a wide
variety of human beings, and from there I get to see all these
bubbles and echo chambers across the world.
2:25:16
It's fascinating to see how
we humans form them.
2:25:16
And I think it's fair to say that Silicon Valley is a
kind of echo chamber, a kind of silo and bubble.
2:25:27
I think bubbles are
actually really useful and effective.
2:25:31
It's not necessarily a negative
thing because you could be ultra-productive.
2:25:34
It could be the Steve Jobs reality distortion field, because
you just convince each other that breakthroughs are imminent, and
by convincing each other of that, you make the breakthroughs imminent.
2:25:48
- Byrne Hobart wrote a book classifying
bubbles.
2:25:48
One of them is financial bubbles, which is like speculation, which is bad,
and the other one is for build-outs, because it pushes people to build
these things.
2:25:56
And I do think AI is in this, but I worry about it transitioning
to a financial bubble, which is - Yeah, but also in the space of ideas, that bubble—you are doing a reality distortion field, and that means you are
deviating from reality.
2:26:11
And if you go too far from reality while
also working, you know, 996, you might miss some fundamental
aspects of the human experience, including beyond Silicon Valley.
2:26:26
This
is a common problem in Silicon Valley: it's a very specific geographic
area.
2:26:30
You might not understand the Midwest perspective, the full experience of all the other humans in
the United States and across the world, and you speak a certain way to each other,
you convince each other of a certain thing, and that can get you
into real trouble.
2:26:44
Whether AI is a big success and becomes
a powerful technology or it's not, in either trajectory
you can get yourself into trouble.
2:26:55
So you have to consider
all of that.
2:26:55
Here you are, a young person trying to decide what
you want to do with your life. - The thing that is...
2:27:02
I don't even
really understand this, but the SF AI memes have gotten to the point
where "permanent underclass" was one of them, which was the idea
that the last six months of 2025 was the only time to
build durable value in an AI startup or model.
2:27:17
Otherwise, all
the value will be captured by existing companies and you will
therefore be poor, which...
2:27:24
that's an example of the SF
thing that goes so far.
2:27:24
I still think for young people going to be
able to tap into it, if you're really passionate about wanting to have an
impact in AI, being physically in SF is the most likely place where
you're going to do this.
2:27:36
But it has has trade-offs.
2:27:41
- I think SF is an incredible
place, but there is a bit of a bubble.
2:27:45
And if you go into
that bubble, which is extremely valuable, just get out also.
2:27:49
Read history books, read literature, visit
other places in the world.
2:27:57
Twitter and Substack are
not the entire world.
2:28:01
- I would say, one of the people I worked
with is moving to SF, and it's like, I need to get him a copy of Season of the
Witch, which is a history of SF from 1960 to 1985, which goes through
the hippie revolution, like all the gays taking over the city and that
culture emerging, and then the HIV/AIDS crisis and other things.
2:28:20
And
it's just like, that is so recent, and so much turmoil and hurt,
but also love in SF.
2:28:28
And it's like, no one knows about this.
2:28:28
It's a
great book, Season of the Witch. I recommend it.
2:28:32
A bunch of my SF friends who get out recommended it to me.
2:28:35
And
I think that's just like living there...
2:28:39
I lived there and I didn't
appreciate this context, and it's just so recent. - Yeah. Okay, let's...
2:28:46
We talked a lot
about a lot of things.
2:28:46
Certainly about the things that were exciting
last year.
2:28:54
But this year, One of the things you guys mentioned
that's exciting is the scaling of text diffusion models, and just a different
exploration of text diffusion.
2:29:02
Can you talk about what that is and what
the possibility it holds?
2:29:09
So, different kinds of
approaches than the current LMs?
2:29:13
- Yeah, so we talked a lot about the transformer
architecture and the autoregressive transformer architecture specifically,
like GPT.
2:29:17
And it doesn't mean no one else is working on anything else.
2:29:21
So, people are always on the, let's say, lookout for the next big
thing.
2:29:25
Because I think it would be almost stupid not to.
2:29:28
Because sure,
right now, the transformer architecture is the thing, and it works
best, and there's, right now, nothing else out there.
2:29:36
But, you know, it's always a
good idea to not put all your eggs into one basket.
2:29:40
So, people are developing
other alternatives to the autoregressive transformer.
2:29:44
One of
them would be, for example, text diffusion models.
2:29:48
And listeners may
know diffusion models from image generation, like Stable Diffusion
popularized it.
2:29:52
There was a paper on generating images.
2:29:56
Back then, people
used GANs, Generative Adversarial Networks.
2:30:00
And then there was this
diffusion process where you iteratively denoise an image, and that resulted
in really good quality images over time.
2:30:07
Stable Diffusion was a company.
2:30:07
Other
companies build their own diffusion models.
2:30:11
And then people are now like, "Okay,
can we try this also for text?"
2:30:15
Doesn't, you know, make intuitive sense
yet, because it feels like, okay, it's not something continuous like a pixel that we can
differentiate.
2:30:19
It's discrete text, so how do we implement that denoising process?
2:30:23
It's
kind of similar to the BERT models by Google.
2:30:30
Like, when you go back to the original
transformer, they were the encoder and the decoder.
2:30:34
The decoder is what we
are using right now in GPT and so forth.
2:30:38
The encoder is more like a parallel technique where you
have multiple tokens that you fill in in parallel.
2:30:45
GPT models,
they do autoregressive generation, completing the sentence one
token at a time.
2:30:49
And in BERT models, you have a
sentence that has gaps.
2:30:57
You mask them out, and then
one iteration is filling in these gaps.
2:31:01
Text diffusion is
kind of like that, where you are starting with some random
text, and then you are filling in the missing parts or refining
them iteratively over multiple iterations.
2:31:12
And the cool thing here
is that this can do multiple tokens at the
same time.
2:31:16
It's like the promise of having it more efficient.
2:31:19
Now, the trade-off is, of course, how good is the quality?
2:31:23
It
might be faster, and now you have this dimension of the denoising process.
2:31:27
The more steps you do, the better the text becomes. And people...
2:31:34
I mean, you can scale in different
ways.
2:31:34
They try to see if that is maybe a valid alternative to
the autoregressive model in terms of giving you the same
quality for less compute.
2:31:46
Right now, there are papers that
suggest if you want to get the same quality, you have to crank up
the denoising steps, and then you end up spending the same compute you
would spend on an autoregressive model.
2:31:58
The other downside is, while
being parallel sounds appealing, some tasks are not parallel.
2:32:01
Like
reasoning tasks or tool use, maybe where you have to ask a
code interpreter to give you an intermediate result.
2:32:09
That is tricky
with diffusion models.
2:32:09
So, there are some hybrids, but the main idea is: how
can we parallelize it?
2:32:13
It's an interesting avenue.
2:32:17
I think
right now, there are mostly research models out there, like
LaMDA and some other ones.
2:32:24
I saw some by startups, some
deployed models.
2:32:24
There is no big diffusion model at scale yet,
like on the Gemini or ChatGPT level.
2:32:32
But there was an announcement by Google, a site where they said they are
launching Gemini Diffusion, and they put it into context of their Gemini Nano 2 model, and they said basically:
for the same quality on most benchmarks, we can generate
things much faster.
2:32:47
You mentioned what's next.
2:32:51
I don't think the
text diffusion model is going to replace autoregressive LLMs, but it
will be something maybe for quick, cheap, at-scale
tasks.
2:32:58
Maybe the free tier in the future will
be something like that.
2:33:08
- I think there are examples where it's already
being used.
2:33:08
To paint an example of why this is better, for example, when
GPT-5 is taking 30 minutes to respond, it's generating one token at
a time.
2:33:15
And this diffusion idea is essentially to generate all of
those tokens and the completion in one batch, which is why it could be
way faster.
2:33:23
And I think it could be suited for...
2:33:26
the startups I'm hearing
about are code startups where you have a code base, and you have somebody that's
effectively "vibe coding," and they say, "Make this change."
2:33:34
And a
code diff is essentially a huge reply from the model, but
it doesn't have to have that much external context, and you can get it
really fast by using these diffusion models.
2:33:45
One example I've heard is that they
use text diffusion to generate really long diffs, because doing it
with an autoregressive model would take minutes, and that time for a user-facing
product causes a lot of churn.
2:33:57
Every second, you lose a lot of users.
2:33:57
So, I think it's going to be this thing where it's going to— ...
2:34:02
grow and have some applications, but I
actually thought that different types of models were going to be used for different
things much sooner than they have been, so I kind of trade off.
2:34:10
I think the
tool-use point is the one that's stopping them from being most general purpose
because, for Claude Code and ChatGPT search, the autoregressive chain is interrupted
with some external tool, and I don't know how to do that
with the diffusion setup.
2:34:28
- So what's the future of tool use this
year and then in the coming years?
2:34:28
Do you think there's going to be a lot of developments
there, and how that's integrated into the entire stack?
2:34:37
- I do think right now, it's mostly
on the proprietary LLM side, but I think we will see more of that
in the open-source tooling.
2:34:41
And I think it is a huge unlock
because then you can really outsource certain tasks from
just memorization to actual— you know, instead of having the
LLM memorize what is 23 plus 5, just use a calculator.
2:34:58
- So do you think that can
help solve hallucination?
2:35:01
- Not solve it, but reduce it.
2:35:01
So the LLM still needs to know when to ask for a tool call.
2:35:06
And the second one is, well, it doesn't mean the
internet is always correct.
2:35:09
You can do a web search, but let's say I
asked who won the World Cup in, let's say, 1998; it still needs to
find the right website and get the right information.
2:35:21
You can still go to the
incorrect website and give me incorrect information.
2:35:24
So I don't think it
will fully solve that, but it is improving it in that sense.
2:35:31
And so another cool paper
earlier this year—I think it was December 31st, so it's
not technically 2026, but close—the recursive language model.
2:35:43
That's a cool idea to kind of
take this even a bit further.
2:35:47
Just to explain, Nathan,
you also mentioned earlier, it's harder to do cool research in academia
because of the compute budget.
2:35:51
If I recall correctly, they did everything with
GPT-5, so they didn't even use local models, but the idea is, let's say you have a
long-context task; instead of having the LLM solve all of it in one
shot or even in a chain, you break it down into sub-tasks.
2:36:06
You have the LLM decide what is a good sub-task, and then recursively call an LLM to solve
that.
2:36:14
And I think something like that, adding tools—you know,
each one maybe you have a huge Q&A task, so each one
goes to the web and gathers information, and then you pull it together
at the end and stitch it back together.
2:36:29
I think there's going to be a lot of
unlock using things like that where you don't necessarily improve the LLM
itself; you improve how the LLM is used and what the LLM can use.
2:36:39
One
downside right now with tool use is you have to give the LLM
permission to use tools.
2:36:43
And that will take some trust, especially
if you want to unlock things like having an LLM answer emails for
you—or not even answer, but just sort them for you or select them for you or
something like that.
2:36:54
I don't know if I would today give an LLM access to my emails,
right?
2:36:58
I mean, this is a huge risk.
2:37:03
- I think there's a cool...
2:37:03
one last point
on the tool use thing.
2:37:03
I think that you hinted at this, and we've both come at
this in our own ways, is that the open versus closed models use tools in very different
ways, where open models, people go to Hugging Face and download the model, and then the
person's going to be like, "What tool do I want?"
2:37:20
I don't know, Exa is my preferred search
provider, but somebody else might care for a different search startup.
2:37:24
Where you release
a model, it needs to be useful for multiple tools, for multiple use cases, which
is really hard because you're making a general reasoning engine model, which
is actually what gpt-oss-120b is good for.
2:37:35
But on the closed
models, you're deeply integrating the specific tool into your experience,
and I think that open models will struggle to replicate some of the
things that I like to do with closed models, which will be like, you can
reference a mix of public and private information.
2:37:51
And something that I keep
trying every three to six months, I try Claude Code on the web,
which is just prompting a model to make an update to some GitHub repository
that I have.
2:37:59
And it's just like that set of secure cloud
environments is just so nice for just sending it off to do this
thing and then come back to me, and these will probably help define
some of the local open and closed niches.
2:38:18
But I think initially, because
there was such a rush to get tool use working, the open models were on the back foot,
which is kind of inevitable.
2:38:22
I think there's so much research, so many resources in
these frontier labs, but it will be fun when the open models solve this
because it's going to necessitate a bit more flexible and potentially interesting
model that might work with this recursive idea to be an orchestrator
and a tool use model, so hopefully the necessity drives
some interesting innovation there.
2:38:45
- So, continual learning—this
is a longstanding topic, important problem.
2:38:50
I think
that increases in importance as the cost of training the models
goes up.
2:38:53
So can you explain what continual learning is and how important
it might be this year and in the coming years to make progress?
2:39:03
- This relates a lot to this kind
of SF zeitgeist of, what is AGI, which is Artificial General
Intelligence, and what is ASI, Artificial Superintelligence, and
what are the language models that we have today capable of doing?
2:39:13
I think
the language models can solve a lot of tasks, but a key milestone
among the AI community is essentially when AI could
replace any remote worker, taking in information and solving
digital tasks and doing them.
2:39:25
And the limitation that's
highlighted by people is that a language model will not learn
from feedback the same way that an employee does.
2:39:36
So if you hire an
editor, the editor will mess up, but you will tell them.
2:39:40
And if you hired a good editor,
they don't do it again.
2:39:40
But language models don't have this ability to modify themselves and learn very
quickly.
2:39:45
So the idea is, if we are going to actually get to something that
is a true, general adaptable intelligence that can go into any remote
work scenario, it needs to be able to learn quickly from feedback
and on-the-job learning.
2:40:00
I'm personally more bullish on language
models being able to just provide them with very good context.
2:40:04
You said, maybe offline, that you can write extensive documents
to models where you say, "I have all this information.
2:40:11
Here are
all the blog posts I've ever written.
2:40:15
I like this type of writing.
2:40:15
My voice is
based on this."
2:40:15
But many people don't provide this to models, and the models
weren't designed to take this amount of context previously.
2:40:23
Agentic
models are just starting.
2:40:23
So it's this kind of trade-off: do we need
to update the weights of this model with this continual learning thing
to make them learn fast?
2:40:31
Or the counterargument is we just need to provide
them with more context and information, and they will have the appearance of learning
fast by having a lot of context and being smart.
2:40:43
- So we should mention the terminology
here.
2:40:43
Continual learning refers to changing the weights continuously so
that the model adapts and adjusts based on the new incoming information, doing so continually,
rapidly, and frequently.
2:41:00
And then the thing you mentioned
on the other side of it generally will be referred to
as in-context learning.
2:41:04
As you learn stuff, there's a
huge context window.
2:41:08
You can just keep loading it with extra information
every time you prompt the system, which I think both legitimately
can be seen as learning.
2:41:21
It's just a different place
where you're doing the learning.
2:41:24
- I think, to be honest with you,
continual learning — updating weights — we already have that in
different flavors.
2:41:28
If you think about how...
2:41:31
I think the
distinction here is: do you do that on a personalized custom
model for each person, or do you do it on a global model
scale?
2:41:39
I think we have that already, going from GPT-5 to 5. 1 and 5. 2.
2:41:43
It's maybe not immediate, but it is a
curated update, a quick curated update where there was feedback about things
they couldn't do, feedback by the community.
2:41:55
They updated the weights,
next model, and so forth.
2:41:55
So it is a flavor of that.
2:41:58
Another even
finer-grained example is like RLVR; you run it, it
updates.
2:42:06
The problem is you can't just do that for each person because
it would be too expensive to update the weights for each person, and I think
that's the problem. Unless you get...
2:42:17
Even at OpenAI scale, building
the data centers, it would be too expensive.
2:42:21
I think that is only
feasible once you have something on the device where the cost is on the
consumer.
2:42:25
Like what Apple tried to do with the Apple Foundation models, putting them
on the phone, where they learn from experience.
2:42:33
- A bit of a related topic, but this kind
of, maybe anthropomorphized term: memory.
2:42:42
What are different ideas for the mechanism of
how to add memory to these systems as we're increasingly seeing?
2:42:46
Personalized memory especially?
2:42:49
- Right now, it's mostly basically stuffing things into the context
and then just recalling that.
2:42:56
But again, I think it's expensive because
you have to—you can cache it, but still you spend tokens on that.
2:43:04
And the
second one is you can only do so much.
2:43:08
I think it's more like a preference
or style.
2:43:08
I mean, a lot of people do that when they solve math
problems.
2:43:12
You say it's way so you can add previous knowledge and stuff,
but you also give it certain preference prompts: "do what I preferred last
time," or something like that.
2:43:20
But it doesn't unlock new capabilities.
2:43:23
So
for that, one thing people still use is LoRA adapters.
2:43:31
These are basically,
instead of updating the whole weight matrix, there are two smaller weight
matrices that you kind of have in parallel or overlays like
the delta.
2:43:38
But yeah, you can do that to some extent, but then again, it is economics.
2:43:45
There
were also papers, for example, LoRA learns
less but forgets less.
2:43:53
It's like, there's no free lunch.
2:43:53
If you
want to learn more, you need to use more weights, but it gets more expensive.
2:43:57
And
then again, if you learn more, you forget more, and you have to find that
Goldilocks zone basically.
2:44:04
- We haven't really mentioned it much,
but implied in this discussion is context length also.
2:44:08
Is there a lot
of innovation that's possible there?
2:44:13
- I think the colloquially accepted
thing is that it's a compute and data problem where you can...
2:44:17
and sometimes small architecture things like attention
variants.
2:44:21
We talked about hybrid attention models, which
is essentially if you have what looks like a state space model within
your transformer.
2:44:28
And those are better suited because you have to
spend less compute to model the furthest along token.
2:44:37
I think that, those aren't free because they have to
be accompanied by a lot of compute or the right data.
2:44:47
How many sequences
of 100,000 tokens do you have in the world, and where do you get
these?
2:44:51
It just ends up being pretty expensive to scale them.
2:44:54
We've gotten
pretty quickly to a million tokens of input context length.
2:44:58
I would
expect it to keep increasing and get to 2 million or 5 million this year,
but I don't expect it to go to 100 million.
2:45:06
That would be like a true
breakthrough, and I think those breakthroughs are possible.
2:45:10
I think of the continual
learning thing as a research problem where there could be a breakthrough that
just makes transformers work way better at this and it's cheap.
2:45:18
These things
could happen with so much scientific attention.
2:45:22
But turning the crank, it'll
be consistent increases over time.
2:45:27
- Looking at the extremes, I think there's,
again, no free lunch.
2:45:27
So, the one extreme to make it cheap: you have, let's
say, an RNN that has a single state where you save everything from
the previous stuff.
2:45:35
It's like a specific fixed-size thing,
so you never really grow the memory because you are stuffing
everything into one state, but then the longer the context gets, the
more information you forget because you can't compress everything
into one state.
2:45:50
Then on the other hand, you have the transformers, which
try to remember every token, which is great sometimes if you want to look up specific
information, but very expensive because you have the KV cache that grows, the
dot product that grows.
2:46:02
But then, like you said, the Mamba layers—I
mean, they kind of have the same problem.
2:46:10
Like an RNN, you try to compress everything
into one state; you're a bit more selective there.
2:46:14
But then I think it's like
this Goldilocks zone again.
2:46:14
With Nemotron 3, they found a good
ratio of how many attention layers do you need for the global
information where everything is accessible compared to having these compressed
states.
2:46:25
And I think that's how we will scale more—by finding better, let's say, ratios in the
Goldilocks zone, like between making computing cheap enough
to run, but then also making it powerful enough to be
useful.
2:46:41
And one more plug here, the Recursive Language Model
paper, that is one of the papers that tries to kind of address the
long context thing.
2:46:48
So what they found is essentially instead of stuffing
everything into this long context if you break it up into
multiple smaller tasks, so you save memory by having multiple
smaller cores, you can actually get better accuracy than having the
LLM try everything all at once.
2:47:08
I mean, it's a new paradigm.
2:47:08
We will see,
you know, there might be other flavors of that.
2:47:12
So I think with that, we
will still make improvement on long context, but then also, like
Nathan said, I think the problem is for pre-training itself, we don't have as
many long context documents as other documents.
2:47:24
So it's harder
to study basically how LLMs behave and stuff
like that on that level.
2:47:31
- There are some rules of thumb where essentially
you pre-train a language model, like OLMo.
2:47:35
we pre-trained at like 8K context
length and then extended to 32K with training.
2:47:38
And there are some rules
of thumb where you're essentially doubling the training context length, it takes
like 2X compute, and then you can normally like 2 to 4X the
context length again.
2:47:50
So I think a lot of it ends up being
kind of compute bound at pre-training, which is in this...
2:47:54
Like we talked about,
everyone talks about this big increase in compute for the top labs this year, and
that should reflect in some longer context windows.
2:48:01
But I think on the post-training side,
there are some more interesting things.
2:48:01
As we have agents, the agents are gonna manage
this context on their own, where now agents, people that use Claude Code a lot dread
the compaction, which is when Claude takes its entire full 100,000 tokens of work
and compacts it into a bulleted list.
2:48:17
But what the next
models will do—and I'm sure people are already working on this—is
essentially the model can control when it compacts and how.
2:48:25
So you can essentially
train your RL algorithm where compaction is an action- ...
2:48:30
where it shortens the history and
then the problem formulation will be, "I want to keep the maximum evaluation
scores that I have gotten while the model compacts its history to
the minimum length."
2:48:39
Because then you have the minimum amount of tokens that
you need to do this kind of compounding autoregressive prediction.
2:48:47
So there are actually
pretty nice problem setups in this, where the...
2:48:51
Like these agentic models
learn to use their context in a different way than just plow forward.
2:48:56
- One interesting recent example
would be DeepSeek-V3.
2:48:56
2, where they had a sparse attention
mechanism where they have essentially a very efficient,
small, lightweight indexer.
2:49:04
And instead of attending to all tokens, it
selects: "What tokens do I actually need?"
2:49:12
I mean, it almost comes back
to the original idea of attention where you are selective, but
attention is always on, you have maybe zero weight on some of them, but you use them all.
2:49:20
But
they are even more like, "Let's just mask that out or not even do that."
2:49:28
And even with sliding window attention
in OLMo, that is also kind of that idea.
2:49:32
You have a rolling window where you keep it
fixed, because you don't need everything.
2:49:35
Occasionally, some layers you might, but
it's wasteful.
2:49:35
But right now, I think, if you use everything, you're on the safe
side—it gives you the best bang for the buck because you never miss information.
2:49:43
And I think this year will be more about figuring out, like you
said, how to be smarter about that.
2:49:51
Right now, people want to have
the next state-of-the-art, and the state-of-the-art happens to
be the brute-force, expensive thing.
2:49:58
And then once you have
that, as you said, keep that accuracy, but let's see how we can
do that cheaper now, with tricks. - Yeah. All this scaling thing.
2:50:08
The reason we get the Claude
4.
2:50:08
5 Sonnet model first is because you can train it faster and you're
not hitting these compute walls as soon.
2:50:16
They can just try a lot more things and get
the model faster, even though the bigger model is actually better.
2:50:22
- I think we should say that there's a lot of
exciting stuff going on in the AI space.
2:50:25
My mind has recently
been really focused on robotics.
2:50:29
Today, we almost
entirely didn't talk about robotics.
2:50:32
There's a lot
of stuff on image generation, video generation.
2:50:36
I think it's fair to say that the most exciting
research work in terms of the amount, intensity, and fervor is in the LLM space, which is why I think it's
justified for us to focus on the LLMs that we're discussing.
2:50:51
But it'd be nice to bring in certain things that might be
useful.
2:50:55
For example, world models— there's growing excitement on that.
2:50:59
Do you think there will be any use in this coming year for
world models in the LLM space? Yes, I do think so.
2:51:06
Also with LLMs, what's interesting here is
that if we unlock more LLM capabilities, it also automatically
unlocks all the other fields because it makes progress faster.
2:51:18
A lot of
researchers and engineers use LLMs for coding.
2:51:24
So even if they work on
robotics, if you optimize these LLMs that help with coding,
it pays off.
2:51:28
But then, yes, world models are interesting.
2:51:32
It's basically where you have the model run a simulation
of the world in a sense, like a little toy thing of the real
thing, which can, again, unlock capabilities
regarding data the LLM is not aware of. It can simulate things.
2:51:47
And I think LLMs happen to work well by pre-training and
doing next-token prediction.
2:51:59
But we could do this even
more sophisticatedly in a sense.
2:52:03
I think there was a paper by Meta, a paper called World
Models.
2:52:07
So where they basically apply the concept of
world models to LLMs again, where instead of just having
next-token prediction and verifiable rewards, checking the answer
correctness, they also make sure the intermediate variables are correct.
2:52:21
You know, it's kind of like the model is learning basically a
code environment in a sense.
2:52:25
And I think this makes a lot of sense.
2:52:28
It's just expensive to do, but it is making things more
sophisticated, like modeling the whole thing, not just
the result.
2:52:40
And so it can add more value.
2:52:44
I remember when I
was a grad student, there is a...
2:52:51
competition called CASP, I
think, where they do protein structure prediction.
2:52:55
They predict
the structure of a protein that is not solved yet
at that point.
2:52:59
So in a sense, this is actually great, and I think
we need something like that for LLMs also, where you do the benchmark, but no one
does.
2:53:07
You hand in the results, but no one knows the solution.
2:53:11
And then after
the fact, someone reveals that.
2:53:11
But, AlphaFold, when it came
out, it crushed this benchmark.
2:53:19
I mean, there
were also multiple iterations, but I remember the
first one.
2:53:23
I'm not an expert in that subject, but the first one
explicitly modeled the physical interactions of the...
2:53:31
You know,
the physics of the molecule.
2:53:34
Also the angles, impossible angles.
2:53:34
And then
in the next version, I think they got rid of this, and just with brute force, scaling
it up.
2:53:38
And I think with LLMs, we are currently in this brute force scaling because
it just happens to work.
2:53:42
But I do think at some point it might make
sense to bring back this thing.
2:53:49
And I think with world
models, I think that is where I think that might be
actually quite cool. I mean, yeah.
2:53:57
And of course, also for robotics,
which is completely unrelated to LLMs. - Yeah.
2:54:03
And robotics is very explicit.
2:54:03
So there's the problem of locomotion or manipulation.
2:54:08
Locomotion is much more solved, especially
in the learning domain.
2:54:08
But there's a lot of value, just like with the
initial protein folding systems, bringing in the traditional model-based
methods. So you don't...
2:54:14
it's unlikely that you can just learn the
manipulation or the whole body, local manipulation problem end to end. That's the dream.
2:54:28
But then you realize when
you look at the magic of the human hand and the complexity of the real world,
it's really hard to learn this all the way through, the way I
guess AlphaFold 2 didn't.
2:54:40
- I'm excited about the robotic learning
space.
2:54:40
I think it's collectively getting supercharged by all the
excitement and investment in language models generally, where
the infrastructure for training transformers, which is a general
modeling thing, is becoming world-class industrial tooling, where wherever there was a limitation
for robotics, it's just way better.
2:55:03
There's way more compute.
2:55:03
And then on top
of that, they take these language models as kind of central units where you can
do interesting explorative work around something that already
works.
2:55:11
And then I see it emerging as, kind of like we talked
about, Hugging Face transformers and Hugging Face.
2:55:18
I think when I was at Hugging Face, I was
trying to get this to happen, but it was too early.
2:55:22
It's like these open
robotic models on Hugging Face, and having people be able to
contribute data and fine-tune them.
2:55:29
I think we're much closer now that the
investment in robotics and self-driving cars is related and it enables this,
where once you get to the point where you can have this sort of ecosystem where
somebody can download a robotics model and maybe fine-tune it to their robot
or share datasets across the world.
2:55:45
There's some work in this
area like RTX, I think it was a few years ago, where people are starting to
do that.
2:55:49
But once they have this ecosystem, it'll look very different.
2:55:53
And then this whole post-ChatGPT boom is putting more
resources into that, which I think is a very good area for doing research.
2:56:02
- This is also resulting in much
better, more accurate, and more realistic simulators being built, closing
the sim-to-real gap in the robotic space.
2:56:10
But, you know, you mentioned
a lot of excitement in the robotics space and a lot of
investment.
2:56:14
The downside of that, which happens in
hype cycles, I personally believe, and most robotics
people believe, that it's not...
2:56:24
Robotics is not going to be solved at the time scale as being
implicitly or explicitly promised.
2:56:32
And so what happens when
there's all these robotics companies that spring up and then they don't have
a product that works?
2:56:36
Then there's going to be this kind of
crash of excitement, which is nerve-wracking.
2:56:45
Hopefully something
else will come in and keep swooping in so that the continued development
of some of these ideas keeps going.
2:56:53
- I think it's also related to the continual
learning issue, essentially, where the real world is so complex.
2:56:57
With LLMs, you don't really need to have
something learn for the user, because there are a lot of things
everyone has to do.
2:57:05
Everyone maybe wants to, I don't know, fix their
grammar in their email or code or something like that.
2:57:13
It's more constrained,
so you can kind of prepare the model for that.
2:57:17
But preparing the robot for the
real world is harder.
2:57:17
I mean, you have the robotic foundation
models, and you can learn certain things like grasping
things.
2:57:24
But then again, everyone's house is different.
2:57:28
It's so different, and that is, I think, where the robot would have
to learn on the job, essentially.
2:57:36
And that, I guess, is the
bottleneck right now: how to, customize it on the fly, essentially.
2:57:42
- I don't think I can
possibly understate the importance of the thing that doesn't
get talked about almost at all by robotics folks or
anyone, which is safety.
2:57:53
All the interesting complexities we talk
about learning, all the failure modes and failure cases, everything we've been talking
about with LLMs—sometimes they fail in interesting ways.
2:58:00
All of that is
fun and games in the LLM space.
2:58:06
In the robotic space, in people's homes,
across millions of minutes and billions of interactions, you really
are almost allowed to fail never.
2:58:17
When you have
embodied systems that are put out there in the real world,
you just have to solve so many problems you never thought you'd have to
solve when just thinking about the general robot learning problem.
2:58:32
- I'm so bearish on in-home learned robots for consumer purchase.
2:58:36
I'm very
bullish on self-driving cars, and I'm very bullish for robotic automation,
e. g.
2:58:40
, like Amazon distribution where Amazon has built whole new
distribution centers designed for robots first rather than humans.
2:58:48
There's a
lot of excitement in AI circles about AI enabling automation and mass-scale manufacturing, and I do think that
the path to robots doing that is more reasonable, where it's a thing that is designed and optimized to do a
repetitive task that a human could conceivably do but doesn't
want to.
2:59:05
And then I'm much, but it's also going to take
a lot longer than people probably predict.
2:59:13
I think the leap from AI singularity to we can now scale
up mass manufacturing in the US because we have a massive
AI advantage is one that is troubled by a lot of political
and other challenging problems.
2:59:31
- Let's talk about timelines,
specifically timelines to AGI or ASI.
2:59:37
Is it fair, as a starting point,
to say that nobody really agrees on the definitions of AGI and ASI?
2:59:46
- I kind of think there's a
lot of disagreement, but I've been getting pushback where a lot of people
kind of say the same thing, which is like a thing that could reproduce
most digital economic work.
2:59:57
So, the remote worker
is a fairly reasonable example.
3:00:00
And I think OpenAI's
definition is somewhat related to that, which is like
an AI that can do a lot of economically valuable tasks—which I don't
really love as a definition, but I think it could be a grounding point, because language models today,
while immensely powerful, are not this remote worker
drop-in.
3:00:19
And there are things that could be done by an AI that
are way harder than remote work, which are like finding an unexpected scientific discovery that you couldn't even posit,
which would be an example of something that somebody says is an artificial
superintelligence problem.
3:00:34
Or, taking in all medical records and
finding linkages across certain illnesses that people didn't know, or
figuring out that some common drug can treat some niche cancer.
3:00:49
They would
say that that is a superintelligence thing.
3:00:52
So these are kind of natural
tiers.
3:00:52
My problem with it is that it becomes deeply entwined with the quest
for meaning of AI and these religious aspects to it.
3:01:01
So there's
different paths you can take it.
3:01:06
- And I don't even know if the
remote worker is a good definition because what exactly is
that?
3:01:10
I actually, I mean, I like...
3:01:14
I don't know if you
like the originally titled AI27 report.
3:01:17
They focus more on code and
research taste, so the target there is the superhuman coder.
3:01:25
So they have
several milestone systems: Superhuman coders,
superhuman AI researcher, then superintelligent AI
researcher, and then the full ASI, artificial superintelligence.
3:01:35
But after you develop the superhuman coder,
everything else follows quickly.
3:01:45
There, the task is to have fully
autonomous, automated coding.
3:01:52
So any kind of coding you need
to do in order to perform research is fully
automated.
3:01:55
And from there, humans would be doing AI research
together with that system, and they will quickly be able to develop a system that
can actually do the research for you. That's the idea.
3:02:07
And initially their
prediction was 2027, 2028, and now they've pushed it back by three to four years to 2031 (mean prediction).
3:02:18
Probably my prediction is even beyond 2031, but at least you can, in a concrete way, think about how difficult
it is to fully automate programming.
3:02:31
- Yeah, I disagree with some of their
presumptions and dynamics on how it would play out, but I think
they did good work in the scenario-defining milestones that are
concrete and tell a useful story, which is why the reach for this AI 2027 document transcended Silicon Valley.
3:02:46
It's
because they told a good story and they did a lot of rigorous work to do this.
3:02:53
I think the camp that I fall
into is that AI is so-called "jagged," which will be excellent at some
things and really bad at some things.
3:03:00
I think that when they're close to this automated software engineer, what it will
be good at is traditional ML systems and frontend, the model is
excellent at; but distributed ML, the models are actually quite bad at
because there's so little training data on doing large-scale distributed learning.
3:03:16
This is something we already see, and I think this will just get amplified.
3:03:20
And then it's kind of messier in these trade-offs, like how you think
AI research works and so on.
3:03:28
- So you think basically a superhuman
coder is almost unachievable, because of the jagged nature of the
thing, you're just always going to have gaps in capabilities?
3:03:38
- I think it's assigning completeness to
something where the models are already superhuman at some types of code.
3:03:42
I think that will continue.
3:03:46
And people are creative, so they'll
utilize these incredible abilities to fill in the weaknesses of the models
and move really fast.
3:03:50
There'll always be this dance for a long
time between the humans enabling the thing that the model
can't do.
3:03:57
And the best AI researchers are the ones that can
enable this superpower.
3:04:04
And I think those lines lead to what we already
see.
3:04:04
I think like Claude Code for building a website, you can stand up a beautiful
website in a few hours or do data going to keep getting better, and
we'll pick up some new coding skills along the way.
3:04:15
Linking to
what's happening in big tech, this AI 2027 report leans into the singularity idea, whereas
I think research is messy, social and largely in the data
in ways that AI models can't process.
3:04:34
But what we do have today
is really powerful and these tech companies are all collectively buying
into this with tens of billions of dollars of investment.
3:04:42
So we are going to
get some much better version of ChatGPT, a much better version of
Claude Code than we already have.
3:04:50
I think it's just hard to predict
where that is going, but the bright clarity of that future is why some of
the most powerful people in the world are putting so much money into this.
3:04:59
And I think it's just kind of small differences between like—we don't
actually know what a better version of ChatGPT is, but also, can it automate AI
research?
3:05:06
I would say probably not, at least in this timeframe.
3:05:13
Big tech
is going to spend $100 billion much faster than we get an automated AI researcher
that enables an AI research singularity.
3:05:22
- So you think your prediction
would be— if this is even a useful milestone or
more than 10 years out?
3:05:30
- I would say less than that on the software
side, but I think longer than that on things like research.
3:05:36
- Well, let's just for fun try to imagine
a world where all software writing is fully automated.
3:05:42
Can
you imagine that world?
3:05:46
- By the end of this year, the amount of
software that'll be automated will be so high.
3:05:50
But it'll be things like
trying to train a model with RL and you need to have
multiple bunches of GPUs communicating with each other.
3:05:57
That'll
still be hard, but it'll be much easier.
3:06:02
- One way to think about this—
the full automation of programming—is just thinking of lines
of useful code written, the fraction of that to the number of humans in the
loop.
3:06:12
So presumably there'll be for a long time humans in the loop of
software writing.
3:06:16
It'll just be fewer and fewer relative to the amount
of code written. Right?
3:06:24
And the superhuman coder—I think
the presumption there is it goes to zero, the number of humans
in the loop.
3:06:28
What does that world look like when the number of
humans in the loop is in the hundreds, not in the
hundreds of thousands?
3:06:39
- I think software engineering will
be driven more to system design and goals of outcomes, where I do
think software is largely going to be.
3:06:46
I think this has been happening
over the last few weeks, where people have gone from a month ago saying, "Oh
yeah, agents are kind of slop," which is a famous Karpathy quote, to
what is a little bit of a meme—the industrialization of software when anyone
can just create software with their fingerprints.
3:07:04
I do think we are closer
to that side of things, and it takes direction and understanding
how systems work to extract the best from the language
models.
3:07:12
I think it's hard to accept the gravity of how much is going to change
with software development and how many more people can do things without
ever looking at the code.
3:07:22
- I think what's interesting is to think
about whether these systems will be independent—completely independent in
the sense that, while I have no doubt that LLMs will kind of at some
point solve coding in a sense, like calculators solve calculating,
right?
3:07:33
So at some point, humans developed a tool where you never need
a human to calculate that number.
3:07:41
You just type it in, and it's
an algorithm.
3:07:41
You can do it in that sense.
3:07:45
And I think that's
the same probably for coding.
3:07:45
But the question isn't...
3:07:49
I think what will happen
is, you will just say, "Build that website."
3:07:53
It will make a really good website, and
then you maybe refine it.
3:07:53
But will it do things independently where...
3:07:57
Will
you still be having humans asking the AI to do something?
3:08:03
Like will
there be a person to say, "Build that website?"
3:08:07
Or will there be AI that
just builds websites or something?
3:08:12
- I think talking about
building websites is— - Too simple.
3:08:16
- The problem with websites and
the problem with the web, you know, HTML and all that kind of stuff,
it's very resilient to just ... slop.
3:08:25
It will show you slop; it's good
at showing slop.
3:08:25
I would rather think of safety-critical systems,
like asking AI to end-to-end generate something that manages
logistics, or manages cars, a fleet of cars—all that kind of stuff.
3:08:41
So it end-to-end generates that for you.
3:08:45
- I think a more intermediate example is
take something like Slack or Microsoft Word.
3:08:49
I think if organizations
allow it, AI could very easily implement features end-to-end and
do a fairly good job for like things that you want to try.
3:08:58
You want to add a new tab in Slack that you want to use, and I think
AI will be able to do that pretty well.
3:09:06
- Actually, that's a really great
example.
3:09:06
How far away are we from that? - Like this year. - See, I don't know. I don't know.
3:09:15
- I guess I don't know how bad production
codebases are, but I think that within...
3:09:18
on the order of a few years, a
lot of people are going to be pushed to be more of a designer and product
manager, where you have multiple of these agents that can try things for you and they
might take one to two days to implement a feature or attempt to fix a bug.
3:09:30
And you have these dashboards, which I think Slack is actually a good dashboard
where your agents will talk to you and you'll then give feedback.
3:09:38
But things like, if I make a website, like, "Do you
want a passable logo?"
3:09:42
I think these cohesive design things and the
style is going to be very hard for models and deciding on
what to add the next time. - I just... Okay.
3:09:54
I hang out with a
lot of programmers and some of them are a little bit on the skeptical side
in general. That's just their vibe.
3:09:57
I just think there's a lot of complexity
involved in adding features to complex systems.
3:10:09
Like, if you look
at the browser, Chrome.
3:10:13
If I wanted to add a feature,
if I wanted to have tabs as opposed to up top, I
want them on the left side. Interface-wise, right? I think we're
not...
3:10:21
This is not a next-year thing.
3:10:26
- One of the Claude releases this year, one
of their tests was: we give it a piece of software and leave Claude to run to
recreate it entirely, and it could already almost rebuild Slack from
scratch, just given the parameters of the software and left
in a sandbox environment to do that.
3:10:41
- So the "from scratch" part,
I like almost better.
3:10:44
- So it might be that smaller and newer
companies are advantaged, and they're like, "We don't have the bloat and complexity,
and therefore this feature exists."
3:10:53
- And I think this gets to the point you
mentioned, that some people you talk to are skeptical.
3:10:57
I think that's
not because the LLM can't do X, Y, Z.
3:11:01
It's because people
don't want it to do it this way.
3:11:05
- Some of that could be a skill issue
on the human side.
3:11:05
We have to be honest with ourselves.
3:11:09
And some of that
could be an underspecification issue.
3:11:13
So, programming, it's like you're
just assuming...
3:11:13
This is like an issue with communication in relationships
and friendships.
3:11:20
You're assuming the LLM is supposed to read your mind.
3:11:24
This is where spec-driven design is really important.
3:11:27
Using natural
language to specify what you want.
3:11:32
- If you talk to people at the labs,
they use these in their training and production code.
3:11:36
Claude Code is
built with Claude Code, and they all use these things extensively.
3:11:40
Dario
talks about how much of Claude's code...
3:11:44
It's like these people are slightly
ahead in terms of the capabilities they have and what they
probably spend on inference.
3:11:53
They could spend 10 to 100x as
much as we're spending on a lowly $100 or $200 a month plan. They
truly let it rip.
3:11:57
And I think that, with the pace of progress
that we have, it seems like- a year ago we didn't have Claude Code
and we didn't really have reasoning models.
3:12:11
The difference between sitting
here today and what we can do with these models is significant,
and there's a lot of low-hanging fruit to improve them.
3:12:18
The
failure modes are pretty dumb.
3:12:18
Like- "Claude, you tried to use a CLI command I
don't have installed 14 times, and then I sent you the command to run."
3:12:27
That, from a modeling perspective, is pretty fixable. So I don't know. - I agree with you.
3:12:34
I've been
becoming more and more bullish in general.
3:12:38
Speaking to what you're
articulating, I think it is a human skill issue.
3:12:42
Anthropic is leading the way, along with other
companies, in understanding how to best use the models for programming;
therefore, they're effectively using them.
3:12:54
There are a lot of
programmers on the outskirts who don't...
3:12:58
I mean, there's not a really good guide on how to use them.
3:13:01
People
are trying to figure it out, but- - It might be very expensive.
3:13:04
The
entry point might be $2,000 a month, which is only for tech companies
and rich people. That could be it.
3:13:13
- But it might be worth it.
3:13:13
If
the final result is a working software system, it might be worth it.
3:13:17
By
the way, it's funny how we converged from the discussion of timeline to AGI
to something more pragmatic and useful.
3:13:25
Is there anything
concrete, interesting, useful, and profound to be said
about the timeline to AGI and ASI?
3:13:32
Or are these discussions a bit
too detached from the day to day?
3:13:39
- There are interesting bets.
3:13:39
A lot
of people are trying to do RLVR— Reinforcement Learning with Verifiable
Rewards—in real scientific domains, where startups with hundreds of millions
of funding have wet labs where they're having language models propose
hypotheses that are tested in the real world.
3:13:54
I would say that they're early,
but with the pace of progress, it's like- ...
3:14:00
maybe they're early by six months
and they make it because they were there first, or maybe they're
early by eight years; you don't know.
3:14:08
That type of moonshot to branch
this momentum into other sciences would be very transformative.
3:14:15
If, AlphaFold moments happen in all
sorts of other scientific domains by a startup solving this.
3:14:23
I think
there are startups—maybe Harmonic is one—where they're going all in
on language models plus Lean for math.
3:14:31
You had another guest where
you talked about this recently, and we don't know exactly
what's going to fall out of spending $100 million on that
model.
3:14:38
Most of them will fail, but a couple might be big
breakthroughs that are very different than ChatGPT or
Claude Code type software experiences.
3:14:50
A tool that's only good for a PhD mathematician but makes
them 100X effective... - I agree.
3:14:58
I think this will happen
in a lot of domains, especially domains that have a lot of resources, like finance, legal, and
pharmaceutical companies.
3:15:09
But then again, is it really
AGI?
3:15:09
Because we are now specializing it again.
3:15:13
Is it really
that much different from back in the day when we had specialized
algorithms?
3:15:17
It's just the same thing, way more sophisticated, but I don't know, is there a threshold for AGI?
3:15:25
I think
the real cool thing here is that we have foundation models we
can specialize.
3:15:29
That's like the breakthrough.
3:15:33
Right now, I think we
are not there yet because, first, it's too expensive, but also, ChatGPT doesn't just give away their
model to customize it.
3:15:39
I think once that's true...
3:15:44
And I can
imagine this as a business model, where OpenAI says at some point, "Hey, Bank of America, for $100
million we will do your custom model," something like that.
3:15:55
I think
that will be the huge economic value-add.
3:15:59
The other thing, though, is also...
3:16:02
Companies, I mean,
what is the differentiating factor?
3:16:06
If everyone uses the
same LLM, if everyone uses ChatGPT, they will all do
the same thing.
3:16:09
Well, if everyone is moving in lockstep,
but companies want to have a competitive advantage, there is
no way around using some of their private data and specializing.
3:16:21
It's gonna be interesting.
3:16:26
- Seeing the pace of progress, it does
feel like things are coming.
3:16:26
I don't think the AGI and ASI thresholds
are particularly useful.
3:16:35
- I think the real question, and
this relates to the remote worker thing, is: when are we going to see a
big, obvious leap in economic impact?
3:16:47
Because currently there's
not been an obvious leap in the economic impact of
LLM models, for example.
3:16:55
And that's, you know, aside
from AGI or ASI, all that stuff, there's a real question of, "When
are we going to see a GDP..." "... jump?"
3:17:06
- Yeah, what is the GDP made up
of?
3:17:06
A lot of it is financial services, so I don't know what this is.
3:17:13
- Right, GDP is a- - It's just hard for me to
think about the GDP bump, but I would say that software development
becomes valuable in a different way, when you no longer have
to look at the code anymore.
3:17:25
When Claude Code will make
you a small business.
3:17:29
Which is essentially, Claude can set up your
website, your bank account, your email, and your whatever else.
3:17:33
And you just have
to express what you're trying to put into the world.
3:17:37
That's not just an enterprise market, but it is hard.
3:17:41
I don't
know how you get people to try doing that.
3:17:45
I guess if ChatGPT can do
it—people are trying ChatGPT.
3:17:49
- I think it boils down to the scientific
question of, "How hard is tool use to solve?"
3:17:53
Because a lot of
the stuff you're implying, the remote work stuff, is tool
use. It's like...
3:17:57
computer use, like how you have an
LLM that goes out there, this agentic system, and does
something in the world, and only screws up 1% of the time.
3:18:11
- Computer use- - Or less. - ...
3:18:12
is a good example of what labs care about
and we haven't seen a lot of progress on.
3:18:16
We saw multiple demos in 2025 of, like, Claude can use your computer, or OpenAI
had operator, and they all suck.
3:18:24
They're investing money in this,
and I think that'll be a good example.
3:18:28
Whereas actually,
something where it just seems like taking over the whole
screen seems a lot harder than having an API that they
can call in the back end.
3:18:39
For some of that, you have to set up a different
environment for them all to work in.
3:18:39
They're not working on your MacBook; they
are individually interfacing with Google and Amazon and Slack, and they
handle all these things in a very different way than humans do.
3:18:51
So some
of this might be structural blockers.
3:18:55
- Also, specification-wise,
I think the problem is for arbitrary tasks, well, you
still have to specify what you want your LLM to do. And how
do you do that? What is the environment? How do you specify?
3:19:07
You
can say what the end goal is, but if it can't solve the end goal...
3:19:11
with LLMs, if you ask it for text, it can always clarify or do
sub-steps.
3:19:15
How do you put that information into a system that, let's say,
books a travel trip for you?
3:19:19
You can say, "You screwed up my credit card information,"
but even to get it to that point, even to get it to that point,
how do you, as a user, guide the model before it can even attempt that?
3:19:31
I think the interface is really hard.
3:19:36
- Yeah, it has to learn a lot
about you specifically.
3:19:39
And this goes to continual learning,
about the general mistakes that are made throughout, and then
mistakes that are made through you.
3:19:48
- All the AI interfaces are getting
set up to ask humans for input.
3:19:51
I think Claude Code we talked about a lot.
3:19:54
It asks feedback and questions.
3:19:54
If
it doesn't have enough specification on your plan or your desired goal, it starts
to ask questions, "Would you rather?"
3:20:02
We talked about Memory,
which saves across chats.
3:20:06
Its first implementation is kind
of odd, where it'll mention my dog's name or something in a chat.
3:20:11
I'm like, "You don't need to be
subtle about this. I don't care."
3:20:15
But things that are emerging,
ChatGPT has the Pulse feature.
3:20:19
Which is like a curated
couple of paragraphs with links to something to look at,
and people talk about how models are going to ask you
questions.
3:20:27
Which I think is a very...
3:20:30
It's probably going to
work.
3:20:30
The language model knows you had a doctor appointment and asks,
"Hey, how are you feeling after that?"
3:20:37
Which again goes into the territory
where humans are very susceptible to this, and there's a lot of
social change to come.
3:20:41
But also, they're experimenting with having the
models engage.
3:20:45
Some people like this Pulse feature, which processes your
chats and automatically searches for information and puts it in the app.
3:20:53
So there are a lot of things coming.
3:20:58
- I used that feature before, and I
always feel bad because it does that every day, and I rarely check
it out.
3:21:02
It's like, how much compute is burned on something I don't even look
at, you know?
3:21:06
It's kind of like, "Oh..."
3:21:11
- There's also a lot of idle compute
in the world, so don't feel too bad. - Okay.
3:21:16
Do you think new
ideas might be needed?
3:21:20
Is it possible that the path to
AGI, however we define that, to solve computer use more generally,
to solve biology and chemistry and physics—sort of the Dario Amodei
definition of AGI?
3:21:30
Do you think it's possible that totally new ideas
are needed? Non-LLM, non-RL ideas?
3:21:45
What might they look like?
3:21:45
We're
going into philosophy land a bit.
3:21:50
- For something like a singularity
to happen, I would say yes.
3:21:54
The new ideas could be architectures
or training algorithms, fundamental deep learning
things.
3:21:58
But in that nature, they're pretty hard to predict.
3:22:02
I think we won't get very far even without those advances.
3:22:06
We might get the software solution, but it might stop at
software and not do computer use without more innovation.
3:22:14
So I think
that a lot of progress will be coming, but if you're gonna zoom
out, there's still ideas in the next 30 years that are
gonna look like that was a major scientific innovation
that enabled the next chapter of this.
3:22:28
And I don't know
if it comes in one year or in 15 years. - Yeah.
3:22:32
I wonder if the bitter lesson holds
true for the next 100 years, what that looks like.
3:22:37
- If scaling laws are fundamental in deep
learning, I think the bitter lesson will always apply, which is
compute will become more abundant, but even within
abundant compute, the ones that have a steeper scaling
law slope or a better offset— like, this is a 2D plot of
performance and compute—and like even if there's more compute available,
the ones that get 100x out of it will win.
3:23:01
- It might be something like literally
computer clusters orbiting Earth with solar panels.
3:23:09
- The problem with that is heat dissipation.
3:23:09
You get all the radiation from the sun and don't have any air to dissipate heat.
3:23:13
But there is a lot of space to put clusters.
3:23:17
There's a lot of solar energy there and
you could figure out the heat dissipation, as there is a lot of energy and there
probably could be engineering will to solve the heat problem— so there could be.
3:23:27
- Is it possible—and we should say
that it definitely is possible— that we're basically going to be
plateauing this year?
3:23:30
Not in terms of— the system capabilities, but
what they actually mean for human civilization.
3:23:43
So on the coding
front, really nice websites will be built. Very nice auto-complete.
3:23:53
Very nice way to understand code bases
and maybe help debug, but really just a very nice helper on the coding
front.
3:24:01
It can help research mathematicians do some
math.
3:24:04
It can help you with shopping. It's a nice
helper. It's Clippy on steroids. What else?
3:24:12
It may be a good
education tool and all that kind of stuff, but computer use turns out
extremely difficult to solve.
3:24:23
So I'm trying to frame the cynical case in all these domains
where there's not a really huge economic impact, but realize
how costly it is to train these systems at every level, both the
pre-training and the inference, how costly the inference is,
the reasoning, all of that. Like, is that possible?
3:24:43
And how
likely is that, do you think?
3:24:47
- When you look at the models, there are
so many obvious things to improve and it takes a long time to train
these models and to do this art, and it'll take us with the ideas
that we have multiple years to actually saturate in terms
of whatever benchmark or performance we are searching
for.
3:25:03
It might serve very narrow niches; like the average
ChatGPT 800 million user might not get a lot of benefit out of this,
but it is going to serve different populations by getting
better at different things.
3:25:18
- But I think what everybody's
chasing now is a general system that's useful
to everybody.
3:25:22
So, okay, so if that's not... That can plateau, right?
3:25:28
- I think that dream is actually kind of
dying.
3:25:28
As you talked about with the specialized models where it's like...
3:25:34
And multimodal is often...
3:25:34
Video generation
is a totally different thing. Thing.
3:25:39
- "That dream is kind of
dying" is a big statement, because I don't know if it's dying.
3:25:42
If you ask the actual frontier lab people, they...
3:25:46
I mean,
they're still chasing it, right?
3:25:48
- I do think they are still rushing
to get the next model out, which will be much better than the...
3:25:52
"Much" is a relative term, but it will be better than the previous one.
3:25:56
And I can't see them slowing down.
3:26:00
I just think the gains
will be made or felt more through not only scaling the model, but now...
3:26:07
I feel like there's a lot of tech
debt.
3:26:07
It's like, "Well, let's just put the better model in there."
3:26:13
Better model, better model.
3:26:13
And now
people are like, "Okay, let's also at the same time improve
everything around it too."
3:26:20
Like the engineering of the
context and inference scaling.
3:26:23
The big labs will still keep
doing that.
3:26:23
And now also the smaller labs will catch up,
because now they are hiring more.
3:26:31
There will be more people and
LLMs.
3:26:31
It's kind of like a circle.
3:26:35
They also make them more productive and
it's just... It's like amplification.
3:26:39
I think what we can expect is
amplification, but not like a change of any...
3:26:43
not like a paradigm change.
3:26:43
I don't think that is true, but everything will be just amplified
and amplified, and I can see that continuing for a long time, you know? - Yeah.
3:26:52
I guess my statement that the
dream is dying depends on exactly what you think it's gonna be doing.
3:26:56
Like, Claude Code is a general model that can do a lot of things,
but it's not necessarily...
3:27:05
It depends a lot on integrations.
3:27:05
I bet
Claude Code could do a fairly good job of doing your email, and the hardest part
is figuring out how to give information to it and how to get it to
be able to send your emails.
3:27:17
But that's just kind of like...
3:27:17
I
think it goes back to what is the "one model to rule everything" ethos,
which is just like a thing in the cloud that handles your entire digital
life and is way smarter than everybody.
3:27:29
It's like it's operating in a...
3:27:34
So it's an interesting leap of faith
to go from "Claude Code becomes that," which in some ways is...
3:27:39
There are some
avenues for that, but I do think that the rhetoric of the industry
is a little bit different.
3:27:49
- I think the immediate thing
we will feel next as a normal person using LLMs will probably
be related to something trivial, like making figures.
3:27:56
Right now, LLMs are terrible at making figures.
3:28:00
Is it because we
are getting served the cheap models with much less inference compute
than behind the scenes? Maybe some.
3:28:09
Like, there are some ways to get
better figures, but if you ask today, ..."
3:28:13
Draw a flowchart of X, Y, Z,"
it's most of the time terrible.
3:28:18
And it is a very simple task for
a human.
3:28:18
I think it's almost easier sometimes to draw something
than to write something.
3:28:25
- Yeah, the multimodal understanding does
feel like something that is odd... ...
3:28:28
that it's not better solved.
3:28:31
- I think we're not saying one obvious
thing that we're not realizing, that's a gigantic thing that's hard
to measure, which is making all of human knowledge accessible—
—to the entire world.
3:28:46
One thing that is hard to articulate
is the huge difference between Google Search and an LLM.
3:28:50
I
feel like I can basically ask an LLM anything and get an answer, and
it's doing less and less hallucination.
3:29:04
And that means understanding my own life, figuring out a career trajectory,
solving the problems all around me, learning about anything
through human history.
3:29:16
I feel like nobody's
really talking about that, because they just immediately take
it for granted that this is awesome.
3:29:25
That's why everybody's using it:
because you get answers for stuff.
3:29:29
Think about the impact across time.
3:29:33
This is not just in the United
States; it's all across the world.
3:29:37
Kids throughout the world being
able to learn these ideas— the impact that has across time is
probably... ... That's the real impact.
3:29:48
Talk about GDP; it won't be
like a leap. It'll be... ...
3:29:52
that's how we get to Mars,
that's how we build these things, that's how we have a million new OpenAIs
and all the innovation from there.
3:30:00
It's this quiet force that permeates
everything: human knowledge. - I agree with you.
3:30:06
In a sense, it
makes knowledge more accessible, but it also depends on what the topic is.
3:30:13
For something like math, you can
ask it questions and it answers, but if you want to learn a topic from
scratch, the sweet spot is still elsewhere.
3:30:28
There are really good math
textbooks laid out linearly, and that is a proven
strategy to learn a topic.
3:30:36
It makes sense, if you start from zero, to use information-dense
text to soak it up, but then you use the LLM to
make infinite exercises.
3:30:43
Like, you have problems in a certain area
or have questions that something's- uncertain about certain
things, you ask it to generate example problems, you solve them, and you need more background
knowledge, you ask it to generate that. But then...
3:31:02
it won't
give you anything, let's say, that is not in the textbook.
3:31:08
It's just packaging it differently, if that makes sense.
3:31:13
But then there are things I feel
like where it also adds value in a more timely sense, where there is no good alternative besides a human doing
it on the fly.
3:31:21
For example, if you're planning to go to
Disneyland and you try to figure out which tickets to
buy for which park when, well, there is no textbook on that.
3:31:32
There
is no information-dense resource.
3:31:36
There's only the sparse internet, and
then there is a lot of value in the LLM. You just ask it.
3:31:40
You have
constraints on traveling these days.
3:31:44
I want to go there and there.
3:31:44
Please
figure out what I need, when and from where, what it costs and
stuff like that, and it is a very customized, on-the-fly package.
3:31:56
And this is like one of a thousand
examples of personalized- Personalization is essentially pulling
information from the sparse internet, the non-information-dense thing where there's no better version that exists. It just doesn't exist.
3:32:07
You make it almost from scratch.
3:32:12
- And if it does exist, it's full of-
speaking of Disney World, full of- what would you call it? Ad slop. It's
impossible.
3:32:16
Take any city in the world, what are the top 10 things to do?
3:32:25
An LLM is just way better to ask than anything on the internet.
3:32:29
- Well, for now, that's because they're
subsidized and they're gonna be paid for by ads. - Oh my goodness. - It's coming. - No. No.
3:32:38
I mean, I'm hoping there's a
very clear indication what's an ad and what's not an ad in that context.
3:32:46
- That's something I mentioned a
few years ago.
3:32:46
If, I don't know, if you are looking for a new
running shoe, well, is it a coincidence that Nike maybe comes
up first? Maybe, maybe not.
3:32:58
But I think there are clear laws.
3:32:58
You have to be clear about that.
3:33:02
I think that's what everyone fears.
3:33:02
It's the subtle message in there, but that also brings us
to the topic of ads, where I think this was a
thing.
3:33:11
Hopefully, I think for- in 2025, just because I think it's they're still not making money
in other ways right now.
3:33:22
Having ad spots in there...
3:33:22
but
the thing is, they couldn't, because there are alternatives without
ads and people would just flock- to the other products.
3:33:31
It's also just crazy how- yeah, how they're one-upping
each other, spending so much money to just get the users. - I think so.
3:33:41
Like, some Instagram
ads— I don't use Instagram, but I understand the appeal of paying
a platform to find users who will genuinely like your product, and that is
the best case of things like Instagram ads.
3:33:56
But there are also plenty of cases
where advertising is very awful for incentives, and I think that a world where the power of AI can integrate with that
positive view of, "I am a person and I have a small business and I
want to make the best, I don't know, damn steak knives in the world,
and I want to sell them to somebody who needs them."
3:34:15
And if AI can make
that sort of advertising thing work even better, that's very good
for the world, especially with digital infrastructure, because
that's how the modern web has been built.
3:34:27
But that's not to say that addicting feeds so that you
can show people more content is a good thing.
3:34:34
So, I think that's even
what OpenAI would say, is they want to find a way that can make
the monetization upside of ads while still giving their users agency.
3:34:45
And I personally would think that Google is
probably going to be better at figuring out how to do this, because
they already have ad supply and if they figure out how to turn
this demand in their Gemini app into useful ads, then they can turn it on.
3:34:56
And somebody will figure it out—I don't know
if it's this year, but there will be experiments with it.
3:35:06
- I do think what holds companies back
right now is really just that the competition is not doing it.
3:35:09
It's
more like a reputation thing.
3:35:13
It's just, I think people are just
afraid right now of ruining or losing their reputation, losing users, because it would make headlines if
someone launched these ads.
3:35:19
But— - Unless they were great, but the first ads
won't be great because it's a hard problem that we don't know how to solve.
3:35:28
- Yeah, I think also the first version of
that will likely be something like on X, like the timeline where you have a
promoted post sometimes in between.
3:35:35
It'll be something where it will say "promoted"
or something small, and then there will be an image.
3:35:39
I think right now the problem
is: who makes the first move?
3:35:43
- If we go 10 years out, the
proposition for ads is that you will make so much money on
ads by having so many users that you can use this to fund better R&D
and make better models, which is why YouTube is dominating the market for any— Netflix is scared of
YouTube.
3:35:58
They have the ads, they make—I pay $28 a month for
Premium.
3:36:02
They make at least $28 a month off of me and many other people.
3:36:06
And they're just creating such a dominant position in video.
3:36:10
So
I think that's the proposition: that ads can make you have
a sustained advantage.
3:36:17
in what you're spending per
user.
3:36:17
But there's so much money in it right now that
somebody starting that flywheel is scary because it's a long-term bet.
3:36:29
- Do you think there'll be some
crazy big moves this year business-wise?
3:36:33
Like Google or Apple
acquiring Anthropic or something like this?
3:36:40
- Dario will never sell, but we are
starting to see some types of consolidation with Groq for $20 billion and Scale AI for almost
$30 billion and countless other deals like this that are
structured in a way that is detrimental to the Silicon
Valley ecosystem, which is this licensing deal where not
everybody gets brought along, rather than a full acquisition that
benefits the rank-and-file employee by getting their stock vested.
3:37:06
That's a big issue for culture to address because the startup
ecosystem is the lifeblood where, if you join a startup, even if it's not
successful, it might get acquired on a cheap premium and you'll get
paid out for this equity.
3:37:24
These licensing deals are taking
the top talent a lot of the time.
3:37:27
The deal for Groq to
NVIDIA is rumored to be better to the employees, but it is
still this antitrust-avoiding thing.
3:37:35
But I think that this trend of
consolidation will continue.
3:37:39
Me and many smart people I
respect have been expecting consolidation to have
happened sooner, but it seems like some of these things
are starting to turn, but at the same time, companies are
raising ridiculous amounts of money for reasons where I'm like, "I don't
know why you're taking that money."
3:37:59
So it's mixed this year, but some
consolidation pressure is starting.
3:38:04
- What kind of surprising consolidation
will we see?
3:38:04
You say Anthropic is a "never."
3:38:08
I mean, Groq is a big
one.
3:38:08
Groq with a Q, by the way. - Yeah.
3:38:12
There's just a lot of startups
and a very high premium on AI startups.
3:38:16
So there could be a lot of - that kind of stuff, yeah.
3:38:19
- $10 billion range acquisitions, which
is really big for a startup that was maybe founded a year ago. I think Manus. ai...
3:38:26
this company based in
Singapore that Meta-founded was founded eight months ago and then had
a $2 billion exit.
3:38:30
I think there will be some other multi-billion dollar
acquisitions, like Perplexity.
3:38:39
- Like Perplexity, right?
3:38:40
- Yeah, people rumor them to
Apple.
3:38:40
I think there's a lot of of pressure and liquidity
in AI.
3:38:43
There's pressure on big companies to have outcomes and- I would guess that a big
acquisition gives people leeway to then tell the next
chapter of that story.
3:38:56
- I guess Cursor—we've been talking
about code—somebody acquires Cursor.
3:39:00
if somebody acquires Cursor...
3:39:02
- They're in such a good position
by having so much user data.
3:39:05
And we talked about continual learning.
3:39:05
They
had one of the most interesting sentences in a blog post, which is that they had
their new Composer model, which was a fine-tune of one of these large Mixture
of Expert models from China.
3:39:13
You can know that by asking it or because the
model sometimes responds in Chinese— ...
3:39:22
which none of the American models do.
3:39:22
And
they had a blog post where they're like, "We're updating the model weights every
90 minutes based on real-world feedback from people using it."
3:39:30
Which is like
the closest thing to real-world RL happening on a model, and it's just
mentioned in one of their blog posts— - That's incredible. - which is super cool.
3:39:38
- And by the way, I should say I use Composer
a lot because one of the benefits it has is that it's fast.
3:39:43
- I need to try it 'cause
everybody says this.
3:39:45
- And there'll be some IPOs potentially.
3:39:45
You think Anthropic, OpenAI, xAI.
3:39:51
- They can all raise so much money so easily that they don't feel a need to.
3:39:53
So long as fundraising is easy, they're not going to IPO because public
markets apply pressure.
3:40:00
I think we're seeing in China that the
ecosystem's a little different with both MiniMax and Z.
3:40:03
ai applying for, filing IPO paperwork, which will be interesting
to see how the Chinese market reacts.
3:40:11
I actually would guess that it's
going to be similarly hypey to the US, so long as all this is going and
not based on the reality that they're both losing a ton of money.
3:40:19
I wish
more of the gigantic American AI startups were public because it would be very
interesting to see how they're spending money and have more insight.
3:40:27
And
also just to give people access to investing in these, because
I think they're some of the most formidable companies—they're the companies of the era.
3:40:37
And the
tradition is now for so many of the big startups in the US to not go public.
3:40:41
It's like we're still waiting for Stripe and the IPO, but Databricks definitely
didn't.
3:40:45
They raised like a Series G or something.
3:40:49
And I just feel
like it's kind of a weird equilibrium for the
market where it's like, I would like to see these
companies go public and evolve in that way that a company can.
3:41:01
- Do you think 10 years from now some
of the frontier model companies are still around? Anthropic, OpenAI?
3:41:08
- I definitely don't see it as a
winner-takes-all unless there truly is some algorithmic secret that
one of them finds that lets this flywheel.
3:41:16
Because the development path
is so similar for all of them.
3:41:16
Google and OpenAI have all the same
products, and then Anthropic's more focused, but when you talk to people it
sounds like they're solving a lot of the same problems. So I think...
3:41:28
and there's
offerings that'll spread out. There's a lot of...
3:41:31
it's a very big cake being made that
people are going to take money out of.
3:41:36
- I don't want to trivialize it, but OpenAI
and Anthropic are primarily LLM service providers.
3:41:43
And some of the
other companies like Google and xAI, linked to X, do other stuff too.
3:41:51
And so it's very possible,
if AI becomes more commodified, that the companies just
providing LLMs will die.
3:42:00
- I think the advantage they have is a
lot of users, and I think they will just pivot.
3:42:04
Like Anthropic, I
think, pivoted.
3:42:04
I don't think they originally planned to work on
code, but they found, "Okay, this is a nice niche, and now we are comfortable
and we push on this niche."
3:42:15
I can see the same thing...
3:42:19
Let's say
hypothetically, I'm not sure if it will be true, but let's say Google
takes all the market share of the general chatbot.
3:42:27
Maybe OpenAI will then
focus on some other sub-topic.
3:42:31
They have too many users to go
away in the foreseeable future.
3:42:37
- I think Google is always ready to
say, "Hold my beer," with AI models.
3:42:40
- I think the question is if
the companies can support the valuations.
3:42:43
I see the AI companies being looked at in some ways like
AWS, Azure, and GCP are, all competing in the same space and all
very successful businesses.
3:42:51
There's a chance that the API market is
so unprofitable that they go up and down the stack to products and hardware.
3:42:59
They have so much cash that they can build power plants and data centers, which is
a durable advantage now.
3:43:03
But there's also a reasonable outcome that
these APIs are so valuable and so flexible for developers that
they become something like AWS.
3:43:15
But AWS and Azure are also
going to have these APIs, so having five or six people competing in the
API market is hard.
3:43:21
So maybe that's why they get squeezed out.
3:43:27
- You mentioned "RIP Llama."
3:43:27
Is
there a path to winning for Meta? - I think nobody knows.
3:43:32
They're moving a
lot, so they're signing licensing deals with Black Forest Labs, which
is an image generation company, or Midjourney.
3:43:42
So I think in some ways on the product and consumer-facing
AI front, it's too early to tell.
3:43:50
I think they have
some people who are excellent and very motivated being close to
Zuckerberg.
3:43:54
So I think there's still a story to unfold there.
3:43:58
Llama is a
bit different, where Llama was the most focused expression of the
organization.
3:44:04
And I don't see Llama being supported to that extent.
3:44:08
I think it was a very successful brand for them.
3:44:12
So they still
might participate in the open ecosystem or continue the
Llama brand into a different service, because people
know what Llama is.
3:44:21
- You think there's a Llama 5?
3:44:24
- Not an open-weight one. - It's interesting.
3:44:26
I think Llama was the
pioneering open-weight model.
3:44:26
With Llama 1, 2, and 3, there was a lot
of love.
3:44:34
But I think then, hypothesizing or speculating,
I think the leaders at Meta, like the upper executives,
they...
3:44:42
I think they got very excited about Llama because they saw how
popular it was in the community.
3:44:46
And then I think the problem was trying to,
let's say, monetize the open—or not monetize the open source,
but use it to make a bigger splash.
3:44:58
It felt almost forced, like
developing these very big Llama 4 models to be on top of the benchmarks.
3:45:08
But I don't think the
goal of Llama models is to be on top of the benchmarks beating, let's say, ChatGPT
or other models.
3:45:12
I think the goal was to have a model that people can use,
trust, modify, and understand.
3:45:20
So that includes having smaller models.
3:45:20
They don't have to be the best models.
3:45:23
And what happened was, these models
were, of course...
3:45:23
the benchmarks suggested that they were better than
they were because they had specific models trained on preferences so that they
performed well on benchmarks.
3:45:32
That's kind of, like, this overfitting thing to force
it to be the best.
3:45:36
But then at the same time, they didn't do the small models
that people could use.
3:45:39
And I think that no one could run these big models
then.
3:45:43
And then there was kind of a weird thing.
3:45:47
I think it's just
because people got too excited about headlines pushing the
frontier. I think that's it.
3:45:54
- And too much on the benchmarking side. - It's too much work.
3:45:57
- I think it imploded under internal political fighting and misaligned
incentives.
3:46:00
The researchers want to build the best models, but
there's a layer of organization— ...
3:46:07
and management that is trying to demonstrate
that they do these things.
3:46:07
And then there are rumors about how, for example,
some horrible technical decision was made.
3:46:19
It just seems like it got so bad
that it all just crashed out.
3:46:24
- Yeah, but we should also
give huge props to Mark Zuckerberg.
3:46:28
I think it comes
from Mark, actually, from Mark Zuckerberg, from the top of the
leadership, saying open source is important.
3:46:35
The fact that
that leadership exists means there could be a Llama 5,
where they learn the lessons from benchmarking and say, "We're
going to be GPT-OSS—" "...
3:46:47
and provide a really awesome
library of open source."
3:46:51
- What people say is that there's
a debate between Mark and Alexandr Wang, who is very bright,
but much more against open source.
3:46:59
And to the extent that he has
a lot of influence over the AI org, it seems much less likely, because it
seems like Mark brought him in for a fresh leadership eye in directing AI.
3:47:06
And if being open or closed is no longer
the defining nature of the model, I don't expect that to be a defining
argument between Mark and Alex.
3:47:18
They're both very bright,
but I just have a hard time understanding all of it because
Mark wrote this piece in July of 2024, which was probably the best blog post at the time, saying
"The Case for Open Source AI."
3:47:36
And then July 2025 came around and it was,
"We're reevaluating our relationship with open source." So it's just kind of...
3:47:42
- But I think also the problem...
3:47:42
Not the problem, but I think, well, we may have been a bit too
harsh, and that caused some of that.
3:47:50
Because I mean, we as open source
developers or the community...
3:47:54
Even though the model was
maybe not what everyone hoped for, it got a lot of backlash.
3:47:58
And I think that was unfortunate because I can see that as a
company, they were hoping for positive headlines.
3:48:05
And
instead of just getting no headlines or positive headlines,
in turn they got negative headlines.
3:48:13
And then it kind of
reflected bad on the company.
3:48:17
I think that is also something
where it's maybe a spite reaction, almost like, "Okay, we tried to
do something nice, we tried to give you something cool, like an
open source model, and now you are kind of being negative about
us, even for the company."
3:48:32
So in that sense, it looks like, "Well,
maybe then we'll change our mind." I guess. I don't know.
3:48:38
- Yeah, that's where the
dynamics of discourse on X can lead us, as a community, astray.
3:48:48
Because sometimes it feels random.
3:48:48
People
pick the thing they like and don't like.
3:48:51
I mean, you can see the same thing
with Grok 4. 1 and Grok Code Fast 1. 0.
3:48:59
I don't think, vibe-wise, people love it
publicly.
3:48:59
But a lot of people use it.
3:49:09
So if you look to Reddit and
X, they don't really give it praise from the programming
community, but they use it.
3:49:17
And the same thing with probably Llama.
3:49:17
I don't understand the dynamics of either positive hype or negative
hype. I don't understand it.
3:49:25
- I mean, one of the stories of 2025
is the US filling the gap of Llama, which is the rise of these Chinese
open-weight models, models- to the point where that was the single
issue I've spent a lot of energy on lately, trying to do policy work to
get the US to invest in this.
3:49:41
- So just tell me the story of ADAM.
3:49:43
- The ADAM Project started as me
calling it the American DeepSeek Project, which doesn't really work for
DC audiences, but it's the story of the most impactful thing I can
do with my career, which is that these Chinese open-weight models
are cultivating a lot of power, and there is a lot of demand for
building on these open models, especially in enterprises in the US that
are very cagey about Chinese models.
3:50:06
- The ADAM Project, American Truly Open Models, is a US-based initiative
to build and host high-quality, genuinely open-weight AI models
and supporting infrastructure explicitly aimed at competing
with and catching up to China's rapidly advancing
open-source AI ecosystem.
3:50:25
- I think the one-sentence summary
would be that... or two sentences.
3:50:29
One is a proposition that open models
are going to be an engine for AI research because that is what people
start with; therefore, it's important to own them.
3:50:37
And the second
one is, therefore, the US should be building the best models so
that the best research happens in the US, and those US companies
take the value from being the home of where AI research is
happening.
3:50:49
And without more investment in open models—we have plots
on the website where it's like, "Qwen, Qwen, Qwen, Qwen"—it's
all these models that are excellent from these Chinese companies
that are cultivating influence internationally.
3:51:04
I think the US is spending way more on AI, and the
ability to create open models that are a generation beyond what
the cutting edge of closed labs costs roughly $100 million,
which is a lot of money, but not a lot of money to these
companies.
3:51:20
Therefore, we need a centralizing force of people who
want to do this.
3:51:24
And I think we got engagement from people pretty
much across the full stack, whether it's policy.
3:51:33
- So there has been support
from the administration?
3:51:36
- I don't think anyone
technically in government has signed it publicly, but I know
people that have worked in AI policy, in both the Biden and Trump
administrations, are very supportive of promoting open-source models
in the US.
3:51:48
I think, for example, AI2 got a grant from the NSF
for $100 million over four years, which is the biggest CS grant the NSF has ever awarded, and it's for AI2
to attempt this. It's a starting point.
3:52:05
But the best thing happens when there are
multiple organizations building models, because they can cross-pollinate
ideas and build this ecosystem.
3:52:13
I don't think it works if it's
just Llama releasing models, because Llama could go away.
3:52:17
The same thing applies for AI2; I can't be the only one building models.
3:52:21
It becomes a lot of time spent on talking to people, whether in
policy...
3:52:29
I know NVIDIA is very excited about this.
3:52:33
I think Jensen
Huang has been talking about the urgency for this, and they've done a
lot more in 2025, where the Nemotron models are more of a focus.
3:52:41
They've started releasing some data along with NVIDIA's
open models, and very few companies do this, especially
of NVIDIA's size, so there are signs of progress.
3:52:52
We hear about
Reflection AI, where they say their two billion dollar fundraise is dedicated
to building US open models, and I feel and their announcement tweet
reads like a blog post, right?
3:53:06
I think that cultural tide is
starting to turn.
3:53:06
In July, four or five DeepSeek-caliber
Chinese open-weight models and and zero from the US.
3:53:16
That's
the moment where I realized, like, "Oh, I guess I have to spend energy on
this because nobody else is gonna do it."
3:53:24
So it takes a lot of people contributing
together, and I don't say that, the Adam Project isn't the thing that's
helping to move the ecosystem, but it's people like me doing this sort
of thing to get the word out.
3:53:35
- Do you like the 2025 America's
AI Action Plan?
3:53:35
That includes open source stuff.
3:53:39
The
White House AI Action Plan includes a dedicated section titled "Encourage
Open-Source and Open-Weight AI," defining such models and
arguing they have unique value for innovation and startups. - Yeah.
3:53:52
I mean, the AI
Action Plan is a plan, but largely, I think it's maybe the most coherent policy document that has come
out of the administration, and I hope that it largely succeeds.
3:54:04
I know people
that have worked on the AI Action Plan and the challenges of taking policy and
making it real.
3:54:08
I have no idea how to do this as an AI researcher,
but largely a lot of things in that were very real, and there's
a huge build-out of AI in the country.
3:54:19
There are a lot of issues that
people are hearing about, from water use to whatever, and we should be able to
build things in this country, but also, we need to not ruin places in
our country in the process of building it, and it's worthwhile
to spend energy on.
3:54:31
I think that's a role the federal
government plays. They set the agenda.
3:54:38
And with AI, setting the agenda
that open-weight should be a first consideration is a large part of what they can do and then
people think about it.
3:54:49
- Also, for education and talent
for these companies, it's very important because otherwise,
if there are only closed models, how do you get
the next generation of people contributing at some point?
3:54:59
Because otherwise, you will point only be able to learn after you joined a company.
3:55:06
But at
that point, how do you hire talented people?
3:55:10
How do you identify
talented people?
3:55:10
I think open source is essential for a lot of
things, but also even just for educating the population and
training the next generation of researchers.
3:55:21
It's the
way, or the only way.
3:55:24
- The way that I could've gotten this
to go more viral was to tell a story of Chinese AI integrating
with an authoritarian state, being ASI and taking over the world, and therefore
we need our own American models.
3:55:31
But it's very intentional why I talk about
innovation and science in the US because I think it's both more
realistic as an outcome, but also it's a world that I
would like to manifest.
3:55:47
- I would say, though, also
even any open-weight model, I do think, is a valuable model. - Yeah.
3:55:55
And my argument is that we should be
in a leading position.
3:55:55
But I think it's worth saying it simply because there
are still voices in the AI ecosystem that say we should consider banning the
release of open models due to safety risks.
3:56:09
And I think it's worth adding
that, effectively, that's impossible without making the US have
its own great firewall, which is also known to not work that well because the cost for training these models, whether
it's one to a hundred million dollars, is attainable to a huge amount of people in the world that want to have influence, so
these models will be trained all over the world.
3:56:32
And we want the models,
especially when, like, I mean, there are safety concerns, but we
want this information and tools to flow freely across the world and into the
US so that people can use them and learn from them.
3:56:45
Stopping that would be
such a restructuring of our internet that it seems impossible.
3:56:51
- Do you think maybe in that case the
big open-weight models from China are actually a good thing in a sense, like,
for the US companies?
3:56:55
Because maybe the US companies you mentioned earlier
are usually one generation behind in terms of what they release open source
versus what they are using?
3:57:03
For example, gpt-oss might not be the cutting-edge
model.
3:57:07
Gemini 3 might not be, but they do that because they know this
is safe to release.
3:57:10
But then when they see, these companies see, for
example, there is DeepSeek-V3.
3:57:14
2, which is really awesome, and
it gets used and there is no backlash, there is no security risk, that
could then, again, encourage them to release better models.
3:57:26
Maybe that, in
a sense, is a very positive thing. - A hundred percent.
3:57:30
These Chinese companies
have set things into motion that I think would potentially not have happened
if they were not all releasing models.
3:57:38
So I think it was like I'm almost sure
that those discussions have been had by leadership.
3:57:45
- Is there a possible future where
the dominant AI models in the world are all open source?
3:57:50
- Depends on the trajectory of progress that
you predict.
3:57:50
If you think saturation in progress is coming within a
few years, so essentially, within the time where financial support
is still very good, then open models will be so optimized and so much cheaper
to run that they'll win out.
3:58:02
This goes back to open source ideas
where so many more people will be putting money into optimizing
the serving of these open-weight common architectures
that they will become standards, and then you could have chips dedicated to
them, and it'll be way cheaper than the offerings from these closed
companies that are custom.
3:58:25
- We should say that the AI27
report kinda predicts one of the things it does from a narrative
perspective is that there will be a lot of centralization.
3:58:32
As the AI
systems get smarter and smarter, national security
concerns will arise, and you'll centralize the labs, and they'll
become super secretive, and there'll be this whole race - ...
3:58:45
from a military perspective of how
do you...
3:58:45
between China and the US.
3:58:48
And so all of these fun
conversations we're having about LLMs...
3:58:52
the generals and the soldiers will come into the room and be like, "All
right.
3:58:56
We're now in the Manhattan Project stage of this whole thing."
3:59:02
- I think in 2025, '26, '27, I don't think something like that is even remotely
possible.
3:59:06
You can make the same argument for computers, right?
3:59:10
You can say,
"Computers are capable and we don't want the general public to get them."
3:59:14
Or chips,
even AI chips, but you see how Huawei makes chips now.
3:59:21
It took
a few years, but...
3:59:21
and I don't think there is a way you
can contain knowledge like that.
3:59:29
I think in this day and age, it is impossible, like the internet.
3:59:33
I
don't think this is a possibility.
3:59:37
- On the Manhattan Project
thing, I think that a Manhattan Project-like thing for open
models would be pretty reasonable, because it wouldn't cost that much.
3:59:45
But I think that will come.
3:59:48
It seems like culturally, the companies
are changing.
3:59:48
But I agree with Sebastian on all of that.
3:59:52
I don't see it
happening nor being helpful. - Yeah.
3:59:58
The motivating force
behind the Manhattan Project was civilizational risk.
4:00:02
It's harder to
motivate that for open-source models.
4:00:08
- There's no civilizational risk.
4:00:10
- On the hardware side,
we mentioned NVIDIA a bunch of times.
4:00:14
Do you think Jensen
and NVIDIA will keep winning?
4:00:18
- I think they have to iterate
and manufacture a lot.
4:00:22
And I think they probably...
4:00:22
what they're doing, they do innovate, but I think there's always the chance that someone
does something fundamentally different, gets very lucky, and
then does something.
4:00:34
But the problem is adoption.
4:00:38
The moat of NVIDIA is probably not just the GPU.
4:00:42
It's
more like the CUDA ecosystem, and that has evolved over two decades.
4:00:45
Even back when I was a grad student, I was in a lab
doing biophysical simulations, molecular dynamics, and we had a
Tesla GPU back then just for the computations.
4:00:57
It was about
15 years ago now.
4:00:57
And they built this up for a long time,
and that's the moat, I think.
4:01:05
It's not the chip itself,
although they have the money to iterate, build, and
scale.
4:01:09
But then it's really about compatibility.
4:01:13
If you're
at that scale, why would you go with something risky where there are only a few chips
they can make per year? You go with the big one.
4:01:21
But
then I do think with LLMs now, it will be easier to
design something like CUDA.
4:01:30
It took 15 years because it was hard,
but now that we have LLMs, we can maybe replicate CUDA.
4:01:35
- And I wonder if there will be a
separation of training and inference compute as we stabilize, and more
compute is needed for inference.
4:01:47
- That's supposed to be the point of
the Groq acquisition.
4:01:47
And that's why part of what Vera Rubin is- where they have a new chip with no
high-bandwidth memory, which is one of the- or very little, which is one of
the most expensive pieces.
4:01:55
It's designed for pre-fill, which
is the part of inference where you essentially do a lot of matrix
multiplications.
4:02:03
And then you only need the memory when you're doing this autoregressive
generation, and you have the KV cache swaps.
4:02:11
So they have this new GPU
that's designed for that specific use case, and then the
cost of ownership per FLOP or whatever is actually way
lower.
4:02:19
But I think that NVIDIA's fate lies in the diffusion
of AI still.
4:02:22
Their biggest clients are still these hyperscale
companies.
4:02:26
Like, Google obviously can make TPUs.
4:02:30
Amazon is making Trainium.
4:02:34
Microsoft will try to do its own things.
4:02:34
And so long as the pace of AI progress is high, NVIDIA's platform is the most
flexible and people will want that.
4:02:40
But if there's stagnation, then creating bespoke
chips, there's more time to do it.
4:02:50
- It's interesting that NVIDIA
is quite active in trying to develop all kinds of different products.
4:02:55
- They try to create areas of commercial
value that will use a lot of GPUs. - Mm-hmm.
4:03:01
But they keep innovating
and they're doing a lot of incredible research, so...
4:03:06
- Everyone says the company's super
oriented around Jensen and how operationally plugged in he is.
4:03:11
And
it sounds so unlike many other big companies that I've heard about.
4:03:15
And so
long as that's the culture, I think that we can expect that to keep progress
happening.
4:03:19
And it's like he's still in the Steve Jobs era of Apple.
4:03:22
So long as that is how it operates, I'm pretty optimistic for their situation because it's like, it is
their top-order problem, and I don't know if making these chips for the
whole ecosystem is the top goal of all these other companies.
4:03:38
They'll do a good
job, but it might not be as good of a job.
4:03:43
- Since you mentioned Jensen,
I've been reading a lot about history and about singular figures in
history.
4:03:47
What do you guys think about the single man/woman view of
history?
4:03:51
How important are individuals for steering the direction
of history in the tech sector?
4:03:58
So, you know, what's NVIDIA
without Jensen?
4:03:58
You mentioned Steve Jobs.
4:04:02
What's Apple
without Steve Jobs?
4:04:02
What's xAI without Elon or DeepMind without Demis?
4:04:11
- People make things earlier and
faster, whereas scientifically, many great scientists credit being in the
right place at the right time and still making the innovation, where eventually
someone else will still have the idea.
4:04:25
So I think that in that way, Jensen
is helping manifest this GPU revolution much faster and much more
focused than it would happen without having a person there.
4:04:37
And this is
making the whole AI build-out faster.
4:04:37
But I do still think that eventually,
something like ChatGPT would have happened and a build-out like this would have
happened, but it probably would not have been as fast.
4:04:48
I think that's the
sort of flavor that is applied.
4:04:55
- These individual people, there are people
who are placing bets on something.
4:04:58
Some get lucky, some don't.
4:04:58
But if you don't
have these people at the helm, it would be more diffused.
4:05:02
It's almost like investing
in an ETF versus individual stocks.
4:05:06
Individual stocks
might go up or down more heavily than an ETF, which is more balanced.
4:05:10
It will eventually go up over time. We'll get there.
4:05:14
But it's just
like, you know, the focus I think is the thing. Passion and focus.
4:05:19
- Isn't there a real case to be made
that without Jensen, there's not a reinvigoration of the
deep learning revolution?
4:05:26
- It could've been 20 years
later, is what I would say.
4:05:30
Or like another AI winter could
have come if GPUs weren't around.
4:05:35
- That could change history completely
because you could think of all the other technologies that could've come in the meantime, and the
focus of human civilization would get...
4:05:44
Silicon Valley would be
captured by different hype.
4:05:48
- But I do think there's certainly
an aspect where it was all planned, the GPU trajectory.
4:05:52
But on the
other hand, it's also a lot of lucky coincidences or good
intuition.
4:05:55
Like the investment into, let's say, biophysical
simulations.
4:05:59
I mean, I think it started with video games and then it just
happened to be good at linear algebra because video games require a lot of linear
algebra.
4:06:07
And then you have the biophysical simulations.
4:06:11
But still, I
don't think the master plan was AI.
4:06:15
I think it happened to be Alex Krizhevsky.
4:06:19
So someone took
these GPUs and said, "Hey, let's try to train a neural network on that."
4:06:23
It happened to work really well, and I think it only happened because
you could purchase those GPUs.
4:06:30
- Gaming would've created a
demand for faster processors if NVIDIA had gone out of
business in the early days.
4:06:37
That's what I would think.
4:06:37
I think that the GPUs would've been different,
but I think GPUs would still exist at the time of AlexNet and at the time
of the Transformer.
4:06:46
It was just hard to know if it would be one company
as successful or multiple smaller companies with worse chips.
4:06:53
But I don't think that's a 100-year delay.
4:06:58
It might
be a decade delay.
4:07:01
- Well, it could be one, two, three,
four, five-decade delay.
4:07:01
I just can't see Intel or AMD doing what NVIDIA did.
4:07:08
- I don't think it would be
a company that exists.
4:07:11
I think it would be a different
company that would rise.
4:07:13
- Like Silicon Graphics or something.
4:07:15
- So yeah, some company that
has died would have done it.
4:07:19
- But just looking at it,
it seems like these singular figures, these
leaders, have a huge impact on the trajectory of the world.
4:07:27
Obviously,
there are incredible teams behind them.
4:07:31
But, you know, having that kind of
very singular, almost dogmatic focus- -is necessary to make progress.
4:07:40
- Yeah, I mean, even with GPT, it wouldn't
exist if there wasn't a person, Ilya, who pushed for this scaling, right?
4:07:47
- Yeah, Dario Amodei was also
deeply involved in that.
4:07:50
If you read some of the histories from
OpenAI, it seems wild thinking about how early these people were like, "We need
to hook up 10,000 GPUs and take all of OpenAI's compute and train one model."
4:07:57
There were a lot of people who didn't want to do that.
4:08:02
- Which is an insane thing to
believe.
4:08:02
To believe in scaling before scaling has any
indication that it's going to materialize. Again, singular figures.
4:08:09
Speaking of which, 100 years from now, this is presumably post-singularity,
whatever singularity is.
4:08:21
When historians look back
at our time now, what technological breakthroughs
would they really emphasize as the breakthroughs
that led to the singularity?
4:08:32
So far we have Turing to today, 80 years.
4:08:36
- I think it would still be computing,
like the umbrella term "computing."
4:08:40
I don't necessarily think that in 100 or 200 years it would be AI.
4:08:44
It could still very well be computers.
4:08:47
We are now taking
better advantage of them, but the fact of computing remains.
4:08:53
- It's basically a Moore's
Law discussion.
4:08:53
Even the details of CUDA and GPUs
won't even be remembered, nor will all this software turmoil.
4:09:01
It'll just be, obviously, compute.
4:09:07
- I generally agree, but is
the connectivity of the internet and compute able to be
merged? Or is it both of them?
4:09:17
- I think the internet will probably
be related to communication.
4:09:21
It could be a phone, the
internet, or satellites.
4:09:25
Compute is more like the
scaling aspect of it.
4:09:29
- It's possible that the internet
is completely forgotten- -that the internet is wrapped into phone
networks, like communication networks.
4:09:38
This is just another manifestation
of that, and the real breakthrough comes from increased compute, or
Moore's Law, broadly defined.
4:09:46
- Well, I think the connection of
people is very fundamental to it.
4:09:50
it's like, you can talk to
anyone.
4:09:50
You want to find the best person in the world for something, they are
somewhere in the world.
4:09:54
And being able to have that flow of information—the
AIs will also rely on this.
4:10:02
I've been fixating on when I
said the dream was dead about the one central model.
4:10:06
The thing that is
evolving is people having many agents for different tasks.
4:10:10
People already
started doing this with different clouds.
4:10:14
It's described as many
AGIs in the data center where each one manages and they
talk to each other.
4:10:18
And that is reliant on networking and the
free flow of information. on top of compute.
4:10:26
But
networking, especially with GPUs, is such a part of
scaling of compute.
4:10:29
The GPUs and the data centers
need to talk to each other.
4:10:36
- Anything about neural networks will
be remembered?
4:10:36
Like, do you think there's something very specific and
singular to the fact that it's neural networks that's seen as a breakthrough,
like a genius, that you're basically replicating, in a very
crude way, the human mind?
4:10:47
The structure of the human
brain, the human mind?
4:10:54
- I think without the human mind,
we probably wouldn't have neural networks, because it just was an
inspiration for that.
4:10:58
But on the other end, I think it's just so,
so different.
4:11:02
I mean, it's digital versus biological, that I do
think it will probably be more grouped as an algorithm.
4:11:11
- That's massively parallelizable... ...
4:11:11
On this particular kind of compute?
4:11:15
- It could have been like genetic
computing; genetic algorithms just parallelized.
4:11:19
It just happens that this
is more efficient and works better.
4:11:23
- And it very well could be that the
LLM, the neural networks, the way we architect them now is just
a small component of the system that leads to singularity.
4:11:33
- If you think of it in 100
years, I think society can be changed more with more compute
and intelligence because of autonomy.
4:11:41
But looking
at this, what are the things from the Industrial Revolution that we
remember?
4:11:45
We remember the engine, which is probably the equivalent
of the computer in this.
4:11:51
But there's a lot of other physical
transformations that people are aware of, like the cotton gin and
all these things, these machines that are still known:
air conditioning, refrigerators.
4:12:04
Some of these things from
AI will still be known.
4:12:08
The word "transformer" could still
be known.
4:12:08
I would guess that deep learning is definitely still known, but
the transformer might be evolved away from in 100 years with AGI researchers
everywhere.
4:12:16
But I think deep learning is likely to be a
term that is remembered.
4:12:28
- And I wonder what the air conditioning and
refrigeration of the future is that AI brings.
4:12:32
If we travel forward 100 years from
now, we transport there right now, what do you think is different?
4:12:36
How do
you think the world looks different?
4:12:40
First of all, do you think there are
humans?
4:12:40
Do you think there are robots everywhere walking around?
4:12:46
- I do think specialized robots,
for sure, for certain tasks. - Humanoid form? - Maybe half-humanoid. We'll see.
4:12:54
I think for certain things, yes, there
will be humanoid robots because it's just amenable for the environment.
4:12:57
But for certain tasks, it might make sense.
4:13:01
What's harder
to imagine is how we interact with the devices and what humans
do with devices.
4:13:05
Well, I mean, I'm pretty sure it will probably
not be the cellphone or the laptop. Will it be implants?
4:13:16
- I mean, it has to be
brain-computer interfaces, right?
4:13:18
I mean, 100 years from now, given
the progress we're seeing now— there has to be...
4:13:22
unless
there's legitimately a complete alteration of how we
interact with reality.
4:13:33
- On the other hand, cars are
older than 100 years, right?
4:13:33
And it's still the same interface.
4:13:37
We haven't replaced cars with something else.
4:13:40
We just made them
better, but it's still a steering wheel, still wheels, you know?
4:13:45
- I think we'll still carry around
a physical brick of compute because people want some ability
to have a private...
4:13:48
Like, you might not engage with it as much as a
phone, but having private information that is yours as an interface
between the rest of the internet, I think that will
still exist.
4:13:59
It might not look like an iPhone and it might be
used a lot less, but I still expect people to carry things around.
4:14:08
- Why do you think the smartphone
is the embodiment of private? There's a camera on it.
4:14:11
There's— - Private for you, like encrypted
messages, encrypted photos... know what your life is.
4:14:22
I guess it's a question of how optimistic
on brain-machine interfaces you are.
4:14:26
Is all that just going to be stored
in the cloud? Your whole calendar?
4:14:30
It's hard to think about processing all the information that
we can process visually through brain-machine interfaces
presenting something like a calendar or something to you.
4:14:44
It's hard to just think about knowing,
without looking, your email inbox.
4:14:49
Like you signal to a computer and
then you just know your email inbox.
4:14:53
Is that something that the human
brain can handle being piped into it non-visually?
4:14:56
I don't know
exactly how those transformations happen.
4:15:03
Humans aren't changing in 100 years.
4:15:05
I think agency and community are
things that people actually want.
4:15:09
- A local community, yeah.
4:15:10
- People you are close to, being
able to do things with them and being able to ascribe meaning to
your life and being able to do things.
4:15:22
In 100 years, I don't think
that human biology is changing away from those on a
time scale that we can discuss.
4:15:30
And I think that UBI
does not solve agency.
4:15:34
I do expect mass wealth, and
I hope that it has spread so that the average life looks
very different in 100 years.
4:15:42
But that's still a lot to happen.
4:15:42
If you think about countries that are early in their
development process to getting access to computing and internet, to
build all the infrastructure and have policy that shares one
nation's wealth with another is...
4:16:00
I think it's
an optimistic view to see all that happening in 100 years- ...
4:16:05
while they are still
independent entities and not just like absorbed into some
international order by force.
4:16:13
- But there could be just better,
more elaborate, more effective ...
4:16:17
social support systems
that help alleviate some levels of basic suffering
from the world.
4:16:21
You know, the transformation of society where a lot
of jobs are lost in the short term, I think we have to really remember
that each individual job that's lost is a human being who's
suffering. That's like a ...
4:16:37
When jobs are lost, the
scale is a real tragedy.
4:16:37
You can make all kinds of arguments about
economics or how it's all going to be okay.
4:16:45
It's good for the GDP,
there's going to be new jobs created.
4:16:49
Fundamentally at the
individual level for that human being, that's real
suffering.
4:16:53
That's a real personal sort of tragedy.
4:16:57
And
we have to not forget that as the technologies are being
developed.
4:17:01
And also my hope for all the AI slop we're seeing is that there will be a greater and
greater premium for the fundamental aspects of the human
experience that are in-person.
4:17:17
The things that we all...
4:17:17
Like seeing
each other, talking together in-person.
4:17:22
- The next few years are definitely
going to be an increased value on physical goods and events— ...
4:17:26
and
even more pressure on slop. So it'll be...
4:17:32
the slop is only starting.
4:17:35
The next few years will be more and
more diverse ... versions of slop.
4:17:38
- They would be drowning
in slop.
4:17:38
Is that what— - So I'm hoping that society
drowns in slop enough to snap out of it and be like, "We can't deal
with it. It just doesn't matter."
4:17:48
And then, the physical has
such a higher premium on it.
4:17:53
- Even like classic examples, I
honestly think this is true, and I think we will get tired of it.
4:17:57
We are already kind of tired of it. I mean, even art.
4:18:01
I don't think art
will go away.
4:18:01
You have paintings, physical paintings.
4:18:05
There's
more value, not just monetary value, but just more
value appreciation for the actual painting than a photocopy of that
painting.
4:18:13
It could be a perfect digital reprint, but there is something when you go to a
museum and you look at that art and you see the real thing and you just think,
"Okay. A human." It's like a craft.
4:18:24
You have like an appreciation for that.
4:18:24
And
I think the same is true for writing, for talking, for any type of experience...
4:18:32
I do unfortunately think it
will be like a dichotomy, like a fork where some
things will be automated.
4:18:40
Like, you know, there are not as many paintings
as there used to be, you know, 200 years ago.
4:18:43
There are more photographs,
more photocopies.
4:18:43
But at the same time, it won't go away.
4:18:47
There will be value in that.
4:18:51
I think the difference will just be,
you know, what's the proportion of that.
4:18:55
But personally, I have a hard time
reading things where I obviously see it's obviously AI generated. I'm sorry.
4:19:00
It might—it might be really good information there,
but I'm just like, "Nah, not for me."
4:19:08
- I think eventually they'll fool
you, and it'll be on platforms that give ways of verifying or
building trust.
4:19:12
So you will trust that Lex is not AI generated, having been
here.
4:19:16
So then you have trust in this channel.
4:19:21
But it's harder for new
people who don't have that trust.
4:19:25
- Well, that will get interesting
because I think fundamentally it's a solvable problem by having trust in certain outlets that they won't do it,
but it's all going to be trust-based.
4:19:37
There will be systems to authorize, "Okay, this
is real. This is not real."
4:19:37
There will be some telltale signs where you can obviously
tell this is AI generated and this is not.
4:19:45
But some will be so
good that it's hard to tell, and then you have to trust.
4:19:49
And well, that will
get interesting and a bit problematic.
4:19:54
- The extreme case of this is to
watermark all human content.
4:19:57
So all photos that we take on our own
have some watermark until they are edited or something like this.
4:20:02
And software
can manage communications with the device manufacturer- device
manufacturer to maintain human editing— which is the opposite of the discussion
to try to watermark AI images.
4:20:11
And then you can make a Google image that has a watermark
and use a different Google tool to remove it. - Yep.
4:20:20
It's going to be an
arms race, basically.
4:20:23
- And we've been mostly focusing
on the positive aspects of AI.
4:20:26
All the capabilities
that we've been talking about can be used to destabilize
human civilization with even just relatively dumb AI applied at scale,
and then further, superintelligent AI systems.
4:20:41
Of course, there's
the sort of doomer take that's important to consider
as we develop these technologies.
4:20:49
What gives you
hope about the future of human civilization, given everything we've been
talking about? Are we going to be okay? - I think we will.
4:20:59
I'm
definitely a worrier, both about AI and non-AI things.
4:21:02
But humans do
tend to find a way.
4:21:02
I think that's what humans are built for: to have
community and find a way to figure out problems.
4:21:13
That's what has gotten us
to this point.
4:21:13
And to think that the AI opportunity and related
technologies is really big.
4:21:22
And I think that there's big
social and political problems to help everybody understand that.
4:21:28
And I
think that's what we're staring at a lot of right now, is like the world
is a scary place, and AI is a very uncertain thing.
4:21:35
And it
takes a lot of work that is not necessarily building things.
4:21:40
It's like telling people and understanding people, that the
people building AI are historically not motivated or wanting to do.
4:21:50
But it is something that is
probably doable.
4:21:50
It just will take longer than people want.
4:21:54
And we have to go through that long period of like hard, distraught AI discussions if we want to
have the lasting benefits. - Yeah.
4:22:04
Through that process, I'm
especially excited that we get a chance to better understand ourselves, us at the individual level as
humans and at the civilization level, and answer some of the
big mysteries, like what is this whole consciousness thing going on here?
4:22:22
It seems to be truly special.
4:22:22
Like, there's a real miracle in our mind.
4:22:26
And AI puts a mirror
to ourselves and we get to answer some of the big
questions about like, what is this whole thing going on here?
4:22:35
- Well, one thing about that is also what
I do think makes us very different from AI and why I don't
worry about AI taking over is, like you said, consciousness.
4:22:43
We
humans, we decide what we want to do.
4:22:47
AI in its current
implementation, I can't see it changing.
4:22:51
You have to tell
it what to do.
4:22:51
And so you have still the agency.
4:22:55
It doesn't
take the agency from you because it becomes a tool.
4:22:59
You can think of
it as a tool. You tell it what to do.
4:23:02
It will be more automatic
than other previous tools.
4:23:06
It's certainly more powerful than a
hammer, it can figure things out, but it's still you in charge, right?
4:23:10
So
the AI is not in charge, you're in charge.
4:23:14
You tell the AI what to
do and it's doing it for you.
4:23:17
- So in the post-singularity,
post-apocalyptic war between humans and machines, you're saying
humans are worth fighting for? - 100%. I mean, this is...
4:23:27
The
movie Terminator, they made in the '80s, essentially,
and I do think, well, the only thing I can see
going wrong is, of course, if things are explicitly programmed to do
the thing that is harmful, basically.
4:23:43
- I think actually in that, in a Terminator
type of setup, I think humans win.
4:23:49
I think we're too clever.
4:23:49
It's hard to
explain how we figure it out, but we do.
4:23:56
And we'll probably be using local LLMs, open
source LLMs to help fight the machines.
4:24:04
I apologize for the
ridiculousness.
4:24:04
Like I said, Nathan already knows I've been a big
fan of his for a long time.
4:24:07
Been a big fan of yours, Sebastian, for a long
time, so it's an honor to finally meet you.
4:24:15
Thank you for everything you put out into the
world.
4:24:15
Thank you for the excellent books you're writing.
4:24:19
Thank you for teaching us.
4:24:19
And
thank you for talking today. This was fun.
4:24:26
- Thank you for inviting us here and having
this human connection, which is actually- - Extremely valuable- human connection.
4:24:33
Thanks for listening to this conversation
with Sebastian Raschka and Nathan Lambert.
4:24:37
To support this podcast,
please check out our sponsors in the description, where you can also
find links to contact me, ask questions, give feedback
and so on.
4:24:44
And now let me leave you with some
words from Albert Einstein.
4:24:52
"It is not that I'm so
smart, but I stay with the questions much longer."
4:24:56
Thank you for
listening, and hope to see you next time.