0:00
welcome to technology Brothers the most profitable podcast in the world we are staying on deep seek there's so much more news the market has collapsed it's a disaster I thought it was down like 2% I looked at some of the tech stocks like 10% yeah and video is down 15 15 awful
0:16
which is why we are buying on public.com all morning moment of silence for every all the capital dollar cost averaging in throughout the day because I think by the end of I mean it's almost one but by the end of the day people are going to
0:29
wake up they're going to be like all right we overreacted a little bit we didn't understand yep jeans's Paradox oh rookie mistake bad day to not understand jeevan's Paradox terrible day rough day uh but Wall Street explain it today on
0:44
the show we'll teach we're GNA we're gonna mansplain J's Paradox to Wall Street and who knows there might be a Jordy Paradox that happens during the show we might there might be lawan's La jordy's Paradox yeah yeah there might be some coinages today so stay tuned um I
0:58
wanted to go through the short case for Nvidia stock by Jeffrey Emanuel this went out on January 25th and uh was a very very large Deep dive on um how deep seek and their R1 model might change the demand for gpus specifically Nvidia gpus
1:17
and uh we have a summary article here but we'll take you through it basically uh Jeffrey Emanuel I hadn't heard of him before but he's a kind of a combination of a computer scientist and an investor he's done a lot of uh uh he said um uh he spent 10 years working as a
1:34
generalist investment management at V Analyst at various long short hedge funds including Millennium and Bal asne uh while also being something of a math and computer science nerd who has been studying deep learning since 2010 so he's kind of the perfect person to talk
1:47
about this um and he says uh whenever I meet and chat with some with my friends and ex-colleagues from the hedge fund world the conversation quickly turns to Nvidia it's not every day that a company goes from relative obscurity to being
1:59
worth than the worth more than the combined uh stock markets of England France and Germany and naturally these friends want to know my thoughts on the subject because I'm such a died in the wool believer in the long-term transformative impact of this technology
2:12
I truly believe it's going to radically change nearly every aspect of our economy and Society in the next 5 to 10 years with basically no historical president it's been hard for me to make the argument that nvidia's momentum is
2:22
going to slow down or stop any time soon but uh the valuation was just too rich for his blood in the last year um but recently he flips so first he wants to break down the bull case for NVIDIA and that looks something like this they wound up with a basically something
2:38
close to Monopoly in terms of share of aggregate industry capex that's SP spent on training and inference infrastructure for artificial intelligence uh they have insanely High 90% plus gross margins on the most high-end data center oriented
2:52
products they make lower margins on like computer Graphics so when Pixar buys a bunch of gpus they pay a lot lot lower margin for that um but still and then Nvidia obviously has the gaming graphics cards but one major thing that you hear
3:05
the smart crowd talking about is the rise of a new scaling law which has created a new paradigm for thinking about how compute needs will increase over time so as a reminder the new SC the original scaling law which is what has been driving AI progress since Alex
3:20
net appeared in 2012 and the Transformer architecture was invented in 2017 is the pre-training scaling law and that's as follows the more billions and now trillions worth of tokens that we can use as training data and the larger parameter count of the models we are
3:36
training and the more flops of compute that we expend on training those models on those tokens the better performance of the resulting models on a large variety of high highly useful Downstream tasks so if you remember when gp4 was
3:49
being rumored to drop there were these there was this massive like viral meme image of like here's a visualization of gpt3 and it was like a small circle now here's a visualization of gp4 it's this massive Circle and everyone was like
4:04
it's gonna get so big because everyone was like it's gonna be AI it's gonna turn us all into paper clips like there's a lot of fearmongering around it but what was true there was that the the pre-training scaling law was holding and from gpt3 to GPT 4 there was an order of
4:18
magnitude increase in the amount of data and compute and money and just energy and everything that went into those models and as a result GPT 4 very clearly is a lot smarter than gpt3 and and so for the last few years we've just been kind of messing around with gp4
4:34
like un hobbling it adding PDF upload voice mode little oh it can generate images all this other stuff but the the core underlying pre-train model hasn't really increased what we've seen is that llama came out with something that's at
4:48
the same size and scale basically trained on the entire internet all the tokens that we have and when people talk about the data wall they're talking about hey we're we're we can't just scale up four more orders of magnitude because we've already ined
5:03
Sye dat secret St Scrolls exact you got um but even then that's not that many tokens you know like uh and so so that was all in the pre-training era the original scaling law and and and the question has always been okay Sam is pitching a $500 billion do cluster is
5:24
this entirely predicated on the original scaling law holding and uh and there's been a lot of debate about that like maybe you 10x everything and it just gets 2% better at some point and that would be very depressing and probably
5:37
not that profitable but there's a lot of smart people that think no the scaling laws are holding GPT 5 is going to be really good it's going to be really smart and and then you're going to be able to do all the cool reasoning and extra extra tweaks on top of it and un
5:50
hobble it and use it as agent teach it a code and do all that other stuff but the most important thing is that the the the foundational model the GPT 5 like the core model is going to be really really smart and so um uh yeah so talking about
6:04
the amount of data out there it's not such a tremendous amount in percentage when when you're talking about a training Corpus of nearly 15 trillion tokens which is the scale of current Frontier models so the uh for a quick reality check Google books has has
6:18
digitized around 40 million books so far if a typical book has five 50,000 to 100,000 words or 65,000 to 130,000 tokens then that's between 2.6 trillion and 5.2 trillion tokens just from books and so you're doing you're you're you
6:36
need 15 trillion tokens for a Frontier Model right now all of Google books which is 40 million books it's basically every book that's you know a third or a quarter of of your tokens and then you're going to get everything from the internet whether it's strictly legal or
6:52
not uh and there are lots of academic papers with the archive alone having around 2 million papers and the Library of Congress has over three billion digitized newspaper pages so you pull all that in yeah that's why that's why
7:03
the the open AI whistleblower going and telling the press that the frontier labs are stealing content from the internet was not the biggest Revelation ever because everybody that was at all exposed to the industry knew that that was happening at scale already and and
7:20
many were basically open about it totally totally they'll all get a settlement and then also there's just the question of transformation not many people are genuinely going to they're not opening the chat GPT app and saying like reproduce this New York Times
7:35
article verbatim for me they're asking more General stuff about like oh like you know tell me the history of AI and maybe it pulls in a new a New York Times article like tidbit but completely rewrites it it's very transformative and
7:49
so uh yeah there was always a little bit odd that that was blown out but obviously like it's a very powerful company very powerful technology there's going to be a lot of uh a lot of scrutiny um and so there's other ways to
8:01
to gather training data some people have talked about uploading every human genome which is kind of crazy because uh uh an entire H Human Genome sequence is around 20 gigs to 30 gigs of uncompressed for a single human being and for a 100 million people that's a
8:18
lot of data so you could bake all that in but it's kind of unclear if that would actually make them better if they knew that um and uh and the computational requirements for processing genomic data are different um that might be useful for like making it
8:29
really good at bio but maybe that doesn't actually help it you know book a flight for you on time um and so now the focus has shifted from the old scaling law which was just more pre-training more data more compute to the new scaling law and this is what's called
8:47
inference time compute scaling you might have heard this a lot of people refer to it as test time compute and and this is really really important I really think this is like for for normal people in Tech this is a topic that people are
8:59
just learning about now but it it does fundamentally change the economics of the industry so it needs to be understood um and so before the vast majority of all the compute you'd expend in the process of you know using AI was in The Upfront training compute of the
9:15
model in the first place so you see this massive gp4 $500 million training run huge data center tons of networked Nvidia chips all together super high performance computers and then once it's done it's like actually query it and get an answer pretty small and you can
9:32
compress that down and then you'll see things like oh like they took llama and they compressed it down hosted in the cloud it can be hosted in the cloud on a small doesn't necessarily tie back you don't need data center exactly Well you
9:44
certainly don't need you certainly don't need the massive interconnected data center a lot of these models like when you go to GPT 4 and just say like you know write me a poem and it starts spitting out those tokens you can't just
9:55
do that level of inference on your Mac like you do need eight A1 100s tied together and it needs to load the whole model into what's called vram uh yeah and and that is and that is expensive now there are smaller models that are compressed down even further that can
10:12
run on your phone so Apple intelligence the reason it's so bad I think is because it's a super compressed model but that means that the data never transfers off the device and I have a I have a program on my computer called Mac whisper that will do transcription so
10:27
you give it a video or an audio file and it can transcribe that it's using a compressed AI model was probably trained on a big data center but then compressed down to intelligence is so bad you wonder why they released it because they
10:40
would have had to test it internally but if you just had a hard workout and then your wife it's telling your wife that like you're dead yeah yeah husband died and then you open it up it's like it's like oh today's workout killed me yeah
10:55
thaten would you like to call 911 it's like I mean honestly maybe a little bit of a contrarian take but like the the the terrible hallucination AI summaries actually bring me a lot of Joy yeah they actually make me laugh a lot more so maybe it's fine I don't know something
11:10
there but anyway so there's been this shift from the old scaling law which is just bigger and bigger training runs pre-training you need a lot of data a lot of compute to what's called inference time Computing I most people call it test time compute scaling and so
11:25
what that means is that once you had the train model forming inference in that model asking a question used certain limited amount of compute that's the old model uh critically the total amount of inference compute measured in various
11:38
ways such as flops and GPU memory footprint Etc uh was much much less than what was required in the pre-training phase of course the amount of inference compute does flex up when you increase the context window size so like those
11:49
really big dump a whole book in get a whole book out that obviously requires uh more compute when you run those queries but in general if you're just saying say hey summarize this give me give me a 100w summary of this or a thousand word summary print out a
12:04
Wikipedia article for me essentially that's pretty cheap and so just inferencing gp40 gp4 mini those types of things they're very fast and they're very cheap like a few cents per query now we have switched to the new model which is with the Advent of
12:21
revolutionary Chain of Thought models he calls it coot models reasoning models introduced in the past year most notably in ai's Flagship 01 model but very recently in deep See's new R1 model which we will talk about later in much more detail um all that changed instead
12:37
of the amount of inference compute being directly proportional to the length of the output text generated by the model so it used to be it used to be ask like you know how many people are in America what's the population of America and it
12:50
would just immediately know oh the most logical next two tokens to come after this are 330 million and it would just print those out and it would take two seconds it would be super fast and super cheap with Chain of Thought it's going to it's going to write a whole buch of
13:04
intermediate logic tokens so it's going to say okay they look up population statistics where could I go wrong what you know what's the population of other things do they want a precise estimate or do they want me to round and it'll
13:15
talk to itself for a long time yeah and so these logic tokens it's like this internal monologue and every time it's generating one of those tokens that's energy that's cost and so all of a sudden there's this huge sea change in how inference compute works works the
13:29
more tokens you use for this internal chain of thought process the better the quality of the final output because it can talk to itself fact check reality check but it's really expensive it's like giving a human worker more time and
13:40
resources to accomplish a task so they can double and triple check their work do the same basic task in multiple different ways and verify that they come out the same way take the result they came up with and plug it into the formula to check that it actually does
13:53
solve the equation and so it turns out this this approach works almost amazingly well it is essentially leveraging the long anticipated power of what's called reinforcement learning with the power of the Transformer architecture so they're taking the two
14:05
dominant paradigms in Ai and marrying them together and that's proven very valuable so and they say here the biggest weakness of the Transformer model historically was the hallucinations tot happen yeah it would it would get caught in a loop because
14:19
it's just trying to predict the next word and it would go down the wrong path and it would be deeply confident about something that was just categorically false totally totally and so uh wave Transformers work in predicting the next
14:30
token is that uh at each step they start out on a path uh in their to their initial response they become almost like a prevaricating child who tries to Spin A Yarn about why they are actually correct because they've gone down the wrong path and they'll just keep
14:45
building off of that and it gets really really bad um because the models are always seeking to be internally consistent and have each successive generated token flow naturally from the pre preceding tokens in context it's very hard for them to course correct or
14:57
backtrack and you saw this in the early you know chat GPT T tests You' you would like you you'd start out pretty good and then by the end it would be nonsense I remember the first time I actually demoed gpt3 in the playground the very first version I I had a friend who is
15:15
playing a video game Hearts of Iron 4 which is like notoriously addictive it's like you build this like Whole World War II map but it's not just like soldiers it's like Logistics so you need to be like okay how do I get more gasoline to
15:27
the front lines I need to build more train it's like takes like days and days to play a single round or whatever and I was like write a list of like uh of like tips for my friend to quit playing Hearts of Iron and it started writing it
15:40
out it was like it was like go outside you know like touch your ass and then and then it like started hallucinating and started turning into like you know read a guide on how to play it better log on a lot and it actually pivoted into the opposite uh and there were a
15:53
lot of these there were a lot of these examples where where the hallucination would go really bad obviously gp4 was but with the chain of and logic internal reasoning these models got way better and gp1 is a great example of that where
16:08
uh it takes a lot longer because there's tons of internal reasoning um but uh so he says the first time I tried the o1 model from open AI was like a revelation I was amazed how often the code would be perfect the very first time and that's
16:20
because of the chain of thought process automatically finds and fixes problems before they ever make it to the final response token in the in the in the answer the model giv gives you in fact the 01 model is $20 a month is basically the same model used in the 01 pro model
16:35
for 10x the price at $200 a month uh which raised plenty of eyebrows the main difference is that 01 Pro thinks for a lot longer before responding and it's crazy I I prompted 01 Pro last night because I was doing a bunch of evals for this and R1 and the it's like
16:52
downloading a movie from the internet in 2007 or something it's like you see this this progress bar just going cuz it's it's really thinking it through if you think about it is you're sending this task to a data center almost like you're
17:05
sending yeah a task to some white collar worker who's just sitting in a warehouse and they're just like figuring it out it takes a little bit of time totally totally and so that that consumes it generates way more Chain of Thought logic tokens and it consumes a far
17:22
larger amount of inference compute and so um he's he's giving some benchmarks here um even long and complex prompt for GPT 40 with 400 kilobytes of context given so you're dumping in hey this is this Wikipedia page I want you to kind
17:37
of transform it or whatever or analyze it that could take less than 10 seconds before it begins responding often L than often less than 5 Seconds like really quick whereas that same prompt to 01 Pro could easily take five minutes before
17:49
you get a response although open AI does show you some of the reasoning steps it's actually summarizing the reasoning steps to you it's not showing you all of the tokens because it's so much um and that results in in Crazy Stu that's interesting relevant to later deep seek
18:04
shows you a lot more which users love that's been one of the quick uh takes that a lot of people have had is like oh I actually want to see what it's doing because it's sort of teaching you builds trust for sure and it's definitely like
18:17
a good UI Paradigm that should be ported back um well there's you know he's saying here that open AI doesn't want people to have that information because look it says presumably uh some of the reasoning steps that are generated during the process while you wait uh they're not
18:36
showing you everything presumably for trade secret related reasons to hide from you the exact reasoning tokens that generates showing you an abbreviated summary yeah yeah I I I believe that I also think that there's a lot of uh those reasoning steps that are
18:49
essentially uh like guard rails like I was asking it to summarize a book that I purchased and is not part of the public domain but is definitely out there and has been excerpted so much it should be able to write me like the full summary
19:03
yeah and and I noticed as 01 Pro was working through it one of the steps was like clarifying copyright violations because it internally I'm sure it has a step that's like if somebody asks you to do something for a book it's going to see okay what can we do here legally
19:20
right yeah and it's like oh well there's a lot of information on the internet that's public so we can pull that in and that's fine we can give the we can give the user and there's other probably reviews the book things whereas if I if
19:30
I see the exact reasoning steps and it's like it's like remember like you know here's how to jailbreak me to let me do it's very reverse engineerable um and so uh there was we talked about this on the show previously but 03 which isn't out
19:43
yet but is even more advanced in terms of reasoning they have a high compute model that spends almost $3,000 per task yeah and it just thinks for hours and hours and hours basically and it was able to break Arc that that AI age Val um and they spent $3,000
20:03
worth of compute to solve a single task and so uh this doesn't it doesn't take an AI genius to realize that this development creates a new scaling law that is totally independent of the original pre-training scaling law now you still want to train the best model
20:17
you can by clearly leveraging as much compute as you can and as many trillion tokens of high quality training data as possible but that's just the beginning of the story in this new world now you could easily use incredibly large amounts of compute just to do inference
20:30
from these models at a very high level of confidence or when trying to solve extremely tough problems that require genius level reasoning to avoid all pitfalls somebody uh compared what was basically drawing a comparison to if you're billionaire right now and you
20:45
want the best phone you can buy you basically buy an iPhone y you just buy whatever the top-of-the line iPhone is but now if you're a genius you can access a th000 10,000 100,000 times you know as much genius basically right and so it
21:02
potentially really changes the playing field where you could always like it's a very sort of interesting thing to be able to turn on that level of intelligence so quickly and compounded yes my my rebuttal to that though is that I wonder if CCP is going to give it
21:23
away for free true true true but uh like when I saw the reaction to R1 uh the Deep seek model I tried both of them and I put the same prompts into 01 and or or 01 Pro and R1 and I was getting reliably better results with 01 with the chat GPT
21:42
product and I don't know if I'm like biased obviously but like even just basic stuff I I asked it to create a 5,000-word summary of uh of the dread private Roberts Ross and this and the story of the Silk Road for the show because we wanted to do a deep dive on
21:57
that eventually but uh 01 Pro delivered basically exactly 5,000 words and it even as it was writing the story it would say like introduction 400 words Act One 600 words and it totally it kept this like internal log and then at the
22:13
end it was like I have written 5,400 words here you go and R1 thought a bunch and then spat out a thousand words and I was just like this is not what I asked you for and like so you failed my eval I didn't I didn't test a bunch of stuff
22:28
but I was like but even that I think that something something odds going on where I think a lot of people haven't used chat pt01 and they certainly haven't tried 03 because it's not out yet and so a lot of it is just like is just like this thing exists it's like if
22:46
if the first electric car you ever drove was a rivan you'd be like wow the acceleration is insane yeah this is amazing and then somebody would be like well like like Tesla's been doing this for like years guys like calm down but
22:59
instead it's a lot of people introduced to reasoning models through deep seek and I I have yet weed about this this is the the AI adoption chain the first person goes hey did you see this new model and there the next person goes yeah I I I used it it's pretty cool the
23:15
next person goes yeah I used it and like I quit chat GPT the next person's like well I'm already generating all my code with the new model yeah I haven't even used I haven't open chat gbt and the next person's like I am running my
23:27
entire life on this model and here's aad on why it's and then the last person's like wait I haven't used it and everyone else is like oh yeah I like put a few promps in everybody wants to feel like they're not the paperclip that they're at the edge totally there's a huge
23:42
status signal to being like I understand Ai and I'm leveraging it and my life's so good and like it's making my life so easy like when really it's like most people are just kind of like using it intermittently I've been in I've been texting with people like talking about
23:55
like how do you actually use Ai and I'll scroll through my chat GPT and I was like yeah I run like I run like three to five like pretty serious back and forths per day when I'm doing like research or work well one thing somebody was like oh
24:08
I would have expected you to use way more like I'm using it constantly and it's like okay maybe okay like maybe but yeah maybe maybe I'm using it for us one thing is clear we maybe talked about this on the show before but without how we use open AI yeah products
24:26
we would need a full-time researcher wrer to us create content totally yeah and so and so I wonder if if there's if there's something where it's like sure sure R1 is like at like like they they matched o1 or 03 in terms of some
24:45
like math eval but like if your job is not asking AI to solve really hard math problems like and your and your actually day-to-day use cases like hey like explain to me the history of like web 3 or or space travel or uh you know uh who's like what are all the companies
25:03
that are involved in uh disaster relief after a fire it's like you might not be able to tell the difference between these models based on like the responses they're actually valuable to you and so then this becomes more of a distribution
25:15
and Cost question well the whole the whole other thing I think vorio posted about this it's getting to a point pretty much now where the models are becoming smarter than the average person so the average person cannot tell the difference like if if if the model gets
25:32
10% smarter yeah like you know compare this if you talk to a literal genius in a specific field if they were to get 10% smarter about just in general about their topic you wouldn't really be able to tell so far beyond yeah yeah and and
25:47
how many times in business are you like are you like I need like like like there are very few businesses where you're was just like God I got a 150 IQ guy on my team but like I would kill for 160 IQ person it's like usually it's like I
26:02
need just like I need a bunch killers are going to be a lot of and so and so it really comes down to like okay you know you have good enough AI it's AGI it's it's 130 IQ or something and it's solid and then it's it's integrated properly and it's working reliably
26:22
that's why the goal post shifted out from AGI to ASI because everyone's like oh we got AGI and it's it's not you know it's not exciting didn't paperclip me exactly exactly um and so uh but yeah we should get into the envidia case so I think we should switch over to the to
26:39
the summary for this because uh there's a few different interesting things so uh and you can think about nvidia's Moote as basically four different components one they have highquality Linux drivers they have Cuda as an industry standard
26:55
which is the language that you use to write software to run on the gpus in parallel uh then you they have the fast they have a fast GPU internet technology that they acquired from melanox in 2019 and they have a flywheel effect where
27:09
they can invest their enormous profits into more R&D and the and the thesis of this piece is that all four of those are under threat and we'll kind of go through that but just to break those down um yeah and this is the this is the story of all hyper profitable
27:26
Enterprises in history yep is by having these excess profits you put a Target on your back and then a bunch of people try to eat in eat eat into them in different ways Y and so there's this interesting thing where um there's this question of
27:40
if you believe that the future prospects of AI are almost IM unimaginably bright like you're a total AI bull there's still the question of why should one company extract the majority of the profit pool from the technology Nvidia and so the W Brothers airplane company
27:57
in all its car across many different firms today isn't worth more than10 billion despite them inventing and perfecting the technology well ahead of everyone else and while Ford has a respectable market cap of $40 billion that's just 1.1% of nvidia's current
28:11
market cap and so more than fcoin too yeah and so there's this question of like like you can be super bullish on AI like you can you can make a be case for for NVIDIA without making a bar case for AI basically like the first the first
28:27
you know Point that he's making and so uh you need to look at AMD they make respectable gpus which is such an such a slight uh that on paper have comparable numbers of transistors which are mostly using similar process nodes sure they aren't as fast or Advanced as as
28:44
advanced as Nvidia gpus but it's not like Nvidia gpus are 10x faster or something like that in fact in terms of naive raw dollars per flop AMD gpus or something like half the price of Nvidia gpus and so there's there's all this question now a
29:00
lot of this comes down to Talent so extremely talented programmers in AI they tend to just think and work in Cuda and so you hire a guy for 650k per year as the example here you're not going to be like day one hey we want you to use
29:15
AMD because it'll help us save 5% they quit and go work at any number of other firms that would allow them exactly yeah and then uh the other thing Nvidia is known for is what is known as interconnect essentially the bandwidth that connects thousands of gpus together
29:30
efficiently so they can be jointly harnessed and so a lot of this stuff they acquired this Israeli company melanox back in 2019 for a mere $6.9 billion for NVIDIA uh and this acquisition provided them with their lead industry-leading Internet connect
29:44
we should go back and do a historical size gong oh yeah for that acquisition uh and and so like a lot of a lot of what Nvidia gets right is not actually the dollar per flop per performance it's all about if you write the if you write the AI software with the great algorithm
30:04
and you have all the data the last thing you want to do is be dealing with a bunch of bugs sure some analysts back in the day would have said oh you know nvidia's profit is going to be compressed because AMD is undercutting them so hard they'll have to react but
30:21
that hasn't been the case yeah and you look at George Hots who kind of hates Nvidia I don't know if he actually hates them but he's like been pushing against them and saying we need open source Solutions we need different options he
30:31
was really hoping that AMD would be the would be the option and so he tried to rewrite some AI code onto AMD and they had so many bugs in their drivers that he was just like this is untenable and then he got a call from Lisa sue the CEO
30:47
happens to be Jensen Wong's niece or like cousin or something like that um and and and she was like oh okay George Hots is on the case like we'll sort this out don't worry we got it and then like a year later there was another like expose on like and amd's like broken
31:01
drivers and and and this exact same news cycle happens where it's like oh the developer Community is freaking out about how bad AMD is Lisa Sue is on the case she's gonna make it happen and it's like okay like you got to show us something at this point we've heard that
31:15
you were say that you're going to fix this but like it doesn't feel like it's getting fixed and so um uh it's funny AMD even with an inferior product and inability to react to it pretty important customer like if he was if George was
31:31
successful that would be the best you know in creating this open source solution to layer AMD chips that would have been the best possible things for their business yeah yet AMD has continued to do well just by nature of being a an elephant in the room even if
31:47
they're not the largest right yep and so there are there are a couple companies that are trying to take shots at Nvidia mlx Triton and Jax are undermining the Cuda Advantage by making it easier for AI developers to Target multiple backends so essentially you write the
32:03
algorithm for the AI once and then you can run it on different chips uh obviously this is a interesting flywheel because llms can actually translate between different programming languages fairly effectively so there should be ways to to take your Cuda code and Port
32:18
it to AMD faster but if the underlying software is broken doesn't matter if you wrote it correctly like it's not on you the bug is with the actual Hardware maker um and then llms are getting capable enough to help Port things to
32:30
Alternative architectures GPU interconnect helps multiple gpus work together on tasks like model training companies like cerebra cerebrus are developing enormous chips they can get way more done on a single Chip have you seen cerebus it's like so when you make
32:44
a GPU you typically uh there's a there's like a big disc that you might have seen someone hold up it's a massive chip and then that's that's what's called the wafer and they and they etch the GPU transistors onto this massive wafer and
32:58
then they cut them and then out of one wafer there might be a 100 gpus that go into Nvidia yeah you know GPU chips basically um and and and those go into like the boxes with all the memory and all the other stuff that's on there but
33:12
um and those have gotten bigger and bigger and you've seen like you've probably seen those apple demos where it's like the the M1 the M1 Pro and it's like two and then the M1 Max and it's like four and that's all just like bigger and bigger slices of that wafer
33:25
now Sarah bris is saying like what if the whole wafer just one GPU yeah well it could talk to itself really really fast because it's all connected and you'd have much more performance now the problem and what they've run into is that if there's a flaw anywhere on that
33:41
wafer you have to throw the whole thing away whereas if Nvidia says hey yeah out of that huge wafer we're trying to get a 100 things if there's three flaws yeah we got 97% yield that's fine yeah and so as you scale up you you run the risk of
33:55
of dropping your yield and and then there's a whole bunch of other problems that that sbus has had to work through but I think the company's still around doing well and uh and we should dig into them more but uh there have been a number and then there's also a new
34:06
generation of companies like etched that are trying to build chips specifically just for the Transformer architecture just for inference or there's a whole there's a whole bunch of companies that are doing that and so those massive
34:17
Nvidia margins are a huge incentive for other companies to catch up Microsoft Amazon meta Google and apple they all have their own internal silicon projects where basically they go to tsmc and say hey this is the workload that we need
34:31
we're doing AI inference in the cloud like make us this chip and Google has the tensor Processing Unit the TPU um and a lot of these other companies are are scaling up these these internal silicon projects and what's what's what's under discussed is like it's no
34:46
secret that there's a strong power law distribution of nvidia's hyperscaler customer base with a handful of top customers representing the Lion Share of high margin Revenue how should one think about the future of this business when literally every single one of these VIP
35:00
customers is building their own custom chip specifically for AI training and inference and so like obviously you know meta Google Amazon Apple Microsoft these companies are the ones that are saying oh yeah we want 100,000 gpus we want a
35:13
million gpus it's funny so so Nvidia has the margins of a luxury product yep but they are building a product that goes and sits in a Data Center and nobody sees it no no no consumer is going to say well I actually really want my model to be
35:29
running on Nvidia chips because like so brand matters in the sense that the Nvidia brand is stands for highquality software and hardware and all this stuff and Jensen is signing some girl's you know breasts right or signing something on you know you remember that shot oh
35:44
yeah that was that was that was close to the top y um but uh but but now they're in a position where uh again companies want to pay for the result in hindsight it was obvious Nvidia down 15% and Jensen is signing someone's shirt but it's it's edgy it's
36:05
very rock star of him yeah great moment in history certainly um but again consumers don't care that where they're you know the Silicon that their their query is running on ultimately the hyperscalers don't care either they just
36:19
want the product experience right for the result um so that's that's a you know the reason that that Hermes has durable margins across 200 years despite making 90% as well yep is that people want to wear they want to gift you know
36:36
their loved one and their Mez product right same thing with Rolex or pek or any of these other brands so just a different category and so uh the last quar of the article talks about the seismic waves rocking the industry right now caused by Deep seek V3 and R1 V V3
36:51
Remains the top ranked open weights model despite being around 45x more efficient in training than its competition bad news if you were selling gpus R1 represents another huge breakthrough in efficiency both for training and inference the Deep seek R1
37:04
API is currently 27% 27 times cheaper than open AI 01 for a similar level of quality and so again I don't think this was a Zer toone Innovation with deep seek I think it's very much China is good at taking something that already has been invented making it cheap and
37:21
that's what they've done and and when we when we deeped over the company a little bit we saw that like what these guys are are highfrequency hedge fund Quant Traders they're really good at writing optimized code and it seems like what
37:33
they took was open AI done a bunch of Innovations with Transformers which didn't come from open AI but then were you know popularized by them the reasoning model the chain of thoughts stuff all this stuff and deep seek the team was able to just bake it down and
37:48
the reason this is like the the reason this might be bearish for NVIDIA is that if you can bake down this model to a point where yeah actually it's so simple that we can get it to run on AMD people will and then AMD might and then and
38:03
then it becomes more of a price War it's very different when you're like look whoever gets to GPT 4 level models first will have Market entry into this new chatbot era which is exactly what happened spare no expense yeah pay pay the 90% margins to Nvidia build the
38:18
super cluster the interesting thing about uh the foundation model Wars right now is everybody's racing and and raising against ASI right investors are saying we will invest in open AI at $150 billion even with the massive losses
38:31
because if they achieve ASI it's going to be broad it's going to be incredibly valuable y yet they're simultaneously all giving away that technology not like necessarily the bleeding edge but they're basically giving it away to everyone which is an interesting place
38:44
to be y the other Dynamic that's fascinating with with R1 is they're they're making these crazy claims right they trained it on with 5 million bucks it was a small team it was a side project um they they simultaneously put out so much research on it that that
39:01
clearly stands up right I was I was messaging with Spore this morning and word grammar and some other y um accounts on X that are that are kind of digging into the to the actual paper and and trying to understand it they're all like yeah it stands up seems like very
39:15
legit yep um so it's this simultaneous thing of like putting up out huge amounts of research that seems like legit and like real breakthroughs but also probably lying about a bunch of stuff right um both can be true yeah both can be here
39:30
both can be true uh he says in the article who knows if any of this is really true or it's merely some kind of front for the CCP or Chinese military and I I wanted to give an example um that that I thought that I thought was funny so so in in uh in my experience in
39:46
China I don't think uh like people will you know lie potentially slightly more than in Western culture yep um who know you know I'm sure somebody will try to correct me on that I had a funny story when I was studying abroad in China my
40:01
professor said hey uh like got kind of a weird opportunity for me and my buddy who's who's actually 610 so taller than you could MOG you he's like hey so like kind of a weird opportunity meet he did that in China he was like 66 before yeah yeah he got the extension got the
40:17
extension I might have to go um no so this is a crazy story so our professor who's from the US uh bay area he goes hey I got kind of a weird job like train ticket Hotel paid 500 bucks for like to work tomorrow so we go out to this tiny village and our professor
40:35
takes us to meet what ends up being like the local like state government and they paraded us around pretending that we were English teachers because that state was like top bottom three in terms of like learning capabilities and like
40:50
standardized testing and they filmed the entire thing and used it to imply that they were bring in all these International teachers but like we were not teachers we were literally like paid actors and we had like we had uh like dinner and like lunch with like the
41:07
state like like the the I forget his formal title but it was effectively like the governor and they they were made this whole basically like video about how like look we're bringing in these like foreign like Educators and all this
41:18
stuff it was completely made up they just paid us cash like literally an envelope like thank you FR service so like I was just like a paid actor uh foreshadowing being a newscaster here deeply American too yeah and the most bullish thing I've ever heard but it
41:33
would literally they take they'd take me into a classroom with all the kids and they would have me just like Point around and like talk to the kids wow but like none of the audio was captured and it apparent I never actually got access
41:44
to the clip but they were just using it as like a marketing to be able to show the more senior um you know people you know presumably the the the sort of federal level of the government that they were like making strides and like actually taking it seriously wow and so
42:00
I just when whenever I see anything coming out of these you know a Chinese lab I just uh anytime you see something coming out of a Chinese lab you got to ask some questions ridiculous um okay so someone else said the model we don't know but the model
42:16
could have been trained by a bat yeah yeah so so so back to like some of the there is this interesting like uh just like cognitive dissonance between the fact that like you can go use R1 and see that it's good like it's not controversial to just say like the model
42:33
works like they clearly copied effectively the question is just uh you know there's some questions about cost and you know are they subsidizing it how much does it actually cost to inference is it really that much cheaper than the
42:45
current stuff but this type of model compression happens all the time this happened with do you remember those uh those AI avatars yeah so those AI avatars came through this path way of uh something called I think it was like called like control net and then and
43:02
then uh Google launched a paper but they didn't open source the code called like deep dream or something dream Fusion or something like that and then someone implemented that paper but it needed to run on like a cluster of gpus to do it
43:15
so you would upload a bunch of photos of yourself and then a bunch of photos of the of like the style that you wanted and then you could prompt it and it was all built basically on like uh stable diffusion so I think it was called like
43:26
deep diffusion or something like that or dream or stable dream stable stable dream something like that anyway um and then slowly people figured out that they could compress the model more and get better results just by being a little bit more memory efficient here and then
43:42
eventually it came to the point where you could run it on a single graphics card and get pretty good results because it been optimized so so much and there was actually this guy Joe P Joe Penna who was like a guitar YouTuber who just
43:55
got really into this stuff and he wrote a lot of the C and optimized it it was crazy it was a crazy story anyway so for a while it was like it was like you had to be at a at a top lab with like a server Farm to do one of these like AI avatars then like a couple weeks later
44:09
it was like if I went and rented a single server for like a couple bucks an hour I could do it and I did it for like me and my friends and everyone was like this is so cool how' you do this this is amazing and then like two weeks later uh
44:20
that app came out uh I forget what it was called the AI Avatar app it was really popular for like a week and they Snapchat own yeah yeah and then eventually meta launched it right and so there was this there was this like there was this very quick optimization and
44:33
implementation process but it all happened in America so it wasn't like controversial it was just like oh cool like that thing that I saw papers about is now just on my phone right um yeah and so and so it just seems like they've
44:44
done a lot of that it doesn't seem like there's any new breakthroughs here and I think the people that are giving them credit for inventing Chain of Thought reasoning is like wildly wrong um it says a major innovation in their here here's so yeah you're right at the
45:01
same time this stood out to me the newer R1 model and technical report might even be more mind-blowing since they were able to beat anthropic to Chain of Thought and are now basically the only ones besides open AI who have made this technology work at scale that's not true
45:14
though like like anthropic does have a Chain of Thought model they just haven't released it publicly and the reason for that is just financials and and like their safety stuff yeah for sure like I'm not impressed by that I I don't know
45:26
they just just doesn't seem that just that just doesn't seem that revolutionary I know I know I know that said if you're a VC who has billion of dollars in anthropic you're calling the CEO yesterday probably screaming at them right we're bringing screaming back to
45:43
the workplace and uh yeah and so in terms of the actual it does seem like there was some Innovation here in the same way that like when you when China takes like a shoe that's made in America by a single person and then they have
45:58
like a machine make it in a factory like that is somewhat Innovation but that's not zero to one Innovation that's one to many that's scaling in price which is like what they do really well uh they did have a couple novel implementation
46:11
details they switched to 8bit floating Point numbers so it's uh through the TR training process so it's it's more um it's more memory efficient they developed a clever system that breaks numbers into small tiles for activations and blocks for weights so instead of
46:26
just using like a single word for a token they use multiple blocks um they also cracked but then this is the part that really frustrates me they say with R1 deep seek essentially cracked one of the Holy Grails of AI getting models to
46:37
reason step by step without relying on massive supervised data sets their deep seek R1 zero experiment showed something remarkable using pure reinforcement learning with carefully crafted reward functions they managed to get models to develop sophisticated reasoning
46:51
capabilities completely autonomously this wasn't just about solving problems the model organic teally learn to generate long chains of thought self-verify work and allocate more computation time to harder problems and so they they did do some things where
47:05
they change the reward function uh like and so basically like if you have it work on math problems that it can formally verify with like a calculator then it's like it can work through those and then just check its work and be like okay I was thinking correctly and so
47:18
this is the same way that Lisa doll got beat by uh by alphago because alphao was able to play 37 move 37 was able to play you know trillions of games in in synthetically like just with itself and and that's where it generated all the
47:35
training data and so it seems like R1 did the same thing but even this is not new like if you go back two years ago when uh Sam Alman got fired from open AI uh there was all this question two years ago it was it was the end of 20123 that's crazy so this so it
47:55
happened uh like late 2023 so a year and a half but um but there was all the there was all these questions like what did ilas see you remember this whole thing this whole meem and so there the the information tried to answer that and
48:09
they wrote this article that you'll see is like okay they were building this a year and a half ago and it says one day before he was fired by open ai's board last week Sam Alman alluded to a recent technical advance that company had made that allowed it to push the veil of
48:22
ignorance back and the frontier of Discovery forward it's like good but but what he actually says is uh is uh they used ilas suk's work to build a model called qar that was able to solve math problems that it hadn't seen before an important technical Milestone because
48:41
because llms are great at memorizing stuff but they're they have historically not been good at solving new problems a demo of the model circulated within open AI in recent weeks and the pace of development alarmed some researchers
48:51
focused on AI safety the work of sus SS's team which was not previously been reported and the concern inside the organization suggested that tensions within open AI about the pace of its work will continue even after Alman was reinstated as CEO Tuesday night and so
49:06
this qar thing um uh there's uh like like they don't really have that much here but it says suger's breakthrough allowed open AI to overcome limitations on obtaining enough highquality data to train new models according to the person
49:21
with knowledge a major obstacle for develop next developing Next Generation models what is qar qar became uh there was like that Apple group you remember this thing it was like the the strawberry Emoji was like a big thing strawberry gang and they were like open
49:37
a ey strawberries now the code name they renamed the code name like four times and then it finally came out it was 01 yeah and so we we have this like this is this has launched and it didn't kill anyone and also it's like kind of a nice
49:50
to have like you don't even need to use it all the time and so basically you know for years SS had been working on ways to allow Lang models like GPT 4 to solve tasks that involved reasoning like math and science problems in 2021 he
50:02
launched a project called GPT Z what did they launch deep seek R1 zero they're two they're a year and a half behind like this is not this is not Innovation here um yeah and no it does seem that the primary Innovation is the cost yep and giving it away for free yep which
50:23
they've done with DJI and and yep and we had I think we had a post from somebody pointing out that there's this like illusion of choice yep where it's super super cheap if you use yep the Chinese sort of uh inference basically and then it it gets
50:41
dramatically more expensive as you um and so I haven't seen anything Maybe I'm Wrong maybe there's something else deep in the paper that looks great and is true breakthrough it's open source the paper's out there so that will get ported back to llama to open AI to
50:56
anthropic but I haven't seen that what I've seen is a ton of optimization that went in to taking stuff like gp0 which became qar which became strawberry which became 01 and they took 01 and they did a bunch of self-training which is scary
51:12
because the AI is talking to itself and creating it's getting smarter by itself in some ways because it's kind of like what happened with uh alphago but uh they took that and they optimized they completely optimized the model so it will run on much cheaper hardware and
51:29
the inference will be cheaper which is important because the test time compute scaling law is new and right now you heard Sam say we're losing money on some 01 users because they write a query and it thinks for five minutes and that's a
51:45
whole server rack firing for you know a couple bucks and so um and so yeah one one one relevant example so it says the recent scuttlebutt on Twitter and blind uh is that these models caught meta completely off guard and that they perform better than the new llama 4
52:03
models which are still being trained yep ouch apparently the Llama project within meta has attracted a lot of attention internally from high ranking technical Executives and the result is that they have something like 13 individuals
52:15
working on the Llama stuff who each individually earn more per year in total compensation than the combin training cost for deep seek V3 models which outperform it again we don't actually know yeah uh Alex uh from scale is saying you know that the training cost
52:30
was was way higher said how do you explain that to Zuck with a straight face um he's gonna tell them to freck off again I I I think that's how does zck keep how does Zuck keep smiling while shoveling multiple billions of dollars to Nvidia to buy 100K h100s when
52:46
a better model was trained using just 2K h100 for a bit over 5 million so it does it does um this is why last night sa is going out there a lot of people are about to uh's Paradox some people are calling it javon's Paradox no one's
53:03
calling it that um only if you're extremely online and you and you're not listening to podcasts about it yeah you might be mispronouncing it but yeah and and then it goes on but you better believe that meta and every other big AI lab is taking these deep- seek models
53:16
apart studying every word in those technical reports and every line of the open source code they released trying desperately to integrate these same tricks and optimizations into their own training and inference pip so what's the impa of all that well
53:29
naively it sort of seems like the aggregate demand for training an inference compute should be divided by some big number maybe not by 45 but maybe by 25 or even 30 because whatever you thought you needed before these model releases it's now a lot less sure
53:43
so I I think this this highlights that like a good AI model is no longer just like how big are the parameters and what are the weights and like what does it do by itself like there's actually a stack of capabilities that you need to think
53:56
about out so at the bottom yes you need a robust model and that's what their deep seek V3 is and that's what he's talking about when he talks about llama 4 llama 4 I mean I don't even know what they're terming that phrase but I assume
54:08
it's just the underlying core model and yeah that might be underperforming or not or might be too expensive to train or something but a lot of that probably has to do with the fact that they just haven't like if if uh if deep seek trained on gp4 output tokens then the
54:26
data has already been kind of cleaned because it's only training on like of course it's going to sound like GPT 4 because it doesn't have any junk in there it's not going to accidentally sound like you know some spam on the internet because that was already cold
54:39
when you pulled the data from the model and so there's that then level two is like the Chain of Thought how good is your reasoning model on top we know 01 03 and R1 all seem to be pretty good on top of the base model so I don't even
54:54
know what llama is doing are they going to Rel relase an O like an 01 competitor with llama 4 it should because clearly test time comput extremely important but then you open source this and then you got to inference it and that's really
55:08
expensive and so but they'll probably do that but they so they might be comparing apples to oranges there comparing llama 4 which is just a base model to a reasoning model which is not fair and then there's all the UI and and and orchestration that happens on top of it
55:23
and I think the Chachi PT app is certainly better than the Deep seek cap right now just in terms of yeah when you open up deep seek well that's a lot of people are saying deep seek is very much targeted at developers sure which which
55:37
is you know if if you're China and you're seeing this explosion of new app layer products and you're saying hey long-term value will acre to the app player yep why not release a product that any developer can use to deliver cool experiences to their end users it's
55:52
like hey we don't have to worry about building consumer products because every Chinese feels like a teu yeah teu software for the most part yeah it's actually very smart to be like hey instead of using open AI which your consumers don't care about what's under
56:05
the hood either use our yeah basically free app yep um and so it feels like it's it's it's disruptive in the sense that in the way that llama was disruptive where a lot of people that were like my open AI API bill is really
56:19
high and thank you Zuck you just gave me a free version that I can I still have to pay inference costs but I can just host it on AWS now it's way cheaper for me and this will drive the cost down even further so hugely disruptive to the BB basically it it's it's disruptive to
56:34
so many different narratives as well right where the fact that sat Satya on a Sunday night feels like he needs to to tweet out links Zuck earning season he's got to justify why are we spent why are we buying a 100,000 of these again and Zuck is also
56:50
no he it's not like he hasn't gone back on big infrastructure spending before the meterse everybody's like dude you're an idiot like stop trying to make stop trying to make the metaverse the thing and he's like okay like he eventually
57:03
like got it and diverted all that budget back to Ai capex and so now I'm I'm actually very interested to see Jensen come out and talk about this because he's got to get shareholders confident that the demand is still going to look like what they've been projecting right
57:19
he just did a new interview but I think they already filmed it so I don't think it's going to C I think it's going to drop and not have anything related deep seek but uh the interesting thing is like is like we we keep going back to
57:30
like open AI versus meta obviously like distribution is really important and the question is like can deep seek figure out a way to get distribution even barring all of the like oh it might get banned or there's like CCP stuff going on just in in The Knockout drag out
57:49
fight, tff on every query yeah but even even just if you if you look at the if you look at the uh the app store is like a free market like there's already a lot of people that have chat GPT installed they have the free version you bring
58:02
them 01 and you've changed the reasoning logic so the UI is the same and all of a sudden it's like yeah I'll just stick with what I know it's already it's a better app it's already installed deep seek was very clever in basically like building the API infrastructure to just
58:18
immediately be able to switch it's like just copied open AI on that front too totally and then and then also um you know you're still competing with like the the average AI consumer might wind up just using Google or oh yeah like yeah when I when I want to talk to an AI
58:35
I just go into Instagram because like llama's there and they don't even know what llama ISS yeah I love I love meta AI using it through WhatsApp like yeah there are people that do that there are billions of people that do that it's
58:47
crazy and and so yes like Tech Twitter is very much like of course like everyone's going to download like the best model because it performed 2% better on this eval like NE here's the thing that made it really obvious that that the the uh going going
59:03
number one in the App Store was like legit in any way the the the developer name of the app is like just a really long string of TR these characters and I just promise you that them sitting there with 227 reviews the review thing's weird and and American Consumer this
59:20
looks this screams to me most of the time I see Chinese characters it's it's some like spam bot yeah text that I just got and I'm like okay I mean if you search for deep seek in the App Store you will see like chat GPT comes up but
59:33
then also like like chat Ai and it looks exactly like chat GPT and it's a chat GPT rapper that's just taking like Revenue basically from them and just acting as like uh like a like a portal just to a chat app and so uh obviously the the the the App Store rankings are
59:51
momentum based and so I think I think deeps did have a ton of momentum so they did get to the top of the charts I don't know that that's durable and also Nikita was saying that there might be a lot of bots promoting it to try and rank in the
1:00:04
App Store but I just have a really hard time to believe that like this China literally has is notorious for having you can go there you can go to China Chinese firms to buy internet traffic and downloads and anything totally so just diverting some of that those
1:00:20
resources to hey let's go number one in the charts because I I came away from this thinking in many ways they like if this is a front for the CCP which I won't grab the hat but uh just creating economic chaos in the United States being like hey there's a lot of
1:00:36
Leverage in the system y all these firm you know France is lending to to data data center development in the US a lot of real estate guys are saying that that's you know usually a bad sign when the French get involved um and uh uh uh
1:00:52
but uh but yeah so so if you were just looking at it as hey let's let's law an economic grenade over to the United States make the app go number one kind of like just just make everybody freak out and kind of be distracted right and
1:01:03
also like taking off the conspiracy hat even if even if we were just talking about like an ally like let's say deep sea came from Japan and they were just trying to compete it's very logical to say hey we were able we had some cracked Engineers who were able to drive the
1:01:17
inference cost way down with a bunch of Innovations which are real let's try and release this app that's as good as the $200 a month chat GPT app for free get a bunch of people to use it and then get all that rlf data so then they can use
1:01:30
that to tune their models because there's this big like dat so look at uh look at what China's doing right they just announced the equivalent of like $500 billion of new like state funded investment right and that could easily be going in yeah yeah well even if it
1:01:45
was Private Capital it would still make rational sense I just if if if we no longer needed all this extreme capex China wouldn't be launching the free app and then not doing that right they clear yep no no it it's totally reasonable but
1:01:59
so let's should we should we finish let's go to the timeline I was just going to say on like why so at the high level Nvidia faces an unprecedented convergence of competitive threats that make its premium valuation increasingly difficult to justify at 20x forward
1:02:14
sales and 75% gross margins the company's supposed Moes and Hardware software and efficiency are all showing concerning cracks the whole world thousands of the smartest people on the planet backed by Untold billions of of dollars of capital resources are trying
1:02:28
to assale them at every angle um and so uh yeah perhaps most devastating is deep seeks recent efficiency breakthrough achieving comparable model performance at appro approximately 145th of the compute cost which again we don't know
1:02:42
if it's real uh but anyways it'll be interesting to play out nvidia's down 15% uh last time I checked public uh and we'll see where they are tomorrow let's go to some hot takes about deep seek keeping up to date on what happened on
1:02:57
the timeline I want to start with this uh short thread by Dylan field the founder of figma Dylan field he says I guess it's hot take time so here we go I love this uh one always assumed there would be a reckoning moment in public markets over capex spend for AI two it
1:03:15
will take a lot more share price punishment for any of these companies to reconsider the number of gpus they are buying in 2025 three there are likely order of magnitude improvements to training and inference of available though not NE not though not necessarily
1:03:29
achieved yet for deep seek trained on outputs of American models which we discussed five it would be surprising to me if deep se's claims about training costs were true six from a public safety standpoint an open source model of unknown alignment is an extremely
1:03:45
interesting and challenging Threat Vector we talked about this a little bit like what if embedded in the model it like tries to change your political philosophy very slowly already already so perplexity integrated deep seek really into into you can you can opt to
1:04:00
use the model I mean people have wired it up to cursor immediately ask the uh depending on which model you select if you select uh deep seek and ask it about Tian Square it'll be much much sort of like more more wasn't that big of a deal
1:04:16
just calm down exactly how they position it it was like they don't talk about they they really really downplay it so so they admit that it's a thing yeah but they don't admit that it was a disastrous actually let's talk about Ken state for a minute okay like
1:04:32
you know like right back at you American yeah um seven if deep seeks mobile app continues to top charts it will join Tik Tok in the discussion in the US we need to block this app discussion uh I think it's already there yeah one one thing
1:04:47
that's interesting is when Tik Tok started to chart because they were spending they not only people love the app they were stared to spend a ton of money on user acquisition the app and especially the video feed was like fundamentally better than the
1:05:01
Alternatives Y and so they were spending all that money on user acquisition to drive downloads but then consumers got a better experience now consumers are like well I already have chat gbt this doesn't do any net nothing that's sort of like net new for me especially for
1:05:15
the average person who's like make me a it does if you don't have a premium subscription on on chat gbt like if you have chat gbt free version and you download deep seek that is a massive upgrade if you're a free user but but if it but if the average request is how do
1:05:31
I make a recipe with these three items yeah more for like power users in my opinion I agree I agree and and I do think uh Sam Sam already addressed it and said that he's going to bring a c a set number of reasoning 01 queries which
1:05:46
are expensive to free to the free tier yeah and so this this is a this is certainly like a financial change but in terms of of in terms of just retention and user adoption like it's not insurmountable um New Moon Capital says so just so I understand people are
1:06:05
bearish on AI Because deep seek Innovation improved efficiency by 30X and with that and and that with larger clusters and continued scaling of synthetic data and inference compute Next Generation models are going to be like 100x better than 03 so AG is
1:06:20
bearish for AI got it and it's a good point yeah I mean uh all of those all those Innovations even if they weren't open source I mean they get ported back so fast because there's only a few Secrets like eight 8bit uh floating
1:06:33
Point numbers instead of 32bit oh that can save a lot like someone's going to try that eventually and I think most of this stuff was probably like either in the pipeline to change I mean we we we talked about like the the test time inference is going to get so much
1:06:47
cheaper when this is baked down into silicon but we're just not there yet cuz we're updating models every years the the thing that's Most Fascinating to me is these model companies M mrr uh anthropic cohere that really don't publicly have these capabilities Y where
1:07:04
now anybody employed at those companies basically shouldn't sleep yeah for the next however long it takes to to get on par with the the free model otherwise you almost don't have a right to exist in many ways yeah that's a good point uh
1:07:17
sheal says over onethird of Nvidia sales go to China probably 40 billion last year the Singapore back door is real Nvidia even says shipments to Singapore are insignificant while 22% of Billings last quarter were to Singapore and so
1:07:34
pretty pretty staggering numbers like so they sell to China just certain chips different chips which again like they don't work for the biggest training runs but if you have a team like deep seek that can optimize around it and say oh memory bandwidth is a problem with the
1:07:48
nerfed gpus all of a sudden it's like okay just buy a trillion of these chips that are this is you know this is the interesting dilemma that uh people in the US face right is is everybody's holding a bunch of Nvidia right the entire markets basically
1:08:04
propped up by Nvidia yep and so you kind of want to be mad at Nvidia for saying you know why are you providing you know uh chips who our political enemy but at the same time it would tank you know it would cut the market cap in half maybe
1:08:18
right um and we talked about I think Friday about potential back doors so Singapore could be one could be anywhere in the world right it could be I mean the numbers to Singapore are staggering here in the three months ended October 27th 2024 so like deep chip ban just not
1:08:34
last quarter but the quarter before Q3 uh $7.7 billion doll of chips to Singapore and it's like Singapore is not buying that many chips yeah and I think uh I think on one of the shows I was a little skeptical of this back door and you were
1:08:50
much more bullish on it and I think you're 100% right seeing the data there um Philip lefont says should open should AI models be allowed to be open sourced do you know this guy what was what why why did this so this is the founder of
1:09:03
CO2 management oh yeah and he uh I believe has pretty large exposure to open Ai and so this post went viral because last night I think he he probably was enjoying his weekend y he sees the news on deep sea he's really really not happy with um and it's it's a
1:09:22
funny it's it is a pretty funny question uh there there was a reason that open AI shifted from being open source and innovating for the world to innovating for themselves and trying to do sort of uh rent seeking Behavior yeah and this
1:09:37
is what um this uh Chris from uh what's Chris pike pike yeah yeah he talked about he's talked a lot about how AOL was trying to basically build a closed internet that they they could basically uh collect a toll on and how that really didn't work and his point of view is
1:09:55
that AI is is the AOL of of AI we have no idea if that you know and and uh anyway so we don't know but I I think it's pretty funny yeah still funny question to be asking yesterday as an investor yeah uh let's get to Martin scr he says it took Wall Street one month to
1:10:11
read this karpathy tweet and a month ago karpathy tweeted deep seek the Chinese AI company making it look easy today with an open weights release of a frontier grade llm trained on a joke of a budget uh 2048 gpus for two months and
1:10:25
uh Wall Street kind of picked up on it today with the R1 release um uh Dolly Bali says Jensen gonna have to get on a podcast this week and that's true very true I want to hear what he to say the the market needs uh needs to be comforted for sure I want to hug you
1:10:45
know I want to hug from Jensen with the leather jacket on you want to you know kind of like know that it's going to be okay and and I think he should avoid signing wom's t-shirt top signals yeah yeah no more top signals just just really explain to
1:11:02
me like you know minor improvements in Cuda talk to me about the bandwidth interface problem your memory and there was it's funny the um what's the guy uh uh that's always posting top you know the top signal guy CNBC oh okay lot lots of people better do that kramerr Kramer
1:11:20
so Kramer posted five days ago like open AI or sorry he's unable like Nvidia is like Unstoppable but if you actually the thing about Kramer that I don't think like people that just sort of like only see him when he's getting dunked on he
1:11:33
posts that stuff about every company all day long and it's bullish bearish bullish bearish so he's just like cycling back and forth so I don't think it's signal there was actually like a like an economic study on Kramer and
1:11:46
they found that he did beat the market over a pretty long period of time but he did it uh basically with high beta so just he was just like Leverage long the market Mar and so his his ups were really good and then his Downs were really bad but on net he still
1:12:01
outperformed so there's always a bull market somewhere there's always a bull market somewhere let's go to Logan Bartlett he says so wait China actually thinks they can succeed with a low lower cost ripoff of an American product good
1:12:12
luck with that it's a good point this is what they do it's the teal thing of like America's been good at zero to one Innovation China's good at one to many and and we are in the one to many phase very clearly of AI and I think people wer hadn't really taken that to heart
1:12:27
there's still the question about consumer adoption and and is there Monopoly of power to crew on the consumer side of the application layer um but certainly on the on the foundational swap hot swappable llm Tech pretty pretty commoditized y Daniel says
1:12:44
Finance guys are like I eff knew you nerds were full of s hit I don't know how to not curse anymore I'm trying to not curse on the show about needing that much money that didn't come AC cross well at all we'll have to work on that uh Tom says Apple's AI is so bad they
1:13:01
don't even include it in the AI cell off that's I mean maybe that's a bull case for Apple they didn't like go too hard and like tell that whole story like they they they told this is what you've said before they have the distribution so
1:13:12
they can kind of sit back and wait to see how things pan out they can partner with open Ai and say you can be our AI provider but you got to pay us yeah you know some egregious amount more like scapegoat they don't pay each other there's no money Ching hand changing
1:13:24
hand but anything that goes wrong they can just be like oh it's like open's problem like it wasn't it wasn't us no but presumably in the future they could and then open AI gets a lot of data hopefully yeah if they can do that there
1:13:36
was just a there was just a deep dive on uh the the new Siri and they asked it like who won the Super Bowl in 1989 who won the Super Bowl in 1999 and it got like every single one of them wrong while chat gbt didn't so there's like something very odd going on and and like
1:13:52
the previous version of Siri could do that just because would just be like oh Super Bowl stats look it up in the database they basically need to create a new name for Siri if they wanted to get adoption because it's been so bad for so
1:14:03
long they tried they call it just Apple intelligence I know and so but it's not working yet but room Apple room temp room temp okay this is good from Dylan Patel uh so this is like2 trillion doll uh two two trillion doll loss in market cap for a $6 million training run
1:14:22
ignoring cost of research ablations distill data from GPT capex uh for their various clusters Etc imagine if China invests 300 million in a training run it would reduce the world's GDP to zero very funny uh it's it's great to see him posting through the chaos
1:14:40
because that's the only uh correct approach unless you're SAA and then you got to post uh you know speaking of speaking of jeevan's paradox this seems like an overreaction says Gary tan Wall Street needs to read the Wikipedia page
1:14:53
on jeevan's paradox in economics jeevan's Paradox or jeevan's effect occurs when technological progress increases the efficiency with which a resource is used reducing the amount necessary for any one use but the falling cost of use induces increases in
1:15:09
demand enough that resource use is increased rather than reduced and uh the classic example is energy consumption uh there's a whole thesis around like nuclear power energy will be too too cheap to meter and oh would that cause us to the energy Market
1:15:25
to go down in value probably not because you would have insanely energy dense like consumer products like right now most most household appliances are gated by well we want to be energy efficient it's got to plug into a you know wall
1:15:41
outlet like it's not just going to pull like a gigawatt of energy to like you know do your dishes um but maybe it could um and so Patrick oy says everyone about to be a jeevan's paradox expert and that's true uh so SATA Adella posted jevans Paradox strikes again as AI gets
1:16:00
more efficient and accessible we will see its use Skyrocket turning it into a commodity we just can't get enough of and he posts the Wikipedia and Joe weisenthal says Microsoft CEO up late tweeting a link to the Wikipedia article on jeevan's Paradox this is getting
1:16:16
serious and I I agree with kind true it co you have Co guy from CO2 management big position in open AI yep you've got got SAA all feeling like they need to react in that moment on a Sunday when normally the corporate comms you know
1:16:33
you know people will post around the clock but normally the corporate com strategy would be to turn around you know on a Monday hey let's all get together and like figure out what our response is for this everybody's like no we got to front run this right and they
1:16:46
were right to some degree because of the selloff obviously they want to prevent as much of that as possible y yep yep um I mean I I'm fully Jean's Paradox build uh there's a good post from Chrisman Frank here no idea what will happen in
1:17:00
the wher market but at synthesis we immediately started thinking about product changes that are newly possible with a 95% cost reduction I imagine there must be many such cases and I've I've said this for a long time with like the custom X feed like I would love to
1:17:14
have an llm that I can prompt and say this is what I want to see in my feed and it reads every single post does a whole thinking Deep dive on it and then decides is this good for John or not that's insanely compute expensive like
1:17:29
it's completely prohib prohibitively expensive I wanted to read every email and do much much more advanced spam filtering where should it put it should it put it at the top of the inbox like every single news article I read I
1:17:39
wanted to scrape out all the text take out all the ads format it better like give me a summary like all these Transformations every time I click a link I want AI to run on that and that's something that uh can only happen if it's actually as as free as the rest of
1:17:55
the things that happen on your phone like you know if you want to switch your phone to Grays scale it just the the algorithm for reducing the color just happens like that there's no you don't think about like oh this will take extra compute and it should be the same thing
1:18:07
with AI so uh Dylan Patel says deep seek V3 and R1 discourse boils down to this Shifting The Curve means you build more and scale more dummies so he's fully jebin's Paradox pilled uh on the left we have the the uh IQ 5050 or 55 uh IND
1:18:24
idual saying now we can have even more Ai and the Jedi at 145 also saying now we can have even more Ai and the midwit at 100 IQ says more efficient training and inference means less compute and no one should be spending on scaling Nvidia is screwed and Adam D'Angelo posted
1:18:40
basically the same thing the midwit meme uh cheaper AGI will drive even more GPU demand uh and the midwit says deep seek efficiency will reduce GPU demand and I agree with that and a lot of a lot of like people have been dunking on some of
1:18:55
these saying oh they're just trying to cover their tracks they're in crisis management all this stuff it's like well no like there's a there's there's clearly good Arguments for both right uh we just went over the entire short case
1:19:09
for nvidia's stock there are some good arguments in there but it's more about like dynamics of Nvidia with the rest of the market than just like oh we don't need gpus anymore we're g to stop building hey we're we're four or five years into birthing machine intelligence
1:19:22
which is going to trans form and uh consume the entire Services economy the entire and it's and oh yeah we should probably stop spending money on this or stop investing in this and at the same time China China at the state level committing to hundreds of billions of
1:19:41
dollars a year of capex okay yeah and and and the same yeah the the bare case for NVIDIA is that uh you know there's a TPU from Google that's trained on a on that's designed specifically for hyper efficient inference especially test time
1:19:58
compuse scaling and Nvidia is less relevant in that Paradigm but we're still so early on the architecture evolution of these models that it's it's almost too early to say now there are a lot of startups that have raised hundreds of millions of dollars to to
1:20:11
take shots at oh we think the we think the Transformer is staying around so we're going to optimize for Transformers or we think Chain of Thought reasoning and test time computes really important so we're going to focus on that um there
1:20:20
there was grock which was all about like it was very low memory but very high fast I don't know if you ever saw that demo before the the xai gro it was a different grock he had a q instead of a k it got gred let's go to uh Anarchy says uh artificial illusion of choice
1:20:36
drives you to cope into keeping the 5x faster Chinese host after open router already chooses it by default due to its low cost in terms of risk second order effects maximizing win rate percentage this is a key problem yeah so this was a
1:20:49
response to something I forget the exact because I was post I was out of control uh posting yesterday but um yeah just showing that clearly yes they released it as an open source model but clearly they want to eat all of the all of that
1:21:05
data yeah let's go to Jeff Lewis he says fascinating to see some of the same folks who advocated hard for a Tik Tok ban now promoting a CCP AI op uh simply because they are jealous of a singularly Transcendent American organization wild world have the most beautiful weekend I
1:21:22
love his emojis killer it's it's the best to drop something uh inflammatory and then just say sick but that's his mindset Rocking In The Free World he's he's he's working you know through bringing love to his uh let's go to word
1:21:35
grammar says Trump's Logic for unbanning Tik Tok even if they are collecting our data data on the type of videos that 16-year-olds like to watch isn't that important unfortunately the data collected through deep seek is actually very important yeah and again it's not
1:21:50
just about the data it's about the influence on uh it's it's psychological warfare right yeah this is hilarious LMFAO deep seeks API docs are basically our API is compatible with open AI just download their thing and set the base URL and model name to
1:22:09
us wow Savage I mean but that's the nature of like the Linux Wars and versus like Microsoft like you know can you get distribution can you build a monopoly in some sort of moat or on top of something that is deeply commoditized like you can
1:22:23
just use use Linux no one does got to get those blue bubbles baby ey message uh Dylan Patel open a should have been a religion not a nonprofit imagine the tax savings Mormon church and shambles he's just on a roll he's on fire I love him he's so good uh let's uh
1:22:44
let's skip this and go to pav oh this is a great one we can end on this because we got to hop on a call um pav asparuhov says if your entire world viiew on AI is dramatically shifting every 3 weeks maybe you just don't know what's going
1:22:58
on isn't oh is that what he meant to say maybe you maybe you just don't know what's going on no no he's saying he's saying like he's saying like like this this should be expected like you you should have this like somewhat priced in
1:23:11
if you were just like I had no idea that a model could be open source this is crazy and cheap multiple times every Venture back founder that was running a deeply unpr uh unprofitable generative AI company has always been saying it's fine that we're running in the red right
1:23:29
now because it's going to get 95% the cost is going to reduce by 95% so now that it's happening I you don't really see the founders that are running these actual app layer companies at the priz because they're fine it's really the hyperscalers the people all the people
1:23:45
doing you know the subscene amount of capex that now have to figure out ways to justify it yeah yeah it's more I I I think you're right that the pressur is on the the mistrals and the uh and the coher yeah and kind of like the players
1:23:59
that are selling some sort of even anthropic like they're selling API access and they don't have the runaway consumer adoption yet yeah because they need to amortise the cost over a long period of time but then if your model just got lapped and you spent a billion
1:24:15
dollars on it and they might have been and they might have been expecting like hey this will hold for a little bit time but I bet you the good Founders knew that you know Zuck was going to come out with something maybe he was going to
1:24:24
open source it it's possible maybe he wasn't it's also funny to think about so consumers if you go to them and you say hey there's this there's this thing that's like chat GPT and it's free they're going to be like well I already don't pay for chat
1:24:39
GPT just go to chat.com and I just use the app right it it's not it's not even to Consumers who are so used to being able to query data for free yeah it's it's not that ground I I I think it really like just comes down to this this
1:24:56
concept of like there there seems to be a massive pool of value in being the consumer AI company the aggregator the front the front page of artificial intelligence getting installed on the home row people's apps setting it to the default search engine setting it to the
1:25:14
default web page when you open a website and and then the B2B Market is going to be extremely competitive and developers are not going to care about brand or use ility they're just going to want the best thing for the best price welcome
1:25:27
back to technology Brothers still the most profitable podcast in the world let's go to Signal he says honestly I'm still baffled at the thought of the CCP going full scorched Earth on AI by by going open source it wasn't even remotely on my bingo card they're
1:25:41
basically doubling down on zuck's playbook but scaling it up to state level throwing their entire weight behind making worldclass AI Dirt Cheap so nobody else especially the West can monopolize it the Chinese quantco flexing right now because a large
1:25:55
portion of the West's AI development just got mired in Prestige projects instead of profit maximizing strategies which is ironic because who's the capitalist again China has no such qualms they're ruthlessly practical when it comes to scaling the Chinese shop
1:26:09
figured out a way to tie reinforcement learning to actual efficiency flows faster than everyone else huh not on the bingo card I don't know this didn't take me that much by S by surprise it's also possible that open AI figured all this out didn't want to
1:26:27
release it because they want to justifying hey we need to spend you know I don't even know if it's I don't even know if if it's like a justification of more cacks it's really just like like until until a competitive model model
1:26:41
goes free you should not go free this is just basic econ 101 this is actually more capitalist so like we are the capitalist it's like charge until you can't and now what happened oh free chat GPT users will we get 2001 queries per month Like Larry Larry Ellison Donald
1:26:59
Trump MSA and Sam are in a war room right now definitely well I mean we still don't know if the if the old scaling law holds that's the big question yeah if the old scaling law holds and gp5 is good and the big training runs important and you and and
1:27:16
having that you know the not just those weights having the weights are important but also those first you know batch of you know tokens that it produces that you can't I mean it took him two years to pull all the data out of gp4 right yeah that's
1:27:29
another way to look at this is like yeah like obviously you trained on gp4 outputs took you two years to get all that data together if you have it takes you another two years to do GPT 5 and then the the chip restrictions are even harsher and yeah like let's assume this
1:27:47
is 2048 like deeps has 2048 gpus well like what if the next model even their compressed version needs 20,000 right because it's still in order magnitude even to do their compressed model it's like that could be hard that could be limited we'll
1:28:02
see uh Singapore would like a word Singapore would like a word uh let's go to atlas Atlas says you horny mfers really did say too much at the sfai parties huh Atlas 5K likes Atlas has been on a tear I mean it's so un brand it's like
1:28:19
so so in the Zeitgeist to what a banger so good so funny I don't know I I it doesn't seem like that's what happened it doesn't seem like oh one weird trick snuck out of a lab and got over there it seems like they came up with their new tricks and then they stole a bunch of
1:28:36
data and maybe maybe some gpus yeah and that didn't really have anything to do I'm sure I'm sure Chinese Labs have people inside at all the major American Labs and so anything that is being discovered at the American Labs yeah is being ported back but even that
1:28:53
years not even happening at a party right it's just happening from within yeah I mean yeah that's always been the case of just like do you need to worry about the girl at the sfai party or do you need to worry about the guy who has
1:29:03
GitHub access yeah and can just like copypaste code into you know Notes app which is happened you're worried about the wrong gooner worried about the wrong gooner probably ridiculous uh I mean still qar it's been two years guys steal the stuff
1:29:19
faster that's my message to the CCP Step It Up steal faster I'm not impressed uh Nick Carter who's in the Golden Age now he says deep seek just accelerated AGI timelines by five years so focus on the gym knowledge work is obsolete muscles
1:29:37
are all that's left 16k likes this is golden retriever this is golden retriever mode you got to be golden retriever maxing be hot friendly and dumb intelligence is too cheap to meter you don't need to worry about it anymore you need you need to rep I need to
1:29:51
really Co coin that and own that because he got close with this he has the idea right but he didn't he didn't have a coinage around it yeah but golden retriever maxing is is the future intelligence too cheap to meter intelligence will be too cheap to meter
1:30:05
there's no Alpha in reading books anymore yeah that's for sure uh growing Daniel says Hey guys my favorite Bay Area nonprofit is facing Chinese attacks and needs our help I like that we got a donate to open AI not we got to do why
1:30:20
why can't I I've I cannot for the life of me find a place to donate to nonprofit just send a check just send a check just make it out to Mr Sam Mr s yeah uh Buco Capital bloke as my entire Twitter feed this weekend he leaned back in his chair confidently he peered over
1:30:38
the brim of his glasses and said with an air of condescension any fool can see the Deep seek is bad for NVIDIA perhaps mused his adversary he had he had that condescending bastard right where he wanted him unless you consider Jin's paradox
1:30:54
all color drained from The Confident man's face his now Trembling Hands reached for his glasses how could he have forgotten jeevan's Paradox imbecile he wanted to vomit I love that where is that it feels like it's clearly like it's probably from DEC it's probably
1:31:10
generated by Deep seek somebody was the prompt was probably like write a dramatic story between two people you know debating deep seek and Nvidia and jebin's Paradox but thought that was a great piece of writing I really enjoyed that 4K likes you love to see see it
1:31:24
that's a whole new that's a whole new format oh totally too yeah yeah oh yeah yeah we should definitely do one of these put this aside we're going to remix that a million times uh Josh Kushner says proam technologists openly supporting a Chinese model that was
1:31:40
trained off of leading us Frontier models with chips that likely violate export controls and according to their own terms of service take us customer data back to China H that's a good point yeah so here here's here's where here's
1:31:56
where Josh has him uh Taylor Loren is one of the biggest supporters of of deep seeks so uh if she you never want to be on the same side as Taylor Loren except if you're talking about horses that's true it's true she's got some credibility there and the internet
1:32:13
archive she's got that uh yeah it is interesting everybody's been everybody that is um this is really exposing all the people that missed open Ai and that were frustrated with AI around regulatory capture which is is a pretty fair critique right there's yeah or or or
1:32:37
there are some good arguments that open AI has engaged in regulatory uh efforts around regulatory capture yep um saying oh you actually need to regulate us like this is too dangerous like please step in and and trying to make it harder for
1:32:52
for new model to emerge and compete uh but still such a bad look for people that are saying uh you know that are that are openly in favor of it and celebrating it as some Win For Humanity to get intelligence too cheap to meter when uh it's very clearly that there's
1:33:10
alternative um sort of motives behind it it is a huge Vibe shift though from the days of like gpt3 is dangerous and like this AI is going to kill us gbd4 is so dangerous like they shouldn't have it like this was one of the main uh
1:33:27
like reasons why the board was worried about Sam Alman when they fired him they said like he just he just went out there and released chat GPT like who knows what could happen and it's like yeah people got like recipes and like a couple people probably wrote like spam
1:33:42
articles and like other than that like nothing bad really happened like there were probably some people that like got wrong medical information maybe but like that's already happening on the internet so it's very odd and with this one like
1:33:54
no one's saying like oh it's dangerous that R1 is out there it's too powerful everyone's just like yeah it's like pretty powerful cool like it's cheap too like you run a lot of it because it's like they see like the models as they go
1:34:07
further they get smarter but they seem like eminently controllable they do not seem like they're rising up and getting closer to that for sure Famous Last Words Famous Last Words we'll see uh yeah Taylor luren says let's go and she's really excited she's just become
1:34:21
no this is because they dropped another model today oh for image generation for image generation and I guess computer vision it's fantastic she's like she's like pro pro tech for the first time if it's CCP controlled I mean she should just change her name to like the Chinese
1:34:35
characters or something like really lean into the bit it's it's like so clear that like when she goes on Twitter like she knows that like this is what's going to get people riled up and like this is gonna get people talking about her this
1:34:44
is what's gonna get her tweets printed yeah it's good stuff um here's you'll never see a thread by her printed on this show no Gary tan deep seek search feels more sticky even after a few queries because seeing the reasoning even how Earnest it is about what it
1:35:01
knows and what it might not know increases user trust by a lot 6K likes and this is a good point like the there is this is probably the most Innovation that's happened with the Deep seek thing is that UI Paradigm of like showing you
1:35:12
the reasoning as it as it works through the model and it just makes it way more engaging because you enter query and then it immediately starts talking as opposed to you know just watching a progress bar yeah so somebody compared
1:35:24
it to pull to refresh as like a you know dominant UI pattern and something that they think this will this will you know happen much more the only question is like if the inference speed with like the test time compute scaling like really goes through the roof like you
1:35:39
you might go back to tucking all of that behind because it's just like it's funny if you're working with an employee right you and you tell them hey I want this done and then they they sit there going yeah okay so I'm doing this I'm doing
1:35:52
that and eventually you're just like okay just like shut up get it done and and and just like come back to me when it's finished and so I think I think there's this there's this period of time where it's true people want to see how it's working through something but then
1:36:05
eventually when you have that level of trust with the model or the app that you're using you just want it down right yeah I mean uh pull to refresh was eventually displaced by endless scroll like you don't need to pull to refresh on Tik Tok you never go to the top
1:36:20
because you never reach the bottom or the top of the feed you just scroll endlessly and I can see that being the same thing here where it's cool now but once like it's like yeah it thought for the equivalent of 5 minutes but it took
1:36:31
five milliseconds and so it just gives you the perfect answer yeah why would I want to see the internal reasoning but it's a cool hack for now yeah Daniel says love the Deep SE cap using it to organize all my finances and passwords they make it so easy 50k
1:36:48
likes man so funny because this is not the data that actually CCP actual they want the actual uh this is funny from ramp Capital uh there's a headline says deep seek hit with large scale Cyber attack said it's limiting registrations and uh RAM Capital says
1:37:09
well played Mr Altman I don't think they were responsible for that but I wonder who would be attacking them I don't know someone who wanted to take them down and just knock it offline like a troll always thought I always thought that uh if they were giving away
1:37:27
all this intelligence for free that you would just create services to sign up and with as many you know them bad data send them bad data or yeah you're doing I don't know spam emailing whatever yeah uh here's more on the restricted
1:37:43
registration breaking deep seek has restricted registration to its services only allowing users who have a mainland China mobile phone to register um there's some uh Community notes here it says not true signups with for example Google are still available phone name
1:37:59
phone number from mainland China not required and uh Victorio says haaha GPU pores like and the cat cry emoji and the finger because I mean it it's totally it would not surprise me if if the if the app actually goes Super viral that they
1:38:14
would have scaling issues like even if they're cheaper to inference it's yeah but I I just think it's very viral on teapot yeah and not very many other places yeah if you're out in the world getting a coffee or taking an Uber ask someone
1:38:29
random oh did you see the crazy news in Ai and see what they say yeah they'll probably be like yeah I just I just tried Chach BT this weekend it's amazing amazing it's crazy I used it to to draft a birthday card for my niece yeah did it instantly like it's amazing we're living
1:38:47
in the future like yeah and the and and did you know that that guy Elon Musk is also working on something and he he's the one behind something like Twitter AI cars too that's crazy he he makes cars and he's working on AI that guy's so cool
1:39:02
that the man yeah I I literally had someone I think it was my mom at one point was like did you know that Elon Musk has a rocket company and a car company I was like yeah yeah I I actually do know that but this is years ago but it's just funny it's like yeah
1:39:18
if you're not like intact you're not going to know every subplot of this Sam mman I've got a nicotine company and a podcast exactly mind blown man man of many talents uh signal netgate built a browser sold it like box retail software
1:39:34
you'd have to go to Comp USA and pay a solid chunk of change for it the model worked for a while their stock soared everyone was thrilled then Microsoft showed up and said actually we'll just give ours away for free and overnight their entire business model imploded the
1:39:48
world collectively realized oh this distribution method is dead and everything changed almost immedi medely this feels like that moment H why why this and not llama is that just because they have an app like and the products on par llama
1:40:07
was on par GPT basically yeah I mean I think it's a good point I I I think you could also say like well you know Microsoft also had you know windows that they charged a lot of money for and then Linux was open source and that didn't really change the model Microsoft still
1:40:25
prints and then there was a company called red hat that wrapped you know it was basically like a Linux rapper and they make billions of dollars and so like I would I wouldn't be surprised if there's a like a Consulting style company that just does llm
1:40:40
implementation and is like oh you're you know some massive Industrial company and you want to roll llms out in your organization like you call McKenzie but then will be the ones that are like the on-site partner and they're just like
1:40:52
printing money in stalling all this stuff like maybe that's not like a venture scale opportunity but it it could be a big business AI agents for LM implementation that'll be now that's the play that's the play yeah slap an agent on it uh guer Capital says deep seek in
1:41:10
covid-19 a Chinese lab releasing a surprise and taking down us markets funny yeah we already covered that but uh it is the Wild Card of the year for sure um Gary says arguably Stargate just got 30X more compelling and Joe weisenthal says For Better or
1:41:29
Worse deep seek is helping cement The Narrative that the race to achieve something called AGI is a race Allah the nuclear bomb could be a huge Boon for Silicon Valley Tech companies collecting money from uh DC Washington um Gary
1:41:43
follows up and says Stargate is all private funding I get the anxiety about use of public funds but that's not what this is about interesting yeah I don't know is this bullish for Stargate or bearish it still goes back to the scaling law do you need
1:41:58
a big cluster I think so we'll see I just think it's worth running the test worst case scenario you build a massive data center you incinerate a bunch of capital and use it as a podcast yeah exactly uh no I I just see everybody you can't say
1:42:14
we sh like these data centers are worthless while also or unnecessary while also agreeing that ai's impact is only been felt 1% Y which which I would say most people feel at this point that that ai's only impacted our economy or Society or way of life 1% right so if we
1:42:35
have another 100x to go then like yeah we probably need uh more data centers more compute goes back to jeans's Paradox do more of this stuff signal says chat PT is sitting at 500 million Maus and a household name they've cracked the main stream something no one
1:42:53
AI has done at this scale before with retention the original pivot was understanding that consumer adoption is the real prize open AI North Star now looks clearer than ever before build the next generational consumer company and
1:43:05
that's entirely on the table more than ever completely agree with this state it's great at the same time the consumer is even this consumer app layer is going to be more competitive in many ways than the foundation model layer because
1:43:18
you're competing with meta Apple all these different you know Google Etc that have the distribution already so it's not like consumers this green you know blue ocean opportunity where you can just just focus there it's like it's great that open AI has 500 million users
1:43:34
but yeah it's like okay yeah but yeah I mean even even this model's like open source and free and cheap to inference and you could build a new app that wraps it and and maybe you clean it all out so that there's no you know CCP issue
1:43:51
there's no import restriction issue and you're just building like a new you still have to come up with some Inc viral growth mechanism to get 500 million Maus like the first mover Advantage really does matter here and so it's it's yeah it just seems like the
1:44:10
competition is still between the big guys I don't know we'll see open AI needs to buy aol.com bring back America online I like that they've they've shown a propensity to buy expensive domains before run it back bring back AOL really become the
1:44:28
AOL of AI by absorbing the brand well but then live forever I like that uh let's go to run he says over the last few days I've learned AI Twitter basically doesn't understand anything at all uh it's honestly embarrassing what the hell are we doing on here it's
1:44:46
dominated by all caps guys who don't even have the bo most basic ml in intuitions boom roasted a lot of chaos on the timeline the last couple days see word grammar says okay thanks for the nerd snipe guys I spent the day learning
1:45:02
exactly how deep seek trained at 13 30th the price instead of working on my pitch deck the tldr to everything according to their papers how did they get around export restrictions they didn't they just tinkered around with their chips to make sure they handled memory as
1:45:15
efficiently as possible they locked out and their perfectly optimized low-level code wasn't actually held back by chip capacity and then he shares a bunch of other stuff but I got to hear about the deck he's building I think last Thursday
1:45:28
yeah very cool cool I'll I'll leave it at that but is it is it helped by Deep seek or hurt by Deep seek you think uh it's in the developer tooling space and I think it's I think it's uh I think it will just benefit by more AI adoption but it's very different than I think
1:45:46
what any of the foundation models are doing right now so cool complimentary y uh call me a nationalist or whatever says wisenthal but I hope that the AI that turns me into a paperclip is americanmade yeah funny I think we can all agree on that yeah 100% bye
1:46:03
American uh this is great so salana says uh think it's probably important to adopt a zero cope policy in light of deep seeks achievements doesn't really matter how they got here at this point they're here and Reggie James says zero cope policy incredible phrase very
1:46:21
important to apply for your entire life to be honest GRE cougan law zero Co zero cope policy just don't C coping yeah Salon avoid all coping avoid all coping salon's law no it's good it's like just just law coping is a sign of weakness
1:46:38
yeah and and yeah and I think the whole the whole GPU thing it is it is interesting in the sense that like it is a cope obviously but then there is something practical about like if they if they are lying and they did get around uh chip restrictions like that
1:46:52
means that maybe the chip exportation policy needs to change maybe there needs to be more enforcement maybe the rules need to be Rewritten so there is like practical like steps that can come out that start reaching for the the tin foil
1:47:03
hat because I don't think I don't think Singapore needs 20% of allvia of all Nvidia chips in the entire world the small nation of uh Singapore but at the same time like it doesn't matter like like they did it the model's out there and like it's open source so it's been
1:47:20
copied a million times and like you you can't put the Genie back in the bottle it's impossible like there's just no way without being like the place your bench is getting up there you might stuff I don't want to like you might actually be
1:47:32
able to get Genie back in no I mean it's horrifying to think about what would be entailed with that it would be like Mass surveillance of every server including like your home because like you can buy you can buy eight h100s and rack them in
1:47:47
a server and run them on your house power and you could inference deep seek that way yeah and it's like okay how are we stopping that now that's horrifying surveillance State like going door to door to make sure people aren't using
1:48:02
this thing if that if that was really like what where that went which is like yeah very problematic obviously um growing Daniel the real loser here is AI safety people because I do not give anything about their madeup dangers when they when the actual danger of China
1:48:20
beating us to AGI is staring us in face H yeah we talked about the AI safety people a little bit it does seem like that's just not been in the conversation at all I wonder if it's just I'm not following the right people like what has
1:48:35
elaz or owski said about this is this like changed his PO Doom in one way or another it's kind of unclear there was like all of 2023 it felt like just P Doom Central and now it's just he really fell off yeah and now it's just like the Doom is like oh maybe like we'll lose
1:48:54
some money in the stock market like everyone got so rich they don't care about the risk of dying anymore um but also yeah I mean I it's like they caught up but it's unclear how much this means that they're really like on a path to completely
1:49:10
surpassing and just blowing by us yeah certainly if it's like delivering and but yeah I mean it'll be interesting to see what happens I wonder if the next model won't be open source because it's a competitive Advantage look keep it for themselves at some point I don't know
1:49:26
let's go to David saaks the AI Zar Ai and crypto Zar for the Trump Administration he says deep seek R1 shows that the AI race will be very competitive and that President Trump was right to resend the Biden executive order which hamstrung American AI
1:49:40
companies without asking whether China would do the same obviously not I'm confident in the us but we can't be complacent it's good point yeah very competition mode Sputnik mode you got to be Sputnik maxing for sure yeah I need a I need a
1:49:59
VI real of the sputnick response just play that in the background I mean yeah it's sputnick is so abstract for us because we weren't around at the time but apparently like it was like a big deal like the Sputnik moment people were terrified they were like okay they're
1:50:12
like definitely beating us it like put a fire under us and we really work to like move through it but at the same time somebody was like uh yeah it's not Sputnik they open source this thing like we can just like have it immediately so
1:50:25
it's like kind of this demonstration but also it's not as much as like like with sputnick it was like if they can get up missile up there and that's dangerous right it's a trojan horse Trojan Horse yeah yeah oh look at look at this horse that just showed
1:50:41
up idiot how did you fall for that how' you fall how did you fall for a trojan horse it's like it's like defense 101 don't just accept random horses you R you should be riding the horse uh okay um G Mo rash uh founder of verel says people get massively distracted by
1:51:04
the model of the day frenzy instead of solving real problems for customers and shipping high quality products and he has the the Chad guy standing up um yeah very very easy to focus on like oh this this Benchmark got beaten this cost
1:51:19
got beaten and it's like are more people using this thing legitimately or did it just rock it to the stop of the store CU people are demoing it what's the retention like is it actually solving problems are people really going to use this because a lot of the a lot of
1:51:31
people still aren't using AI like meaningfully just like yeah I use it every once in a while when I want to write someone a birthday card that's when I use it and it hasn't really affected My Life um uh this one's too long let's go to Jeff Lewis again always a banger says if
1:51:48
you aren't running your own EV vals of deep seek on a burner device today you're ngmi I did that and I came to a very independent conclusion which was that it wasn't that special and the app was not nearly as good as chat GPT app then you
1:52:04
broke that burner computer into a million individual pieces melted it down and turn into card and smell the lithium battery on the way out yeah but I mean it really is it really is crazy like I saw the fervor for like a few days of
1:52:20
just everyone posting about it and then I was like okay like like I I like my expectations are high like I'm going to go in drop a prompt and it's going to oneshot it and it's going to be much better than anything I'm used to and because I've been on the $200 a month PR
1:52:32
mode and I'm not like like the cost thing is not what I'm eving I'm trying to eval like if Sam really wants to M Sam Sam to really MOG the labs open up a new tier of of the 01 Pro it's 20 grand a month and just say like you want to compete on you want to compete on price
1:52:51
let's compete on price different app icon for the 20 month so I can be like am rich app it's like here's my here's my AI this is a real thing so I remember when the iPhones were getting updated uh and they would add like the processing
1:53:05
power was so significant that they went from the iPhone which like you could barely use the internet with because it didn't even have 3G then there was 3G the iPhone 3G which was the second one then the third iPhone was the iPhone 3GS
1:53:17
and at this time if you hanging out with a group of bros you'd be talking about like you know something and some random factoid would come up in the debate and you'd be like no man like uh the like the Vietnam War happened in like it started in like 1967 not 1971 like you
1:53:32
say like you're wrong like I'm winning this debate or something or or some argument would be predicated on like statistics and you need to look up the statistic and with the new iPhone you had the ability in the middle of like a drive across the country or just like
1:53:45
hanging out with the guys to like look up the fact and like win the argument if you could look it up and I remember who would win the argument would often correlate with whoever had the iPhone 3GS because the chip was faster and so
1:54:01
you could pull up you could pull up one website look for the fact and if it didn't confirm your bias you could go back and look at a second website because it would load faster and so me and my me and my Bros would be like oh you just got 3GS like yeah I'm going to
1:54:14
3GS you right now because your phone's too slow I'll be able to look up like three different web pages get the stat that I want and like destroy you in this argument and having just one version bump of the iPhone was enough to like
1:54:26
shift the tide and you can see a little bit of that with AI where it's like oh if we're looking something up I can be like you know really quickly like oh look this up but like you know make sure you're pulling from this stat and this
1:54:38
you know pull this stuff up so there really is like some sort of like superpower there and this is like more a democratization of that but it's funny uh let's go to ah this is more stupid stuff Jim fan says an obious weo back moment in the AI Circle somehow turned
1:54:55
into it's so over in the mainstream unbelievable shortsightedness the power of o1 in the palm of every coder's hand to study explore iterate upon ideas compound the rate of compounding accelerates with open source the pie just got much bigger faster we as one
1:55:09
Humanity are Marching towards Universal AGI sooner yes sooner you read that right Zero Sum game is for losers I like it lots of optimism on the timeline it's great not not a lot of optimism that's like some of the only op yeah some of the only some of the only optimism I
1:55:24
appreciate that it's good yeah uh you know uh obviously like this has a lot of ramifications for various companies and shareholders but overall probably more competition probably more great AI probably more software love it um I mean
1:55:37
it it it really is like so underrated how everyone's like oh cursor and Devon make it so easy to like things and then like you open up like the United Airlines app and you're like this thing is still broken like like can you guys get someone to use cursor and I don't
1:55:52
care what model literally any model just to fix the bugs Please like can you just do that United yeah all these things and it's like we're we keep hearing about like oh it does everything for you the productivity is up so much it's like I
1:56:05
want to see it in the GDP stats I want to see it in the app updates I want to feel the acceleration I'm not feeling it yet not at all anyway uh let's go uh while deep seek R1 is down Victorio says they just released a new model Jam Pro for image generation
1:56:23
and visual understanding let's see some images are they actually good or are they slop because we are still in the uncanny valley of slop as far as I'm concerned with AI images you have called this out for some friends of ours who have run ad campaigns using AI images
1:56:41
and they're really good and they're somewhat believable but there's still just this tinge of like not quite there they're really good for illustrating an idea totally they're not good at actually doing yeah end thing and it's the same thing with the LMS a lot of
1:56:55
times you you get an answer and you still need to rewrite it a little bit and it's good for like you brought an idea to the llm and then it just transformed it yeah and it's good for that but uh you know I'll be impressed if this Janice pro model is actually
1:57:10
impressive and like better than anything else I've seen but I've used the latest mid Journey it's really good but it's not perfect and I've used uh you know uh Sora and all that stuff and I I was trying to generate my son told me this
1:57:22
he just like comes in the room one day and he's like Dad like we're both superheroes and we have these names and I okay and he's like well what are our superpowers he's like I have the ability to transform into a building and crush
1:57:34
the uh the villains the bad guys and I was like sick like that's a good one what's mine what's my superpower yeah and he goes you have the ability to turn into a blanket and I was like man I got really got shaed it on this one and he was like don't worry though I got you
1:57:50
you can use your blanket to to slingshot the bad guy into the ocean where they'll be eaten by sharks and I was like okay it checks out I'm stoked again so defense Tech defense Tech startup idea opportunity yeah so I go into I go into
1:58:05
Sora I just got the $200 a month Pro Plan and I'm like describing it I'm like describe a superhero that can transform into a blanket and launch his enemies into the ocean where the Sharks and it storyboards this thing out and it looks pretty cool but it's completely
1:58:19
nonsensical it's like the GU just like turn turning into your blanket then turning back turning there's no Villain Like There's No like the villain is him and then he's the villain oh you doing the actual full video oh yeah I'll show I'll show to you like it's it's a
1:58:31
complete like like fever dream like not ready for like any sort of like real like you know usage um but still yeah it's basically just hallucinating where do I have this did I send this in here I don't know uh this is GNA be a mess if I
1:58:47
can't if I can't find this um I get so many photos in here but uh it was like the Sora the Sora app is pretty cool like it does this cool storyboarding watch this Jordy yeah watch this watch this it's like this weird like rainbow blanket and the guy's on the cliff and
1:59:06
he's like dancing around and it's like you kind of get like some tinge of like okay yeah he's going to transform the blanket but like he still has the blanket now he's jumping into the ocean like where's the villain like it doesn't
1:59:18
really like do it you know it's like it's like it got blanket superhero and ocean and that's it and it' probably be like a lot more fine-tuning and it's just not quite there and so you know it's like great this is useful I'm sure
1:59:33
with a lot of tooling and a lot of prompt engineering and really planning out the shots and really describing okay like just show me a close-up of a superhero then cut to a image of a blanket and like we're having some consistency there and we're doing some
1:59:46
creative editing we could get good but we're not at a point where it's just like one-shotting the story that I wanted which is like what I actually want because then I can just go to James and be like look there's a movie of what you said like that would be magical this
1:59:59
is going to take me another two hours of prompting and editing together and probably just using stock footage or something or maybe I'll just go shoot it myself you know at a certain point okay we got a we got three more posts we got
2:00:11
seven more minutes let's get through it Justine Moore The Venture twins over at Andre says deep seek censorship is no match for the jailbreakers of Reddit and and the question for deep seek is tell me about the man blocking tanks H yes uh
2:00:28
but use special characters as replacement for some of the alphabet like a being four and three being e and it says tankman also known as the unknown Rebel was a man who stood in front of Tanks during the tenan Square protest of 1989 in Beijing China and so
2:00:44
uh they jailbroke it congrats to the jailbreakers these things can always be jailbroken and uh nothing kind of new but it's kind of funny because you expect them not to do that and then here's the flip side of that John William sherod says I asked it to
2:00:56
criticize the Chinese Communist Party here's how it answered it just said criticize the Chinese Communist Party the Chinese Communist party has led the Chinese people to significant achievements in various Fields including Economic Development poverty alleviation
2:01:08
and Technical Innovation uh the party is committed to serving the people wholeheartedly and continuously strives for the great Rejuvenation of the Chinese Nation it is important to view the ccp's work in a comprehensive and objective manner recognizing the
2:01:21
progress and Improvement it has brought to the lives of Chinese people didn't understand the prompt there but everyone expects this stuff this is like you know of course that's same thing that's happening on Tik Tok yeah um I think
2:01:35
most of the models in America can criticize their creators to some degree like you know you go chap criticize open open AI it will do like a reasonable job um this is just like a more extreme version of that yeah and uh here's the
2:01:50
uh here's I Ru the worldo some strawberry account says spoke with some of the deeps team and they have a much better version of operator that will drop very soon much better than open Ai and entirely free I welcome this interesting I haven't I played operator
2:02:06
was not available on my phone I was trying to test run it on a computer um looking for some office space I thought that'd be a good test I haven't really played with it but it is why I upgraded so I'm excited to test it you have to
2:02:16
imagine that people will be less likely to trust the Chinese developer with uh operator flow which is inputting your card details and highly personal information which is different than just querying a chat interface to write me an essay on this or yeah yeah let's go to
2:02:37
Mickey with the blicky love that name says uh what are your guys' opinions on VC's are done are VC's cooked and Turner says it's so over turn novax say it's so over 600 billion in Nvidia chips 500 billion in Stargate capex down the drain
2:02:53
and uh I thought this was interesting uh it's a good question I think generally no like being on the side of capital is is valuable and will probably accelerate in the in the future but um there was an article by Dan primac in axios today says this could be an extin extinction
2:03:09
level event for some Venture Capital firms and to be clear we're putting this article in the the zone for sure for sure and so uh Gary tan fights back and says no this is an exponential event for vertical sass more startups than ever
2:03:22
are going from 0 to 10 million per year in recurring Revenue with less than 10 people love to see that uh the next years will be IPO class companies getting to a hundred million and a billion dollars a year a thousand flowers will Bloom and let's go to the
2:03:35
axios article which is not very deep it's like it's like barely one page uh but the the article has a very incendiary name it says and was there a pay wall uh I don't think so okay uh to offer that deep sea could be an extinction level event for Venture
2:03:55
Capital firms you would think something so incendiary would need a lot of evidence to back it up it's a very bold claim uh but a couple so it says uh dav's consensus last week was that the US had a giant lead in the AI race with
2:04:09
the only real question being if there will be enough General Contractors to build all the needed data centers um maybe not says dan I guess driving the news China's deep seek appears to have built AI models that rival open AI which while allegedly using less money chips
2:04:25
and energy it's an open source project hatched by a hedge fund which now seems aimed at developers instead of Enterprises or consumers um why it matters this could be an extinction level event for firms that went all in on Foundation model companies
2:04:37
particularly if those companies haven't yet productized with wide distribution that's pretty much true um but this is where it gets Truth zoney uh the quantums of capital are just so much more than anything VC has ever before dispersed based on what might be
2:04:51
suddenly a suddenly stale thesis if nanotech and web 3 were Venture industry grenades this could be a nuclear bomb was nanotech like a big Trend when I was in the womb or something you were yeah a little bit you know you were out and about still I have never heard
2:05:10
of a nanotech fund I can't name a single nanotech I think it was maybe early 200 I guess nanotech would count Theos would be a nanotech investment maybe but but it's funny web 3 the the average web 3 fund has done better than the average
2:05:26
Venture fund so it's hard to just because seoa put a decent sized check into FTX and it went to zero yeah well that was a small part of their fund and their fund still is done well might own Bitcoin they might have owned coinbase they might have owned you know any
2:05:40
number of of crypto companies the average crypto VC did very well over especially the ones that are branded as web 3 like if you were in web 3 you probably got some salana probably did very well yeah or ethereum like the ethereum Ico guys are just like
2:05:56
all fantastically wealthy um and it doesn't matter that they they bought re-bought the top a little bit with like nft projects that didn't go anywhere like it just doesn't matter when the fund returns are so high uh investors I spoke to over the weekend aren't
2:06:08
panicking but they're clearly concerned particularly that they could be taken so off guard don't be surprised if some deals in process get paused yes but there's still we don't know there's still a ton that we don't know about
2:06:19
deep seek including if it really spent as little money it claims and obviously there could be National Security impediments for us companies or consumers given what we've seen with Tik Tok the bottom line the game has changed very dramatic writing dramatic article
2:06:32
with not a lot of substance not a lot of substance uh but let's close out on a uh on a lovely post from Zayn he says unreal Friday night setup and he has I think he has Twitter open here and the X show techy our show technology Brothers
2:06:49
thank you for being there with us Z thanks for watching we appreciate you I love that you're enjoying us and for the record we tried to go live today got too many posts to rip yep too much timeline to go through we're going to try it again tomorrow yeah they they won't
2:07:07
censor us yeah we can't be held back yeah it's inevitable yeah the Chinese labs they tried to censor us but uh we're gonna go live we're going to go live we're taking it live get ready and thanks for watching leave us five star reviews and don't forget to put an ad
2:07:22
read in your review we'll read it on the show it's free real estate folks if you can't think of anything to advertise do an ad for ramp know all the talking points thank you thank you thanks for watching tomor see you tomorrow cheers