0:00
One of those concerns whenever we have new 📍 technologies is, okay, this is gonna, you know, do harm to the world and only the adults in the room.
One of those concerns whenever we have new 📍 technologies is, okay, this is gonna, you know, do harm to the world and only the adults in the room.
Make the technology available and be qualified to put the safeguards on the technology.
But throughout history, we know that's not the case.
What worried me was kind of a panopticon where a few people were controlling a very powerful technology.
It'll be beneficial for the world to have an ecosystem based on an open source foundation.
We will want these models to fit on consumer level hardware.
These models are trained to blabber text to predict the next characters and who knows how large of a scale they need to be until they're great mathematicians.
There's a wonderful quote that I love that is you can't ask a person to make a list of ideas that would never occur to them.
Uncertainty is very important because it's an admission of the fact that, you don't know what you're doing, but this is your very best guess Emergent type models, they are far more adaptive and far more similar to complex adaptive systems.
We have like a hundred billion neurons and after a certain number of neurons, these, these phenomena emerge.
Jim O'Shaughnessy: Hi, I'm Jim O'Shaughnessy and welcome to Infinite Loops.
Sometimes we get caught up in what feel like infinite loops when trying to figure things out.
Markets go up and down, research is presented and then refuted, and we find ourselves right back where we started.
The goal of this podcast is to learn how we can reset our thinking on issues that hopefully leaves us with a better understanding as to why we think the way we think and how we might be able to change that to avoid going in infinite loops of thought.
We hope to offer our listeners a fresh perspective on a variety of issues and look at them through a multifaceted lens — including history, philosophy, art, science, linguistics, and yes, also through quantitative analysis.
And through these discussions help you not only become a better investor, but also become a more nuanced thinker.
With each episode we hope to bring you along with us as we learn together.
Thanks for joining us, now please enjoy this episode of Infinite Loops.
Jim O'Shaughnessy: Well, hello everyone.
It's Jim O'Shaughnessy with yet another Infinite Loops.
Today, I cannot tell you how excited I am to have my guest and colleague at Stability AI, David Ha, who is the head of strategy and our lead architect for our model design at Stability AI, formally of Google Brain.
And according, David, to many of the experts that I spoke to after I heard that you were joining us, one of the very top minds working in AI today. Welcome.
David Ha: Oh, thanks, Jim. I'm humbled to be here.
I think this field is so large that there's no top minds, and as we discussed many times, there's many minds and there's many brains that work together and we are all standing on the shoulders of giants in this field.
Jim O'Shaughnessy: Yes, absolutely.
And we're going to get to that, because that's a mutual interest of both of ours, sort of the collective mind.
I call it the human colossus.
Now that we have AI, I think I'm incredibly excited for what the future holds with the advances that we are making, standing on the shoulders of giants.
And one of the first things that I just have to ask you, you could work anywhere, you could work anywhere.
Why'd you pick Stability AI?
David Ha: That's a good question.
I mean, I'm really passionate about not just the technology, machine learning technology, but also just making them broadly available to the world.
When I think about what happened in the last few years, there's an explosion of generative AI tech.
As these models scale to larger sizes, you get more capabilities.
But one of the problems there is, just as these models scale, they require much more infrastructure and work in order to serve them.
It's not like the consumer or the hobbyist can, even two years ago, one or two years ago, can download a state-of-the-art model, like a language model, and have them work on their machines.
These are, at the time, they need to be productionalized and served on large corporate machines.
So to me, it's a double-edged sword.
As you scale up the model, you get better capabilities, but at the same time you get more control from the people who were working on these.
And being in the industry, being at Google, have lots of friends at open AI and so on, you can see the writing on the wall.
David Ha: And when I decide to leave the company, I had no plans to join Stability AI.
I wanted to just work by myself on open-source machinery models, be a tinkerer.
I was talking to Andrej Karpathy as well, who left Tesla around that time, and he became a tinkerer.
Anyway, I was like, Hey, let's form a group of tinkerers of ex tech people and we can all work on open-source projects.
And I was really passionate about text to image models last year.
The first time when I saw DALL-E 1 came out, I was following some of the efforts on the open-source reproductions.
And when DALL-E 2 came out, I was one of the first beta testers of DALL-E 2 from Open AI.
And I knew this is going to be a game changer, because I think language models were well for a while, but there's something special about the visual domain.
I think we've been evolved to understand vision more intuitively instantly, than processing text.
David Ha: So I became totally obsessed with text to image models.
And when I decided to leave Google to work on open-source efforts, I focused on open-source models like DALL-E Mini, which went viral last year and Stable Diffusion.
And that was the time when I saw some of the efforts from Emad when he started Stability AI, with a company whose mission is to open-source or back open-source efforts in the development.
So he approached me to help out with the mission, and after some thought, I decided, yeah, why not?
David Ha: I think Stability AI is a good platform that he set up to push this efforts.
I think, given the world needs an open-source alternative to closed source model, so just how decades ago there was open-source operating systems and closed source operating systems.
Now there are open-source models and closed source models, and there needs to be that spectrum.
And as you said, I could have just started a small company with a few ex Googlers and raised some VC funding, worked on some model to do a product market fit and try to bootstrap from there, get a few customers, make them happy and so on.
David Ha: But that's not so fun.
You want to be at a platform that is at the center of the open-source movement, and then to push things forward.
And I think that's what I signed up for.
It's going to be challenging, it's not going to be easy, because just like how companies have developed business models and directions around open-source operating systems, we're going to be developing an entirely new set of business directions, business models, to make sustainable developments in open-source models.
So I think it's not going to be easy, it's going to be challenging.
But it's a challenge, that I think is worth pursuing.
Compared to doing the typical X big company tech employee startup.
So it is just more fun, in my opinion.
Jim O'Shaughnessy: Yeah, I agree.
And I am passionate about open-source too.
That's why I invested in the company.
I'm ruled by a Taoist-Stoic philosophy, which is go with the flow.
Don't let things that you can't have any effect on, disturb you.
But also, when there are things where you can have an effect by participating, you should.
And when I saw Emad and the release of the total model with weights, I went, that's what I'm looking for.
Because I believe that, just philosophically, I'm all for both.
I'm all for coexisting, and I'm all for everyone making things better through cooperation, through different business models, et cetera.
Jim O'Shaughnessy: But what worried me was a panopticon where a few people were controlling a very powerful technology.
And I have four grandkids and I don't want them to grow up in that kind of world.
I want them to grow up in a world where open-source at least has a place at the table.
Plus, on top of it, when you look at the history of open-source, it runs the internet now.
And in studying open-source versus closed source, I just did a recording earlier today with a guy who wrote a book about von Neumann.
And one of the things that I learned when I was researching that is, his kind of the father of open-source in that he wanted all of his work to be put into the public domain.
And I was like, I like this guy. David Ha: Yeah, exactly.
I like Von Neumann is a huge inspiration to the field.
I have his books as well.
And I think I read somewhere along the lines of, I think even after his death, some of his partners tried hard to document his last works to publish memoirs of his work.
I think it's throughout, he started it, but throughout history, open-source or the general hacker culture of tinkering things, not Black Hat hacker, but more like the hacker culture of tinkering things, that's when innovation really starts.
Back in the days when you had the hobbyist tinkering around with building their machines.
When I was a kid, people didn't buy a shiny MacBook.
You buy a box and then you put your PCI cards in the box, you install the RAM by yourself, you tinker with it, and that led to a generation of tech talents.
David Ha: When Linux came out, people were tinkering with the operating system, building things on top of it.
And then now you get mobile devices, toys that run Linux.
And it's going to happen the same way with machine learning models as well.
You're going to have closed source closed models inside big companies for specific business reasons.
But I think it'll be beneficial for the world to have an ecosystem based on an open-source foundation. So in the very...
Let's say fundamentally, people can tinker and pull apart the models to understand what's going on, build things on top of it, and even collectively identify ways to work on malicious use of these models as well.
David Ha: Because one of the concerns whenever we have new technologies is, okay, this is going to do harm to the world and only the adults in the room should make the technology available and be qualified to put the safeguards on the technology.
But throughout history, we know that's not the case.
With this technology, it's the same.
The more eyes we have on it, the more understanding we have for these technologies.
David Ha: Language models is a good example.
The there's a sufficient community in academia and also industry, looking at language models.
And there's a whole body of literature built up on identifying malicious uses of language models, what can be done, how to analyze them, biases and so on.
So yeah, it is just that's one direction.
David Ha: The other direction I feel, is when these models are open, in the sense that people can tinker and develop them collectively, you get all sorts of innovations that you will never get at big companies.
Like Stable Diffusion is an example where I would say, we're not really optimizing for pure state-of-the-art performance.
We're optimizing for the best performance given the constrained resources that we have.
David Ha: For example, say something like Stable Diffusion would fit on a consumer level GPU.
And because of that, you suddenly have a million people that could develop on it, and you get efforts, you get developments on the research side too.
People can develop fine-tuning algorithms so that you can generate images from Stable Diffusion model with an extra 30 training samples of ourselves.
And that's something that is simply not possible for large models, because you'd need the model to be orders of magnitude, maybe three to four orders of magnitude, and you'd need an entire data center to store such a model for it to know everyone.
And maybe you don't even want the model to know everyone, maybe you know, want the model to fit on your iPhone and know you, right?
So, I think you get all sorts of these innovations when the model is out there, that the model fits on even a higher end, consumer level, hardware.
David Ha: The parallel I was thinking of is, back in the day there were innovations in craze super computers in the mini computer, which is a mega computer and all of these computers.
But the real innovations come when there's the personal PC computer that people can play around with, develop games on.
And once you have people like John Carmack, who really worked on getting the maximum bang for the buck on this hardware.
And I think we're going to see the most important innovations is going to happen, not necessarily at the huge data center level, but at the individual GPU level.
David Ha: So for me, as someone who's running the strategy and the research org at Stability, one of the things we're trying to do differently, is not simply work on just scale, raw scale, and bigger is better, but we want to have developed algorithms in a clever way, so that they can do a lot with very little.
And because of that, Stable Diffusion now can work on a smartphone.
Future models we develop, we have this constraints in mind, so that we would want these models to fit on consumer level hardware.
David Ha: And it goes back to our ambitions in AI as well.
What's human intelligence?
It's not like we don't need a gigawatts to power our brains, we need 30 watts, what's in here.
So just how we've been evolved to do maximum with minimum resources, I think, it makes sense when we're developing AI, to do a lot with reasonable resources.
Jim O'Shaughnessy: Yeah, because constrained resources, and that was one of the things that I also was drawn instantly to the open architecture, because of the idea when you study innovations, when you study all of that, cognitive diversity is such a huge advantage.
And I was actually talking to an investor friend and he was asking me, well, are you worried that the United States government is going to shut down AI research?
And my response was, and I need to give it more thought, it was more off the cuff, but my response was, absolutely not.
Because the ability to allow open architecture, the ability to allow innovations, because we can allow hundreds of millions of people around the world to tinker and make these things better, will make us much more robust, if you will, in terms of all of the innovations, than say a closed system would ever be able to do.
Jim O'Shaughnessy: I mean, there's a wonderful quote that I love that is, you can't ask a person to make a list of ideas that would never occur to them.
And so I often find when I talk to even very smart people, they miss this.
I'm a huge fan of David Deutsch, ‘The Beginning of Infinity’, and they miss this idea that's a common trap even brilliant people fall into, is this notion that we know what the innovations of tomorrow are going to be.
We have absolutely no idea.
And he's got a wonderful piece in the book where he says, what were the smartest people in 1900 commenting about the internet and quantum physics?
And he says they weren't commenting at all because nobody knew about them.
And at the same time, when you go back and read the greatest minds then, well, we might as well just close the patent office.
We might as well close physics down, because we know everything.
Jim O'Shaughnessy: And so it's this idea that you share with me that, God, no, we don't know everything.
And the ability to allow many, many minds of varying interests, to use the base technology, I think, is going to be able to allow far greater innovations than a closed system.
Jim O'Shaughnessy: And that leads me to a question.
What was your biggest surprise when you saw?
So we released the model with weights.
I was blown away by the tsunami of innovations and everything else.
I'd lost the bet I'd made with a friend.
A number of different kinds of innovations.
What was your reaction when you saw all of this coming online?
David Ha: Yeah, I mean, I was blown away as well.
I think, as you mentioned, no one can predict with sufficient accuracy, where this technology will develop.
You mentioned that the military may have funded the development of the internet, but whoever developed the internet, Tim Berners, he would not predict where the internet is now, as someone who made the worldwide web back then.
And just like now people are even on Twitter and so on, you have all of these...
The thought leadership saying, "Okay, this is how we're going to get AGI.
We're going to scale, we're going to do this."
We know, we don't need to bet, in 30 years time, we're still going to be around and things will move in a different way than what people have predicted today with absolute certainty.
And I think it goes to what you said, there's a quote by a politician, "There's known knowns and there's known unknowns and there's unknown unknowns," and I think in this technology is all unknown unknowns.
We don't know what's going to come.
And the best way to maximize the innovation in this unknown unknown is to admit that you simply don't know the list of all of the great things that's going to happen and you want to put technology in a place where these things can be discovered.
You want to maximize, to quote my friend Ken Stanley and Joel Lehman, evolutionary researchers who wrote the book, ‘Greatness Cannot be Planned’.
Jim O'Shaughnessy: That's a great book. David Ha: Yeah.
You just have to maximize the stepping stones and I think something...
Getting back to what we're doing, an open-source generative model, a large language model that's trained on the world's knowledge, you put it out there, people would find things to do with it that you have just simply no idea. I was just blown away.
I track some of these developments and there's the obvious, the people making images or certain artistic styles and so on, that's a given.
David Ha: What I thought was really notable to me is when you pair up these models with creativity in terms of letting kids draw.
So I was blown away with the example of a father playing around with the Stable Diffusion model and feeding in his toddler's doodle of, I think of little monster robots and then they use the model to convert it into a high quality robot monster that made chocolate cake or something like that.
And I feel like the father will show that to his son and say, "Hey see, look son, this is what you made with the model."
And to me that is very satisfying to see this sort of feedback loop because we're not viewing these models as some sentient machine with...
That's going to do its own things.
I really value the interaction loop that people are using these models for because they enable a stronger bond or deeper communication between people like the case of the kid feeding in their doodles into this model to create newer paintings and they can use that to feedback into their own idea generation process.
David Ha: We're going to see more of this.
It's going to be gradual.
Just like with any new technologies, when my kid grows up, maybe in a few years, they will take for granted a world where you can type in some texts and produce some nifty image from it.
And all of their kids will be using at school will be using these to do their homework assignments and it's just going to be a given.
And I think the one thing, I think, also our founder Emad is interested and passionate about is on the educational aspect.
And when you ask me what kind of things that I was blown away with is...
I think with all things with our founder Emad, he thinks five steps ahead, whereas I will say, "Okay, to get there we have to think 1, 2, 3, 4, 5, how to get there."
David Ha: I think it's not like we're going to have a curriculum using language models and Stable Diffusion, whatever, at school.
It's going to be very gradual.
At parents, at teachers, kids, they're going to slowly use these technologies, give feedback to make better versions of them and I think we're going to...
That's one of the aspects that I'm quite excited about in enabling deeper human to human communication.
Related to that, for me, some of the use cases that I found most fascinating because both of us are avid Twitter users and we see what...
We see so many generative AI images every day.
When I honestly speak at this point, offends some people, but I'm actually not that excited about generating art, per se, using these models.
I'm more excited about people using these models to generate what might not be considered art. More like memes.
David Ha: I think last year when these models become open-source, people were not able to use them for their full extent with Open AI because of some of the censorship issues around it.
I cannot generate a Picasso image of Trump, for example, last year.
But what I noticed was people started to use these open-source models to communicate things that are culturally relevant at the time of the moments.
In the US there's obviously lots of discussion and controversy around guns, weapons, and the right to bear arms.
And I think last year we saw some unfortunate shootings at school and people...
Actually, I saw people use something like an open-source text-to-image model to just let out steam and they created a Fisher-Price gun or something like that.
David Ha: And I thought that that went viral and I thought, "Yeah, this is art.
This is what contemporary art is.
There's a message behind something like this," that's a Fisher-Price grenade.
It's not a very fancy painting. It doesn't look so nice.
But people can use these models to create something that has a deeper meaning and communicate or vent their thoughts at the height of cultural issues of the moments to the rest of their peers and I think this is one of the casual creativity moments that really impacted me.
Jim O'Shaughnessy: And you and I are both, it seems big believers in the what they're calling the centaur approach, man plus AI coming up with all sorts of incredibly creative, innovative things and that's one of the things that I see happening.
Particularly with artists, I mentioned to you when we were off mic, that I've been collecting art for a long time and so I got a lot of worried phone calls from some of the artists that we collect because we tried to collect living artists and when I spoke to them about
the iterative ability, that they could iterate on their own work and almost have a companion there giving them feedback, the feedback loop being so rapid, they became actually quite excited. I believe it's a crawl, walk, run and it's understandable when people are frightened
I believe it's a crawl, walk, run and it's understandable when people are frightened and confused.
Jim O'Shaughnessy: And one of the things that I'm trying to do and why I'm delighted to have you here now with me is to help them understand and not simply dismiss their concerns.
It's one of the reasons why Stability AI has an opt-out from our training models for artists if they don't want to participate.
I think O'Shaughnessy Ventures, which is my company, which is the company that invested in Stability AI, is also right now negotiating with an AI tool company that allows the artists to actually work with their hand but then train it on their own dataset and keep a growing library of all the iterations.
And when I tell my artist friends about it, they're like, "Wait, like a paintbrush?"
And I went, "Yeah, exactly.
Or your finger or whatever you want to use," and they get really excited.
So I definitely see this idea and I read your paper, Teaching Machines to Draw, which you wrote in 2017, and it immediately resonated with me because of the way that you approached it.
Jim O'Shaughnessy: It was like, "Let's have a process that is very much a collaboration between machines and we humans."
So I'm very, very excited about all of that coming down.
But I also understand and try to help people, are trying to come along and people get afraid of new tech. They always have.
You go all the way back to radio and people thought radio waves were dangerous and you go back before then, the telegraph was going to destroy society, novels were going to destroy society and when you do a deep dive, you see a bunch of polemics against, who?
William Shakespeare who was corrupting the youth of England and should be banned.
And so that seems to be something that we face time and time again.
And so that leads me into one of the things that I wanted to do with you.
I have a lot of friends who say, "I really would love to just have some basic understanding of what..."
Jim O'Shaughnessy: They're willing to say to me, "What's a large language model?
And why is everyone so excited by GPTChat and large language models?
And what's the difference between a large language model and a stable or a diffusion model and what's a multimodal type thing?"
So if you wouldn't mind and you would indulge me, could you give us just a little bit of a tutorial for your average?
Most of our watchers and listeners are quite bright, so you don't have to dumb it down, but large language models are great for certain purposes.
They're not so great for other purposes.
Same with generative models.
So if you wouldn't mind, just a 101 on large language models then versus the other models that we're working with. David Ha: Yeah, sure.
Let's take a step back before we talk about large language models, language models or just models in particular.
At the end of the day, these models, they're statistical models.
They're prediction models.
They model the statistics of the world and the data we feed them to train them.
Like we were talking about von Neumann earlier, if we even stepped back another century when these models started...
When people started to use these for things like when Bayes I think he was a priest or a prior. Jim O'Shaughnessy: Yeah. You're right. That's right. He was. David Ha: Yeah. Yeah.
And he had to figure out a way to model how long people would live so that the widows and orphans can gets receive an annuity to support their... To live from.
So even back at that time, there was a concept of a model that the idea that you cannot predict everything with absolute certainty.
You have to model the statistics from a bunch of data.
So the simplest model that people use a few hundred years ago is, okay, you have a bunch of data, you wanted to model, say, your heights given your age or your weight given your age, then you would fit a linear model to that.
It's not going to be perfect, but here's a bunch of data points.
To the data point, you draw the best line that fits it and you observe some uncertainty around it, then this is a prediction model.
Given X, you predict what likely Y is with some uncertainty.
David Ha: So that uncertainty is very important because it's an admission of the fact that you don't know what you're doing, but this is your very best guess and this is the error bar that you're going to get.
And that's basically the foundation of machine learning is you have data and you try to find the relationships between the data sets and you make a prediction model with an uncertainty.
And then when machine learning started taking off is when we have larger sets of data.
So we no longer have a hundred or a thousand samples of human height versus weight or age versus weight.
We would have a hundred million samples of all sorts of different characteristics.
Then we can no longer think in one dimension anymore.
Once you extend beyond one dimension, you get a two dimension, you get a plane and beyond that is a hyperplane, you're thinking about tens of thousands of dimensions.
David Ha: You're fitting models of tens of thousands of dimensions, which...
And they may not be linear models, they could have models with curvature or it is very common in statistics there.
There's like, how do you say?
The sigmoidal models where you're not predicting one thing, you're predicting the probability of that thing happening, like a zero or one threshold.
So machine learning, I think the advent of deep learning really became popular around exactly 10 years ago when researchers, some of my previous colleagues and Geoff Hinton has demonstrated that's these neural network models that can be trained to understand the statistical properties of really large data sets.
Back then you had a data set developed at Stanford by Fei-Fei Li's group called ImageNet and for a while it was all of these handcrafted traditional computer vision methods.
But what deep learning has shown is you can have less handcrafting of features, just give the model the data and let it learn the rules by itself from the data in order to get the best prediction accuracy.
David Ha: So I think from 10 years ago you started to see the rise from simple linear regression or logistic regression models to more of a deep neural network models that can have a...
Rather than having a pure linear or logistic regression, you can can have more curvature in hyper-dimensional space.
This is something that neural networks do.
And in a sense, maybe some people think this is what our brain does, as well.
We have a hundred billion neurons and after a certain number of neurons, these phenomena emerge.
So I'm going to talk about that next.
So we're at the stage where neural network models are starting to be really good at prediction given, because it can model lots of data.
David Ha: Then the interesting thing is, sure, you can train things on prediction or even things like translation.
If you have paired English to French samples, you can do that.
But what if you train a model to predict itself without any labels?
So that's really interesting because one of the limitations we have is labeling data is a daunting task and it requires a lot of thought, but self-labeling is free.
Like anything on the internet, the label is itself, right?
So what you can do is there's two broad types of models that are popular now.
There's language models that generate sequences of data and there's things like image models, Stable Diffusion you generate an image.
These operate on a very similar principle, but for things like language model, you can have a large corpus of text on the internet.
And the interesting thing here is all you need to do is train the model to simply predict what the next character is going to be or what the next word is going to be, predict the probability distribution of the next word.
David Ha: And such a very simple objective as you scale the model, as you scale the size and the number of neurons, you get interesting emerging capabilities as well.
So before, maybe back in 2015, '16, when I was playing around with language models, you can feed it, auto Shakespeare, and it will blab out something that sounds like Shakespeare.
David Ha: But in the next few years, once people scaled up the number of parameters from 5 million, to a hundred million, to a billion parameters, to a hundred billion parameters, this simple objective, you can now interact with the model.
You can actually feed in, "This is what I'm going to say," and the model takes that as an input as if it said that and predict the next character and give you some feedback on that.
And I think this is very interesting, because this is an emergent phenomenon.
We didn't design the model to have these chat functions.
It's just like this capability has emerged from scale.
David Ha: And the same for image side as well.
I think for images, there are data sets that will map the description of that image to that image itself and text to image models can do things like go from a text input into some representation of that text input and its objective is to generate an image that encapsulates what the text prompt is.
And once we have enough images, I remember when I started, everyone was just generating tiny images of 10 classes of cats, dogs, airplanes, cars, digits and so on.
And they're not very general.
You can only generate so much.
David Ha: But once you have a large enough data distribution, you can start generating novel things like for example, a Formula 1 race car that looks like a strawberry and it'll do that.
This understanding of concepts are emergent.
So I think that's what I want to get at.
You start off with very simple statistical models, but as you increase the scale of the model and you keep the objectives quite simple, you get these emergent capabilities that were not planned but simply emerge from training on that objective.
David Ha: It's similar as a researcher, I think both of us are interested in things like civilization and developments.
We ourselves, we only have a very simple optimization objective to survive and to maybe to pass on our genes to our descendants.
And somehow throughout this simple objective, like human civilization has emerged with all the goodness.
And I find it fascinating that we have a parallel universe where you have a simple objective of, "Let's predict the next character," and you get this vast understanding.
So yeah, I think that's the high level description of what's going on and what we could see from these principles.
Jim O'Shaughnessy: One of the quotes that you like that I also love compares emergence to engineering.
And the quote was, "Bridges are designed to be indifferent to their environment and withstand fluctuations," whereas with emergent type models, they are far more adaptive and far more similar to complex adaptive systems, which I'm fascinated by, which we both are, I know since chatting with you offline.
Jim O'Shaughnessy: And one of the things that you made me start thinking about a lot was the idea of resource constraints.
And you use our own evolution as you just mentioned, that goodness, we have two primary objective functions, live and pass our genes on.
And out of those two simple objective functions we tried to maximize, came this incredible world of 8 billion sentient beings.
So I love the connections between emergence and what's happening right now in all of the models.
But I also wonder, are these causative or are they correlative, do you think?
David Ha: That's a good question.
It's tricky, because when people...
Let's take a step back and say your job as an engineer is to design a system, to identify whether an image is a cats or not a cats.
And before a neural networks or machine learning, you would have to come up with all of these rules for figuring out, okay, let's put in the whisker detector and let's put in... Does it have two eyes?
Is the cats a full cats, with the body or just the head of the cats and so on.
David Ha: Like these expert systems back in the '70s or '80s, you would have maybe come up with 2000 rules.
And that is an example of a hand engineered bridge rather than an emergent system.
An emergent system would be like, okay, here are a million pictures of cats, figure out what's a cat. And it'll do that.
David Ha: So it is very tricky because there's also the question of correlation versus causation in this.
And one of the examples that I like the most is, that's why the neural networks also it's a double-edged sword because sometimes your model might treat a correlation as causation.
There's some examples of ImageNet classification that I find hilarious.
There's a general category inside ImageNet called ants, like the insect ants.
David Ha: Some models, I think they optimize too hard to get state-of-the-art accuracy that people were feeding in an image of the side, the corner of a wall at home, and it would classify that as an ant, because maybe there were lots of examples of an ant around the corner of a wall that's where they hangout.
So it would just think this is an ant because well, that's what the data is suggesting.
David Ha: So I think one of the arguments is, well maybe there's not enough data to suggest that.
There's not enough examples for it to do that.
But this is debatable as well.
Maybe just purely scaling on data and having the rules learned is not the only way forward.
One of my hypothesis is, it's a combination.
Well, some works I did before, there's a paper called ‘Weight Agnostic Neural Networks’, where I tried to find...
David Ha: My collaborators and I we did a project where we trained neural networks without training the neural networks.
We only found the architecture of the neural networks.
But that architecture still had to work, not great, but still had to work kind of, if we randomized the parameters of the model.
So you actually try to find the neural network model that needs to work for some task, even if the parameters were randomized.
David Ha: And I think one of the intuition or the parallels or the inspirations for that research is the structure of our own brains.
It's not this random like a neuro network that have no structure.
There's a lot of structure in the brains and how it's developed.
And the architecture of that is optimized for particular task of survival on Earth.
We will not survive at the bottom of the ocean or on another planet, is on Earth.
We're not general intelligences.
We're very good narrow intelligences for Earth.
Jim O'Shaughnessy: Lines to Earth, right? David Ha: Yeah, exactly.
So I feel like getting back to causation versus correlation.
A lot of the cause of structures, they could be learned, but sometimes maybe through evolution or there's an outer loop of how the system is defined or how the rules were set up so that the systems can learn may influence ultimately what it can learn and the understanding that it can do.
One thing for example, when people are training now these large language models, in addition to doing things like language, they can also generate computer code.
David Ha: So now GitHub has a pretty good business.
There's a product called Copilots, as you may have heard of that, where the programmer can use it as an assistance and say, hey, I want to have this JavaScript code for this use case, have a little form, or I want to sort a list of numbers, generate some code to do that and it'll do that.
And in order to do that, the data set of the large language model obviously needs to contain lots of programming code.
Otherwise, if you're only going to train a language model on the best works of Shakespeare, it's not going to go C++.
That needs to be in a data set.
David Ha: So there has been some papers that found when you mix the data set, if you put in a decent chunk of human languages and computer code, even when the language model is trying is prompted to have a natural language conversation, there is some evidence that the conversations are more structured and more methodical because maybe it knows how to code.
So the way it talks to you would be a very difference than if it didn't know how to code.
David Ha: So these are as an example of things that you can take into the structure or you specify the structure of the data sets and that will guide the model to ultimately learn certain properties of the world.
So maybe that will help just how I view how to approach things like causation versus correlation.
But that's a really good question and it's unsolved.
No one knows the answer and just theories, but this is one of the biggest questions in AI, how do you view this. Jim O'Shaughnessy: Yeah.
My company, O'Shaughnessy Ventures, it has several verticals, but what we're probably best known for is the investing side.
And we are right now looking at a company that the founders are maintaining that large language models are good at very many things, but the current ones at least are not trained to be really good with numbers.
They're good at, two plus two equals four but the business use cases, which is a massive market as you know, are going to require a very different probabilistic model that this company is bent on building one of those foundational models, but it's got the accuracy.
Jim O'Shaughnessy: So the challenge of course is no CFO at any company is going to use any model where there's a chance that when the CFO goes, "Yeah, what was the earnings from our Bangladeshi subsidiary between March and February of last year?"
He's going to want that query to give an exactly correct answer, and that is a not trivial pursuit.
David Ha: This is exactly correct because these models are trained to blabber, text to predict the next characters and who knows how large of a scale they need to do until they're great mathematicians.
As we mentioned earlier, if you want image model to generate every single human on Earth, it's going to require data centers the size of an entire country.
And we don't know right now what extra scale we need for these models to be able to process numbers perfectly or whether even that's feasible given the energy requirements in running these.
But what I can say though, this is an interesting question you raised.
This is very interesting because if you think about it, we suck at math too.
Jim O'Shaughnessy: Yes we do.
David Ha: If you ask me- Jim O'Shaughnessy: There's a great line that we humans go, we're trained to do 1, 2, 3 a lot. David Ha: Yeah, exactly.
If the CFO ask me, "Hey David, what's your profit projections to the cents at the end of the year?" I go, "I don't know."
What I'm trying to say though, what humans are good at is, okay, if the CFO asked me this question, I got to generate a report.
I take out my calculator and I bring out my spreadsheets and I start typing the numbers in and I will present the CFO with the spreadsheets, with the formulas in there.
This is how I arrived at that result.
And I showed them this calculator, this is how I arrive at that result.
And surely a calculator is supposed to perform calculations will be much better than a human, let alone a language model at doing these things.
David Ha: So getting back to the language model on the research and development side, there is quite a bit of work on not necessarily training these models only to generate text, but to train them to use tools as well.
So imagine if your machine learning system is interfaced with a calculator or even a programming language or a database, then in theory, your language model would able to output a query for the calculator and get the result from that query, or it can produce something like a SQL expression to query a database and get output from the database.
And I think that is much more approachable compared to simply scaling things up.
So, once you throw in a neural interface for a calculator or a neural interface for a database, this is what we need to do as well.
Then your algorithm can be trained to use these tools and to get back to the users with results based on the tools that they used and also, the user will be able to figure out what queries they submitted to the database and also record what has been typed on a virtual calculator.
David Ha: So, the user, the CFO can have confidence that, "Okay, this is the steps taken to arrive at this result." And we're not that far.
There are companies out there, started by my friends, who are doing exactly this.
There's a company called Adept AI where they're actually training these language models, not necessarily to output text, but to navigate webpages and to give results based on navigation.
David Ha: This is very powerful because you could put anything on a webpage, you could put a calculator on a webpage as well.
So, you can put a scratchpad, you could put a programming language on a webpage.
There's another company called Perplexity AI, I'm not sure you've heard of it.
They made a language model tool that can do Twitter search or something like that.
And what they try to do, from my understanding, is they try to train a language model to not just generate text, but to output a SQL query into a database.
David Ha: So, you can ask Twitter, "Who is the most annoying user that follows me?"
And it will have some sort of a weird SQL language, supposedly, that has some certain neural traits like that.
We can come up with that.
And it might be able to queries their Twitter database and based on some special filtering, they come up with using the language model, give you, "Here's your top 10 annoying Twitter followers." Or something.
Jim O'Shaughnessy: I fear that if I queried that database, it would just come back with my account. David Ha: Yeah.
So, you can see how the magic of language models can be used in conjunction with so-called very strict ways, very methodical ways, mechanical ways of using traditional interfaces.
And I think that's going to be very powerful and we're going to see more of that.
Getting back to what we discussed earlier, it's not just about scale, it's about getting your maximum capabilities when you have minimum resources.
And given that we have decades of database research and computers were designed to compute, to add and multiply numbers, we should let our models use these capabilities naturally rather than relearn how to do them from emergence, sometimes.
David Ha: It's like, even humans, we're going to have to wait like a million generations to evolve biological brains that are good at multiplying numbers, whereas our ingenuity is not that, but the fact that we can use tools.
So, I think we're going to need for our models to learn how to use tools rather than just doing everything by themselves, to address some of these challenges that you outlined just now.
Jim O'Shaughnessy: And that is very much in line with the way I see this proceeding.
I agree completely, and I always try to get people to understand that really what we're doing is just we're going to need to train them to use tools because we understand how to use tools and that is how we have this cumulative cultural evolution, because we can time-bind what we find through printing or some other way of capturing our learning and pass it forward in time.
Robert Anton Wilson calls that time-binding. David Ha: Yep.
Jim O'Shaughnessy: And we see that this idea of the human colossus taking shape, where you move from strictly zero-sum games to positive-sum games, you move from just pure cutthroat competition to collaboration, you move to a higher level of understanding with that built up knowledge.
Jim O'Shaughnessy: You said at the beginning, "We're standing on the shoulders of giants."
And that is very much the way I see this proceeding.
I know that we are working on an open-source solution to GPT Chat, those types of large language models, and just the sheer number of parameters sometimes take my breath away.
I'm not going to list them all, but they're in the billions, which is an ass-load.
I looked that up by the way, an ass-load is an actual term of measurement.
Actually, it's a butt-load. David Ha: Butt-load. Yeah. Or a shit-load. Jim O'Shaughnessy: Yeah. Exactly. Exactly.
But I also agree that this idea of compression, this ability to use 70 billion different parameters to get down to a better model as an open-source model.
But I think it's important for everyone to understand that different needs have different model solutions. Right?
Like the one we were just discussing, this company we're looking at is very bullish on the fact that they could write a new foundational model that is SQL-sourced and able to train it up on that as opposed to the way traditional large language models are doing.
Jim O'Shaughnessy: What do you think about though, we see, and I use it too, this idea, I always look to evolution, not just our evolution, evolution in general on this planet.
So, N=1, we're on the third rock from the sun, but at least on this planet, complex adaptive systems seems to rule everything around us, from spores all the way up to massive human organizations.
Jim O'Shaughnessy: So, it does seem to me to be at least reasonable to assume that going down that path of emergent-type models that can use tools is going to get us further along then this brute force.
And I love Taoism, and one of the things that is always ever present in the Tao Te Ching, is that which is about to die is stiff and doesn't adapt or bend.
And that which is going to live is supple, it is movement, movement forward.
Stasis is death, movement is life.
And I think that if we use those as our operating assumptions when we're looking at how to develop these models better and how to make them serve us better, which I think is what we want here, I definitely think that your idea about, "Resource constrained emergence, let's look at evolution."
That just makes a tremendous amount of sense to me. Right? David Ha: Yeah.
I think that's a great quote that you just quoted as well.
How living things move and adapt and change and reinvent themselves in general, that's the idea.
And dead things or just stale, they don't move.
When I look at how these technologies are going to develop in the world, in the same way that's complex systems or emergent phenomena has ruled the world, has shaped human society, it's going to be messy.
This is not going to be a clean process.
David Ha: Even if the large companies want to make it very clean, like, "Okay.
All of these technologies is going to be under this API that we're going to be in control of and everyone is going to get their technology development through us."
The reality is definitely not going to move that way.
With or without us, our company, people are going to make open-source solutions.
What I think is going to happen is you're going to have open-source language models.
They can be developed by Stability AI, they can be developed by other companies, other countries and other governments or academic institutions.
And they're just going to be out there roaming the world and people are going to use them in a gazillion different ways.
David Ha: When I think about developing a language model, as we discussed earlier, I would prefer such a model to somehow function with reasonable compute requirements, maybe a few GPUs that they can order on Amazon and somehow people can serve them.
And once you get into a world where most people who can afford to run their own language model, even maybe rent one on a cloud for a dollar an hour or something, you get an exponentially larger types of use cases that we will never see.
David Ha: You're going to see maybe certain professions have their own models for that profession, like the legal profession, and so on.
You're going to see play writers have their own language models that's customized for their uses.
This is going to be the ugly as well.
You're going to see malicious uses of language models and you're going to see people who can find ways to detect malicious uses of language models.
So, you're going to see all of these systems out there in the world, and just in complex theory, there's going to be predator/prey cycles, there's going to be equilibrium.
And where we end up with will be that new equilibrium, once everything is out of the box. Yeah. Exciting times ahead. Jim O'Shaughnessy: Yeah.
Jim O'Shaughnessy: Okay, so let me ask you the final question that I always ask all of my guests, and that is, David, you are obviously an incredibly brilliant guy.
I'm going to make you the emperor of the world for one day.
You can't kill anyone, you can't incarcerate anyone.
But what you can do is I'm going to hand you a magic microphone and you can say two things into it, and you are going to incept all 8 billion people and all the language models and generative models out there too. Why not?
We're going to incept all the models as well.
Jim O'Shaughnessy: And you can say two things that all of these ...
The humans will wake up, the models will turn on, and they'll think, "I just had the most brilliant thought and I'm going to pursue it."
What two things are you going to whisper and incept in humanity? David Ha: Oh yeah.
I mean, that's a great question.
I mean, there's so many things we can do, but one of the beliefs I have is we need more communication in our world.
Understand people with different perspectives, just in general, just talking to people is a good thing.
So, one thing would be, "Okay, find someone in your neighborhood that you've never talked to before and talk to them for an hour to understand their life." Right?
Jim O'Shaughnessy: That's fantastic. I love that one. David Ha: Yeah.
The other thing is if people have kids, spend more time with your kids.
If you don't have kids, go to your neighborhood school to spend time with other people's kids.
And I think that's going to just make the world better. These two things.
Jim O'Shaughnessy: Totally agree.
And I have four grandchildren ranging in age from nine to newborn, and I was musing while watching, we were all together, and I was watching the older kids and even my two-and-a-half year old granddaughter, we should maybe think about modeling how children go about learning and bring it into our conception of how we might make our training these models better, because- David Ha: Yeah. Definitely.
Yeah, I mean, one of the use cases is we can use these models to not educating kids but understanding ourselves.
Jim O'Shaughnessy: Exactly. Exactly.
And that is, I think, both of our aims here, better communication, better collaboration, help us understand ourselves better.
That to me, is a north star that's worth aiming at, and I'm delighted- David Ha: Exactly. Jim O'Shaughnessy: ...
to be your colleague and aiming at it together. David Ha: Yeah. Likewise, Jim. Always a pleasure.
Jim O'Shaughnessy: Great pleasure and thanks for coming on.