Vladimir Vapnik: Predicates, Invariants, and the Essence of Intelligence | Lex Fridman Podcast #71

0:00

- The following is a conversation with Vladimir Vapnik, part two, the second time we spoke on the podcast.

0:07

He's the co inventor of support vector machines, support vector clustering, VC theory and many foundational ideas in statistical learning.

0:15

He was born in the Soviet Union, worked at the Institute of Control Sciences in Moscow, then in the U. S.

0:20

, worked at ATT&T, NEC Labs, Facebook AI Research, and now is a professor at Columbia University.

0:29

His work has been cited over 200,000 times.

0:33

The first time we spoke on the podcast was just over a year ago, one of the early episodes.

0:39

This time we spoke after a lecture he gave titled: "Complete Statistical Theory of Learning," as part of the MIT series of lectures on Deep Learning and AI that I organized.

0:49

I'll release the video of the lecture in the next few days.

0:53

This podcast and the lecture are independent from each other so you don't need one to understand the other.

0:59

The lecture is quite technical and math heavy.

1:03

So if you do watch both, I recommend listening to this podcast first, since the podcast is probably a bit more accessible.

1:11

This is The Artificial Intelligence Podcast.

1:14

If you enjoy it, subscribe on YouTube, give it five starts on Apple PodCast, support it on Patreon, or simply connect with me on Twitter @LexFridman, spelled: F-R-I-D-M-A-N.

1:24

As usual, I'll do one or two minutes of ads now, and never any ads in the middle that can break the flow of the conversation.

1:31

I hope that works for you and doesn't hurt the listening experience.

1:36

This show is presented by Cash App, the number one finance app on the App Store.

1:40

When you get it, use code: LexPodcast.

1:43

Cash App lets you send money to friends by BitCoin and invest in the stock market with as little as $1.

1:49

Broker services are provided by Cash App Investing, a subsidiary of Square, and member S. I. P. C..

1:56

Since Cash App allows you to send and receive money digitally peer to peer, and security in all digital transaction is very important, let me mention that PCI data security standard.

2:06

PCI DSS Level One, that Cash App is complaint with.

2:12

I'm a big fan of standards for safety and security and PCI DSS is a good example of that.

2:18

Where a bunch of competitors got together and agreed that there needs to be a global standard around the security of transactions.

2:25

Now we just need to do the same for autonomous vehicles and A. I. systems in general.

2:31

So again, if you get Cash App from the App Store or Google Play, and use the code: LexPodcast, you get $10, and Cash App will also donate $10 to FIRST, one of my favorite organizations that is helping to advance robotics and STEM education for young people around the world.

2:49

And now, here's my conversation with Vladimir Vapnik.

2:55

You and I talked about Alan Turing yesterday, a little bit. - Yes.

2:59

- And that he, as the father of artificial intelligence may have instilled in our field an ethic of engineering in that science.

3:07

Seeking more to build intelligence rather than to understand it.

3:11

What do you think is the difference between these two paths of engineering intelligence and the science of intelligence?

3:21

- It's a completely different story.

3:23

Engineering is imitation of human activity.

3:28

You have to make a device which behaves as a human behaves.

3:35

You have all the functions of human.

3:39

It does not matter how you do it.

3:42

But to understand what is intelligence about, is quite different problem.

3:48

So I think, I believe, that it's somehow related to predicated talk yesterday.

3:57

Because, look at Vladimir Propp's idea.

4:05

He just found such a one here, predicates. He called it units.

4:16

Which can explain human behavior, at least in Russian tales.

4:22

You look at the Russian tales and derive from that.

4:24

And then people realize that they're more violent in Russian tales.

4:29

It is in TV, in movie serials and so on and so on.

4:33

- So you're talking about Vladimir Propp, who in 1928 published a book, "Morphology of the Folk Tale." - Exactly.

4:42

- Describing 31 predicates that have this kind of sequential structure that a lot of the stories' narratives follow in Russian folklore and in other content.

4:55

We'll talk about it; I'd like to talk about predicates in a focused way, but let me, if you'll allow me, to stay zoomed out on our friend Allen Touring.

5:04

And you know, he inspired a generation with the imitation game. - Yes.

5:11

- Do you think, if we can linger on that a little bit longer do you think we can learn?

5:18

Do you think learning to imitate intelligence can get us closer to understanding intelligence?

5:26

Why do you think imitation is so far from understanding?

5:32

- I think that it is different between you have different goals.

5:39

Your goal is to create something, something useful.

5:44

And that is great, and you can see how much things was done and I believe that it will be done even more.

5:53

Self-driving cars and all sorts of this business.

5:56

It is great, and it was inspired by Turing's vision.

6:03

But understanding is very difficult.

6:05

It's more or less a philosophical category.

6:08

What means understandable?

6:11

I believe in things which start from Plato.

6:16

That there exists world of ideas.

6:19

I believe that intelligence, it is world of ideas.

6:22

But it is world of pure ideas.

6:26

And when you combine that with reality things, it creates as in my case, in the variants, which is very specific.

6:38

And that I believe, the combination of ideas and a way to constructing the variant is intelligence, but first of all a predicate.

6:53

If you know predicate, and hope for this is not too much predicate exists.

7:01

For example, sort of unpredicted for human behavior is not a lot.

7:04

- Vladimir Propp used 31 (sighs) you could even call 'em predicates, 31 predicates to describe stories, narratives.

7:17

Do you think human behavior, how much of human behavior, how much of our world, our universe, all the things that matter in our existence can be summarized in predicates of the kind that Propp was working with?

7:32

- I think that we have a lot of forms of behavior.

7:37

But I think the predicate is much less.

7:41

Because even in these examples which I gave you yesterday, you saw that predicate can be, one, predicate can construct many different invariants, depending on your data.

7:59

They're applying to different data and they give different invariants.

8:04

But pure ideas, maybe not so much. - Not so many.

8:09

- I don't know about that.

8:11

But my guess, I hope, that's my challenge about digit recognition, how much you need.

8:19

- I think we'll talk about computer vision and 2-D images a little bit in your challenge.

8:24

- [Vladimir] That's exact about intelligence.

8:28

- That's exactly about, no, that hopes to be exactly about the spirit of intelligence in the simplest possible way. - Yeah, absolutely.

8:38

You should start the simplest way, otherwise you will not be able to do it.

8:42

- There's an open question whether starting at the MNIST digit recognition is a step towards intelligence or it's an entirely different thing.

8:51

- I think that to build records using say, 100, 200 times less examples, you need intelligence. - You need intelligence.

9:01

So let's, because you use this term, and it would be nice, I'd like to ask simple, maybe even dumb questions.

9:10

Let's start with a predicate.

9:12

In terms of terms and how you think about it, what is a predicate? - I don't know.

9:19

(laughs) I have a feeling, formally they exist.

9:25

I believe that predicate for 2-D images. One of them is symmetry. - Hold on a second.

9:33

Sorry, sorry to interrupt and pull you back.

9:36

At the simplest level, we're not being profound currently.

9:40

A predicate is a statement of something that is true. - [Vladimir] Yes.

9:46

- Do you think of predicates as somehow probabilistic in nature or is this binary?

9:54

This is truly constraints of logical statements about the world.

9:59

- In my definition, the simplest predicate is function.

10:04

Function and you can use this function to make inner product, that is predicate.

10:10

- What's the input and what's the output of the function?

10:14

- Input is X, something which is input in reality.

10:18

Say, if you considered digit recognition, in pixel space.

10:25

But it is function which in pixel space.

10:29

But it can be any function, pixel space.

10:35

And you choose, and I believe that there are several functions, which is important for understanding of images.

10:46

One of them is symmetry, it's not so simple construction, as I described is the derivative, it's all this stuff.

10:55

But another I believe, I don't know how many, is how well-structurized is picture. - Structurized? - Yeah.

11:04

- What do you mean by structurized?

11:06

- It is formal definition, say, something heavy on the left corner, not so heavy in the middle and so on.

11:17

You describe in general, concept of what you assume.

11:21

- Concepts, some kind of universal concepts. - Yeah.

11:26

But I don't know how to formalize this.

11:30

- This is the thing, there's a million ways we can talk about this, I'll keep bringing it up.

11:35

We humans have such concepts, when we look at digits.

11:41

But it's hard to put them, just like you're saying now, it's hard to put them into words.

11:45

- You know, that is example.

11:48

When critics in music, trying to describe music, they use predicate.

11:58

And not too many predicate but in different combination.

12:02

But they have some special words for describing music.

12:08

And the same should be for images.

12:12

But maybe there are critics who understand essence of what this image is about.

12:21

- Do you think there exists critics who can summarize the essence of images, human beings?

12:30

- [Vladimir] I hope that, yes.

12:32

- Explicitly state them on paper?

12:37

The fundamental question I'm asking is do you (chuckles) do you think there exists a small set of predicates that will summarize images?

12:48

It feels to our mind like it does that the concept of what makes a two and a three and a four-- - No, no, no.

12:56

It's not only on this level.

13:01

It should not describe two, three, four.

13:04

It describes some construction which allows you to create an invariants.

13:11

- And invariants, sorry to stick on this, but terminology.

13:14

- Invariants, it is, it is projection of your image.

13:24

Say, I can say, looking at my image, it is more or less symmetric and I can give you a variable of symmetry.

13:35

Say, level of symmetry using this function which I gave yesterday.

13:43

Then you can describe that your image has these characteristics exactly in the way how musical critics describe music.

13:59

But this is invariant applied to specific data, to specific music, to something.

14:07

I strongly believe in this Plato idea, that exists world of predicate and world of reality and predicate and reality are somehow connected and you have to figure out that.

14:22

- Let's talk about Plato a little bit.

14:24

So you draw a line from Plato to Hegel to Wigner to today. - Yes. - So Plato has forms.

14:33

The theory of forms; there's a world of ideas and a world of things, as you talk about.

14:39

And there's a connection, and presumably the world of ideas is very small, and the world of things is arbitrarily big, but they're all, what Plato calls them like, it's a shadow, the real world is a shadow from the world of form.

14:54

- Yeah and you have projection - Projection. - Of a world of idea. - Yeah, very poetic.

14:59

- And in reality you can, realize this projection using these invariants because it is a projection for only specific examples which create specific features of specific objects.

15:18

- So the essence of intelligence is, while only being able to observe the world of things, try to come up with a world of ideas.

15:27

- Exactly, like in this music story.

15:30

Intelligent musical critics knows this world and have a feeling about them.

15:33

- I feel like that's a contradiction; intelligent music critics.

15:39

I think music is to be...

15:44

enjoyed in all its forms.

15:47

The notion of critic, like a food critic.

15:50

- [Vladimir] No, I don't want attach emotion.

15:52

- That's an interesting question.

15:53

Does emotion, there's certain elements of the human psychology, of the human experience which seem to almost contradict intelligence and reason. Like emotion, like fear.

16:07

Like love, all of those things.

16:11

Are those not connected in any way to the space of ideas? - This I don't know.

16:20

I just want to be concentrating on very simple story. On digit recognition.

16:27

- So you don't think you have to love and fear death in order to recognize digits?

16:32

- I don't know because it's so complicated.

16:36

It involves a lot of stuff which I never considered.

16:41

But I know about digit recognition.

16:45

And I know that for digit recognition to get the records from small numbers of observations, you need predicate but not special predicate for this problem.

17:03

But universal predicate which understand world of images. - Of visual information. - Visual, yeah.

17:11

But on the first step, they understand say, a world of 100 digits or characters or something simple.

17:21

- [Lex] So like you said, symmetry is an interesting one.

17:23

- No, that's what I think one of the predicate is related to symmetry, the level of symmetry.

17:31

- Okay, degree of symmetry. - Yeah.

17:32

- So you think symmetry at the bottom is a universal notion and there's, there's degrees of a single kind of symmetry?

17:41

Or is there many kinds of symmetries?

17:43

- Many kinds of symmetries.

17:46

There is a symmetry/anti-symmetry say, letter S.

17:52

So it has vertical anti-symmetry.

17:58

It could be diagonal symmetry, vertical symmetry.

18:02

- So when you cut vertically the letter S-- - Yeah, and then the upper part and lower part are in different directions.

18:16

- Inverted along the Y-axis.

18:18

But that's just like, one example of symmetry right?

18:21

Isn't there like-- - Right, but there is a degree of symmetry.

18:26

If you play all this lineated stuff to do tangent distance, whatever I described, you can have a degree of symmetry.

18:40

And that is what is describing reason of image.

18:45

It is the same as you will describe this image saying about digit S has anti-symmetry, digit three is symmetric more/less. Look for symmetry.

19:04

- Do you think such concepts like symmetry predicates, like symmetry, is it a hierarchical set of concepts?

19:14

Or are these independent, distinct predicates that we want to discover a subset of?

19:23

- There is a degree of symmetry.

19:26

And this idea of symmetry made very general, like degree of symmetry, the degree of symmetry can be zero, no symmetry at all.

19:40

Or degree of symmetry of let's say, more or less symmetrical.

19:47

But you have one of these descriptions, and symmetry can be different.

19:52

As I told, horizontal, vertical, diagonal.

19:57

Anti-symmetry also is a concept of symmetry.

20:01

- What about shape in general?

20:03

I mean, symmetry is a fascinating notion, but it-- - No, no, I'm talking about digits, I would like to concentrate on all I would like to know predicate for digit recognition.

20:14

- Yes, but symmetry is not enough for digit recognition, right?

20:18

- It is not necessarily for digit recognition; it helps to create invariant, which you can use when you will have examples for digit recognition.

20:35

You have regular problem of digit recognition.

20:38

You have examples of the first class and second class.

20:41

Plus you know that there exists this concept of symmetry.

20:45

And you apply when you're looking for decision rule, you will apply concept of symmetry, of this level of symmetry which you estimate from.

21:01

Everything is, comes from weak convergence. - What is convergence?

21:07

What is weak convergence?

21:09

What is strong convergence?

21:11

I'm sorry I'm gonna do this to you.

21:13

What are we converging from and to?

21:16

- You're converging, you would like to have a function.

21:20

The function which say, indicate a function which indicate your digit five, for example.

21:28

- A classification/task-- - Let's talk only about classification task.

21:33

- So classification means you will say whether this is a five or not, or say which of the 10 digits it is.

21:40

- Right, right, I would like to have these functions.

21:46

Then, I have some examples.

21:56

I can consider property of these examples.

22:01

Say, symmetry, and I can measure level of symmetry for every digit.

22:08

And then I can take average from my training data.

22:16

And I will consider only functions of conditional probability, which I'm looking for in my decision level.

22:27

Which applying to the digits will give me the same average as I observed on training data.

22:41

So actually this is different level of description of what you want.

22:48

You want, not just, you show not one digit.

22:54

You show this predicate, show general property of all digits which you have in mind.

23:03

If you have in mind digit three, it gives you property of digit three and you select as admissible set of functions, only function which keeps this property.

23:16

You will not consider the other functions.

23:20

So you're immediately looking for smaller subset of functions.

23:24

- That's what you mean by an admissible function.

23:25

- And admissible function. Exactly.

23:28

- Which is still a pretty large, for the number three, that's a large-- - It's large but, if you have one predicate.

23:36

But according to, there is a strong and weak convergence.

23:42

Strong convergence if convergence in function.

23:46

You're looking for the function, on one function, and you're looking on another function.

23:52

And square difference from them should be small.

23:59

If you take difference at any points, make a square, make an integral, and it should be small.

24:05

That is convergence in function.

24:08

Suppose you had some function, any function.

24:11

So I would say, I say that some function converged to this function.

24:17

If integral from square difference between them is small-- - That's the definition of strong convergence.

24:24

- That's the definition of-- - Two functions, the integral of the difference is small.

24:28

- Yeah, it is convergence in functions. - Yeah.

24:31

- But you have different convergence in functions.

24:35

You take any function, you take some function Fe and take inner product this function, this F function.

24:46

F zero function which you want to find, and that gives you some value.

24:52

So you say that a set of functions converge in inner product to this function if this value of inner product converge to value F zero. That is for one Fe.

25:12

But three converges requires that it converges for any function of field of space.

25:20

If it converge for any function of field of space, then you will say that this is a weak convergence.

25:28

You can think that when you take integral, that is property, integral property of function.

25:35

For example, if you will take sine or cosine, it is coefficient of say, free expansion.

25:45

If it converged for all coefficients or free expansion, so under some condition it converged to function you're looking for.

25:58

But weak convergence means any property.

26:02

Not convergence, not point-wise.

26:05

But integral property of function.

26:09

So weak convergence means integral property of functions.

26:13

When I'm talking about predicate, I would like to formulate which integral properties I would like to have for convergence.

26:28

And if I will take one predicate, predicate its function which I measure property.

26:37

If I will use one predicate and say, I will consider only function which give me the same value as this predicate, I select a set of functions from functions which is admissible, in the sense that function which I'm looking for in this set of functions.

27:01

Because I'm checking in training data, it gives the same.

27:08

- [Lex] Yeah, so it always has to be connected to the training data in terms of-- - Yeah, but property, you can know independent of training data.

27:18

And this guy, Propp, says that there is formal property.

27:24

31 property and if you-- - A fairy tale, the Russian fairy tale. - Right.

27:27

But the Russian fairy tale is not so interesting.

27:30

It's more interesting that people applied this to movies, to theater, to different things and the same works. They're universal.

27:42

- So I would argue that there's a little bit of a difference between the kinds of things that would apply to, which are essentially stories and digit recognition.

27:54

- [Vladimir] It is the same story.

27:55

- You're saying digits, there's a story within the digit? - Yeah.

27:59

(laughing) So but my point is, well, I hope that it's possible to beat a record.

28:08

Using not 60,000 but say, 100 times less.

28:13

Because instead you will give predicates.

28:17

And you will select your decision not from right set of functions.

28:24

But from set of function which gives this predicate.

28:28

But predicate is not related just to digit recognition.

28:32

- Right, so-- - Like in Plato's case.

28:36

- (laughs) Do you think it's possible to automatically discover the predicates?

28:42

So, you basically said that the essence of intelligence is the discovery of good predicates. - [Vladimir] Yeah.

28:51

- Now the natural question is, you know, that's what Einstein was good at doing in physics.

28:59

Can we make machines do these kinds of discovery of good predicates?

29:04

Or is this ultimately a human endeavor? - That I don't know.

29:09

I don't think that much it can do.

29:11

Because, according to theory about weak convergence, any function from Hilbert Space can be predicate.

29:23

So you have infinite number of predicate in upper.

29:27

And before you don't know which predicate is good in which.

29:32

But whatever Propp showed, and why people call it breakthrough, that there is not too many predicate which cover most of situation that happen in the world.

29:51

- So there's a sea of predicates.

29:54

Only a small amount are useful for the kinds of things that happen in the world?

30:01

- I think that, I would say only a small part of predicates very useful. Useful, all of them.

30:08

- Only a very few are what we should, let's call them good predicates. - Very good predicates. - Very good predicates.

30:19

Can we linger on it, what's your intuition?

30:21

Why is it hard for a machine to discover good predicates?

30:27

- Even in my talk described how to do a predicate.

30:30

How to find new predicates, I'm not sure that it is very good predicate.

30:35

- [Lex] What did you propose in your talk?

30:37

- In my talk I gave example for diabetes. - Diabetes, yeah.

30:42

- When we achieve some percent so then we're looking from area, where some sort of predicate which I formulated, does not... keep invariant.

31:03

So if it doesn't keep, I retrain my data.

31:07

I select only functions which keep this invariant.

31:10

In the way I did it, I improved my performance.

31:14

I came looking for this predicate.

31:16

I know technically how to do that.

31:19

You can of course do it using a machine, but I'm not sure that we will construct the smartest predicate.

31:30

- Well this is the, allow me to linger on it.

31:34

Because that's the essence, that's the challenge.

31:36

That is artificial, that's the human-level intelligence that we seek, is the discovery of these good predicates.

31:43

You've talked about deep learning as a way to, the predicates they use and the functions are mediocre.

31:53

Or you can find better ones.

31:55

- Let's talk about deep learning. - Sure, let's do it.

31:57

- I know only Jans LaComb, convolutional network, and what else?

32:05

And it's very simple convergence.

32:07

- [Lex] There's not much else to know.

32:09

- To fix it left and right. - Yes.

32:10

- I can do it like that with one predicate.

32:14

- [Lex] Convolution is a single predicate.

32:16

- It's single, it's single predicate.

32:19

- Yes, but-- - You know exactly, you take the derivative for translational, and predicate should be kept.

32:31

- So that's a single predicate, but humans discovered that one or at least-- - Not that that is the least not too many predicates.

32:39

And that is big story because Jan did it 25 years ago, and nothing so clear was uttered to deep network.

32:50

And then I don't understand why we should talk about deep network instead of talking about piece-wise linear functions which keeps this predicate.

33:03

- You know, a counter argument is, that maybe the amount of predicates necessary to solve general intelligence, say in the space of images, doing efficient recognition of handwritten digits is very small.

33:23

So we shouldn't be so obsessed about finding, we'll find other good predicates like convolution, for example.

33:30

There has been other advancements like, if you look at the work with attention, there's attentional mechanisms, especially used in natural language focusing the network's ability to, to learn at which part of the input to look at.

33:47

The thing is, there's other things besides predicates that are important for the actual engineering mechanism of showing how much you can really do, given such these predicates.

34:00

I mean, that's essentially the work of deep learning is constructing architectures that are able to be given the training data to be able to converge towards a function that can approximate, that can generalize well.

34:22

It's an engineering problem. - Yeah, I understand.

34:26

But let's talk not on an emotional level, but on a mathematical level.

34:31

You have set of piece-wise linear functions, it is all possible neural networks.

34:42

It's just piece-wise linear functions. It's many, many pieces.

34:44

- Large, large number of piece-wise linear function. - Exactly. - Very large. - Very large.

34:50

- [Lex] Almost, feels like too large.

34:51

- It's still simpler than say, convolution, which is reproducing Hilbert's Space and we have a Hilbert's set of functions. - What's Hilbert Space?

35:03

- It's space with infinite number of coordinates.

35:09

A function for expansion, something like that. So it's much richer.

35:15

And when I'm talking about closed-form solution, I'm talking about this set of functions.

35:20

Not piece-wise linear set, which is particular case of (chuckles) it is small part of it.

35:30

- So neural networks is a small part of the space you're, of functions you're talking about?

35:36

- Small set of functions. - Yeah. - Let me take it that.

35:40

But it is fine, it is fine.

35:42

I don't want to discuss the small or big we take and what.

35:47

So you have some set of functions.

35:51

So now when you're trying to create architecture, you would like to create admissible set of functions, which all your tricks, to use not all functions, but some subset of this set of functions.

36:07

Say when you're introducing convolutional network.

36:10

It is a way to make this subset useful for you.

36:16

But from my point of view, convolutional, it is something, you want to keep some invariance.

36:24

Say, translation invariance.

36:27

But now if you understand this, and you cannot explain on the level of ideas what neural network does, you should agree that it is much better to have a set of functions.

36:46

As I say, this set of functions should be admissible, it must keep this invariant, this invariant, and that invariant.

36:55

You know that as soon as you incorporate new invariants, set of functions becomes smaller and smaller and smaller.

37:02

- [Lex] But all the invariants are specified by you the human.

37:06

- Yeah, but what I am hope, that there is a standard predicate like, Propp showed.

37:15

That's what I want to find for digit recognition.

37:19

If we start, it is completely new area of what is intelligence about on the level, starting from Plato's idea.

37:28

What is the world of ideas?

37:32

And I believe there is not too many.

37:36

But you know, it is amusing that mathematicians are doing something with neural network in general function.

37:44

But people from literature, from art, they use this all the time. - That's right. - Invariants.

37:52

Say, it is great how people describe music.

37:57

We should learn from that. Something on this level.

38:02

So why, Vladimir Propp who was just theoretical.

38:09

Who studied theoretical literature, he found that.

38:13

- You know what, let me throw that right back at you.

38:15

Because there's a little bit of a, that's less mathematical and more emotional, philosophical Vladimir Propp.

38:22

I mean, he wasn't doing math. - No.

38:26

- And you just said another emotional statement, which is you believe that this Plato world of ideas is small. - I hope. - I hope.

38:38

Do (chuckles), do you, what's your intuition, though, if we can linger on it? That bothers me.

38:44

- You know, because not just small or big. I know exactly.

38:51

That when I'm introducing some predicate, I decrease set of functions.

38:59

But my goal to decrease set of functions much.

39:04

- By as much as possible.

39:04

- By as much as possible.

39:07

Good predicate which does this, then I should choose next predicate which does this, which decreases set as much as possible.

39:17

So, set of good predicates, it is such that the decrease, amount of admissible functions-- - So if each good predicate significantly reduces the set of admissible functions that there naturally should not be that many predicates.

39:35

- No, but, if you reduce very well the VC dimension of the function of admissible set of functions, it's small.

39:46

And you need not too much training data to do well.

39:53

- [Lex] And VC dimension, by the way, is some measure of capacity of this set of functions.

39:57

- Right, roughly speaking, how many function in this set.

40:02

So you're decreasing, decreasing, and it makes it easier for you to find the function you're looking for.

40:10

But the most important part, to create a good admissible set of functions.

40:15

And it probably is that there are many ways but, a good predicate is such that it can do that.

40:26

For this duck, you should know a little bit about duck.

40:31

- What are the three fundamental laws of ducks?

40:35

- Looks like a duck, swims like a duck, and quacks like a duck. - And quacks.

40:38

You should know something about ducks to be able to-- - Not necessarily. Looks like say, horse. It's also good.

40:46

- [Lex] So it's not (chuckles), it generalize from ducks.

40:47

- Yes, and talk like it, and make sound like horse or something.

40:54

And run like horse and moves like horse.

40:57

It is general, it is general predicate that this applies to duck.

41:04

But for duck you can say, play chess like duck.

41:09

- [Lex] You cannot say, play chess like a duck. - Why not?

41:11

- So you're saying you can but that that would not be a good-- - [Vladimir] No, you will not reduce that function.

41:18

- Yeah, you would not reduce the set of functions.

41:21

- So you can, the story is formal story and it's a magical story is that you can use any function you want as a predicate.

41:31

But some of them are good, some of them are not.

41:33

Because some of them reduce a lot of functions.

41:37

The admissible set, some of them (mumbles) - So the question is, and I'll probably keep asking this question, but how do we find such predicates? What's your intuition?

41:47

Handwritten recognition, how do we find the answer to your challenge?

41:52

- Yeah, I understand it like that. I understand what. - What defined?

41:57

- What that means, a new predicate. - Yeah.

42:02

- Like, guy who understands music can say these words he describes when he listens to music.

42:09

He understands music, he'll use not too many different.

42:13

Or you can do it like Propp.

42:15

You can make collection, what you're talking about music. About this, about that.

42:20

It's not too many different situations you describe.

42:25

- Because we mentioned Vladimir Propp a bunch, let me just mention, so there's a sequence of 31 structural notions that are common in stories.

42:36

And I think-- - They're called units.

42:38

- Units, and I think they resonate.

42:40

It starts, just to give an example, Absention: a member of the hero's community or family leaves the security of the home environment, then goes through the interdiction, a forbidding edict or command that's passed upon the hero.

42:54

Don't go there; don't do this.

42:56

The hero's warned against some action.

42:58

Then step three: Violation of interdiction.

43:03

You know, break the rules, break out on your own.

43:07

Then reconnaissance, the villain makes an effort to attain knowledge needed to fulfill their plot, so on. It goes on like this.

43:16

Ends in a wedding, number 31, happily every after.

43:21

- He just gave description of all situations.

43:25

He understands his world-- - Of folk tales.

43:28

- Yeah, not folk-- - Stories.

43:31

- Stories, and these stories not just in folk tales.

43:36

These stories in detective serials as well.

43:40

- [Lex] And probably in our lives.

43:42

We probably live-- - Read this.

43:46

They're all, this predicate is good for different situations. For movie, for theater.

43:55

- By the way, there's also criticism, right?

44:00

There's another way to interpret narratives. From... Claude Levi Strauss.

44:07

- I am not in this business.

44:12

- I know, that's theoretical literature, but it's looking at paradigms behind them.

44:15

- [Vladimir] It's always, this discussion-- - Philosophers argue. - Yeah. - Yeah.

44:20

- But at least there is units.

44:23

It's not too many units that can describe.

44:27

But describe probably gives the other units.

44:30

Or another way for description.

44:31

- Exactly, another set of units.

44:33

- [Vladimir] Another set of predicates, yes.

44:36

It doesn't matter how, but they exist probably.

44:41

- My question is, whether given those units, whether without our human brains to interpret these units, they would still hold as much power as they have.

44:53

Meaning, are those units enough when we give them to the alien species?

44:58

- Let me ask you, do you understand digit images?

45:05

- No, I don't understand. - No, no, no.

45:07

When you can recognize these digit images, it means that you understand. - Yes, I understand.

45:13

- You understand characters, you understand.

45:15

- Nope, nope, nope, nope.

45:22

It's the imitation versus understanding question.

45:25

Because I don't understand the mechanism by which I understand-- - No, no.

45:29

I'm not talking about, I'm talking about predicates.

45:32

You understand that it involves symmetry, maybe structure, maybe something else.

45:37

I cannot formulate, I just was able to find symmetries, or degree of symmetries.

45:43

- So this is a good line.

45:46

I feel like I understand the basic elements of what makes a good hand recognition system my own.

45:54

Like, symmetry connects with me.

45:56

It seems like that's a very powerful predicate.

45:59

My question is, is there a lot more going on that we're not able to introspect?

46:05

Maybe I need to be able to understand a huge amount in the world of ideas.

46:14

Thousands of predicates, millions of predicates in order to do hand recognition.

46:20

- [Vladimir] I don't think so.

46:23

- Both your hope and your intuition are such that very few-- - No, no, let me explain. You're using digits.

46:31

You're using examples as well.

46:34

Theory says that if you will use all possible functions from Hilbert Space, all possible predicates, you don't need training data.

46:49

You just will have admissible set of functions which contain one function. - Yes.

46:57

So the trade off is, when you're not using all predicates, you're only using a few good predicates, that you need to have some training data. - Yes, exactly.

47:05

- The more good predicates you have, the less training data you need. - Exactly.

47:11

That is intelligent learning.

47:14

- Okay, I'm gonna keep asking the same dumb question, handwritten recognition, to solve the challenge, you kind of propose a challenge that says we should be able to get state of the art MNIST error rates by using very few, 60 maybe few examples per digit.

47:32

What kind of predicates do you think you'll-- - [Vladimir] That is the challenge.

47:37

(laughs) So people who will solve this problem. - They will answer. - They will answer it.

47:41

- Do you think they'll be able to answer it in a human explainable way?

47:48

- They just need the right function, that's it.

47:50

- But so, can that function be written I guess by an automated reasoning system?

47:58

Whether we're talking about a neural network learning a particular function, or another mechanism.

48:05

- No, I'm not against neural network.

48:08

I am against admissible set of function which creates neural network. You did it by hand.

48:16

You don't do it by invariants, by predicate, by reason.

48:24

- But neural networks can then reverse, to the reverse step of helping you find a function.

48:30

The task of a neural network is to find disentangled representation, for example is what they call, is to fine that one predicate function that's really captures some kind of essence.

48:45

Not the entire essence, but one very useful essence of this particular visual space.

48:52

Do you think that's possible?

48:55

Listen, I'm grasping, hoping there's an automated way to find good predicates, right?

49:00

So the question is, what are the mechanisms of finding good predicates, ideas, that you think we should pursue?

49:08

A young grad student listening right now. - I gave example.

49:13

So find situation where predicate, which you're suggesting, don't create invariant.

49:27

It's like in physics, find situation where existing theory cannot explain it.

49:36

- [Lex] Find a situation where the existing theory can't explain it. - Cannot explain this.

49:40

- [Lex] So you're finding contradictions.

49:42

- Find contradiction, and then remove this contradiction.

49:46

But in my case, what means contradiction, you find function which, if you will use this function, you're not keeping invariants.

49:57

- [Lex] So really the process of discovering contradictions. - Yeah. It is like in physics.

50:05

Find situation where you have contradiction for one of the property, for one of the predicate, then include this predicate making invariants.

50:18

And solve, again, this problem, now you don't have contradiction.

50:23

But it is not the best way probably, I don't know, to looking for predicate. - That's just one way. Okay. - That, no, no. It is brute force way. - The brute force way.

50:37

What about the ideas of what, big umbrella term of symbolic AI?

50:45

In the '80s with Xper Systems, sort of logic, reasoning-based systems.

50:53

Is there hope there to find some, through sort of, deductive reasoning, to find good predicates?

51:05

- [Vladimir] I don't think so.

51:09

I think that just logic is not enough.

51:12

- Kind of a compelling notion, though.

51:14

That when smart people sit in a room and reason through things, it seems compelling.

51:20

And making our machines do the same is also compelling.

51:26

- Everything is very simple when you have infinite number of predicate.

51:34

You can choose the function you want.

51:38

You have invariants and you choose the function you want.

51:43

But you have to have not too many invariants to solve the problem.

51:56

And how from infinite number function to select finite number and hopefully small finite number of functions.

52:09

Which is good enough to extract some small set of admissible functions.

52:17

So they will be admissible, it's for sure, because every function just decrease set of function and leaving it admissible. But it will be small.

52:27

- But why do you think logic-based systems can't help?

52:34

Intuition, not-- - Because you should know reality, you should know life.

52:39

This guy like Propp, he knows something and he tried to put in invariants his understanding.

52:48

- So but that's the human, yeah, yeah.

52:50

But see, you're putting too much value into Vladimir Propp's knowing something.

52:57

- No, it is-- - Am I being misunderstanding?

53:01

- What means you know life? What it mean? - You know common sense.

53:07

- No, no, you know something.

53:10

Common sense, it is some rules. - You think so?

53:14

Common sense is simply rules?

53:17

Common sense is every, it's mortality, it's fear of death, it's love, it's spirituality, it's happiness and sadness.

53:30

All of it is tied up into understanding gravity which is what we think of as common sense.

53:37

- I don't credit or discuss of that.

53:39

I want to discuss understand digit recognition.

53:45

- Any time I bring up love and death you bring it back to digit recognition. I like it.

53:50

(laughs) - No, you know, it is doable because there is a challenge. - Yeah.

53:55

- Which I still have to solve it; if I will have a student concentrate on this work, I will suggest something or so.

54:04

- You mean handwritten recognition?

54:06

Yeah, it's a beautifully simple, elegant, and yet-- - I think that I know invariants which will solve this. - You do?

54:13

- I think that, I think that. But it is not universal.

54:21

I want some universal invariants which are good not only for digit recognition, for image understanding.

54:30

- So let me ask, how hard do you think is 2-D image understanding?

54:38

If we can kind of intuit handwritten recognition, how big of a step, leap, journey is it from that?

54:48

If I gave you good, if I solved your challenge for handwritten recognition, how long would my journey then be from that to understanding more general, natural images? - Immediately.

54:59

You will understand it as soon as you will make a record. - You think so?

55:05

- Because it is not for free.

55:07

As soon as you will create several invariants which will help you to get the same performance that the best neural net did using more than 100 times less examples.

55:27

You have to have something smart to do that. - And you're saying?

55:31

- That's not an invariant.

55:33

It is predicate because you should put some idea how to do that. - But okay.

55:41

Let me just pause, maybe it's a trivial point, maybe not, but handwritten recognition feels like a 2-D, two-dimensional problem.

55:51

And it seems, like how much complicated is the fact that most images are a projection of a three-dimensional world onto a 2-D plane?

56:03

It feels like for a three-dimensional world, we need to start understanding common sense in order to understand an image.

56:12

It's no longer visual shape and symmetry.

56:17

It's having to start to understand concepts of, understand life. - Yeah.

56:24

You're talking that there are different invariant, different predicate, yeah.

56:28

- [Lex] And potentially much larger number.

56:32

- You know, maybe, but let's start from simple.

56:36

- [Lex] But you said that it would be immediate.

56:38

- No, you know, I cannot think about things which I don't understand yet. This I understand.

56:44

But I'm sure that I don't understand everything there.

56:48

- [Lex] Yeah, that's the difference-- - It's like they say, do as simple as possible, but not simpler, and that is exact case.

56:56

- With handwritten rec-- - With handwritten.

56:58

- Yeah, but that's the difference between you and I.

57:02

(laughs) I welcome and enjoy thinking about things that I completely don't understand.

57:10

Because to me it's a natural extension, without having solved handwritten recognition to wonder how, how difficult is the next step of understanding 2-D and 3-D images, because ultimately, while the science of intelligence is fascinating, it's also fascinating to see how that maps to the engineering of intelligence.

57:34

And recognizing handwritten digits is not, doesn't help you, it might, it may not, help you with the problem of general intelligence.

57:46

We don't know; it'll help you a little bit. We don't know how much. - It is unclear. - It's unclear. - Yeah. - It might very much.

57:50

- But I would like to make a remark: I start not from very primitive problem like, challenge problem; I start with very general problem, with Plato, so you understand.

58:09

And it comes from Plato, digit recognition.

58:13

- So you basically took Plato and the world of forms and ideas, and mapped, and projected it into the clearest, simplest formulation of that big world into handwritten recognition.

58:27

- I will say that I did not understand Plato until recently, and until I considered weak convergence and then predicate and then, oh, this is what Plato thought.

58:45

- Can you linger on that?

58:47

How do you think about this world of ideas and world of things and Plato? - It is a metaphor.

58:53

- It's a metaphor for sure.

58:55

It's a compelling, it's a poetic and a beautiful metaphor.

58:58

But what can you-- - But it is a way how you should try to understand how a duck ideas in the world.

59:07

So from my point of view, it is very clear, but it is a line all the time people looking for that.

59:18

Say, Plato then Hegel, whatever reasonable it exists, whatever exists it is reasonable.

59:26

I don't know what he have in mind, reasonable.

59:30

- [Lex] Right, these philosophers again.

59:31

- No, no, no, no, no, no, it is, it is next stop of Wilner.

59:37

That which we might understand something of reality.

59:41

He did the same Plato line.

59:43

And then it comes suddenly to Vladimir Propp.

59:48

Look, 31 ideas, 31 units, and describes everything.

59:54

- There's abstractions, ideas that represent our world, and we should always try to reach into that.

1:00:03

- Yeah, but what you should make a projection on the reality, but understanding is, it is abstract ideas.

1:00:11

You have in your mind several abstract ideas which you can apply to reality.

1:00:17

- And reality in this case so if we look at machine learning is data. - This example, data. - Data.

1:00:24

Okay, let me put this on you, because I'm an emotional creature.

1:00:28

I'm not a mathematical creature like you.

1:00:30

I find compelling the idea, forget the space, this sea of functions.

1:00:36

There's also a sea of data in the world.

1:00:39

And I find compelling that there might be, like you said, teacher, small examples of data that are most useful for discovering good, whether it's predicates or good functions, that the selection of data may be a powerful journey, a useful mechanism.

1:01:02

You know, coming up with a mechanism for selecting good data might be useful, too.

1:01:07

Do you find this idea of finding the right data set interesting at all?

1:01:14

Or do you kinda take the data set as a given?

1:01:17

- I think that it is, you know, my thing is very simple.

1:01:22

You have huge set of functions.

1:01:26

If you will apply, and you have not too many data.

1:01:32

If you pick up function which describes this data, you will do not very well. - Like randomly pick?

1:01:41

- Yeah (mumbles) It will be irritating.

1:01:46

So you should decrease set of function from which you're picking out one.

1:01:53

So you should go somehow to admissible set of functions.

1:01:59

And this, what about weak convergence?

1:02:04

From another point of view, to make admissible set of function, you need just the deal, just function which you will take an inner product.

1:02:20

Which you will measure property of your function. That is how it works.

1:02:31

- [Lex] No, I get it, I get it.

1:02:32

I understand it but do you, the reality is-- - But let's discuss, let's think about examples.

1:02:40

You have huge set of function and you have several examples.

1:02:45

If you just trying to keep, take function which satisfies these examples, you still will not have it.

1:02:56

You need decrease, you need admissible set of functions.

1:02:59

- Absolutely, but what, say you have more data than functions.

1:03:07

I mean, maybe not more data than functions, 'cause that's-- - That's impossible.

1:03:11

- Impossible, I was trying to be poetic for a second.

1:03:15

I mean, you have a huge amount of data, a huge amount of examples.

1:03:19

- But amount of function can be even bigger.

1:03:22

- Even bigger, I understand.

1:03:24

- Everything is (chuckles) - [Lex] There's always a bigger boat.

1:03:27

- Whole Hilbert Space of functions. - I got you, but okay.

1:03:31

But you don't, you don't find the world of data to be an interesting optimization space?

1:03:38

Like, the optimization should be in the space of functions.

1:03:45

- In creating admissible set of functions.

1:03:46

- [Lex] Admissible set of functions.

1:03:48

- You know, even from the classical, this is so.

1:03:54

From structured reasoning, you should organize function in the way they will be useful for you. - Right.

1:04:07

- And that is admissible step.

1:04:10

- But the way you're thinking about useful is you're given a small set of examples.

1:04:17

- Small set of functions which contain functions by looking for it.

1:04:21

- Yeah, but as looking for it based on the empirical set of small examples. - Yeah.

1:04:28

But that is another story, I don't touch it.

1:04:31

Because I believe that these small examples, it's not too small.

1:04:37

Say 65%, law of large numbers works.

1:04:41

I don't need uniform law.

1:04:43

The story is that in statistics there are two laws.

1:04:46

Law of large numbers and uniform law of large numbers.

1:04:51

So I want a situation where I use law of large numbers but not uniform law of large numbers.

1:04:58

- [Lex] Right, so 60 is law of large. It's large enough. - I hope, I hope.

1:05:02

It still needs some evaluation, some balance, et cetera.

1:05:08

What I did is the following that, if you trust that say, this average gives you something close to expectation so you can talk about that, about this predicate. - Yeah.

1:05:26

- [Vladimir] And that is basis of human intelligence.

1:05:30

- Good predicates, the discovery of good predicates is the basis of human intelligence.

1:05:34

- No, no, it is discovery of your understanding of world.

1:05:39

Of your total logic of understanding the world.

1:05:45

Because you have several functions which you will apply to reality.

1:05:51

- Can you say that again?

1:05:52

So you're-- - You have several functions, predicates, but they're abstract.

1:06:01

Then you will apply them to reality, to your data, and you will create in this weak predicate, which is useful for your task.

1:06:12

But predicate are not related specifically to your task, to this here task.

1:06:17

It is abstract functions, which being applied, applied to-- - Many tasks that you might be interested in.

1:06:24

- It might be many tasks. I don't know. Different tasks.

1:06:29

- [Lex] Well they should be many tasks. Right? - Yeah, I believe. Like in Propp case.

1:06:35

It was for fairy tales, but it's happened everywhere.

1:06:40

- Okay, we talked about images a little bit but, can we talk about Noam Chomsky for a second?

1:06:46

(laughing) - I don't know him very well. - Personally?

1:06:54

- Not personally I don't know. - His ideas. - His ideas.

1:06:58

- Well let me just say, do you think language, human language is essential to expressing ideas, as Noam Chomsky believes?

1:07:08

So like, language is at the core of our formation of predicates. Human language.

1:07:16

- In all the story of language is very complicated.

1:07:20

I don't understand this and I thought about-- - Nobody does.

1:07:24

- I'm not ready to work on that because it's so huge.

1:07:30

It is not for me, and I believe not for our century. - The 21st century. - Not for 21st century.

1:07:39

We should learn something, a lot of stuff from simple tasks like digit recognition.

1:07:45

- So you think digital recognition, 2-D image, how would you more abstractly define it, digit recognition?

1:07:56

It's 2-D image, symbol recognition essentially?

1:08:05

I'm trying to get a sense, sort of thinking about it now, having worked with MNIST forever, how small of a subset is this of the general vision recognition problem and the general intelligence problem? Is it?

1:08:24

Is it a giant subset, is it not?

1:08:27

And how far away is language?

1:08:30

- You know, let me refer to Einstein.

1:08:34

Take the simplest problem, as simple as possible, but not simpler, and this is challenge, is simple problem.

1:08:44

But it's simple by idea, but not simple to get it.

1:08:50

When you will do this, you will find some predicate which helps it.

1:08:57

- Yeah, with Einstein you can, you look at General Relativity, but that doesn't help you with quantum mechanics.

1:09:06

- And that's another story.

1:09:08

You don't have any universal instrument.

1:09:11

- Yeah, so I'm trying to wonder if which space we're in.

1:09:17

Whether handwritten recognition is like General Relativity and then language is like, quantum mechanics, are you still gonna have to do a lot of mess to universalize it but, I'm trying to see.

1:09:35

What's your intuition why handwritten recognition is easier than language?

1:09:42

Just, I think a lot of people would agree with that, but if you could elucidate sort of, the intuition of why.

1:09:51

- I don't know, no, I don't think in this direction.

1:09:56

I just think in the direction that this is problem, which if you will solve it well, we will create some abstract understanding of images. Maybe not all images.

1:10:19

I would like talk to guys who doing in real images in Columbia University. - What kind of images? Unreal you said? - Real images. - Real images.

1:10:29

- Yeah, what their idea is, there are predicate what can be predicate.

1:10:35

I say symmetry will play a role in real-life images.

1:10:41

In any real-life images, 2-D images, let's talk about 2-D images. Because... that's what we know.

1:10:52

And neural network was created for 2-D images.

1:10:55

- So the people I know in vision science, for example, for people who study human vision, that they usually go to the world of symbols and like, handwritten recognition but not really.

1:11:06

It's other kinds of symbols to study our visual perception system.

1:11:11

As far as I know, not much predicate-type of thinking is understood about our vision system.

1:11:17

- [Vladimir] They did not think in this direction.

1:11:19

- They don't, yeah, but how do you even begin to think in that direction?

1:11:24

- That is, I would like to discuss this.

1:11:27

Because if you'll be able to show that it is what's working, and theoretical thing, it's not so bad.

1:11:40

- So if we compare it to language, language has like letters, a finite set of letters and a finite set of ways that you can put together those letters so it feels more amenable to kind of analysis.

1:11:53

With natural images, there is so many pixels-- - No, no, no, letter, language is much, much more complicated.

1:12:03

It involves a lot of different stuff.

1:12:08

It's not just understanding of simple class of tasks.

1:12:15

I would like to see list of task where language is involved. - Yes.

1:12:20

So there's a lot of nice benchmarks now in natural language processing, from the very trivial, like, understanding the elements of a sentence, to question/answering, so much more complicated where you talk about open domain dialogue.

1:12:36

The natural question is, will handwritten recognition, it's really the first step of understanding visual information. - Right.

1:12:48

But even our records show that we're going wrong direction.

1:12:54

Because we need 60,000 digits.

1:12:56

- So even this first step, so forget about talking about the full journey.

1:13:01

This first step should be taking in the right direction.

1:13:03

- No, no, in wrong direction because 60,000 is unacceptable.

1:13:07

- No, I'm saying it should be taken in the right direction because 60,000 is not acceptable.

1:13:14

- You can talk, it's great we have 1/2 percent of error.

1:13:18

- And hopefully the step from doing hand recognition using very few examples, a step towards what babies do when they crawl and they understand their physical environment.

1:13:29

- I don't know what baby do.

1:13:29

- I know you don't know about babies, but-- - If you will do from very small examples, you will find principals which are different.

1:13:38

- That will apply to babies.

1:13:40

- From what we're using now.

1:13:45

Theoretical it's more or less clear.

1:13:48

That means you will use weak convergence, not just strong convergence.

1:13:54

- Do you think these principals are, will naturally be human interpretable? - [Vladimir] Oh yeah.

1:14:02

- So we'll be able to explain them and have a nice presentation to show what those principals are?

1:14:08

Or are they very, going to be very kind of, abstract kinds of functions?

1:14:14

- For example, I talk yesterday about symmetry.

1:14:18

And I gave very simple examples.

1:14:20

The same will be like that.

1:14:22

- You gave like, a predicate of a basic for-- - For symmetries. - Yes.

1:14:26

For different symmetries and you have for-- - For degree of symmetries.

1:14:30

That is important, not just symmetry exist and does not exist; degree of symmetry.

1:14:38

- [Lex] Yeah, for handwritten recognition.

1:14:41

- It's not for handwritten, it's for any images.

1:14:45

But I would like to apply it to handwritten.

1:14:47

- Right, in theory it's more general. Okay, okay.

1:14:55

So a lot of the things we've been talking about falls, we've been talking about philosophy a little bit, but also about mathematics and statistics.

1:15:05

A lot of it falls into this idea, a universal idea of statistical theory of learning.

1:15:11

What is the most beautiful and sort of, powerful or essential idea that you've come across, even for yourself just personally, in the world of statistics or statistic theory of learning?

1:15:25

- Probably uniform convergence which we do with (mumbles).

1:15:32

- [Lex] Can you describe universal convergence?

1:15:36

- You have law of large numbers.

1:15:40

So for any function, expectation of function, average of function converged expectation.

1:15:48

But if you have a set of functions, for any function it is true.

1:15:53

But it should converge similar to anywhere therefore all set of functions.

1:16:01

For learning, you need uniform convergence; just convergence is not enough.

1:16:11

Because when you pick up one which gives minimum, you can pick up one function which does not converge and it will give you the best answer for this function.

1:16:31

So you need the uniform convergence to guarantee learning.

1:16:35

Learning does not really enter your law of large numbers, really universal.

1:16:46

The idea of universal convergence exists in statistics for a long time.

1:16:55

It is interesting that as I think about myself, how stupid I was for 50 years, I did not see weak convergence.

1:17:08

I work only on strong convergence.

1:17:10

But now I think that most powerful is weak convergence because it makes admissible set of functions.

1:17:18

And even in old proverbs, when people tried to understand recognition about duck law, looks like a duck and so on, they used weak convergence.

1:17:32

People in language they understand this.

1:17:36

But when we're trying to create artificial intelligence if we want invent in different way.

1:17:46

Just consider strong convergence, artificial intelligence.

1:17:50

- So reducing the set of admissible functions you think there should be effort put into understanding the properties of weak convergence?

1:17:59

- You know, in classical mathematics, in Hilbert Space, there are only two forms of convergence, strong and weak. Now we can use both.

1:18:16

That means that we did everything.

1:18:21

And it so happened, that when we used Hilbert Space, which is very rich space, space of continuous functions which has an integral and square.

1:18:38

So we can apply weak and strong convergence for learning and have closed-form solution.

1:18:45

So for computation it is simple.

1:18:47

For me it is sign that it is right way.

1:18:52

Because you don't need any this theory.

1:18:56

Yes, do whatever you want.

1:18:59

But now the only way left is the concept of what is predicate? - Of predicate.

1:19:04

- But it is not statistics.

1:19:08

- By the way, I like the fact that you think that heuristics are a mess that should be removed from the system, so closed-form solution is the ultimate goal.

1:19:16

- No it so happens that when you're using right instrument, you have closed-form solution.

1:19:28

- Do you think intelligence, human-level intelligence, when we create it will, will have something like a closed-form solution?

1:19:43

- Now I'm looking on bounds, which I gave bounds for convergence.

1:19:50

And when I'm looking for bounds, I'm thinking, what is the most appropriate kernel of this bound would be.

1:20:02

So we know that in say, all our businesses we use radial basis function.

1:20:12

But looking for the bound I think that I start to understand that maybe we need to make corrections to radial basis function, to be closer to what's better for these bounds.

1:20:28

So I'm again trying to understand what type of kernel have best approximation, not an approximation, best fit to these bounds.

1:20:43

- Sure, so there's a lot of interesting work that could be done in discovering better function than the radial basis functions for the kinds of bounds you would find.

1:20:53

- It still comes from, you're looking to match and trying to understand.

1:21:00

- From your own mind looking at the-- - Yeah but-- - I don't know.

1:21:03

- Then I'm trying to understand what will be good for that.

1:21:11

- Yeah, but to me there's still a beauty, again, maybe I'm descending value toward heuristics.

1:21:17

To me, ultimately intelligence will be a mess of heuristics.

1:21:23

And that's the engineering answer, I guess. - Absolutely.

1:21:27

When you're doing say, self-driving cars, the great guy who will do this.

1:21:35

It doesn't matter what theory behind that.

1:21:40

Who has a better theory have to apply.

1:21:46

It is the same story about predicate because you cannot create rule for, situation is much more than you have rule for that.

1:22:01

Maybe you can have more abstract rules, then it will be less literal.

1:22:08

It is the same story about ideas and ideas applied to specific cases.

1:22:16

- But still you should-- - You cannot avoid this.

1:22:18

- [Lex] Yes of course, but you should still reach for the ideas to understand the science. - Yeah, yeah.

1:22:23

- Let me kind of ask, do you think neural networks or functions can be made to reason?

1:22:34

What do you think, we've been talking about intelligence, but this idea of reasoning.

1:22:39

There's an element of sequentially disassembling, interpreting the images.

1:22:48

When you think of handwritten recognition, we kind of think that there will be a single, there's an input and an output; there's not a recurrence. - Yeah.

1:23:00

- What do you think about, sort of, the idea of recurrence?

1:23:04

Of going back to memory and thinking through this sort of, sequentially, mangling the different representations over and over until you arrive at a conclusion?

1:23:20

Or is ultimately all of that can be wrapped up in a function?

1:23:22

(chuckles) - You're suggesting, that let us use this type of algorithm.

1:23:29

When I started thinking, I first of all, starting to understand what I want.

1:23:36

Can I write down what I want?

1:23:41

And then I try to formalize.

1:23:44

And when I do that, I'm thinking how to solve this problem.

1:23:57

Till now I did not see situation where-- - Where you need recurrence. - Recurrent.

1:24:04

- But do you observe human beings? - Yeah.

1:24:07

- Do you try to, it's the imitation question, right?

1:24:12

It seems that human beings reason, this kind of sequentially, sort of, does that inspire in you a thought that we need to add that into our intelligence systems?

1:24:30

You're saying, okay, you've kind of answered saying, until now I haven't seen a need for it.

1:24:37

And so because of that, you don't see a reason to think about it?

1:24:41

- You know, most of things I don't understand.

1:24:45

In reasoning, in humans, it is for me too complicated.

1:24:52

For me, the most difficult part is to ask questions, good questions.

1:25:03

How it works, how people asking questions. I don't know this.

1:25:11

- You said that machine learning's not only about technical things, speaking of questions, but it's also about philosophy.

1:25:20

What role does philosophy play in machine learning?

1:25:23

We talked about Plato, but generally thinking in this philosophical way, how does philosophy and math fit together in your mind?

1:25:37

- Just ideas, and then their implementation.

1:25:39

It's like predicate, say, admissible set of functions.

1:25:49

It comes together, everything.

1:25:51

Because, the first declaration of theory was done 50 years ago, all that necessary, so everything there.

1:26:02

If you have data you can, and you, in your set of functions has not big capacity.

1:26:13

So law of inter-dimension, you can do that.

1:26:15

You can make structuralist minimization, control capacity.

1:26:21

But there was not table to make admissible set of function with.

1:26:29

Now when suddenly we realize that we did not use another idea of convergence, which we can, everything comes together.

1:26:41

- But those are mathematical notions.

1:26:43

Philosophy plays a role of simply saying that we should be swimming in the space of ideas.

1:26:52

- Let's talk, what is philosophy?

1:26:54

Philosophy means understanding of life.

1:26:59

Understanding of life, say people like Plato, they understand on very high, abstract level of life.

1:27:09

And whatever I'm doing, it's just implementation of my understanding of life.

1:27:16

But every new step, that is very difficult.

1:27:22

For example, to find this idea that we need weak convergence, was not simple for me.

1:27:40

- So that required thinking about life a little bit.

1:27:44

Hard to trace, but there was some thought process.

1:27:49

- You know, when I'm thinking about the same problem for 50 years now, and again and again and again, I'm trying to understand that, this is very important, not to be very enthusiastic.

1:28:06

But concentrate on whatever that was not able to achieve. - Patient. - Yeah. And understand why.

1:28:14

And now I understand that, because I believe in math, I believe that, in this idea.

1:28:23

But now when I see that there are only two ways of convergence, and we're using loss.

1:28:32

That means that we must as well as people do it.

1:28:40

But now, exactly in philosophy and what we know about predicate, how we understand life can be described as a predicate. I thought about that.

1:28:54

And that is more or less obvious level of symmetry.

1:29:00

But next, I have a feeling it's something about structures.

1:29:09

But I don't know how to formulate, how to measure and measure a structure and all this stuff.

1:29:17

The guy who will solve this challenge problem, then when they will look at how he did it, probably just only symmetry is not enough.

1:29:31

- [Lex] But something like symmetry will be there. Structures of that kind. - Oh yeah, absolutely. Symmetry will be there.

1:29:37

A level of symmetry will be there.

1:29:40

And level of symmetry, anti-symmetry, diagonal, vertical, I even don't know how you can use in different direction the degree of symmetry; that's very general. But it will be there.

1:29:55

I think that people are very sensitive to the idea of symmetry.

1:29:58

But there are several ideas like symmetry.

1:30:04

As I would like to learn.

1:30:07

But you cannot learn just thinking about that.

1:30:11

You should do challenging problems and then analyze them.

1:30:16

Why it was able to solve them. And then you will see.

1:30:22

Very simple things, it's not easy to find.

1:30:26

(Lex laughs) Even talking about this, every time. - Yes.

1:30:32

- I was surprised, I tried to understand, these people describe in language strong convergence mechanism for learning.

1:30:44

I did not see it, I don't know.

1:30:47

But weak convergence, the duck story and story like that when you will explain, you will use weak convergence argument.

1:30:57

It looks like a (mumbles) but when you try to formalize, you're just ignoring this. Why? Why 50 years?

1:31:08

From start of machine learning.

1:31:10

- [Lex] And that's the role of philosophy, thinking about life. - I think that might be. I don't know.

1:31:18

Maybe this is theory also, we should blame for that because empirical risk minimization and now just starting, if you read now textbooks, they just about bound about empirical risk minimization.

1:31:34

They don't look for another problem like admissible set.

1:31:41

- But on the topic of life, perhaps we, you, could talk in Russian for a little bit.

1:31:50

What's your favorite memory from childhood?

1:31:53

Okay, I want you to be my (speaks in Russian) - Music.

1:31:59

- How about, can you try and answer in Russian?

1:32:02

(speaking in Russian) What kind of musica?

1:32:10

(speaking in Russian) (speaking in Russian) Now that we're talking about Bach, let's switch back to English, 'cause I like Beethoven and Chopin so.

1:33:18

- [Vladimir] Chopin is another amusing story.

1:33:20

I was-- - But Bach, if we talk about predicates, Bach probably has the most sort of, well defined predicates that underlie it.

1:33:31

- You know, it is very interesting to read what critics writing about Bach, which words they're using, they're trying to describe predicates. - Yeah. - And then Chopin.

1:33:47

It is very different vocabulary.

1:33:52

Very different predicate.

1:33:55

And I think that, if you will make collection of that.

1:34:02

So maybe from this you can describe predicate for digit recognition.

1:34:07

- [Lex] From Bach and Chopin.

1:34:10

- No, no, no, not from Bach and Chopin.

1:34:12

- [Lex] From the the critic interpretation of the music, yeah.

1:34:14

- They're trying to explain you music, what they use this?

1:34:23

They describe high-level ideas of Plato's ideas behind this music. - That's brilliant.

1:34:31

Art is not self-explanatory in some sense.

1:34:34

So you have to try to convert it into ideas.

1:34:39

- It is insulate problems.

1:34:40

When you go from ideas to, to the representation.

1:34:46

It is easy way, but when you're trying to go back, it is you will pose problems but, nevertheless, I believe when you're looking from that, even from art, you will be able to find predicate for digit recognition.

1:35:03

- That's such a fascinating and powerful notion.

1:35:08

Do you ponder your own mortality? Do you think about it? Do you fear it?

1:35:13

Do you draw insight from it? - About mortality? Oh yeah.

1:35:21

- [Lex] Are you afraid of death?

1:35:25

- Not too much, not too much.

1:35:29

It is pity that I will not be able to do something which I think I was a feeling to do that.

1:35:39

For example, I will be very happy to work with guys, take tradition from music.

1:35:50

To write this collection of description, how they describe music, how they use a predicate. And from art as well.

1:36:00

Then take what's in common, and try to understand predicate which is absolute for everything.

1:36:08

- [Lex] For visual recognition and see that there is a connection. - Yeah, yeah, exactly.

1:36:13

- [Lex] There's still time; we've got time.

1:36:16

(laughing) We've got time.

1:36:21

- It takes years and years and years. - You think so? - It's a long way.

1:36:26

- See, you've got the patient mathematician's mind.

1:36:30

I think it could be done very quickly and very beautifully.

1:36:34

I think it's a really elegant idea. Some of many.

1:36:36

- Yeah, you know, the most time it is not to make this collection, to understand what is in common, to think about that once again and again and again.

1:36:48

- Again and again and again.

1:36:50

But I think sometimes, especially when you just say this idea now, even just putting together the collection and looking at the different sets of data.

1:37:03

Language, trying to interpret music, criticize music, and images, I think there will be sparks of ideas that will come.

1:37:11

Of course, again and again you'll come up with better ideas but even just that notion is a beautiful notion.

1:37:15

- I even have some example.

1:37:21

So I have friend, who was specialized in Russian poetry.

1:37:30

She is professor of Russian poetry.

1:37:35

She did not write poems, but she know a lot of stuff.

1:37:44

She make books, several books and one of them is a collection of Russian poetry.

1:37:54

She has images of Russian poetry.

1:37:57

She collected all images of Russian poetry.

1:38:00

And I asked her to do following.

1:38:05

You have Nip's digit recognition.

1:38:09

And we get 100 digits, less than 100, I don't remember, maybe 50 digits.

1:38:18

And try from practical point of view, describe every image you see, using only words of images of Russian poetry. And she did it.

1:38:34

And then we tried to, I call it learning using privileged information.

1:38:44

I call it privileged information.

1:38:45

You have on two languages.

1:38:48

One language is just image of digit.

1:38:53

And another language by it a description of this image.

1:38:57

And this is privileged information.

1:39:02

And there is an algorithm when we are working with privileged information, you're doing well. Better, much better.

1:39:08

- So there's something there. - Something there.

1:39:13

And there is and the thing, she unfortunately died.

1:39:20

The collection of digits and poetic descriptions of those digits.

1:39:29

- [Lex] So there's something there in that poetic description.

1:39:33

- I think that there is an abstract ideas on the Plato level of ideas.

1:39:39

- Yeah, that are there, that could be discovered, and music seems to be a good entry point.

1:39:44

- But as soon as we start this, here's this challenge problem.

1:39:50

- The challenge problem-- - It immediately connected to all this stuff.

1:39:55

- Especially with your talk and this podcast, I'll do whatever I can to advertise.

1:40:00

Such a clean, beautiful, Einstein-like formulation of the challenge before us. - Right.

1:40:05

- Let me ask another absurd question.

1:40:10

We talked about mortality, we talked about philosophy of life; what do you think is the meaning of life?

1:40:17

What's the predicate for mysterious existence here on Earth? - I don't know.

1:40:33

It's very interesting how we have, in Russia, I don't know if you know the guy Strogatski?

1:40:44

They're writing futures they're thinking about, Hume, what's going on.

1:40:51

And they have an idea that there are developing two type of people: Common people and very smart people. They just started.

1:41:07

And these two branches of people will go in different directions very soon.

1:41:13

So that's what they're thinking about next.

1:41:16

- (laughs) So the purpose of life is to create two (chuckles) two paths. - Two paths. - As human societies. Yeah.

1:41:25

- Yes, simple people and more complicated people.

1:41:29

- Which do you like best?

1:41:31

The simple people or the complicated ones?

1:41:34

- I don't know, Strogatski, he's just, he's fantasy but you know, every week we have guy who is just writer, and also Soletskoff literature.

1:41:51

And he explained how he understands literature and human relationship. How he sees life.

1:42:02

And I understood that I'm just small kid comparing to him.

1:42:09

He is very smart guy in understanding life.

1:42:14

He knows this predicate, he knows big blocks of life.

1:42:20

I'm amused every time I listen to him.

1:42:24

And he's just talking about literature.

1:42:27

And I think that I was surprised.

1:42:33

So the managers in big companies, most of them are guys who study English language and English literature. So why?

1:42:52

Because they understand life. They understand models.

1:42:57

And among them, maybe many talented creatures, which are just analyzing this.

1:43:06

And this is big science like Propp did. This is his blocks. Yes, very smart.

1:43:17

- It amazes me that you are and continue to be humbled by the brilliance of others.

1:43:23

- I'm very modest about myself.

1:43:25

I see so smart guys around.

1:43:28

- Well let me immodest for you.

1:43:31

You're one of the greatest mathematicians/statisticians of our time, it's truly an honor. Thank you for talking.

1:43:36

- No, no, no, okay, okay. - And let's talk. - (laughs) It is not. - Yeah, let's talk. - I know my limits.

1:43:45

- [Lex] Let's talk again when your challenge is taken on and solved by a grad student. - Let's talk again.

1:43:51

- Especially when-- - [Vladimir] I hope that this happens.

1:43:57

- Maybe music will be involved.

1:43:58

Vladimir, thank you so much. It's been an honor. - Thank you very much.

1:44:02

- Thanks for listening to this conversation with Vladimir Vapnik.

1:44:05

And thank you to our presenting sponsor Cash App.

1:44:08

Download it, use code: LexPodcast.

1:44:10

You'll get $10 and $10 will go to FIRST, an organization that inspires and educates young mind to become science and technology innovators of tomorrow.

1:44:20

If you enjoyed this podcast, subscribe on YouTube, give it five stars on Apple PodCast, support it on Patreon, or simply connect with me on Twitter @LexFridman.

1:44:30

And now let me leave you with some words from Vladimir Vapnik.

1:44:35

"When solving a problem of interest, "do not solve a more general problem "as an intermediate step." Thank you for listening.

1:44:44

I hope to see you next time.