
57 segments available
Jitendra Malik is a professor at Berkeley and one of the seminal figures in the field of computer vision, the kind before the deep learning revolution, and the kind after. He has been cited over 180,000 times and has mentored many world-class researchers in computer science. Support this podcast by supporting our sponsors: - BetterHelp: http://betterhelp.com/lex - ExpressVPN at https://www.expressvpn.com/lexpod EPISODE LINKS: Jitendra's website: https://people.eecs.berkeley.edu/~malik/ Jitendra's wiki: https://en.wikipedia.org/wiki/Jitendra_Malik PODCAST INFO: Podcast website: https://lexfridman.com/podcast Apple Podcasts: https://apple.co/2lwqZIr Spotify: https://spoti.fi/2nEwCF8 RSS: https://lexfridman.com/feed/podcast/ Full episodes playlist: https://www.youtube.com/playlist?list=PLrAXtmErZgOdP_8GztsuKi9nrraNbKKp4 Clips playlist: https://www.youtube.com/playlist?list=PLrAXtmErZgOeciFP3CBCIEElOJeitOr41 OUTLINE: 0:00 - Introduction 3:17 - Computer vision is hard 10:05 - Tesla Autopilot 21:20 - Human brain vs computers 23:14 - The general problem of computer vision 29:09 - Images vs video in computer vision 37:47 - Benchmarks in computer vision 40:06 - Active learning 45:34 - From pixels to semantics 52:47 - Semantic segmentation 57:05 - The three R's of computer vision 1:02:52 - End-to-end learning in computer vision 1:04:24 - 6 lessons we can learn from children 1:08:36 - Vision and language 1:12:30 - Turing test 1:16:17 - Open problems in computer vision 1:24:49 - AGI 1:35:47 - Pick the right problem CONNECT: - Subscribe to this YouTube channel - Twitter: https://twitter.com/lexfridman - LinkedIn: https://www.linkedin.com/in/lexfridman - Facebook: https://www.facebook.com/LexFridmanPage - Instagram: https://www.instagram.com/lexfridman - Medium: https://medium.com/@lexfridman - Support on Patreon: https://www.patreon.com/lexfridman
In this introduction, Lex Fridman introduces Jitendra Malik, a prominent professor at Berkeley and a key figure in computer vision. Malik's extensive contributions to the field, including his influence on deep learning and mentorship of researchers, set the stage for a deep dive into the complexities of computer vision.
"the following is a conversation with jitendra malik a professor at berkeley and one of the"
Jitendra Malik discusses why computer vision is often underestimated. He explains that much of human visual processing occurs subconsciously, leading researchers to mistakenly believe that replicating this in machines is straightforward. The complexity of visual processing in the human brain highlights the challenges faced in AI.
"in 1966 seymour papper at mit wrote up a proposal called the summer vision project to be given as far as we know to 10 students to work on and solve that summer so that proposal outlined many of the c..."
Malik introduces the concept of the 'fallacy of the successful first step' in computer vision. He explains how initial successes can mislead researchers into thinking that solving complex vision problems is easier than it truly is, emphasizing the significant challenges that remain in achieving high accuracy.
"you said the higher level parts are the harder parts i think vision appears to to be easy because uh most of what visual processing is subconscious or unconscious right so we underestimate the difficu..."
In this segment, Malik contrasts the confidence seen in natural language processing with the hesitance in computer vision. He discusses the necessity of common sense reasoning in vision tasks and how this understanding is often overlooked, making the challenges of vision interpretation more complex.
"yeah i think in the early days it could have been excused because in the early days all aspects of ai were regarded as too easy but i think today it is much less excusable and i think why people fall ..."
Malik elaborates on the cognitive challenges inherent in vision tasks, particularly in autonomous driving. He argues that while some aspects of driving can be automated, the need for sophisticated reasoning in edge cases presents significant hurdles that current systems struggle to overcome.
"still an open problem but in your sense how much understanding is required to solve vision like this put another way how much something called common sense reasoning is required to really be able to i..."
This segment focuses on the complexities of autonomous driving, where Malik discusses the limitations of current vision-based systems like Tesla's Autopilot. He highlights the need for robust solutions that can handle diverse driving conditions and the critical importance of understanding human behavior in driving scenarios.
"in the near future and the reason is because i think there will be that 0.01 percent of the cases where quite sophisticated cognitive reasoning is called for however there are tasks where you can firs..."
Malik explains how vision tasks extend beyond mere perception to include predictive modeling of behaviors in driving. He emphasizes the importance of understanding the actions of other agents on the road, such as pedestrians and skateboarders, and how this requires a higher level of cognitive processing.
"is much more important so just for the for the fun of it since you mentioned let's go there briefly about autonomous vehicles so one of the companies in the space tesla is work with andre karpathy and..."
In this segment, Malik discusses the data requirements for training computer vision systems. He contrasts the data needs of machines with those of humans, arguing that current systems require significantly more data to learn the same capabilities, which poses a challenge for developing effective autonomous driving solutions.
"work under all kinds of driving conditions at that point it's not just a question of vision or perception but really also of control and dealing with all the edge cases so where do you think most of t..."
Malik speculates on the future of autonomous driving technology, questioning whether it will rely on existing architectures like neural networks or evolve into something fundamentally different. He emphasizes the importance of learning from human experiences and the limitations of current deep learning approaches.
"and as much as it informs what's going to happen in the future so i think we have to build predictive models of of of behaviors of people and and those can get quite complicated so uh uh i mean uh i i..."
Malik elaborates on the differences between human learning and current machine learning techniques, particularly in the context of driving. He explains that humans accumulate visual knowledge from a young age, allowing them to learn control in driving quickly, while machines rely on a more simplistic, data-driven approach that lacks the richness of human experience.
"uh i will tell you what i would bet on uh so and this is at my general philosophical position on how these uh learning systems have been uh what we have found currently very effective in computer visi..."
In this segment, Malik highlights how teenagers entering driver education are already visual experts due to their early experiences. He discusses how this foundational knowledge allows them to handle novel driving situations effectively, contrasting it with the limitations of current computer vision systems that lack such innate understanding.
"because from 0 to 16 they have built a certain repertoire of vision in fact most of it has probably been achieved by age 2 right in in this period of age up to age 2 they know that the world is three-..."
Malik reflects on the evolution of learning techniques in artificial intelligence, particularly neural networks. He argues that while there is potential for neural networks to accumulate knowledge like humans, significant advancements in learning methodologies are necessary to achieve this level of sophistication.
"they're learning a sense of typical traffic situations now the the that education process can be quite short because they are coming in as visual geniuses and of course in their future they're going t..."
This segment explores the complexity of human learning compared to supervised learning in machines. Malik emphasizes that human learning involves exploration and manipulation of the environment, leading to a richer dataset that machines currently cannot replicate. He advocates for developing models that encompass the various aspects of human learning.
"same i think uh i don't see any in principle problem with neural networks doing it but i think the learning techniques would need to evolve significantly so the current uh the current learning techniq..."
Malik discusses the differences in computational power between the human brain and modern computers. He notes that while advancements in GPU technology have brought us closer to the brain's capabilities, the efficiency and style of computation remain vastly different, posing challenges for real-world applications.
"learning and what we would need to do is to develop models of all of these and then train our systems in that with that kind of uh protocol so new new methods of learning yes some of which might imita..."
In this segment, Malik articulates the overarching problem of computer vision, linking it to biological perception and action. He argues that perception must be coupled with action to be meaningful, drawing parallels between the evolution of perception in animals and the current challenges faced in computer vision.
"human brain or referencing marvik hans marvel the so do you do you think there's something interesting valuable to consider about the difference in the computational power of the human brain versus th..."
Malik emphasizes the importance of connecting perception to action in the context of computer vision. He discusses how early multicellular organisms relied on this connection for survival, suggesting that modern computer vision systems should similarly integrate perception with actionable outcomes.
"but it's not in the of the same style it's of a very different style so i mean for example the the style of computing that we have in our gpus is far far more power hungry than the style of computing ..."
Malik explains how the evolution of visual systems has led to sophisticated capabilities in humans. He discusses how scientists impose structures on these systems for analysis, but in reality, various processes operate simultaneously, highlighting the complexity of understanding and modeling vision.
"of large scale let me ask sort of the high level question step taking a step back how would you articulate the general problem of computer vision does such a thing exist so if you look at the computer..."
In this segment, Malik addresses the imperfections of human perception models compared to the external world. He notes that while our visual systems have evolved significantly, they are not perfect, and psychologists have demonstrated various illusions that reveal these limitations.
"different places but you need to know where to go and that's really about perception or seeing i mean i mean vision is perhaps the single most perception sense but all the others are equally are also ..."
Malik concludes that the fundamental purpose of vision is to facilitate action. He discusses how our hyper-evolved visual systems serve various purposes beyond mere survival, including aesthetic appreciation, and how these evolved capabilities influence modern applications of computer vision.
"and they have create various illusions to show the ways in which it is imperfect but it's amazing how far it has come from a very simple perception action loop that you exists in you know an animal 50..."
Malik contrasts the challenges of static image analysis with those of dynamic video processing in computer vision. He reflects on historical limitations that led to a focus on static images and discusses the ongoing challenges in video computing, particularly for researchers and companies alike.
"uh that's the most fundamental purpose we have by now hyper evolved so we have this visual system which can be used for other things for example judging the aesthetic value of a painting and this is n..."
Jitendra Malik discusses the historical limitations in computer vision that led to a focus on single images instead of video. He explains how past computational constraints forced the community to prioritize edge detection over retaining full image data, impacting the evolution of computer vision techniques. Malik emphasizes that while video processing remains underexplored, advancements are on the horizon as computational power increases.
"sometimes we we can simplify our problem so much that we essentially lose part of the juice that could enable us to solve the problem and one could reasonably argue that to some extent this happens wh..."
Malik highlights the current state of video recognition, asserting that it lags approximately ten years behind object recognition. He quantifies this gap by referencing performance metrics from challenging video datasets, suggesting that the field is poised for significant progress in the coming years as researchers tackle the complexities of video data.
"we are no longer detecting edges right we process images with convnets because we don't need to we don't have that those compute restrictions anymore now video is still under studied because video com..."
In this segment, Malik explores the necessity of integrating knowledge bases and reasoning into dynamic scene understanding for improved action recognition. He draws parallels between human cognitive processes and AI, advocating for the development of schemas that can enhance long-term video understanding, similar to how children learn through observation.
"already asked but once again so for dynamic scenes do you think do you think some kind of injection of knowledge basis and reasoning is required to help improve like action recognition like if if if u..."
Malik discusses the importance of mimicking child-like learning in AI systems, particularly in the context of computer vision. He emphasizes the need for interactive learning environments that allow AI to conduct experiments and build causal models, drawing on the analogy of children learning through real-world experiences and controlled experiments.
"that when we are going to do long-form video understanding we are going to need to do this i think the kinds of technology that we have right now with 3d convolutions over a couple of seconds of clip ..."
This segment delves into the concept of active learning in AI, where systems learn through interaction with their environment. Malik argues that this approach can help bridge the gap between correlation and causation, enabling AI to develop a deeper understanding of the world, akin to how children conduct experiments to refine their knowledge.
"to simulate the adult mind why not rather try to produce one which simulates the child's so that's a really interesting point if i think about the benchmarks we have before us the the tests of our com..."
Malik expresses optimism about the future of simulation environments in AI, highlighting advancements in computer graphics that could lead to more realistic simulations. He believes that as these environments improve, they will provide valuable opportunities for AI systems to learn and interact with their surroundings in a meaningful way.
"so one of the fundamental aspects of learning like a child is the interactivity so the child gets to play with the data set it's learning from yes it's against the select i mean you can call that acti..."
In this segment, Malik breaks down the components of early vision, including sensation, perception, and cognition. He discusses the significance of image statistics in understanding visual data and how biological systems perform compression to manage the vast amount of information processed by the brain, drawing parallels to advancements in artificial neural networks.
"deep learning as just curve fitting right because it's uh but i i don't quite agree about as a troublemaker he is but uh causality is important but causality is not is not like a single silver bullet ..."
Jitendra Malik discusses the advancements in creating realistic models in computer vision, particularly in simulating both optical and physical interactions. He emphasizes the importance of image statistics and the redundancy in images that allows for effective compression, drawing parallels to biological systems and artificial neural networks.
"but then they could start to do pretty realistic models of that and so on and so forth so the graphics people have shown that they can do this forward direction not just for optical interactions but a..."
Malik elaborates on the concept of early vision, which includes sensation, perception, and cognition. He poses questions about what can be learned from image statistics and discusses the significance of compression in biological settings, highlighting the differences between human vision and artificial neural networks.
"what can we learn from image statistics that we don't already know so at the lowest level what um what can we make from just this the the statistic the basics so there were the variations in the rock ..."
In this segment, Malik contrasts bottom-up and top-down processing in computer vision. He explains how current systems operate in a feed-forward manner, starting from raw pixels, and discusses the implications of feedback mechanisms in biological vision, suggesting that a combination of both approaches could enhance machine vision.
"yeah just just the statistics yeah um how much how much well so i mean the the way to think about it is just how successful is image compression right and we we and there are and that's been done with..."
Malik emphasizes the importance of feedback in biological vision systems compared to the purely feed-forward approach of artificial neural networks. He argues for the need to incorporate feedback mechanisms in machine vision to better handle ambiguous stimuli and improve overall performance.
"yeah so this is this is at the lower level right so we are we are we as i said yeah that's focusing on low level statistics so to linger on that for a little bit uh you mentioned how far can bottom-up..."
This segment focuses on the relationship between segmentation and recognition in computer vision. Malik explains how segmentation allows for the identification of objects without needing to name them, and discusses the implications of weak supervision in human learning compared to current machine learning methods.
"totally feed forward they're trained in a very top-down way so they're trained by saying okay this is a cat there's a cat there's a dog there's a zebra etc and i'm not happy with either of these choic..."
Malik introduces the three R's of computer vision: recognition, reconstruction, and reorganization. He describes how recognition involves labeling objects, reconstruction relates to creating internal models from images, and reorganization focuses on structuring visual information into meaningful entities.
"and the two are functionally equivalent because if you have a feedback network which just has like three rounds of feedback you can just unroll it and make it three times the depth and create it in a ..."
In this segment, Malik discusses how recognition, reconstruction, and reorganization interact in computer vision. He emphasizes the importance of understanding these relationships to improve machine vision systems and highlights the need for a more integrated approach to solving vision problems.
"because for example we're looking at a video stream and the horse moves and that enables me to say that all these pixels are together yeah so the gestural psychologists used to call this the principle..."
Malik elaborates on the significance of segmentation in computer vision, explaining how it enables the identification of objects and facilitates recognition with less supervision. He discusses practical applications, such as medical diagnosis, where precise segmentation is crucial.
"and and and and the question that you can ask is so for me i'm inspired a lot by human vision and i care about that you could be a just a hard-boiled engineer not give a damn so to you i would then ar..."
This segment highlights the concept of weak supervision in the context of segmentation and recognition. Malik argues that humans can learn to identify objects with minimal guidance, suggesting that machine vision could benefit from similar approaches to reduce the need for extensive labeled data.
"uh of drawing outlines around objects versus a bounding box or and then classifying that object what's what's the value of segmentation what is it as a problem in computer vision how is it fundamental..."
Malik discusses segmentation as a fundamental capability in computer vision that allows for the separation of objects from their backgrounds. He emphasizes its role in enabling recognition tasks and how it can be applied in various fields, including medical imaging.
"to uh as we are growing up to acquire uh names of objects with very little supervision so suppose the child lets posit that the child has this ability to separate out objects in the world then when th..."
In this segment, Malik addresses the challenges of segmenting non-rigid objects, such as humans and animals, in computer vision. He explains how the movement and appearance of these objects complicate the segmentation process and the need for advanced techniques to handle such variability.
"than we require today and you think of segmentation as this kind of task that takes on a visual scene and breaks it apart into into interesting entities yeah that might be useful for whatever the task..."
Malik provides definitions for the three R's of computer vision: recognition, reconstruction, and organization. He explains how recognition involves labeling, reconstruction relates to creating models from images, and organization focuses on structuring visual information into coherent entities.
"so it's all sort of a challenge you mentioned the three hours of computer vision are recognition reconstruction reorganization can you describe these three r's sure how they interact yeah so uh so rec..."
This segment explores how recognition, reconstruction, and organization interact within computer vision systems. Malik discusses the importance of understanding these interactions to improve the effectiveness of machine vision and the need for a holistic approach to vision problems.
"inverse graphics i mean that's one way to think about it so graphics is your you have some internal computer representation and uh you have a computer representation of some objects arranged in a scen..."
Malik reflects on the evolution of computer vision in the context of deep learning. He discusses how modern neural networks facilitate the integration of multiple representations and the importance of recognizing the connections between different vision tasks.
"so segmentation is a small part of that so segmentation gets us going towards that yeah and you kind of have this triangle where they all interact together yes so how do you see that interaction in uh..."
Malik explores the concept of end-to-end learning in computer vision, contrasting it with a child development perspective. He highlights six lessons from child development that can inform AI learning, including being multimodal and incremental. Malik argues for a broader view of learning that encompasses lifelong development rather than just supervised learning for specific tasks.
"so speaking of neural networks how much of this uh problem of computer vision of the organization recognition can be um reconstruction how much of it can be learned end to end do you think instead of ..."
In this segment, Malik illustrates the significance of multimodal learning through the example of a child interacting with objects. He explains how combining tactile and visual signals enhances understanding and learning. Malik discusses the potential of using multimodal signals in AI to improve weak supervision and the importance of temporal links between different sensory inputs.
"that we can learn from children uh of be multimodal be incremental be physical explore be social use language can you speak to these perhaps picking one that you find most fundamental toward yeah time..."
Malik argues that vision is a more fundamental ability than language, tracing the evolutionary development of both. He explains how spatial intelligence and the ability to manipulate objects laid the groundwork for language development. This segment delves into the relationship between vision and language, emphasizing that language builds upon the perceptual experiences derived from visual interactions with the world.
"not yeah i mean we have a little bit of work at this but uh but but so much more needs to be done yeah so so so so this this is this is a good example be physical that's to do with uh like the one thi..."
Malik critiques the traditional Turing Test as a measure of intelligence, suggesting that it oversimplifies the complexities of AI. He proposes a more nuanced approach that includes various tasks across different domains, such as visual understanding and manipulation. Malik emphasizes the need for diverse benchmarks to assess AI capabilities beyond mere imitation.
"yeah what you refer to as the spatial intelligence yeah yeah to linger a little bit we mentioned touring and his uh mention of we should learn from children nevertheless language is the fundamental pi..."
Malik identifies long-form video understanding as a significant unsolved problem in computer vision. He discusses the complexities involved in understanding behavior, intentions, and predictions in video content. This segment highlights the need for advanced tracking and modeling of entities in video to achieve a deeper understanding of dynamic scenes.
"amount of work to be done so on the visual understanding side in this intelligence olympics that we've set up yeah what's a good test for one of many of visual scene understanding uh do you think such..."
Malik elaborates on the challenges of achieving rich 3D understanding in computer vision. He contrasts traditional multi-view geometry techniques with modern approaches that rely on single-view predictions. This segment underscores the need for a more unified understanding of 3D representations, emphasizing the importance of learning from diverse visual experiences to build internal models.
"uh clearly clearly unsolved which is what i would call long-form video understanding so so we have a video clip and we want to understand the behavior in there in terms of agents their goals intention..."
In this segment, Malik addresses the explainability of neural networks in achieving visual understanding. He argues that while humans are not fully explainable to each other, there are contexts, such as medical diagnoses, where interpretability is crucial. Malik suggests that high-performing systems can operate as black boxes, raising questions about the balance between performance and transparency in AI.
"uh learning as you move around the world uh notion of 3d so so we we have our succession of visual experiences and from those we so in as part of that i might see a chair from different viewpoints or ..."
Malik shares his insights on the potential for achieving artificial general intelligence (AGI). He distinguishes between the theoretical possibility of AGI and the pragmatic timeline for its realization, expressing skepticism about achieving it within the next two decades. This segment explores the complexities of known and unknown challenges in AI development, particularly in natural language understanding and cognition.
"diagnosis based on data which was data collected of for american males who are in their 30s and 40s and maybe not so relevant to me maybe it is relevant you know et cetera et cetera and we i mean in m..."
Malik emphasizes the importance of addressing AI's current implications rather than waiting for AGI. He discusses the ethical responsibilities associated with deploying AI systems, highlighting real-world consequences such as biases in decision-making and the risks posed by self-driving cars. This segment calls for continuous vigilance in ensuring the safety and fairness of AI technologies.
"i i think that you know donald trump's felt is not a favorite person of mine but one of his lines is very good which is about known knowns known unknowns and unknown unknowns so in the business we are..."
In this concluding segment, Malik reflects on the influence of recommendation algorithms on society. He discusses how platforms like YouTube and Facebook shape our access to information and ideas, raising concerns about the control these algorithms exert over public discourse. Malik warns of the growing dependence on AI systems and the need for awareness of their societal impact.
"lookout for worrying about safety biases risks right i mean the self-driving car kills are pedestrian and they have right i mean the this uber incident in arizona yeah right it has happened right this..."
Malik reflects on the moments in his life that have shaped his happiness and career. He emphasizes the importance of being in the right place at the right time in scientific research and how this has influenced his journey in the field of computer vision.
"could go in a bunch of different directions some of which might be harmful so yeah you're right the the the threats of ai are here already we should be thinking about them on a philosophical notion if..."
Malik discusses the transformation of computer vision from a nascent field to one that is now solving practical problems. He shares his perspective on the progress made over the years and the ongoing challenges that remain in the discipline.
"time and you can you can work on problems at a time when they're just too premature you know you butt your head against them and and nothing happens because it's the prerequisites for success are not ..."
In this segment, Malik shares insights on mentorship, crediting his students' intelligence and creativity. He discusses the importance of guiding students to choose the right problems to tackle in their research, emphasizing the balance between technical skills and problem selection.
"a productive phase yes yeah yeah i think people 500 years from now would laugh are you calling this field mature yeah that is very possible yeah so but you're also lest i forget to mention you've also..."
Malik elaborates on the concept of 'the art of the soluble' in research, explaining how to identify problems that are challenging yet approachable. He stresses the significance of having the intuition to recognize ripe problems for investigation.
"run success is always based on technical competence your you know you're quick with math or you are whatever i mean there's certain technical capabilities which make for short-range progress long-rang..."
Malik discusses the value of intellectual breadth in research, sharing how his background in psychology and neuroscience allows him to connect disparate ideas. He highlights the importance of fostering a broad perspective among students to enhance their research capabilities.
"correct the other thing which i have which i i think i bring to the table uh is i is a certain intellectual breadth i i've spent a fair amount of time studying psychology neuroscience relevant areas o..."
In the conclusion of the podcast, Malik expresses gratitude for the conversation and reflects on the journey of scientific inquiry. He emphasizes the importance of beauty and meaningful connections in life, leaving listeners with a thought-provoking quote from Dostoevsky.
"that's that's of some value well it was beautifully refreshing just to hear you naturally jump to psychology back to computer science and this conversation back and forth i mean that that's uh that's ..."