
22 segments available
François Chollet on June 16, 2025 at AI Startup School in San Francisco. François Chollet is a leading voice in AI. He's the creator of the Keras library, author of Deep Learning with Python, and the founder of the ARC Prize, a global competition aimed at measuring true general intelligence. He's spent years thinking deeply about what intelligence actually is, and why scaling up today’s AI models isn’t enough to reach it. In this talk, he walks through the limits of pretraining and memorized skills, and lays out a path toward true general intelligence— AI that can adapt on the fly, reason in new situations, and invent novel solutions. He explains why abstraction and compositionality matter, how ARC became the benchmark for progress, and what his team at a new research lab called Ndea is building next. Apply to Y Combinator: https://ycombinator.com/apply Work at a startup: https://workatastartup.com Chapters: 00:00 - The Falling Cost of Compute 00:57 - Deep-Learning’s Scaling Era & Benchmarks 01:59 - The ARC Benchmark 03:02 - The 2024 Shift to Test-Time Adaptation 05:01 - What Is Intelligence? 07:12 - Why Benchmarks Matter (and Mislead) 08:57 - ARC 1 Exposes Scaling Limits 10:58 - ARC 2: Compositional Reasoning Arrives 12:55 - Humans vs. Models on ARC 2 14:58 - Previewing ARC 3 & Interactive Agency 17:00 - Kaleidoscopic Hypothesis and Abstractions 22:00 - Type 1 vs. Type 2 Abstractions 26:00 - Discrete Program Search & Inventive AI 29:00 - Fusing Intuition with Symbolic Reasoning 32:00 - Building AGI Through Meta-Learning Systems 33:44 - NDEA: a new AI research lab
François Chollet discusses the historical decline in compute costs and its impact on AI development. He highlights how the affordability of GPU-based compute and large datasets in the 2010s led to significant advancements in deep learning, particularly in computer vision and natural language processing. This segment sets the stage for understanding the scaling laws that have driven AI progress.
"Hi everyone, I'm Francois. I'm super excited to share with you some of my ideas about HGI and how we're going to get there. This chart right there is one of the most important facts about the world. T..."
Chollet critiques the prevailing belief that simply scaling up AI models will lead to general intelligence. He introduces the Abstraction Reasoning Corpus (ARC) benchmark, which reveals the limitations of current models in achieving fluid general intelligence. This segment emphasizes the distinction between memorized skills and the ability to adapt and understand new problems.
"scaling laws that Jared told you about a few minutes ago. So it really seemed like really it all figured out and many people extrapolated that more scale uh was all that was needed to solve everything..."
In this segment, Chollet outlines a pivotal shift in AI research towards test-time adaptation, where models learn and adapt during inference rather than relying solely on pre-trained knowledge. He discusses the significant progress made in ARC benchmarks and the emergence of models demonstrating genuine fluid intelligence, marking a departure from traditional scaling paradigms.
"2024, everything changed. the AI research community started pivoting to a new and very different pattern test adaptation creating models that could change their own state at test time to adapt to some..."
Chollet explores the fundamental question of what intelligence truly means. He contrasts two views: one that equates intelligence with task performance and another that sees it as the ability to handle novel situations. He argues that intelligence is a process, not just a collection of skills, and emphasizes the importance of adaptability in defining true intelligence.
"adaptation, what else might be next for AI? And to answer these questions, we have to go back to a more fundamental question. What is even intelligence? What what do we mean when we say we're trying t..."
In this segment, Chollet discusses the importance of how we measure intelligence in AI. He critiques traditional benchmarks for their focus on static skills rather than fluid intelligence and highlights the need for definitions that capture the essence of adaptability and problem-solving in novel situations. This segment underscores the implications of measurement on AI development.
"are confusing the process and its output. So don't confuse the road and the process that created the road. So to formalize this a bit, I see intelligence as the conversion ratio between the informatio..."
Chollet concludes by addressing the ultimate goal of AI: achieving autonomous invention rather than mere automation of tasks. He argues for a redefinition of intelligence that aligns with the aspirations of AGI, aiming to tackle humanity's most challenging problems and accelerate scientific progress. This segment encapsulates the vision for the future of AI and its potential impact.
"distinction between static skills and fluid intelligence. So between having access to a collection uh of static programs to solve known problems versus being able to synthesize brand new programs on t..."
François Chollet introduces the ARC benchmark, designed to measure intelligence in AI systems. He explains that ARC tasks require general intelligence rather than memorized knowledge, making it a unique tool for assessing AI capabilities. Chollet highlights that ARC is not just a test but a way to direct research towards solving critical bottlenecks on the path to AGI.
"meant to be. And to achieve that, we need a new target. We need to stop targeting fluid intelligence itself, the ability to adapt and invent. So one definition of AGI only enops automation. So it incr..."
Chollet elaborates on the design of ARC tasks, which are based on core knowledge that even young children possess. He emphasizes that solving these tasks requires intelligence rather than rote memorization, making them challenging for AI. This distinction highlights the gaps in current AI capabilities and the need for new approaches to achieve AGI.
"and also humans. So ARK1 contains 1,000 tasks like this one here. And each task is unique. So that means that you cannot cram for ARC. You have to figure out each task on the fly by using your general..."
Chollet clarifies that ARC is not a definitive measure of AGI but a tool to guide research towards understanding intelligence. He discusses how ARC has resisted the pre-training scaling paradigm, demonstrating that fluid intelligence cannot be achieved through scaling alone. This insight is crucial for advancing AI research and development.
"And meanwhile pretty pretty much every other benchmark out there is targeting fixed known tasks. So they can't actually be solved or hacked via memorization alone. That's what makes ARC fairly easy fo..."
François Chollet critiques ARC 1 for being a binary test that does not adequately measure fluid intelligence. He explains that while it can indicate the presence of some intelligence, it does not provide a nuanced understanding of AI capabilities. Chollet emphasizes the need for more sensitive tools to evaluate AI systems effectively.
"and ARC has completely resisted the pre-training scaling paradigm. Even after a 50,000x scale up of pre-trained baselons, their performance on ARC stayed near zero. So we can decisively conclude that ..."
Chollet presents ARC 2, which aims to challenge reasoning systems and improve the evaluation of AI capabilities. He highlights the increased complexity of tasks in ARC 2, which require deliberate thinking and are not easily solvable through brute force. This benchmark is designed to provide a more granular assessment of AI systems compared to its predecessor.
"Well not yet. What you see on this graph is that ARK1 was a binary test. It was a minimal reproduction of fluid intelligence. So it only really gives you two possible modes. Either you have no fluid i..."
In this segment, Chollet discusses the performance of AI models on ARC 2 compared to human participants. He shares results from testing with a diverse group of individuals, emphasizing that while humans can solve the tasks, current AI models struggle significantly. This disparity underscores the ongoing challenges in achieving AGI.
"reasoning systems. It changes the test adaptation pattern. The benchmark format is still the same. There's a much greater focus on probing compositional jarization. So the tasks are still very feasibl..."
Chollet outlines the future direction of ARC with the development of ARC 3, which will assess interactive agency in AI systems. He describes how this benchmark will challenge AI to learn and adapt in novel environments, marking a significant departure from previous formats. This evolution is crucial for advancing towards true AGI.
"with no prior training. So how well do AI models do? Well, if you take basel models like GPT4.5, Lama 4, it's simple. They get 0%. There is simply no way to do these tasks simply via memorization. Nex..."
Chollet introduces ARC 3, a significant evolution from previous benchmarks, focusing on assessing agency and interactive learning. He explains how this new framework will evaluate AI's ability to autonomously explore and achieve goals in unfamiliar environments, setting a strict efficiency standard for task completion.
"And to be clear, I don't think ARK 2 is the final test. We're not going to stop at ARC 2. We've started development on RKGI 3 and AR 3 is a significant departure from the input output pair formats of ..."
François Chollet presents the Kaleidoscope Hypothesis, arguing that intelligence is the ability to identify and recombine fundamental abstractions from past experiences. He explains how these abstractions allow for efficient navigation of novel situations, emphasizing the importance of both abstraction acquisition and on-the-fly recombination in achieving true intelligence.
"release a developer preview so you can start playing with it. What's it going to take to solve AR 2 and we're still very far from it today. Uh then solve AR 3 and we're even further away from that. Ma..."
In this segment, Chollet elaborates on the concept of intelligence as the efficiency of operationalizing past experiences to tackle new challenges. He critiques current AI models for their inefficiency in learning and applying skills, stressing that true intelligence involves both data and computational efficiency.
"And this involves identifying um invariance uh structure uh things that seem to be repeated principles. And these building blocks, these atoms are called abstractions. And whenever you encounter a new..."
Chollet distinguishes between two types of abstraction: type one (value-centric) and type two (program-centric). He explains how both types are essential for human reasoning and cognition, and discusses the limitations of current AI models in effectively utilizing type two abstractions, which are crucial for tasks requiring discrete reasoning.
"were missing a couple of things. First, these models lacked the ability to do on the-fly recombination. So, at training time, they were learning a lot. They were acquiring many useful abstractions, bu..."
François Chollet discusses the role of discrete program search in AI invention, contrasting it with traditional deep learning methods. He highlights how past AI systems have successfully utilized discrete search for creative tasks, emphasizing that true innovation in AI will require moving beyond mere automation to harnessing the power of combinatorial search.
"important. I said that intelligence is about mining abstractions from data and then re combining them. There's really two kinds of abstraction. There's type one and type two. They're pretty similar to..."
In this segment, Chollet compares program synthesis to traditional machine learning, explaining how program synthesis is more data-efficient and relies on combinatorial search over graphs of operators. He discusses the challenges of combinatorial explosion in program synthesis and the need for a balance between type one and type two abstractions.
"does. So what's discrete program search? It's basically combinatoral search over graphs of operators taken from some language, some DSL. And to better understand it, you can try to draw an analogy bet..."
Chollet explores how human intelligence combines type one intuition with type two reasoning. He uses chess as an analogy to illustrate how intuition helps narrow down options for reasoning, emphasizing the importance of integrating both types of abstraction in AI development.
"explosion wall. I said earlier that intelligence is a combination of two forms abstraction. Type one and type two. And I really don't think that you're going to go very far if you go all in on just on..."
Chollet outlines a vision for AI systems that function like programmers, capable of synthesizing solutions on the fly. He describes how these systems will leverage a library of abstractions and deep learning-based intuition to tackle new tasks, aiming for a model that can adapt and improve over time.
"using type one intuition to make type two calculation tractable. So how is the merger between type one and type two going to work? Well the key system two technique is discrete search over a space of ..."
In this concluding segment, Chollet shares insights about his new research lab, Ndea, and its mission to create AI capable of independent invention and scientific discovery. He emphasizes the need for AI that can expand knowledge frontiers and outlines the approach of using deep learning-guided program search to achieve this goal.
"move towards systems that are more like programmers that approach a new task by writing software for it. And when faced with a new task, your programmer like metalarner will synthesize on the fly a pr..."