
16 segments available
In this episode of Decoded, Ankit and Francois walk through the motivation and math behind world models. They cover a number of areas including why sample efficiency is one of the biggest unsolved problems in AI, how deterministic differentiable control and Newtonian physics represent a "perfect world model," and why the action space explosion makes chess tractable but robotics nearly intractable. Apply to Y Combinator: https://www.ycombinator.com/apply Work at a startup: https://www.ycombinator.com/jobs Transcript: https://ycrootaccess.substack.com/p/world-models-an-intuitive-introduction Chapters: 00:00 — Intro 01:45 — What would perfect efficiency look like? 05:10 — World models in the human brain 09:20 — Control theory & the drone example 14:30 — When physics breaks down 17:45 — Chess, Go & the action space problem 24:10 — Why AlphaGo can't scale 28:00 — Monte Carlo tree search explained 34:00 — Self-Driving: state space is infinite 40:30 — Model-Free vs. Model-Based RL 44:00 — Why robotics is the hardest case 48:20 — World models that actually work 54:10 — JEPA & latent space tricks 59:00 — Open problems remaining 01:04:30 — Does this pass the squint test? 01:08:00 — Outro
"One of the biggest open problems in AI right now is how to solve sample efficiency. That is, how do you get models to quickly learn new tasks or skills from relatively small amounts of training data? ..."
"our current state-of-the-art AI systems, what people consider frontier intelligence, basically can't do them, >> right? I mean there we come into new problems with such inductive bias from K through 1..."
"have this crazy good world model. And there's the this uh neuroscientist at Stanford named Shaw Duckman who basically is of the view that the entire point of the growing neoortex for the during the gr..."
"going to be my uh utide by the mass and g. And so that's it. And now I have my transition function. Now how do I get to a policy? And I'm going to apply something called model predictive control or re..."
"rockets and landing rockets in Florida. Let's just say that like there's different if I have my launch pad here and I have a whole bunch of houses here. Let's just say the path going from here to here..."
"for exactly this. Yeah. And then what you can do and this is like the in vogue thing to do since Danar and and uh um the dreamer paper series from V1 to V4 is do action conditioning later like similar..."
">> So the cardality of the state I think is going to be s uh two or three it turnary thing here I guess >> the 19 squar I think it's 361. >> Yeah something like that 361. >> Um my transition same issu..."
"give me uh 361 uh uh numbers that sum to one. And so I'll have some probability of uh of where these things are going to go for the of where my my opponent will play. Um here. >> So these are like the..."
"and the number of uh you know steps you would have to take here presumably have to be way more than 800 in order to get any reasonable uh kind of sampling of this and so you're probably multiplying th..."
">> Let's consider even just like a very simplified >> What do you have? You have a steering wheel that you can turn left, right? You have a a brake pad. >> Yeah. >> And you have the gas. >> Yeah. >> A..."
"based RL I have not just some pi but I have also my uh sigh as well here and so uh by uh by including this I can have a much stronger policy but it would take a lot more time to perform inference beca..."
"bigger action spaces. You have to sell some model. >> Yeah. Uh Lane Macintosh, I played hockey with at Stanford, who now runs Tesla FSD. Um I can ask him, but I would bet money that they shard the dat..."
"where they have this um joint model of um state transitions and actions. They train it by first instantiating it with the open- source one video diffusion model and then it only takes them about 500 h..."
"and a t in pixel space and have this is uh let's say at time t t t t t t t t t t t t t t t t t t t t t t t t t t t t t t t t t t t t t t t t t + 1, t+ etc etc and I have to actually predict now the fu..."
"let's just say for example uh I have you know uh a house here and I want to train the model on you not driving into the house. And so let's say I put I put it into a state right here to drive into the..."
"um and so you need to estimate that very quickly and adapt and that these models just kind of don't have a mechanism to do it. Yeah. And then I guess there's like the practical speed elements of these..."