0:00
The world is changing so quickly.
The world is changing so quickly.
This is probably a little bit obvious, but you should just try things and and like every day do something with AI.
Last summer I took a weekend and used GPT-5 to help me build an iPhone app.
I hadn't done that in a decade.
Yeah, it's so fast and so easy and that was you know an age ago.
That was like 8 months ago.
Now it's even faster and easier.
Don't limit yourself like anything that you imagine you should just try to use AI and see how far you can get with it.
And you'll be you know making the world better.
>> Welcome to another episode of the Light Cone.
Ian Fisher is the co-founder and co-CEO of Poetic, which is building recursively self-improving AI reasoning harnesses for LLMs.
Previously, he spent a decade as a researcher at Google DeepMind and founded a mobile dev tools company through YC years ago. Welcome, Ian. >> Thank you. I'm so happy to be here. >> What is Poetic?
How's it different than RL?
You know, how's it different than context engineering?
>> At Poetic, what we're building is a recursively self-improving system.
And so recursive self-improvement is this you know kind of the holy grail of AI where the AI is making itself smarter.
The core insight that we had is that we could do recursive self-improvement far faster and cheaper than all of the other ways that people had been proposing to do this.
And so obviously I'm I can't go into details about what that what that is what our particular approach is.
But most of the approaches out there involve you know they require you to train a new LLM from scratch.
And training LLMs from scratch costs you know hundreds of millions of dollars and takes months of effort.
And so the >> And then Anthropic or OpenAI will come along and just eat your lunch in the next model release. >> Right, right.
And you know of course Anthropic and OpenAI and Google, they're exploring recursive self-improvement, but typically at that level of having the, you know, having to train a new model for every step of self-improvement that they do.
>> I mean, that seems like actually the, like, defining thing that a startup really, really wants.
Like, I know that I want to take advantage of whatever the next model is, but the second you're in fine-tuning land, I'm spending, you know, millions to hundreds of millions of dollars, and then guess what?
Like, it I just lit it on fire cuz you know, the next version of the frontier model comes out, and I'll never catch up.
Whereas, like, working with your systems means that I will always have the thing that is better than the thing that's out of box, and that's sort of, like, the holy grail.
>> Yeah, we think that this is incredibly valuable to anybody who's building on top of large language models.
We don't view the, you know, the frontier models as competitors.
They're, you know, they're the ones that were using the stilts, you know, building stilts to stand on top of.
But, if we didn't have that that foundational layer, then, you know, Poetic couldn't exist.
>> Yeah, I mean, being the smartest model, you know, it's a game of inches, actually.
And like, so those inches matter a lot.
How do we actually get started?
I mean, you've built something that basically any startup could use that it's sort of like stilts, really.
>> We have built a system that can automatically generate systems for your particular problem that will always outperform the underlying language models.
And without kind of the massive expense, as you're saying about the bitter lesson, where, you know, what would you what would you have done without Poetic?
You probably would have said, "Okay, we're going to first collect a large data set, you know, like, tens of thousands of examples for our particular problem that we're working on, and we're going to fine-tune, you know, the best model we can put get our hands on.
Maybe that's, you know, one of the frontier models, or maybe it's an open weights model.
Doesn't particularly matter.
You're going to spend a lot of money on that fine tuning.
The the computer is so expensive.
And then at the end of it, you have something that you know, works better than the thing that you fine tuned on top of, but by then a new model has come out and it's better than the thing that you fine tuned.
You know, you fine tuned you know, like three years ago on top of GPT 3.
5 or whatever and then GPT 4.
0 four comes out and it just blows you out of the water.
And so are you going to do that again or are you going to go out of business?
And like in some cases the the latter.
With Poetic, what we end up giving you is a you know, people are calling these things harnesses now, but you know, a word agentic system or whatever you want to call it that sits on top of one or more language models and it just performs better than them.
And when the new model comes out, that same harness is perfectly compatible with it.
And you don't need to change anything to get the you know, an even bigger performance bump.
Additionally, we can you know, continue to optimize for this new model, whatever the new model is that you want to use and you know, make it even better.
But you you don't lose out on the you know, hundreds of millions of dollars.
In fact, we do this so much more cheaply than fine tuning would cost as well.
>> And you've done this actually a bunch of times, right?
Like I remember when you first came out with your paper in December of last year, you shot to the top of Arc AGI V2.
And then you've done this a bunch of times for other benchmarks too.
What you know, what was that like?
>> Arc AGI was this was kind of yeah, our us coming out of stealth letting people know that we could tackle these really hard problems.
And in particular, you know, we wanted to show that our system could generate these what we call the you know, we call our system like the Poetic meta system can generate reasoning systems that are highly effective.
Gemini 3 DeepMind had just come out and they were you know, really quite uh, dramatically at the top of the leaderboard at 45% uh, and 2 days later we released our results where, um, we were showing that we could get, uh, a lot higher than that.
>> So they come out with soda and then you come in right above them every single time, which is like wild to see, honestly.
That's what it's like to have stilt, you know, like whatever model comes out, you can be taller than that one with poetic, which is like that's so awesome.
>> Yeah, so the interesting thing is that, uh, we were half the cost of Gemini 3 Deep Think because we were building on top of Gemini 3 Pro, which is a much cheaper model, um, but we still got, uh, in the end a 9% point improvement on the official verification.
So they were at 45% and we were and like $70 something and we were at, uh, 54% and $32 per problem.
>> So recently you guys just announced some incredible results for Humanity's last exam.
Can you tell us more about those?
>> Humanity's last exam is, uh, a set of 2,500 really, really hard questions written by experts in, uh, many different domains.
They're they're meant to be, uh, challenging even for, uh, PhDs in those fields.
AI hasn't passed it yet, uh, but we got to 55%, which is almost 2 percentage points higher than the the previous, uh, state of the art, which came out just last week, uh, from, uh, Anthropic with Claude Opus 4. 6. Uh, they got 53. 1% and we got 55% on it.
>> And one thing that, uh, Humanity's exam doesn't publish is the cost of getting those results.
In your case, this run was done with less than around six figure, how much was it?
>> We didn't publish any, uh, cost for this, but, uh, I can say that the the optimization cost us less than 100K, yeah.
>> Which is impressive because each of these big foundation models train runs are in the hundreds of millions of dollars and you guys as a company you're only seven people? >> That's right. Yeah.
Yeah, it's seven seven research scientists and research engineers. Yeah. >> That's impressive.
And I think the thing that's very interesting about your approach is sort of taking a very scientific approach to the emergent behaviors that a lot of the best founders are doing with models.
I think a lot of founders that get very good results for agents, they treat the underlying model as a common layer that you can switch in between and there's a certain task for example for GPT-5.
2 like very hard to verify bugs get sent to that versus architecture that gets sent to Claude 4.
6 but you're kind of doing this automatically instead of having a human conducting is very impressive.
I think there's something more special going on underneath.
Can you tell us a bit about how it works?
>> Yeah, it sounds magical.
So, what can you tell us?
>> Right, you're you're So, you're getting at a core a really core thing, you know, these harnesses, they are code prompts, data, you know, built on top of one or more language models, right?
And so, this is something that in principle you can build by hand um or with like cloud code or whatever.
But in in practice it takes a lot of work to do these to you know, have all the insights to make this to make these work well.
And so, the core technology that we've developed at Poetic is recursive self-improvement.
So, we we have a recursively self-improving system which we call the Poetic meta system.
The output of that system is systems that solve hard problems where a hard problem is you know, something that you if you gave it to GPT-5.
2 it would struggle to give you reliable robust result, you know, just to use an example.
So, this is a very big advantage for us.
We can generate these systems in a much more automated manner, which means that we can do it much more quickly and much more cheaply than if you hired a team yourself to try to make your own, you know, your own agent to solve your particular task.
solve your particular task. But not only that, um since, you know, this is really an automated optimization process, if you already have done that work, you, you know, you're a you're a startup that's like going after a particular vertical and you've put together, you know, you
think you understand your problem pretty well, you've put together your agent, and you you uh, you know, maybe it's working pretty well, but you know you can get something better or you really need something better, um then you can bring that to us uh, and we can optimize uh, that entire agent or pieces of that agent. We could optimize just the
We could optimize just the prompts, just the reasoning strategies.
Uh, there's a lot of different things that we can do uh, depending on your particular needs.
>> It sounds like this is a complete different paradigm than RL because we went through the S-curve of regular pre-training, RL with when OpenAI released 01, and now this feels like a new one. It sounds special.
It sounds It rhymes a lot with our RNNs, which is a whole different paradigm than our than RL, right?
>> It's going to depend on the particular task, the particular type of problem that we're going after, that we're trying to solve, um and the underlying models that we're working with.
But uh, effectively you could say like each model or each set of models that we're working with will have its their own uh, S-curve.
The Poetic system, the Poetic meta system itself is also going to have its own S-curve.
And so as the Poetic meta system gets better and as the underlying models get better, you you'll find that the uh, you know, the S-curve that you're dealing with keeps shifting higher and higher until ultimately either you saturate or like >> Reach AGI?
>> Yeah, reach AGI, reach super intelligence is, yeah.
>> Given it's still to you might like hit the ceiling first then.
>> That's the goal, right?
You want to hit the ceiling first with Poetic.
>> I think a lot of startups that we work with um and then in my spare time I you know do a bunch of context engineering.
And then the thing is we're sort of like tuning it, tuning evals, tuning like we're context stuffing ourselves.
What does that even feel like to have a you know recursively self-improving version of like prompt engineering context engineering?
>> We don't spend a lot of time looking at the particular data that we're working with.
Uh instead we're letting the Poetic meta system look at that data and and so like the meta system you know if it if it thinks that it needs to put more things in the context, it do more context stuffing or whatever it will it will do that.
If it needs to like generate a bunch of examples um uh to get the get better performance it'll do that for you, right?
It it was pretty interesting to look at the um prompt outputs in particular I'd say for RKG I and that uh you know I think you can read those and say well that's not what a human would have written uh pretty clearly and and it's you know there's some unexpected stuff and you know it made some really simple examples and one of the examples is actually wrong.
Uh but we didn't change it.
We're like well this is you know this is the thing that it output we'll just leave it be.
Um you know we don't want to go in and monkey around with things.
And so uh historically in machine learning you always you know it's like the the rule was you have to know your data set really well.
Um but now we're kind of outsourcing that to the AI itself where the AI is the it's the AI's job to understand the data set and figure out where are the failure modes um and where are the kind of robust reasoning strategies that uh the model that that the agent could uh use um to get better performance.
>> of it is like much you it the output is much better prompts and then how much of it is like the harness itself uh context stuffing or summarizing in the right way or re-ranking in the right way so that like you have some number of like mega LLM calls and then how do you get the most out of um, each of those calls?
>> Yeah, and so that definitely varies per problem, but what we've seen in fact our our last paper at DeepMind was not doing this recursive self-improving stuff, but we were we were showing that you could build these harnesses manually to solve really hard problems.
And what we saw is there is that you know, we manually optimized the prompts really hard for these very hard problems and that got us a little bit of the way.
In this particular case, you know, the the hardest the hardest task we were working on we got like to 5% performance with Gemini 1. 5 flash. This was a while ago.
And then when we added on the the reasoning strategies, we went from 5% to 95%. >> Oh my god.
>> And so this is typically what we see, you know, like everybody's out there kind of doing some amount I wouldn't say everybody, but many people are out there kind of doing some amount of automated prompt prompt optimization.
You know, Jeppa is this very popular paper.
Everybody's kind of re-implementing that.
That will get you some performance improvements, but it's very far from everything that you can get if you actually think about these reasoning strategies that are really going to be written in code rather than in just better prompts.
>> So, if startups want to use Poetic to put their agent on stilts, what should they do?
>> Yeah, so right now we haven't released anything yet, but if you go to poetic.
ai, there is a button you can click to get sign up for early access.
And if you're a startup or a company who has a really hard problem and you've tried everything that you can to make it reliable and robust and you just can't get all the way there, you you need something more, then let us know.
We're looking for problems like that.
So, just tell us tell us what it is that you're working on and we'll reach out.
You'll be the first to know when we're when we're ready to work with you.
>> I mean, if you're at the top of humanities last exam, then I mean, that's that's pretty big.
So, it's You're all the you're already all the way out there at Soda, and then I guess the Stilts basically let any agentic company become Soda. >> That's the idea. Yeah, yeah.
And you know, we view the Arcade GI results and the humanities last exam results as showing kind of two different capabilities that we have.
We can really improve your reasoning, and we can really improve deep knowledge extraction from these models.
>> And then you're just totally vaccinated against the bitter lesson. >> Exactly.
>> YC's next batch is now taking applications. Got a startup in you? Apply at ycombinator. com/apply.
It's never too early, and filling out the app will level up your idea. Okay, back to the video.
>> Slight sort of change change of topic, but something I was curious about.
So, you arrived at Google over a decade ago when they acquired your first YC startup at Portable.
At Portable was porting mobile apps cross-platform, right?
Like Android or whatever.
It's quite different to recursive self-improving AGI.
How did you make that leap?
What happened once you got to Google?
What made you think that you maybe wanted to shift down and do something different?
And just would love to hear that story.
>> The acquisition was this amazing opportunity to reflect on what I really wanted to be doing next, right?
Like Google was in the you know, itself is a place where you can do so many different things.
So, I spent some time thinking about where where I wanted to go next in in in my journey.
I realized that the problems that I was most uh about were really actually AI and and robotics.
Um and it's the best people in the world, many of them in those fields were at Google at the time.
Uh and so I went and talked to them.
They uh let me come join, you know, a new AI robotics team uh in Google research, which was this amazing opportunity for me since that wasn't my background.
My background was like uh computer security and then this cross-platform mobile, you know, it's systems building uh stuff.
I was able to join this team and I'll I'll tell you the truth that I very quickly realized that uh hardware is hard and I didn't really want to be doing robotics.
It was more aspirational at that moment uh but I was really um uh passionate about machine learning.
So I just uh I made a very hard switch into just doing machine learning research uh and did that for, you know, about a decade at Google and and then uh Google and then DeepMind.
>> What's maybe some advice that you have today for engineers who want to get into sort of more of the AI side, probably the applied AI and build startups around AI?
Like how should they think about that?
>> You know, the world is changing so quickly.
I I this is probably a little bit obvious, but you should just try things and and like every day uh do something uh do something with AI.
Always try to push yourself to find the boundaries of what they're capable of.
Uh and uh and build the things that you that you want to build, right?
Um uh even even for me, you know, last summer I took a weekend and used um GPT-5 to help me build an iPhone app.
I hadn't I hadn't done that in a decade.
And yeah, it's so fast and so easy and that was, you know, that was an you know, an age ago.
That was like eight months ago.
Uh now it's even faster and easier. Don't limit yourself.
Like anything that you imagine you should just try to use AI and see how far you can get with it and you'll be, you know, making the world better.
>> That's all we have time for today, but Ian, thank you so much for giving us all Stilts.
We can't wait to use it at YC.
I can't wait to use it for Gary's list.
I mean, there's just so much to do, so.
>> Yeah, thank you for having me. This was a lot of fun.