
5 segments available
ARC-AGI is redefining how to measure progress on the path to AGI - focusing on reasoning, generalization, and adaptability instead of memorization or scale. During this month's NeurIPS 2025 conference, YC's Diana Hu sat down with ARC Prize Foundation President Greg Kamradt to find out why most AI benchmarks fail, how ARC-AGI reveals the limits of today’s models, and why measuring intelligence may be harder than building it. Apply to Y Combinator: https://www.ycombinator.com/apply Chapters: 00:11 — What ARC Prize is and why it exists 00:38 — François Chollet’s definition of AGI 01:48 — What ARC-AGI Actually Tests 02:25 — When LLMs Failed the ARC Benchmark 02:44 — The Reasoning Breakthrough 03:38 — ARC-AGI Becomes the Standard 04:20 — Vanity Metrics 04:49 — False Positives in AI Progress 06:06 — The Evolution of ARC-AGI 07:05 — Inside ARC-AGI v3 08:55 — Measuring Intelligence beyond just accuracy 10:25 — What happens if a model solves ARC-AGI?
Greg Camrad discusses the ARC Prize Foundation's innovative approach to defining and benchmarking intelligence in AI.
"I'm excited today to welcome Greg Camrad who is the president of the Ark Prize. >> That's, right. >> Thanks, for, coming, here, at, Europe's, 2025 in beautiful San Diego. >> Thank, you,, Diana. >> So,..."
This segment explores the evolution of AI performance measurement and the impact of reasoning paradigms on model advancements.
"MMLU plus, and now we have humanities last exam. Those are going super human right? Arc benchmarks, normal people can do these. And so we actually test all of our benchmarks to make sure that um norma..."
This segment discusses the pitfalls of interpreting AI adoption and performance metrics while emphasizing the true mission behind progress in artificial general intelligence.
"all these releases? >> So, it's, it's, going, really, well, that they're adopting it. Um, however, we're mindful of vanity metrics that come from there, too. So just because they use us doesn't necess..."
This segment discusses the upcoming interactive benchmark RGI 3 for measuring AGI through engaging video game-like environments without instructions.
">> three, is, all, about. >> Yes,, absolutely., So, RKGI1, came, out, in 2019. That was France proposed it. I think he made all 800 tasks himself within it which is a huge feat in in and of itself. Um..."
This segment explores the multi-faceted approach to evaluating AI intelligence through efficiency metrics, including data requirements and energy consumption.
"there's something missing still. There's something clearly missing that we need to um need new ideas for research on. >> So, there's, this, big, theme, in, terms, of measuring intelligence with human ..."