Ilya Sutskever discusses the perplexing phenomenon where AI models excel in evaluations yet fail to deliver significant economic impact. He illustrates this with an example of 'vibe coding,' where the model struggles to fix bugs, often introducing new ones instead. This raises questions about the models' training and generalization capabilities, highlighting a disconnect between evaluation performance and real-world effectiveness.
"One of the very confusing things about the models right now they are doing so well on eval but the economic impact seems to be [music] dramatically behind. It's very difficult to make sense of how can..."
Ilya Sutskever discusses the perplexing situation where AI models excel in evaluations yet fail to deliver significant economic impact. He illustrates this with an example of 'vibe coding,' highlighting how models can introduce new bugs while attempting to fix existing ones. This raises questions about the models' training and generalization capabilities, suggesting a disconnect between their evaluation performance and real-world effectiveness.
"One of the very confusing things about the models right now they are doing so well on eval but the economic impact seems to be [music] dramatically behind. It's very difficult to make sense of how can..."