Sundar Pichai on AI deployment frontier

    by Evira: Sundar Pichai

    Sundar Pichai says the AI models we actually use can trail Google’s peak capability by months, not because the research isn’t there, but because the deployment frontier is shaped by cost, latency, and scale. Today’s Pro can become tomorrow’s Ultra once the economics improve, and he notes Gemini Flash may be more impactful than Pro since low latency can beat a small intelligence gap. Benchmarks are getting worse at capturing real-world usefulness, and the next big leap might already exist, just waiting on engineering to make it usable

    Transcript (en)

    Do you feel any of those limitations or is it full steam ahead on all fronts? I think it's compute limited in this sense, right? Like, you know, we can all, part of the reason you've seen is to flash, nano flash and pro models, but not an ultra model. It's like for each generation, we feel like we've been able to get the pro model at like, I don't know, 80, 90% of ultra capability. but ultra would be a lot more like slow and a lot more expensive to serve. But what we've been able to do is to go to the next generation and make the next generation's pro as good as the previous generation's ultra, but be able to serve it in a way that it's fast and you can use it and so on. So I do think scaling laws are working, but it's tough to get at any given time, the models we all use the most is maybe like a few months behind the maximum capability we can deliver, right? Because that won't be the fastest, easiest to use, etc. Also, that's in terms of intelligence. It becomes harder and harder to measure performance in quotes because, you know, you could argue Gemini Flash is much more impactful than pro. Just because of the latency, it's super intelligent already. I mean, sometimes like latency is maybe more important than intelligence, especially when the intelligence is just a little bit less in Flash, it's still an incredibly smart model. And so you have to now start measuring impact. And then it feels like benchmarks are less and less capable of capturing the intelligence of models, the effectiveness of models, the usefulness, the real world usefulness of models.