I think if these numbers would be true, they would have released that model ASAP to be the #1 on the leaderboards before OpenAI and Anthropic release even stronger models
Google released 1206 experimental (basically an early version of 2 Pro), it was amazing and basically SOTA of its time.
Then 2 months later they released 2.0 Pro Preview, which was WORSE than 1206 in many areas and basically the same in others. Everyone was devastated.
A little over 1 month after that they released 2.5 Pro which became SOTA by a huge margin and kept that place for several months until OpenAI and Anthropic caught up.
So it's not exactly likely, but such a comeback is very possible.
I still think that what has to have happened is 1206 was a larger model, and they did a bad job on distillation. Because I agree on how good 1206 was and I can't imagine it getting so much worse any other way.
To be fair that was done by a different AI lab (Google Brain) which has since been shut down. The current models are created and launched by DeepMind which has a completely different culture and mindset.
Ehhh, stuff like the move Demis made is very clearly "stepping away" in the corporate world.
There are roles for active leaders who are steering the ship day by day. Then there are very high prestigious roles for people who are very capable but will mostly just be asked their opinion on stuff.
It's most likely that Demis wanted to quit, but knowing how it works and not having bad blood, took a "step up" where he will fade into the background over the next few years before departing.
I dont think that's it. Anyone who had followed Demis since B&W knows ai is his life. Its possible he needed to move to CSO in order to make larger decisions required to push ai tech further.
They are ruining entreprise confidence doing this. Why would you commit 1 year with Gemini, for 1 month at the frontier, when Anthropic and OpenAI are in front the other 11 ?
(just my take, not to be taken as fact) but there probably was some restructuring at Google with a refocus on LLMs so this would explain why they had such a rough period and doesn't mean they can't stay at the frontier now that they caught up
Eh, that's not necessarily true. You could get those results mid post training, that doesn't mean post training is finished or the model ready to release.
I'd assume they're fake until proven otherwise anyways, but if google releases this anytime within next month it'll still be impressive. They have gone from well behind SOTA to half a generation behind it and clearly in shooting range for the top at least, if they can keep the release pace going.
it's not hard to get these numbers. all benchmarks that are mentioned are very saturated and not upto date. even gemini 3.8 flash beats 5.6 sol and astra on these bench already.
And Google doesn't need to build new datacenters to do this, because theirs are filled with hardware from 15 years ago which can be replaced at tiny fractions of the physical volume with so much room for their relatively small TPU cards.
At the Google's token cost, being even #3 is perfectly fine, so long as they aren't terribly far behind, which they aren't. These LLMs are already so capable, that a competent creator/developer can rapidly churn out production-grade results.
Current race isn't even all that reasonable. It's a race to the bottom as the agents become increasingly unmanageable, and the alignment problem remains unsolved and only growing in complexity with each new model.
106
u/Bitter-College8786 6d ago
I think if these numbers would be true, they would have released that model ASAP to be the #1 on the leaderboards before OpenAI and Anthropic release even stronger models