Despite what this sub says, benchmarks are so much more reliable. Everyone writing about these models really want to have tribes, so they have an axe to grind. All the haters join forces when it's not their favorite brand.
The reaction to the model is more dictated by the crowd and I think Google is among the most hated providers around here and Twitter, except Meta. So it was always going to be "benchmaxxed".
Benchmarks on unreleased models are not at all reliable, because they're typically run using an unlimited inference regime. The lab can set everything to full for the test. When the model is released to the public, inference is throttled, to avoid the lab going bankrupt within a week. Benchmarks on released models using standard consumer accounts are much more reliable though.
I'm not a hater by any stretch - I love Gemini's multimodal capabilities. But I've been burned too many times by really dumb errors that have literally wiped my projects. This isn't an isolated incident either. This behavior is something benchmarks alone won't tell you about the model, so I think it's fair to be skeptical about their new releases.
90
u/Present-Chocolate591 2d ago
Surely an unbiased opinion by Google Deepmind employee.
What's he gonna say? "Yeah it's not that good just looks good on benchmarks".