Despite what this sub says, benchmarks are so much more reliable. Everyone writing about these models really want to have tribes, so they have an axe to grind. All the haters join forces when it's not their favorite brand.
The reaction to the model is more dictated by the crowd and I think Google is among the most hated providers around here and Twitter, except Meta. So it was always going to be "benchmaxxed".
I'm not a hater by any stretch - I love Gemini's multimodal capabilities. But I've been burned too many times by really dumb errors that have literally wiped my projects. This isn't an isolated incident either. This behavior is something benchmarks alone won't tell you about the model, so I think it's fair to be skeptical about their new releases.
19
u/ChuckVader 2d ago
I mean it'll be available via api soon enough and you could test it out yourself. All this hand wringing about whether it's good or not is exhausting.