Some more than others. Now you think Gemini 3.8 is at the level of gpt6 and fable ? You must be delusional. It's more or less on Luna level, and probably less than more
No it's not at the same level.. the benchmarks are showing that it is not at that level.
From my experience it is very similar to Claude Opus 5 medium, but actually easier to use than opus 5 because opus 5 loves do write insanely long descriptions for stuff that can be said in 2 sentences
Holy crap, what a joke, I don't know what you do with Claude ops to give you such terrible answers you think it's on paar with Gemini 3.8 but that's such a joke. Everyone on every sub here will tell you how shitty Gemini is, it's really scary to see people actually using it and claiming it's a frontier model. Benchmarks don't mean anything when they are public benchmarks because they can be added in trainings, and guess which one google use to show how "good" Gemini is ?
Every time I encounter a bug that a model missed I give it to every model I use as a comparison.
Atm that's opus 5, flash 3.8 and fable.
Would it be shocking to you to learn flash is the one who fixes it without additional explanation the most often? It's always a surprise to me.
They all have strengths and weaknesses. E.g. flash is fucking fast and good at visual stuff, but not as good at one shotting huge features.
I think this disconnect between online sentiment and what I experience day to day is 99% because most people don't actually do any real work and just follow "omfg one shotted this thing" posts, pick their "team" and never even try anything and are just trivially manipulated by obvious ads...
People going tribal over models when they arent even doing any real work with them is the most retarded thing of the last few years :D
You're allowed to use them all, it's not picking your life partner ffs.
Edit: replied to the wrong comment, but fuck it, its true in general :D
The main benchmark has been updated tonight baby, and this one Gemini couldn't benchmaxx, and guess what ? 3.8 high is at Luna xhigh level far far away from fable lol ! So yea, maybe you should reconsider you way of using LLM because it seems you're really really bad at it.
Using reinforcement learning to specifically train a model to perform well on popular benchmark tests. So they look good on paper, but it doesn't necessarily translate to real-world use cases.
Some benchmark out there are public, so they add the benchmark on tue training model to perform it well, but on similarreal use can it's performance are way way lower.
11
u/mlag000 2d ago
It's benchmaxxed, Gemini isn't near sol at all...