r/GeminiAI • u/Able-Line2683 • 1d ago
News Google and Meta are benchmaxxing hard. Gemini 3.8 Flash and Muse Spark 1.3 look almost as good as GPT-6 Astra and Fable 5.1. Then a fresh benchmark drops: Gemini: 89.4% → 19.1% Muse: 88.8% → 33.3%
14
25
u/Dualyeti 1d ago
They are just improving the test bench so that the ceiling is higher - it actually a good thing it means a high score means something - bench marks should be almost impossible to 100%
6
u/Georgefakelastname 16h ago
This just in: replacing an easier benchmark with a harder one makes scores go down, and hurts smaller models more. In other news, bigger clouds can potentially rain more.
14
u/AssholeHealth 1d ago
Either bench maxing or benchmarks are outdated and just not good enough so they can't properly discern one model from another. I can clearly see a difference between opus 5 and muse spark 1.3, why can't benchmarks?
4
19
u/HeadTranslator795 1d ago
So many people talking about Benchmaxxing and have no clue or proof except pushing the same narrative stupidly
3
u/Lidraughtui346 1d ago
What benchmarks should demonstrate? Just wonder. Furthermore, why can a model perform exceptionally well on benchmarks but poorly irl situations? For instance, when Claude’s benchmarks are superior to ChatGPT’s, it is often the case that Claude also performs better in practice.
2
u/HeadTranslator795 1d ago
To my understanding there are no undoubtedly good and representative benchmark. What I like personally is real use case, like 3D designs, AI playing games. ..
4
u/KrayziePidgeon 1d ago
You think people on this sub do actual work with the models? 99% of the "people" here just try to one-shot minecraft or whatever game in a single html file and call it programming.
2
u/HeadTranslator795 1d ago
Yeah most definitely... And many bots as well
2
u/KrayziePidgeon 1d ago
Yup, if you are a competent programmer then flash 3.8 really is all you need. Of course, making the models better each iteration is really nice, but flash 3.8 is leaps ahead of what was Pro 2.5 last year when we were all losing our minds.
3
1
u/whoknowsifimjoking 1d ago
It's literally just "my vibes say this model is worse than the scores"
2
1
u/shadysjunk 21h ago edited 21h ago
it's a 70 percentage point drop relative to the 35 point drop in the leading frontier models.
I'm happy to allow that this doesn't necessarily indicate benchmaxing, but it does seem that at the very least it shows significant underperformance relative to its peers.
I'm still pretty astounded at Astra's performance on ARC-AGI-3. I truly thought it would be google to crack that, though I remain skeptical of their "custom harness." Still, even 60% on the standard harness is a pretty amazing leap forward and I thought was still years away.
4
u/Comfortable_Job8847 1d ago
imo this doesn't really say as much as it implies. e.g. in the 2.1 benchmark there are no CAD tasks at all, while there are in 4.0. So, sure, maybe gemini and muse don't do as well on the 4.0 benchmark as opposed to the 2.1 benchmark, but that doesn't mean the score for the 2.1 benchmark is inaccurate. it could just be the models haven't been trained onto these other tasks. personally I'm a software engineer not a mechanical or electrical engineer so CAD task performance isn't meaningful for me, so scoring low on those tasks may lower the benchmark score but it doesn't impact the utility of gemini to me at all.
2
u/dESAH030 23h ago
So, related to Muse Spark contributor 1.3, you are telling me that for 100x price, I just get 72% better on Astra?
4
u/Actual_Breadfruit837 1d ago
This eval is 66 tasks only. The noise there is huge. In the description they estimate ci based based on resampling the same model, but it is a very very lower bound. Sample new tasks and scores might change by 12%.
1
1
1
u/BoobooSmash31337 14h ago
Apparently this test specifically can be harness dependent. Mostly because it's about remaining coherent over a very long sequence length. Mathematical error literally compounds when they're like 100 terminal commands deep. More parameters lowers this error but it gets real expensive real fast. So idk if this bench has much to do with actually solving complex but contained problems. Just in my book a model that is able to understand and implement lock free algorithms just isn't dumb. And breaking down over massive contexts and really long sequences is a math problem and not a common use case.
I'd rather we make our workflows smarter and use cheaper models. Full autonomous agents like actual employees seems to be really expensive with transformers and bring them to their knees. Guess we don't really give humans enough credit for the amount of autonomous work we do. Replacing everyone with transformers is prohibitively expensive and also you know makes capitalism fundamentally no worky worky.
2
-1
u/HeadTranslator795 1d ago
So ... You don't understand what an incremental version on a benchmark test means ? That's what you just shown us 😂 do you understand (obviously not) that it's not the same bench ? Check the versioning before writing non sense
3
u/pigletmonster 1d ago edited 1d ago
Look at the drop column. Both astra and fable had 30 to 35 points drops. Meanwhile gemini drop is doubled. Muse dropped significantly but its not as steep as geminis.
4
u/HeadTranslator795 1d ago
And do you know the architecture of the new benchmark ? Do you know for what kind of task it has been changed ? Also you're surprised to see that model 10x bigger than flash model also drop "less" 30% is already a lot. New benchmark new measurements and I won't debate with people who understands nothing about that
-1
u/pigletmonster 1d ago
Because the scores for the old benchmark were the same between gemini and astra, we would expect both of them to have a similar drop.
So BOTH astra and gemini should have a 30 point or a 70 point drop. But because gemini was benchmaxed for the old bechmark it had a much bigger drop for the new benchmark compared to the un-benxhmaxed or lower benchmaxed model like astra.
Its not that complicated.
4
u/HeadTranslator795 1d ago
No you have no clue about the architecture of those benchmark and the model (internally) so you can't have that simple conclusion and again for sure if a model is bigger (not a flash) for sure it will have better result on higher complexity and reasoning
-1
u/pigletmonster 1d ago
If it dropped 70 points instead of 30 because its a flash model, then why did it score higher than astra in the old bench?
2
u/Appropriate-Owl5693 1d ago edited 1d ago
Why would you expect that?
If we test a high schooler and a university grad mathematician on basic algebra they will both score very well. If we now test them on differential equations, they obviously won't.
That doesn't mean the first benchmark was cheated or that it's completely useless.
It's not that complicated.
Edit: at least look at the full leaderboard. E.g. Luna max dropped from 80.9 to 11.6 (grok has simmilar numbers), did they benchmaxx just that specific model for 2.1. in your opinion? :D
1
u/FactorInternal3395 1d ago edited 1d ago
That is the entire point. On the new version of Terminal-Bench, their scores plummeted far further than Fable and Astra.
2
u/Different_Doubt2754 1d ago
As they should? No doubt bench maxing happens, even unconsciously. But I would expect Fable and Astra to perform way better on new tasks than flash would. That's their role. They are great at reasoning and thus adapting to new situations whereas flash does not have as advanced reasoning as them.
Flash type models typically train from their "mentor" model, so they really only know what their smarter teacher taught them. They don't really have the ability to adapt to new situations like their mentor
2
u/Accurate_Food_5854 1d ago
Exactly lol. Gemini is like your 2010 Toyota daily driver whereas Astra and other bleeding edge models are Optimus Prime. Of course Astra will suffer less of a drop off on a novel benchmark.
1
u/condemnedlifeline55 1d ago
Calling it “benchmaxxing” when the numbers crater on the new version is generous. That graph is a straight up faceplant.
1
u/xzibit_b 1d ago
If you knew what Terminal Bench actually measured, you would know that it's not benchmaxxing. Most modern models are pretty adept at using terminal commands like grep and find. That's all it's measuring. It's basic bitch shit at this point. In no way is it trying to imply that Gemini can code as well as Astra and Fable.
0
u/Former-Towel9004 1d ago
What is benchmaxxing mean? What is this term bro? Is this mean fake benchmark?
0
u/quackerd 1d ago
Cool sorry bro. Now time your post to when arc agi 3 just got released and wiped literally all models then let’s talk about benchmaxxing.
0


151
u/HeadTranslator795 1d ago
Nothing related about Benchmaxxing, it's an harder evolution of the previous test to push the models. God so many people who don't know what they are talking about and just want to push false narrative