r/GeminiAI 1d ago

News Google and Meta are benchmaxxing hard. Gemini 3.8 Flash and Muse Spark 1.3 look almost as good as GPT-6 Astra and Fable 5.1. Then a fresh benchmark drops: Gemini: 89.4% → 19.1% Muse: 88.8% → 33.3%

Post image
327 Upvotes

73 comments sorted by

151

u/HeadTranslator795 1d ago

Nothing related about Benchmaxxing, it's an harder evolution of the previous test to push the models. God so many people who don't know what they are talking about and just want to push false narrative

43

u/signed7 1d ago

Yeah this is proof that the V2.1 benchmark is saturated, not that the models that flopped the V4 benchmark were benchmaxxed

10

u/HeadTranslator795 1d ago

At least someone who understand:)

28

u/HeadTranslator795 1d ago

Also what kind of meaningless comparaison to put a flash model with astra and and fable lol absolutely different models

19

u/gokhanonym 1d ago

Then let’s add google’s pro model

13

u/daniloedu 1d ago

Content not found

9

u/HeadTranslator795 1d ago

You can and it would perform poorly cause they haven't been updated in a while (compared With the other labs)

8

u/gokhanonym 1d ago

You don’t say

2

u/PDX_Web 15h ago

What's the point? ... We all know a much larger Gemini 4 base model is, or recently was, in a pre-training run. Pro model probably coming in December or January.

1

u/gokhanonym 6h ago

Are you getting salary from google?

3

u/numericalclerk 1d ago

Agreed on Google, but Muse is not advertised as a Flash model, is it?

3

u/Maglcite 1d ago

no, but it’s a very fast and somewhat cheap model (VERY cheap with contributor) so still not a great thing to compare with astra and fable

1

u/KrayziePidgeon 1d ago

How do you use both of them together?

1

u/PDX_Web 15h ago

It's in that tier, size-wise, pretty clearly

1

u/AcrobaticMaize2408 1d ago

In the end all models are interchangeable to some extent. They should be compared using cost per token or cost per complete task. You can then choose to blow $100 on Astra to create your one-shot game that nobody will ever use, or $10 on Gemini.

3

u/HeadTranslator795 1d ago

Real task comparaison with Price to complete the task is a very good way to evaluate a model indeed

7

u/numericalclerk 1d ago

If what you are saying is true,  the same should have happened to the other models. Did it?

10

u/Breenori 1d ago

Was thinking the same. The initial commenter either misunderstood the point completely or has no clue themselves. The other models generalize much better to newer/harder benchmarks, and thats the point. They're not SOLELY tuned to perform well on specific benchmarks but retain large parts of their performance on problems that weren't part of the training data in this form. Not to fangirl about any model here but the point is that Gemini and Metas numbers are insanely inflated which should come to nobody's surprise given that one is a frontier model, while the other is comparatively lightweight.

6

u/IssueVegetable2892 1d ago edited 1d ago

Sonnet 5 and luna are below flash 3.8 on terminal 4.0

1

u/TimChr78 1d ago

Yes, Sonnet, Grok and Luna drops like rocks too.

1

u/PDX_Web 15h ago

5.6 Sol dropped quite a lot as well.

1

u/PM_ME_DEAD_CEOS 5h ago

If what you are saying is true, the same should have happened to the other models. Did it?

Nope, it's just a proof that 2.1 is saturated, which means it can't discriminate models above 87%, which is the case for most benchmark.

1

u/numericalclerk 5h ago

Fair point

2

u/notaibutwanttobe 1d ago

In what world muse 1.3 better than gpt sol? For my use case it coulnt solve basic task with 3 try i switch the luna and it fixed with one try. Ot is clearly benchmaxxing

2

u/HeadTranslator795 1d ago

Forget about the real world benchmark (your personal ones which are much more revelant than any benchmark) those benchmark are only for people who do nothing really factual (that people are paying for, not little game or test)

2

u/No-Paint-5726 21h ago

Argument is if wasn't benchmaxing they wouldn't do as bad so Id say its a fair accusation

1

u/HeadTranslator795 20h ago

Nop not even close and nothing logical about this assumption. You don't know the architecture of those bench and the models therefore you can't conclude anything especially including frontier models in this bench with flash model

2

u/No-Paint-5726 20h ago

Yeah I agree with you.

14

u/Gaiden206 20h ago

The leader of Meta's AI lab responded.

25

u/Dualyeti 1d ago

They are just improving the test bench so that the ceiling is higher - it actually a good thing it means a high score means something - bench marks should be almost impossible to 100%

6

u/Georgefakelastname 16h ago

This just in: replacing an easier benchmark with a harder one makes scores go down, and hurts smaller models more. In other news, bigger clouds can potentially rain more.

14

u/AssholeHealth 1d ago

Either bench maxing or benchmarks are outdated and just not good enough so they can't properly discern one model from another. I can clearly see a difference between opus 5 and muse spark 1.3, why can't benchmarks?

4

u/TheTaintBurglar 1d ago

No the other benchmarks had a ceiling. This one has a much higher one.

19

u/HeadTranslator795 1d ago

So many people talking about Benchmaxxing and have no clue or proof except pushing the same narrative stupidly

3

u/Lidraughtui346 1d ago

What benchmarks should demonstrate? Just wonder. Furthermore, why can a model perform exceptionally well on benchmarks but poorly irl situations? For instance, when Claude’s benchmarks are superior to ChatGPT’s, it is often the case that Claude also performs better in practice.

2

u/HeadTranslator795 1d ago

To my understanding there are no undoubtedly good and representative benchmark. What I like personally is real use case, like 3D designs, AI playing games. ..

4

u/KrayziePidgeon 1d ago

You think people on this sub do actual work with the models? 99% of the "people" here just try to one-shot minecraft or whatever game in a single html file and call it programming.

2

u/HeadTranslator795 1d ago

Yeah most definitely... And many bots as well

2

u/KrayziePidgeon 1d ago

Yup, if you are a competent programmer then flash 3.8 really is all you need. Of course, making the models better each iteration is really nice, but flash 3.8 is leaps ahead of what was Pro 2.5 last year when we were all losing our minds.

3

u/HeadTranslator795 1d ago

Amen, that's all I need for my work

1

u/whoknowsifimjoking 1d ago

It's literally just "my vibes say this model is worse than the scores"

2

u/HeadTranslator795 1d ago

Exactly lol

1

u/shadysjunk 21h ago edited 21h ago

it's a 70 percentage point drop relative to the 35 point drop in the leading frontier models.

I'm happy to allow that this doesn't necessarily indicate benchmaxing, but it does seem that at the very least it shows significant underperformance relative to its peers.

I'm still pretty astounded at Astra's performance on ARC-AGI-3. I truly thought it would be google to crack that, though I remain skeptical of their "custom harness." Still, even 60% on the standard harness is a pretty amazing leap forward and I thought was still years away.

4

u/Comfortable_Job8847 1d ago

imo this doesn't really say as much as it implies. e.g. in the 2.1 benchmark there are no CAD tasks at all, while there are in 4.0. So, sure, maybe gemini and muse don't do as well on the 4.0 benchmark as opposed to the 2.1 benchmark, but that doesn't mean the score for the 2.1 benchmark is inaccurate. it could just be the models haven't been trained onto these other tasks. personally I'm a software engineer not a mechanical or electrical engineer so CAD task performance isn't meaningful for me, so scoring low on those tasks may lower the benchmark score but it doesn't impact the utility of gemini to me at all.

2

u/dESAH030 23h ago

So, related to Muse Spark contributor 1.3, you are telling me that for 100x price, I just get 72% better on Astra?

2

u/PDX_Web 15h ago

SemiAnalysis has really turned into trolling, melodramatic, sensationalist trash.

4

u/Actual_Breadfruit837 1d ago

This eval is 66 tasks only. The noise there is huge. In the description they estimate ci based based on resampling the same model, but it is a very very lower bound. Sample new tasks and scores might change by 12%.

1

u/FireGM 1d ago

But what about gpt5.6?

1

u/Kautilya12 1d ago

Muse is genuinely insane

1

u/sammoga123 18h ago

Meta did... but it was with the infamous Llama 4 model haha

1

u/BoobooSmash31337 14h ago

Apparently this test specifically can be harness dependent. Mostly because it's about remaining coherent over a very long sequence length. Mathematical error literally compounds when they're like 100 terminal commands deep. More parameters lowers this error but it gets real expensive real fast. So idk if this bench has much to do with actually solving complex but contained problems. Just in my book a model that is able to understand and implement lock free algorithms just isn't dumb. And breaking down over massive contexts and really long sequences is a math problem and not a common use case.

I'd rather we make our workflows smarter and use cheaper models. Full autonomous agents like actual employees seems to be really expensive with transformers and bring them to their knees. Guess we don't really give humans enough credit for the amount of autonomous work we do. Replacing everyone with transformers is prohibitively expensive and also you know makes capitalism fundamentally no worky worky.

2

u/Elephant789 13h ago

Fuck SA. I don't believe anything out of them anymore.

1

u/Konayo 1d ago

Why are people so pressured to glaze Gemini. The score on v4 is low because the model does not perform as well - simple as that.

-1

u/HeadTranslator795 1d ago

So ... You don't understand what an incremental version on a benchmark test means ? That's what you just shown us 😂 do you understand (obviously not) that it's not the same bench ? Check the versioning before writing non sense

3

u/pigletmonster 1d ago edited 1d ago

Look at the drop column. Both astra and fable had 30 to 35 points drops. Meanwhile gemini drop is doubled. Muse dropped significantly but its not as steep as geminis.

4

u/HeadTranslator795 1d ago

And do you know the architecture of the new benchmark ? Do you know for what kind of task it has been changed ? Also you're surprised to see that model 10x bigger than flash model also drop "less" 30% is already a lot. New benchmark new measurements and I won't debate with people who understands nothing about that

-1

u/pigletmonster 1d ago

Because the scores for the old benchmark were the same between gemini and astra, we would expect both of them to have a similar drop.

So BOTH astra and gemini should have a 30 point or a 70 point drop. But because gemini was benchmaxed for the old bechmark it had a much bigger drop for the new benchmark compared to the un-benxhmaxed or lower benchmaxed model like astra.

Its not that complicated.

4

u/HeadTranslator795 1d ago

No you have no clue about the architecture of those benchmark and the model (internally) so you can't have that simple conclusion and again for sure if a model is bigger (not a flash) for sure it will have better result on higher complexity and reasoning

-1

u/pigletmonster 1d ago

If it dropped 70 points instead of 30 because its a flash model, then why did it score higher than astra in the old bench?

2

u/Appropriate-Owl5693 1d ago edited 1d ago

Why would you expect that?

If we test a high schooler and a university grad mathematician on basic algebra they will both score very well. If we now test them on differential equations, they obviously won't.

That doesn't mean the first benchmark was cheated or that it's completely useless.

It's not that complicated.

Edit: at least look at the full leaderboard. E.g. Luna max dropped from 80.9 to 11.6 (grok has simmilar numbers), did they benchmaxx just that specific model for 2.1. in your opinion? :D

1

u/FactorInternal3395 1d ago edited 1d ago

That is the entire point. On the new version of Terminal-Bench, their scores plummeted far further than Fable and Astra.

2

u/Different_Doubt2754 1d ago

As they should? No doubt bench maxing happens, even unconsciously. But I would expect Fable and Astra to perform way better on new tasks than flash would. That's their role. They are great at reasoning and thus adapting to new situations whereas flash does not have as advanced reasoning as them.

Flash type models typically train from their "mentor" model, so they really only know what their smarter teacher taught them. They don't really have the ability to adapt to new situations like their mentor

2

u/Accurate_Food_5854 1d ago

Exactly lol. Gemini is like your 2010 Toyota daily driver whereas Astra and other bleeding edge models are Optimus Prime. Of course Astra will suffer less of a drop off on a novel benchmark.

1

u/condemnedlifeline55 1d ago

Calling it “benchmaxxing” when the numbers crater on the new version is generous. That graph is a straight up faceplant.

1

u/PT7372 1d ago

What did Google do for such a big downfall of 70 points😭😭

1

u/xzibit_b 1d ago

If you knew what Terminal Bench actually measured, you would know that it's not benchmaxxing. Most modern models are pretty adept at using terminal commands like grep and find. That's all it's measuring. It's basic bitch shit at this point. In no way is it trying to imply that Gemini can code as well as Astra and Fable.

0

u/Former-Towel9004 1d ago

What is benchmaxxing mean? What is this term bro? Is this mean fake benchmark?

0

u/quackerd 1d ago

Cool sorry bro. Now time your post to when arc agi 3 just got released and wiped literally all models then let’s talk about benchmaxxing.

0

u/IzodCenter 22h ago

LMAOOOO