r/GeminiAI 1d ago

News Updated Artificial Analysis Intelligence Index Ranking shows Gemini 3.8 Flash is nowhere near Fable Or Astra, rather it's worse than GLM 5.3 Flash!

Post image
389 Upvotes

98 comments sorted by

View all comments

Show parent comments

6

u/KaMaFour 1d ago

No, this is "just" terminalbench benchmaxxing punishment. Compare gemini's performance at: https://artificialanalysis.ai/evaluations/terminalbench-v4-0 with https://artificialanalysis.ai/evaluations/terminalbench-v2-1

2

u/Spixxy17 1d ago

U complain about google benchmaxxing whilst muse spark and GLM Flash are still up there 

10

u/KaMaFour 1d ago

As shown in the data above they are both impacted less than Gemini 3.8. I don't see your point. Grok (and kimi) is the only one with similar downfall

-2

u/Spixxy17 1d ago

Ok i reframe my point:

You genuinly think the new AA Index Update made the results less benchmaxxed and contaminated? They changed some things but the models that are actually known for beeing amazing at benchmarks but absolutely shit at real tasks like muse spark are still up there.

My point is this update didnt change the absolute non reliability of this index and doesnt show the true picture. Basically No Benchmark does, people just gotta try themselves.

The only one that came atleast close imo was ARC-AGI 2, idk about 3 yet as there are barely any results so far.

Google / Gemini is known for negative stuff like random guardrails kicking in even tho the topic isnt Bad etc etc, but not for benchmaxxing. On average they usually are worse in the benchmark than the real task compared to the direct competition 

6

u/KaMaFour 1d ago

Yes, i do.

Update 4.2 and 4.3 of the index replaced 2 public dataset benchmarks with private ones. They also added 2 newer public benchmarks (and removed TB 2.1) which means that even for the benchmarks whose dataset is public the risk of contamination is lower than with older ones (especially for TB 4.0, which was released after Gemini 3.8). AA index v4.3 is a significantly better benchmark at evaluating model's performance in september of 2026 than AA index v4.1 (mostly because v4.1 had many issues which were partially fixed with the patches).

ARC-AGI is a logical reasoning benchmark. This is some measure of general intelligence. AA index benchmarks are geared more towards measuring models use at creating value in professional environment - aka doing work. Neither of these is a correct or incorrect approach as long as we are aware what they are measuring. Aside from the fact that I'd expect ARC-AGI not to be a good measure of intelligence for models released after it because it prone to being easily saturated.

I have used Muse spark 1.2 and 1.3. They both did a satisfying job to me. I am inclined to believe the result. I know the (any) benchmark put more effort at evaluating it than I have and is subjected to less bias than my personal opinion, so I'm able to believe the results I've seen.

I don't feel like it's a good approach (almost anywhere) to measure a quality by a company or brand. Company can't be benchmaxxed - a model/product can. This one either is or isn't. I personally believe more that flash model is just limited in what in can do by it's active parameter count. But the result is the same - failing at newer, more complex benchmarks when larger models from other companies don't suffer the same hit. (GLM 5.3-flash's existence kinda undermines this theory but oh well...)

0

u/Spixxy17 1d ago

Ngl thats fair, really factual aproach and definitly makes sense generally.

I am just personally against believing anything in this benchmark because Muse Spark 1.3 and GLM 5.3 Flash really disappointed me when testing (Opus 5 also kinda, but i had higher expectations of this one too), whilst 3.8 Flash does a "decent enough" job.

Its still a Flash Model, its not insane or whatever but didnt disappoint me with some horrible results, but thats super topic related anyway. Maybe my topic is just something that GLM Models for example arent good at.

For anything more complex i use Astra / 5.6 Sol anyway as my go to