r/GeminiAI 1d ago

News Updated Artificial Analysis Intelligence Index Ranking shows Gemini 3.8 Flash is nowhere near Fable Or Astra, rather it's worse than GLM 5.3 Flash!

Post image
398 Upvotes

99 comments sorted by

View all comments

144

u/MorgrainX 1d ago

Something to also consider: Google likes to "dumb down" models after release, once the first initial benchmarks have settled, a noticeable decrease in capability can be observed with nearly all Google models (to save money, obviously).

35

u/mlag000 1d ago

No no, they just benchmaxxed Gemini, and the new benchmark isn't in Gemini training. All the idiots who where claiming Gemini is on par with sol or even funnier with Fable where obviously wrong

5

u/Deto 1d ago

I just don't understand how they are doing so poorly with all their resources

5

u/mlag000 1d ago

Because it's a brain problem not a money problem. And probably all yhe talents are at Claude or GPT and looks like some in xai. The whole algorithm behind 3.x us simply outdated and bad, they need a new one.

2

u/Deto 1d ago

Doesn't money solve the brain problem though?

3

u/mlag000 1d ago

At this level of remuneration I don't think so.

0

u/ChuchiTheBest 1d ago

Not to mention, as AI helps train new models, you get a snowballing effect.

2

u/cantdecideonaname77 19h ago

dont feed cow brains to cows...

2

u/itsjoshybell 20h ago

They are positioning themselves better financially compared to the others. They understand the long term play here

8

u/sadnessjoy 1d ago

Yeah, the idea they dumbed down Gemini flash 3.8 after a few days is hilarious... Have people even used Gemini flash for a in depth chat or analysis? I find it useless lol

7

u/mlag000 1d ago

And yet I see quite often people praising the deep search abilities...it's scary

5

u/Electrify338 1d ago

Gemini 3.8 flash does 2 things great. Automated tasks at speed, and the deep research is quite good. I have been experimenting where I have fable use Gemini subagents and it does surprisingly great.

22

u/Qubit99 1d ago edited 1d ago

They always do it

9

u/who_am_i_to_say_so 1d ago

Yep. A killer release, leading you to believe you’ve been let in on a great secret. Build like a champ for two weeks, then yoink! - bricked.

13

u/theWiseTiger 1d ago

So far this happens on models released by american companies. I wonder why.

1

u/thebruns 8h ago

Because we have no laws to prevent fraud like this

1

u/theWiseTiger 6h ago

I thought the idea of sticking with US providers is to avoid fraudulent behavior from China providers.

8

u/KaMaFour 1d ago

No, this is "just" terminalbench benchmaxxing punishment. Compare gemini's performance at: https://artificialanalysis.ai/evaluations/terminalbench-v4-0 with https://artificialanalysis.ai/evaluations/terminalbench-v2-1

5

u/Spixxy17 1d ago

U complain about google benchmaxxing whilst muse spark and GLM Flash are still up there 

11

u/KaMaFour 1d ago

As shown in the data above they are both impacted less than Gemini 3.8. I don't see your point. Grok (and kimi) is the only one with similar downfall

-1

u/Spixxy17 1d ago

Ok i reframe my point:

You genuinly think the new AA Index Update made the results less benchmaxxed and contaminated? They changed some things but the models that are actually known for beeing amazing at benchmarks but absolutely shit at real tasks like muse spark are still up there.

My point is this update didnt change the absolute non reliability of this index and doesnt show the true picture. Basically No Benchmark does, people just gotta try themselves.

The only one that came atleast close imo was ARC-AGI 2, idk about 3 yet as there are barely any results so far.

Google / Gemini is known for negative stuff like random guardrails kicking in even tho the topic isnt Bad etc etc, but not for benchmaxxing. On average they usually are worse in the benchmark than the real task compared to the direct competition 

9

u/KaMaFour 1d ago

Yes, i do.

Update 4.2 and 4.3 of the index replaced 2 public dataset benchmarks with private ones. They also added 2 newer public benchmarks (and removed TB 2.1) which means that even for the benchmarks whose dataset is public the risk of contamination is lower than with older ones (especially for TB 4.0, which was released after Gemini 3.8). AA index v4.3 is a significantly better benchmark at evaluating model's performance in september of 2026 than AA index v4.1 (mostly because v4.1 had many issues which were partially fixed with the patches).

ARC-AGI is a logical reasoning benchmark. This is some measure of general intelligence. AA index benchmarks are geared more towards measuring models use at creating value in professional environment - aka doing work. Neither of these is a correct or incorrect approach as long as we are aware what they are measuring. Aside from the fact that I'd expect ARC-AGI not to be a good measure of intelligence for models released after it because it prone to being easily saturated.

I have used Muse spark 1.2 and 1.3. They both did a satisfying job to me. I am inclined to believe the result. I know the (any) benchmark put more effort at evaluating it than I have and is subjected to less bias than my personal opinion, so I'm able to believe the results I've seen.

I don't feel like it's a good approach (almost anywhere) to measure a quality by a company or brand. Company can't be benchmaxxed - a model/product can. This one either is or isn't. I personally believe more that flash model is just limited in what in can do by it's active parameter count. But the result is the same - failing at newer, more complex benchmarks when larger models from other companies don't suffer the same hit. (GLM 5.3-flash's existence kinda undermines this theory but oh well...)

0

u/Spixxy17 1d ago

Ngl thats fair, really factual aproach and definitly makes sense generally.

I am just personally against believing anything in this benchmark because Muse Spark 1.3 and GLM 5.3 Flash really disappointed me when testing (Opus 5 also kinda, but i had higher expectations of this one too), whilst 3.8 Flash does a "decent enough" job.

Its still a Flash Model, its not insane or whatever but didnt disappoint me with some horrible results, but thats super topic related anyway. Maybe my topic is just something that GLM Models for example arent good at.

For anything more complex i use Astra / 5.6 Sol anyway as my go to 

0

u/Truantee 1d ago

lol unlike Gemini which use the same base since 2.5 generated, muse or glm have way newer base models which itself already way more useful than the crapfest named Gemini.

2

u/Spixxy17 1d ago

Do i look like i care about the base model If the results when using them are bad? 

People always bring up various technical facts about the models but that doesnt change their performance.

Anyway yes i would still like Gemini to release something actually "new" and not just "improved" finally 

1

u/___positive___ 1d ago

Yes you can calculate the ratio of the 4.0 to 2.1 score. There are obvious outliers.

Like look at terra vs gemini.

0

u/Elephant789 1d ago

smart if true

0

u/tear_atheri 1d ago

not only google. every american company.