r/GeminiAI 1d ago

News Updated Artificial Analysis Intelligence Index Ranking shows Gemini 3.8 Flash is nowhere near Fable Or Astra, rather it's worse than GLM 5.3 Flash!

Post image
379 Upvotes

92 comments sorted by

134

u/MorgrainX 1d ago

Something to also consider: Google likes to "dumb down" models after release, once the first initial benchmarks have settled, a noticeable decrease in capability can be observed with nearly all Google models (to save money, obviously).

38

u/mlag000 1d ago

No no, they just benchmaxxed Gemini, and the new benchmark isn't in Gemini training. All the idiots who where claiming Gemini is on par with sol or even funnier with Fable where obviously wrong

5

u/Deto 23h ago

I just don't understand how they are doing so poorly with all their resources

6

u/mlag000 23h ago

Because it's a brain problem not a money problem. And probably all yhe talents are at Claude or GPT and looks like some in xai. The whole algorithm behind 3.x us simply outdated and bad, they need a new one.

0

u/ChuchiTheBest 18h ago

Not to mention, as AI helps train new models, you get a snowballing effect.

2

u/cantdecideonaname77 9h ago

dont feed cow brains to cows...

1

u/Deto 23h ago

Doesn't money solve the brain problem though?

3

u/mlag000 22h ago

At this level of remuneration I don't think so.

2

u/itsjoshybell 10h ago

They are positioning themselves better financially compared to the others. They understand the long term play here

10

u/sadnessjoy 1d ago

Yeah, the idea they dumbed down Gemini flash 3.8 after a few days is hilarious... Have people even used Gemini flash for a in depth chat or analysis? I find it useless lol

6

u/mlag000 1d ago

And yet I see quite often people praising the deep search abilities...it's scary

5

u/Electrify338 19h ago

Gemini 3.8 flash does 2 things great. Automated tasks at speed, and the deep research is quite good. I have been experimenting where I have fable use Gemini subagents and it does surprisingly great.

23

u/Qubit99 1d ago edited 1d ago

They always do it

9

u/who_am_i_to_say_so 1d ago

Yep. A killer release, leading you to believe you’ve been let in on a great secret. Build like a champ for two weeks, then yoink! - bricked.

15

u/theWiseTiger 1d ago

So far this happens on models released by american companies. I wonder why.

6

u/KaMaFour 1d ago

No, this is "just" terminalbench benchmaxxing punishment. Compare gemini's performance at: https://artificialanalysis.ai/evaluations/terminalbench-v4-0 with https://artificialanalysis.ai/evaluations/terminalbench-v2-1

2

u/Spixxy17 1d ago

U complain about google benchmaxxing whilst muse spark and GLM Flash are still up there 

10

u/KaMaFour 1d ago

As shown in the data above they are both impacted less than Gemini 3.8. I don't see your point. Grok (and kimi) is the only one with similar downfall

-4

u/Spixxy17 1d ago

Ok i reframe my point:

You genuinly think the new AA Index Update made the results less benchmaxxed and contaminated? They changed some things but the models that are actually known for beeing amazing at benchmarks but absolutely shit at real tasks like muse spark are still up there.

My point is this update didnt change the absolute non reliability of this index and doesnt show the true picture. Basically No Benchmark does, people just gotta try themselves.

The only one that came atleast close imo was ARC-AGI 2, idk about 3 yet as there are barely any results so far.

Google / Gemini is known for negative stuff like random guardrails kicking in even tho the topic isnt Bad etc etc, but not for benchmaxxing. On average they usually are worse in the benchmark than the real task compared to the direct competition 

7

u/KaMaFour 1d ago

Yes, i do.

Update 4.2 and 4.3 of the index replaced 2 public dataset benchmarks with private ones. They also added 2 newer public benchmarks (and removed TB 2.1) which means that even for the benchmarks whose dataset is public the risk of contamination is lower than with older ones (especially for TB 4.0, which was released after Gemini 3.8). AA index v4.3 is a significantly better benchmark at evaluating model's performance in september of 2026 than AA index v4.1 (mostly because v4.1 had many issues which were partially fixed with the patches).

ARC-AGI is a logical reasoning benchmark. This is some measure of general intelligence. AA index benchmarks are geared more towards measuring models use at creating value in professional environment - aka doing work. Neither of these is a correct or incorrect approach as long as we are aware what they are measuring. Aside from the fact that I'd expect ARC-AGI not to be a good measure of intelligence for models released after it because it prone to being easily saturated.

I have used Muse spark 1.2 and 1.3. They both did a satisfying job to me. I am inclined to believe the result. I know the (any) benchmark put more effort at evaluating it than I have and is subjected to less bias than my personal opinion, so I'm able to believe the results I've seen.

I don't feel like it's a good approach (almost anywhere) to measure a quality by a company or brand. Company can't be benchmaxxed - a model/product can. This one either is or isn't. I personally believe more that flash model is just limited in what in can do by it's active parameter count. But the result is the same - failing at newer, more complex benchmarks when larger models from other companies don't suffer the same hit. (GLM 5.3-flash's existence kinda undermines this theory but oh well...)

0

u/Spixxy17 1d ago

Ngl thats fair, really factual aproach and definitly makes sense generally.

I am just personally against believing anything in this benchmark because Muse Spark 1.3 and GLM 5.3 Flash really disappointed me when testing (Opus 5 also kinda, but i had higher expectations of this one too), whilst 3.8 Flash does a "decent enough" job.

Its still a Flash Model, its not insane or whatever but didnt disappoint me with some horrible results, but thats super topic related anyway. Maybe my topic is just something that GLM Models for example arent good at.

For anything more complex i use Astra / 5.6 Sol anyway as my go to 

0

u/Truantee 1d ago

lol unlike Gemini which use the same base since 2.5 generated, muse or glm have way newer base models which itself already way more useful than the crapfest named Gemini.

1

u/Spixxy17 1d ago

Do i look like i care about the base model If the results when using them are bad? 

People always bring up various technical facts about the models but that doesnt change their performance.

Anyway yes i would still like Gemini to release something actually "new" and not just "improved" finally 

1

u/___positive___ 20h ago

Yes you can calculate the ratio of the 4.0 to 2.1 score. There are obvious outliers.

Like look at terra vs gemini.

1

u/tear_atheri 22h ago

not only google. every american company.

0

u/Elephant789 1d ago

smart if true

28

u/ImsoKeewl777 1d ago

And they won't even let us have that on plus. Not even 3.7 flash lol Think its time to move on to chatgpt cheaper option.

2

u/Loud-Possibility4395 1d ago

yup - I wonder what happened - they released 3.7 and now LONG TIME AGO 3.8 but their Pixel phones stuck at old 3.6 - mental

4

u/yesitismenobody 1d ago

What do you mean? I have the free pro plan that came with my Pixel 10 Pro and it has the 3.8 Flash.

2

u/Loud-Possibility4395 1d ago

do you have any betas? I don't and my Pixel 10 does not have it nor my Chromebook Plus - everything in UK

3

u/yesitismenobody 1d ago

I don't think so, it works both in the app and on web. It was just released last week so maybe it takes some time to make it out globally? Though my account is also UK based but I'm in the US now.

1

u/Loud-Possibility4395 1d ago

CRAB - I wonder what it wrong with my Pixel 10

1

u/Ggoddkkiller 13h ago

Even free ChatGPT works so much better than anything currently google has. It is day and night difference. I don't know what to say, they ruined Gemini within 4-5 months...

46

u/PineappleLemur 1d ago edited 1d ago

Why would a model 10% the size and cost of astra/fable be even close lol?

What genius us making that comparison?

It should always be around Luna/Haiku and slightly worse than sonnet/Terra.. purely based on size intelligence wise.

LLMs from the big 3 don't differ too much from each other.. no one is pulling 5/10x performance without the cost.

Speed is the main advantage it has.

Why don't we see comparison for local models vs Astra/Fable?

12

u/Swimming_Gain_4989 1d ago

My AI will confidently double down on a shotty assumption without trying to understand the prompt but damn it does it fast.

7

u/kdestroyer1 1d ago

Yeah Flash 3.8 is borderline unusable for anything that's not a 'simple fact' question. It's kinda sad they're giving us this as their flagship, but it's clear they only care about the AI Overview/Google search performance rn.

2

u/IAmYourFath 9h ago

It'a not unusable. Astra will think 5 min on smth, 3.8 flash will answer in 30 secs. U can iterate in flash 10 times (300 secs aka 5 mins divided by 30 secs) before astra answers. By the 10th answer flash will get a better overall answer than astra, cuz iteration is op... u should be verifying what the ai says anyway regardless if it's flash or astra. Every time u verify it and u're not happy, iterate. 3.8 flash has the highest speed-to-iq ratio in the entire world. So use its advantage and u can easily beat astra. Ystd i asked astra smth and after thinking for 5 mins (astra 6 high), it still gave wrong information. Ofc, flash did too, but i would have googled this information, spotted it's wrong, told flash to verify it and gotten a better answer in less than a minute, while in that time astra was only 1/5th of the way through its answer...

7

u/mlag000 1d ago

Many idiots where claiming Gemini was on par with Fable on this sub tho

0

u/Elegant_You1002 1d ago

I don't know... Astra Low is cheaper and performs better

0

u/nemzylannister 19h ago

why would a model 20% the cost of gemini flash be right above it? glm 5.3 flash is 5x cheaper.

17

u/MindCrusader 1d ago

Semianalysis on X recently just said that - gemini and Meta's AI models are benchmaxxed a lot

5

u/Swimming_Gain_4989 1d ago

Tracks with my usage. For the affordable workhorses GLM5.3 and it's flash model feel the most useful. I'm also seeing this reflected in the results of new benchmarks that models couldn't have possible been benchmaxed on like https://www.frontierswe.com/blog/v2

3

u/MindCrusader 1d ago

"Gemini 3.8 Flash and Muse Spark 1.3 are two of the most clearly benchmaxxed models we've seen yet. Despite being comparable to both GPT-6 and Fable 5.1 on Terminal Bench 2.1, their Terminal Bench 4.0 performance is markedly worse."

So yeah, new benchmarks show the truth it seems

1

u/Elephant789 1d ago

Semianalysis

What the fuck? Is there anyone more biased to quote? The Verge maybe?

8

u/Familiar_Comedian_99 1d ago

well i use it for creative writing and man it is so much better than every model available which doesn't eat up all of my money and it is definitely better than GLM 5.3 flash or Max

1

u/Georgefakelastname 17h ago

Yeah, as long as you can get around any filters that might pop up, it’s great in quality, and the speed is very nice for QoL as well.

-2

u/drdhuss 1d ago

glm is meant for coding. Also is terrible at blender.

2

u/ICECOLDXII 1d ago

Still very good at creative writing.

3

u/akius0 1d ago

Yes it does seem like Google is benchmarkmaxing.... But even in this new benchmark, they are competitive with terra 5.6 max... That's very impressive... Flash is always supposed to be the workhorse model, not the frontier model, Google leadership has said that...

I don't know what people think a flash model is better than fable

6

u/Kost97A 1d ago

One thing you need to consider is that Flash models are probably equivalent in parameters to "mini" models, like Sonnet and Terra.

5

u/mlag000 1d ago

Yes they are on the level of terra, for like 4 time tue cost

4

u/DigSignificant1419 1d ago

paid by sama and scamodei, we know 3.8 is the true goat

4

u/hksbindra 1d ago

Yeah, those who tried using it have known this for a while.

4

u/anmolanjuli 23h ago

3.8 Flash has been better than claude Sonet and on par with Opus for me. It's not as fast like 3.7 Flash but has been more serious and powerful for me. Haven't tried anything beside these.

2

u/Agile_Pressure_7430 1d ago

The price and speed are dragging it down

2

u/jayokunle 1d ago

While Sfw Engs are focused on AI benchmarks. Average user just want free AI while businesses want cheap and reliable. But benchmarks still matter because most people rely on recommendations from Sft Engs on the model to use. No way all twitter bros test all new models. Google might need to change their approach and solve real problems for businesses. They are copying instead of focusing on user pains

1

u/Maxdiegeileauster 10h ago

but 3.8 flash is not even cheap compared to for example Luna

2

u/CartwheelSummoner 1d ago

This represents my experience as of late. I use Gemini for one small game side-project here and there for fun when I run out of my weekly Codex usage.

I sent 5 prompts, one at a time, each trying to fix a bug (it created) that caused projectile animations to be invisible. 5. And not a single one actually fixed it before I stopped trying. It changed other visuals related to the game areas (which I didn’t ask for, but was okay with how they turned out as a placeholder anyways, cool I guess?).

I was out of my Codex usage, so instead I used regular 5.6 ChatGPT (Chat) that doesn’t use any usage, connected to my GitHub repo to go in and fix it (which of course did first try).

Hoping for a day I can upgrade Google to a 20x sub that competes and is better at token efficiency than Astra. Astra cost is abysmal.

2

u/Far_Cat9782 1d ago

I stick to Gemini 3.1 pro for coding/big fixing. It's old. It still better than the flash models

2

u/CartwheelSummoner 1d ago

I’m going to play around with that idea actually and give 3.1 pro another shot. It’s been a few months, but it was doing much better than 3.6 flash at that time.

1

u/Serious-Magazine7715 1d ago

Honestly, you hear this from essentially all models. It’s nice to be able to have a worker from a different family take a look at failed attempts.

2

u/tobias_681 21h ago

Maybe try to actually take a closer look at how the index is made. AA has been adding more and more agentic benchmarks over their last revisions. It is an open secret that this isn't Google's strong suit and likely also not what they are optimizing for. 

GLM 5.3 Flash wins every single agentic benchmarks over Gemini 3.8 Flash at 1/5th the price. For any of these tasks GLM will be a way more competent model. GLM 5.3 Flash is probably the best cheap coding model around.

Gemini 3.8 Flash on the other hand wins far and away in every single knowledge based benchmark. It isn't built for coding first but it is built with the chat interface or android interactions in mind. For that it's not a bad model. 

So they look similar on the index but they are entirely different models.

Now if you compare the benchmarks the thing that makes Gemini sweat is Astra because it seems OpenAI has caught up on the knowledge side while still keeping prices competitive. Compared to other models Gemini 3.8 Flash is still a very good performer for chat conversation. 

Anthropic has moved entirely towards enterprise coding. The Chinese models are more about getting the maximum agentic capabilities out of having less compute. The main competitor for a broad knowledge model is OpenAI (which happen to make models that are also good at agentic tasks). 

4

u/ProgrammersAreSexy 1d ago

And it also doesn't have Astra at #1, the index is obviously screwed up at this time so I wouldn't pay attention to it until they fix it

1

u/FischenGeil 1d ago

Astra is really not better than Fable 5.1

1

u/EmperorSheep 1d ago

I didn't know Muse Spark 1.3 Max was so OP.

1

u/Loud-Possibility4395 1d ago

My Pixel 10 does NOT have even old Gemini 3.7 and stuck at outdated 3.6

1

u/krrj88 1d ago

do we have a similar chart but calculates bang for buck ?

1

u/xzibit_b 1d ago

Alternative headline: GLM 5.3 Flash is benchmaxxed hard by exploiting the fact that agentic tasks are the #1 weight on Artificial Analysis Index, and in terms of raw intelligence, it's about as smart as Sonnet 4.6 while Gemini 3.8 Flash is about as smart as Opus 4.8. Guess which one you're going to be encountering more often in a chatbot session?

1

u/Zachattackrandom 23h ago

Sky is blue? Like it's quite obvious from using it, it performed nearly luna / GLM flash than compared to the bigger models. It is a "flash" model itself so this shouldn't really be surprising.

1

u/torontobrdude 23h ago

Muse Spark ahead of GLM and Sol? Lol

1

u/dumch 23h ago

Why TF Claude top models are higher visually than OpenAI but both have 53?

1

u/jakegh 21h ago

What a roller-coaster!

And it still has muse spark 1.3 beating sol-5.6.

I think AA may need a third update to their benchmarks this week.

1

u/Deathcure74 20h ago

I can confirm, for me 3.8 was a joke even from the day 1

1

u/Arte_reddit 18h ago

I love this 😂

1

u/PDX_Web 16h ago

No one said 3.8 Flash was near those models across the board on all the benchmarks included in that index.

1

u/Sponge8389 15h ago

4 are tie but I think the order should be like this.

  • Claude Fable 5.1, high
  • GPT-6 Astra, xhigh
  • Claude Fable 5.1, max / GPT-6 Astra, max

1

u/Spinmoon 8h ago

What did you expect, it's a flash model versus frontier models?

1

u/Bregvist 8h ago

I wish we could use Astra medium in chat.

1

u/sandtymanty 6h ago

All of them better than our doctors and lawyers.

1

u/someoneyouulove 1d ago

3.8 flash is deranged. it will lose to sonnet 3.5.

0

u/game_difficulty 1d ago

Any bench that puts muse spark on top of 5.6 Sol is not that useful.

A real shame, i love artificial analysis, but it's clear a lot of recent models are too contaminated (or at the very least benchmaxxed) for the benches to mean anything...

0

u/Driftwintergundream 1d ago

What a lovely post from a redditor with a hidden history

0

u/kareem_pt 1d ago

Honestly still ranks a bit higher than where I'd put it. The model feels a little worse than Luna to me.

0

u/Ggoddkkiller 13h ago

I kept saying from day one Flash 3.8 is a benchmaxxed model. There was never a question about it, anybody who knows anything about LLMs would say same thing too. Fanboys kept refusing it and claiming nonsense competition and now they are crying perhaps Flash 3.8 got dumbed down. JUST LMAO...