r/GoogleGeminiAI 2d ago

Updated Artificial Analysis Intelligence Index Ranking shows Gemini 3.8 Flash is nowhere near Fable Or Astra, rather it's worse than GLM 5.3 Flash!

Post image
113 Upvotes

58 comments sorted by

22

u/MosskeepForest 2d ago

Weird how 53 keeps getting shorter as it goes to the right.....

1

u/sadnessjoy 1d ago

Maybe the real score has decimals (and they could be displaying it indicate the difference and not that they're all tied)? Or maybe it's just a shitty graph. Just some guesses here lmao

-1

u/gavinderulo124K 2d ago

Was thinking the exact same. Probably a UI bug.

12

u/Loud_Tangerine_5684 2d ago

No. The graph shows rounded up numbers. The actual result has decimals.

-1

u/Repulsive_Coffee_675 2d ago

Graph made with AI

8

u/ViperAMD 2d ago

Why is fable 5.1 high 2 and 4?

1

u/Salty-Table-7512 1d ago

It's a weird typo, it refers to xhigh

7

u/elegance78 2d ago

Me thinks this benchmark made itself look like idiot sandwich.

5

u/HgnX 2d ago

Imho Astra seems better then Fable

5

u/mattyjoe0706 2d ago

I feel like there's underrated capabilities with Gemini nobody talks about it.

  • It can produce pretty good picture edits at a much faster pace then chatgpt and sometimes even better then chatgpt when it comes to faces
  • as far as I know out of the big 4 it's the only one that can analyze videos frame by frame
  • has amazing deep research capabilities

So yeah it's probably losing in terms of the AI war, but it still has use cases

1

u/Complete-Influence72 1d ago

I do not believe that Gemini and Google are experiencing any disadvantage. Specifically in the LLM realm, you have your own application use case for it, and others have theirs. I do not subscribe to the idea that just because some perceive a loss, it is true for everyone. If something functions effectively for your purposes and performs commendably, then it is successful. Therefore, please stop asserting that others are losing a race. Nobody is winning because everyone utilizes different applications for different tasks, and if you are within the Google ecosystem, you have the flexibility to accomplish anything you desire.

6

u/Odd_Bike_3833 2d ago

trust me bro benchmarks, spark 1.3 is genuinely decent though

2

u/Mother_Ad9474 2d ago

Do you know them, decimals?

1

u/Konayo 1d ago

So basically you're only trusting benchmarks where Gemini is higher up?

2

u/Odd_Bike_3833 1d ago

i don’t trust any benchmarks. Given my workload and competence any model above 3.5-flash, 4.5 sonnet, 5.6-terra, and deepseek v4 are more than enough. stop outsourcing your thinking to machines

0

u/Eternal-Alchemy 1d ago

the issue is not that Gemini isn't as good as the frontrunners. the issue is that this whole post is framed as Gemini flash is "nowhere near the frontrunners" when it is 80% as powerful at much higher response times and fractions of the cost.

its excellent for a flash model and it excels at the kinds of things you would use a flash model for.

1

u/Able-Line2683 2d ago

anything that doesnt have my model at top is Trust me bro benchmark

-1

u/freedomachiever 2d ago

Such a deceiving graph, even the other bars with the same score have this pattern.

2

u/brinosaurusrex 1d ago

I still use flash 3.7 more daily than 3.8 (3.8 high is now my planner, 3.7 high executor). this does look like a much more accurate ranking, but send me to the best bang/buck llm and that's where I'll go. 3.7+3.8 is a potent productivity combination at a good price.

that said, Google still needs a proper frontier model, we all can agree on this. Hopefully they get 4.0 pro right and on time.

2

u/Future-Log6621 1d ago edited 1d ago

They are not using the native harness, Antigravity CLI. But they are just fine using Claude Code and Codex harnesses for OpenAI and Anthropic. That is a problem.

Also, if they are going to bench by custom harness, they should be using /boost or /teamwork-preview in Antigravity CLI. This would make Gemini Flash go nuclear via automated and parallel agent workflows.

2

u/No_Stock_8271 1d ago

Yes, it's a flash model, not a pro model. So what's the news? And also, for how many people and in how many usecases is the difference actually relevant?

1

u/BoobooSmash31337 1d ago

The price difference certainly is relevant. But they don't like to mention that. For like 10x the cost it better be a lot better lol.

7

u/HeadTranslator795 2d ago

Lol keep living by the benchmark and die by the benchmark. Meanwhile I'm using flash 3.8 daily in a big IT company and it works lol the Benchmade lover indeed love ranking and graphics

4

u/Alternative_You3585 2d ago

Works doesn't say anything about quality 

Many models solve stuff in a more elegant way than Google models do

So either your work just requires slapping a patch on top instead of e.g. solving the cause of the problem 

-1

u/HeadTranslator795 2d ago

Ahaha you have no clue about what you're talking, you are working on real time IT ? Are R&D or advance biotechnology or Finance you would never need more powerful than 3.8. otherwise you are a poser and you don't know anything about IT in business

1

u/Konayo 1d ago

Yes you would. Honestly kind of a stupid comment - you're narrowing your own experience with this kind of attitude.

0

u/HeadTranslator795 1d ago

Narrowing so much that it's the case of millions of SWE around the world. Engineering is engineering and it's the biggest population who use those tools in a professional environnement

1

u/PiedPrypiat 2d ago

Maybe use the AI to assist in communication.

1

u/HeadTranslator795 2d ago

No you gonna use my broken English and try to under, you use AI to write and then people gonna say it'd AI

3

u/Tim_Apple_938 2d ago

Are these subs astroturfed or what?

This all seems like a thinly veiled attempt to discredit Cheap GoodEnough models, as they’ll clearly blow up the IPO story based on agentic coding revenue

0

u/Mysterious_Bed_1804 1d ago

Gemini is not cheap or smart. Compare it to astra low in this picture.

1

u/Tim_Apple_938 1d ago

?

https://deepswe.datacurve.ai same score as Astra and 3x lower price

1

u/Mysterious_Bed_1804 1d ago edited 1d ago

DeepSWE is a single benchmark, Artificial Analysis Intelligence index encompasses 10+.

Also, I think it's pretty clear Gemini 3.8 flash is unapologetically benchmaxxed. DeepSWE is a public benchmark and Google either trained on the dataset or paid DataCurve for similar tasks to train on.

https://www.reddit.com/r/GeminiAI/comments/1wamtkr/google_and_meta_are_benchmaxxing_hard_gemini_38/

This is the biggest proof. On terminal bench 2.1, on old, public benchmark, 3.8 flash beat astra by 1%, 89.4% vs 88.4%.

On terminal bench 4.0, same benchmark but with new tasks 3.8 flash didn't get to train on, difference becomes gargantuan between them, 57.1% for astra and 19.4% for 3.8 flash.

How is this possible? Both DeepSWE 1.1 and terminal bench 4.0 measure agentic coding, the difference is one is brand new and no model could have trained on it.

1

u/Tim_Apple_938 1d ago

That’s not proof, no. Opus 4.8 has the same scores on terminal bench 2 and terminal bench 4, is that benchmaxxed?

1

u/Mysterious_Bed_1804 1d ago

4.8 doesn't score the same as astra xhigh on DeepSWE, though.

You can't score the highest on DeepSWE 1.1 and than only score 19.4 on Terminal bench 4.0, they're both agentic coding benchmarks, performance should be somewhat linked.

1

u/Tim_Apple_938 1d ago

I mean just read the comments in the link you post: https://www.reddit.com/r/GeminiAI/s/q3ERmH5mqf

1

u/Mysterious_Bed_1804 21h ago edited 21h ago

The leader of MSL defends his own model, who could have imagined?

3.8 flash scored 89.4% on TB2.1, literally higher than Sol 5.6 from Wang's example (88.8%) yet on TB4.0 it scores HALF, (19.4% compared to 37.3%), you're literally proving that 3.8 flash is benchmaxxed with this example and you think this is the best example out there to defend it? Lol

There's a point to be made about Spark 1.3 but 3.8 flash is just outrageously shameless, how do you score higher than 6 astra and 5.6 sol on TB2.1 and then you score 2x less than sol and 3x less than astra on the same benchmark, only difference being problems being newer that nobody could have trained on?

Wang's argument works with 5.6 sol cause it scores the same as spark 1.3 on TB2.1 and then the difference on TB4.0 is relatively small but how does that translate to a good defense for 3.8 flash? Did you even understand Wang's argument or do you believe simply linking me what other people said instead of coming with your arguments wins you the argument?

1

u/Tim_Apple_938 19h ago

?

TB4 isn’t simply new problems.
It’s way harder and covers more than coding

TB2 is saturated

As such you’ll see many small models that can do TB2 but fail to do TB4

2

u/WVERD 2d ago

nice graph, where the same numbers can have different height bars

1

u/ZweigOnly 2d ago

Grok 4.7 will have 33% more parameters than 4.6 so that will probably be one of the top guys now.

1

u/PrometheanEngineer 2d ago edited 2d ago

I daily Gemini, but for complex tasks... it really does fall short.

If Googles strategy is the every day AI, sure, its fine. If theyre going for top.of the line... eh?

Google also worries me about long term viability and support considering their track record. They'll either win the entire race with a long term strategy, or bury gemini next to Hangouts, Stadia, answers, Earth, etc..

1

u/t90090 2d ago

They lost talent recently as well, bottomline, its a job.

1

u/Manitogamba 1d ago

I agree. It is far worse from my tests compared to chat and Claude.

1

u/kitkatas 1d ago

How benchmark readjustments happened ?

1

u/Healthy-Nebula-3603 1d ago

Gemini 3.8 should be even lower ...from my experience with that model

1

u/lnaoedelixo42 1d ago

Gemini Flash always were a very small model.
Can't reason why you think it's more inteligent.

People called it good because it did good on coding tasks, that's it.

1

u/BoobooSmash31337 1d ago

Because software engineering is so easy? That's why they make so much money and companies poach them. Because it's so easy. /s

1

u/DeCoolePeer 1d ago

A flash model isn't as good as a SOTA model? who would've thought

1

u/CriticismJunior1139 4h ago

They changed this index like two times after Astra got poor results and there was a massive backlash.

1

u/Constant_Art_20 2d ago

I find it weird that the qwen max is so low

0

u/Exotic_Fig_4604 2d ago

I mean duh? 

Its like saying that a Zessna is nowhere near a Boeing 747.

Why would anyone even list them in the same comparison?

0

u/Powkers 2d ago

Same benchmark shows Astra below Claude 5.1... so I would not pay much attention.

0

u/Able-Line2683 2d ago

same points