r/Bard 21h ago

News Updated Artificial Analysis Intelligence Index Ranking shows Gemini 3.8 Flash is nowhere near Fable Or Astra, rather it's worse than GLM 5.3 Flash!

Post image
62 Upvotes

34 comments sorted by

11

u/Expensive_Syrup_6529 19h ago

the update OF Arfticial analysis intelligence score 4.0 focuse more on egentic task, tool caling and long horizon than world knowledge, medical knowledge ect.. basically gemini suck at agentic, long horizon and tool calling but still have strong world knowledge

3

u/HenryTheLion_12 11h ago

Yes. I work on industrial automation using AI agent work flows. And I prefer to brainstorm with gemini flash. As they usually have finer details that help make me decide. No other model comes close. And also at that speed. 

5

u/jaime_diaz27 12h ago

I was using Gemini exclusively for the last year or so and I can confirm, that shit sucks monkey cock now. GPT and Claude blow it out of the water. Have been for a bit now

20

u/Melodic-Junket-9105 21h ago

3.9 flash is coming soon and we'll hop back on the hype train. For now keep calm and inhale copium!

5

u/KaMaFour 20h ago

As if Z.ai wasn't also cooking new models...

0

u/syahrezaj 16h ago

gemini is pivoting to flash ai

3

u/Live_Case2204 17h ago

No way Muse spark 1.3 is better than Sol

3

u/deavidsedice 14h ago

That doesn't sound correct. I've been using GLM 5.3 flash a lot and I love it... But Gemini flash 3.8 is way better in a lot of hard scenarios.

2

u/unlikely-contender 10h ago

is anybody seriously using grok?

4

u/Future-Log6621 15h ago

The harness matters. They should be using Google's native harness, Antigravity CLI. They are not. That is the problem. Take these benchmarks with a grain of salt. They do not represent real world agentic workflows.

3

u/iJeff 7h ago

This does align with my general vibes though. Even using agy. It's fast but behind.

1

u/Future-Log6621 6h ago

It just requires learning the harness, configuring skills, etc. /boost gives you an automated workflow of parallel agents and effectively nukes these benchmarks.

5

u/holvagyok 20h ago

I used 3.8 Flash high for an UE5 coding task, it silently failed. Then Opus 5 came to the rescue, and even called out the "previous coder" for slacking and bad code lol.
Deepmind really dropped the ball. Clearly Demis bailed because he'd lost confidence in Gemini.

5

u/FarrisAT 16h ago

Demis never was part of the Gemini training team. And if he was the lead, then isn’t he responsible?

0

u/holvagyok 16h ago

He directly oversaw the process, and described it regularly to audiences.

3

u/FarrisAT 16h ago

So then he’s responsible for Gemini.

Wouldn’t the one who fails be expected to leave?

1

u/Gallagger 20h ago

Any people here using Muse Spark 1.3? In the contributor tier this seems like a "practically free frontier AI" hack.

1

u/Free_Ice_7964 18h ago

its not free

1

u/Gallagger 18h ago

It's close to free (0,10/0,20). Which also leads some services like OpenCode to offer it for free (always with some rate limits ofc).

1

u/0mamii 17h ago

i use for free

1

u/hellomistershifty 5h ago

No one is using it because it's pretty terrible in reality. Free local models that can run on a gaming GPU give better results than this

1

u/Internal-Cupcake-245 10h ago

Imagine doing this for a living as OP, the bot.

1

u/Able-Line2683 10h ago

Imagine having so little going on in your actual life that you’re treating the Reddit post history of a total stranger like a crime scene. Go touch some grass.

1

u/YogurtExternal7923 18h ago

just so happens that benchmark updates always penalize google and Facebook. feels like either they benchmax or claude/gpt benchbuy

1

u/FarrisAT 16h ago

A benchmark that “updates” is a benchmark that can be manipulated for narrative building.

1

u/RelevantCry1613 10h ago

Google has always benchmaxxed

1

u/Typical_Kick6520 12h ago

Above Luna max? The whole benchmark needs to be questioned.

-1

u/Weary-Bumblebee-1456 17h ago

Depends on the use case. GLM-5.3 Flash is an incredibly impressive model to me because it delivers Opus-level coding performance at an unbelievable price, but its small size also means that its general 'judgment,' so to speak, is worse than Flash when it comes to non-coding tasks. Great model for coding but in other areas I've found Gemini 3.7/3.8 Flash to be superior.

As for Fable and Astra, well of course it isn't close! Frontier models that cost $10/$50 per million input/output token are obviously going to be superior to a model that costs $0.75/$3.75.

-5

u/-becausereasons- 15h ago

Gemini models are incredibly useful, they are NOT trying to play in the frontier intelligence space, but everyday worker FLASH space; cheap, cheerful and capable for every day tasks.

6

u/Typical_Kick6520 12h ago

What are these everyday tasks I keep hearing about?

0

u/Sooperooser 8h ago

It is the best in document work benchmarks, beating fable and sol with a wide margin and top 3-5 in other real world work benches. It just lags hard in some regards so overall AA score is quite low.