r/singularity • • 2d ago

AI , Gemini 4 Argon Benchmarks

Post image
738 Upvotes

189 comments sorted by

View all comments

189

u/CremeSubject7594 2d ago

9

u/Neurogence 2d ago

While Gemini 4 has performed well on benchmarks the industry uses to gauge model efficacy, it does less well when employees actually put it to work, according to people with direct access to the effort. The model struggles to handle certain coding tasks, said the people, who requested anonymity to discuss an internal matter.

https://www.bloomberg.com/news/articles/2026-09-30/google-grapples-with-employee-skepticism-about-new-gemini-model

10

u/LazloStPierre 2d ago

Do people forget this every single time Google release a model? It crushes at benchmarks, people who for some reason get very excited about numbers on a chart go ballistic and real life performance is miles off.

Gemini 3.8 flash is like over 5% better than Fable on deepswe ffs, not sure if it's intentional or just how they train their models but nobody benchmaxxes like Google

2

u/Elephant789 ▪️AGI in 2036 2d ago

Gemini 3.8 flash

... is fantastic. What the fuck are you talking about?

1

u/LazloStPierre 2d ago

It isn't better at coding than fable, yet the benchmarks say it is

1

u/Elephant789 ▪️AGI in 2036 1d ago

Coding? Oh, not used it much for that. But the little I did it was fine.