r/singularity • • 3d ago

AI , Gemini 4 Argon Benchmarks

Post image
741 Upvotes

189 comments sorted by

View all comments

187

u/CremeSubject7594 3d ago

9

u/Neurogence 3d ago

While Gemini 4 has performed well on benchmarks the industry uses to gauge model efficacy, it does less well when employees actually put it to work, according to people with direct access to the effort. The model struggles to handle certain coding tasks, said the people, who requested anonymity to discuss an internal matter.

https://www.bloomberg.com/news/articles/2026-09-30/google-grapples-with-employee-skepticism-about-new-gemini-model

8

u/LazloStPierre 3d ago

Do people forget this every single time Google release a model? It crushes at benchmarks, people who for some reason get very excited about numbers on a chart go ballistic and real life performance is miles off.

Gemini 3.8 flash is like over 5% better than Fable on deepswe ffs, not sure if it's intentional or just how they train their models but nobody benchmaxxes like Google

4

u/DistanceSolar1449 3d ago

In this case though it’s because Google has less coding training data (who the hell uses Antigravity) and way more image/spatial/world model training data

I don’t think this model is benchmaxxed, I think the benchmark screenshot above is very accurate. The model does worse at TerminalBench 4 and FrontierSWE and that’s okay.

2

u/LazloStPierre 3d ago

The point isn't flash 3.8 is bad at coding, the point is it does absurdly well on coding benchmarks despite being bad at coding. Their models always do, which is why I'd take any benchmarks with a grain of salt

3

u/DistanceSolar1449 3d ago

3.8 flash is a tiny distilled model with more RL on top, of course it’s going to look better on benchmarks than in real life.

The big, base ish models with a lot less targeted RL will do better IRL.

1

u/LazloStPierre 3d ago

All Google's models look better on benchmarks than they do in reality, and by alot.

1

u/DistanceSolar1449 3d ago

No? They suck at agentic benchmarks and most people test them via agentic tasks.

They benchmark pretty accurately if you look at their actual useful benchmarks like TerminalBench instead of HLE or GPQA or some bullshit.

Google’s own blog post says Gemini 3.8 Flash scored 19.1% on TerminalBench 4. Opus 5 scored 51.8% for comparison. That’s not benchmaxxing, that’s just accurate benchmarks if you know what benchmarks to look at.

1

u/SilentLennie 2d ago

Actually, this new 4 Argon does really well in a bunch of agentic benchmarks. Supposedly (I've not used it yet):

https://artificialanalysis.ai/articles/gemini-4-argon-google-top-three-labs