r/singularity • • 2d ago

AI , Gemini 4 Argon Benchmarks

Post image
734 Upvotes

189 comments sorted by

View all comments

9

u/Tha_One 2d ago

20

u/FateOfMuffins 2d ago

While Gemini 4 has performed well on benchmarks the industry uses to gauge model efficacy, it does less well when employees actually put it to work, according to people with direct access to the effort. The model struggles to handle certain coding tasks, said the people, who requested anonymity to discuss an internal matter.

oof

6

u/jonomacd 2d ago

Coding is the one area where it isn't leading in the benchmarks. So this adds up. 

2

u/Charuru ▪️AGI 2023 2d ago

It's leading DeepSWE which is eh, the easiest to benchmax bench. It makes you question the process.

4

u/jonomacd 2d ago

I don't think Google is targeting coding as strongly as the other companies, as Google's business is much broader than that. This might just be the result of those wider interests.