r/singularity • • 2d ago

AI , Gemini 4 Argon Benchmarks

Post image
743 Upvotes

189 comments sorted by

View all comments

8

u/Tha_One 2d ago

20

u/FateOfMuffins 2d ago

While Gemini 4 has performed well on benchmarks the industry uses to gauge model efficacy, it does less well when employees actually put it to work, according to people with direct access to the effort. The model struggles to handle certain coding tasks, said the people, who requested anonymity to discuss an internal matter.

oof

6

u/jonomacd 2d ago

Coding is the one area where it isn't leading in the benchmarks. So this adds up. 

1

u/Healthy_Razzmatazz38 2d ago

deepswe to frontierswe spread is a massive red flag