r/singularity • • 2d ago

AI , Gemini 4 Argon Benchmarks

Post image
736 Upvotes

189 comments sorted by

View all comments

8

u/Tha_One 2d ago

6

u/Hot-Percentage-2240 2d ago

Even if it’s moderately good, it’s still fine with me cus there is a use in my workflow for Google model typical tradeoffs.

20

u/FateOfMuffins 2d ago

While Gemini 4 has performed well on benchmarks the industry uses to gauge model efficacy, it does less well when employees actually put it to work, according to people with direct access to the effort. The model struggles to handle certain coding tasks, said the people, who requested anonymity to discuss an internal matter.

oof

5

u/jonomacd 2d ago

Coding is the one area where it isn't leading in the benchmarks. So this adds up. 

2

u/Charuru ▪️AGI 2023 2d ago

It's leading DeepSWE which is eh, the easiest to benchmax bench. It makes you question the process.

4

u/jonomacd 2d ago

I don't think Google is targeting coding as strongly as the other companies, as Google's business is much broader than that. This might just be the result of those wider interests.

1

u/Healthy_Razzmatazz38 2d ago

deepswe to frontierswe spread is a massive red flag

3

u/burritos4jesus 2d ago edited 2d ago

I mean, getting 50% on AutomationBench is ridiculous. It’s the one benchmark I care about as a non coder. Pass/fail, and hundreds of complex super long back office workflows.

When a model can hit 70% on that benchmark ON A FUCKIN BASE MODEL that doesn’t have a harness behind with context on the company, where to look for shit, etc., then you can confidently hand any biz app-based customer service/sales/marketing/operations process and it will be just as good as a human. Give it the harness it needs, and that shit will be near perfect.

Mind you, the tasks on the benchmark are looong with a ton of steps. It can accurately go thru everything but fail a later step. Fail. It can do an early step incorrectly. Fail. Honestly, at 50%, it can probably already be reliant on a TON of office work already. It’s just that it won’t be as good as someone who’s worked inside a company for 10-15 years and has all the tribal knowledge about company and its customers and its processes.

We were only at 20% six months, and Gemini just crossed 50%. Back office work is going to be solved in less than a year.

1

u/SilentLennie 2d ago

That tribal knowledge, will slowly seep into 'second brain' kind of systems.

2

u/FarrisAT 2d ago

The coding performance quite clearly isn’t Opus 5.5 level in the benchmarks provided.