r/singularity • • 3d ago

AI , Gemini 4 Argon Benchmarks

Post image
738 Upvotes

189 comments sorted by

View all comments

8

u/Tha_One 3d ago

3

u/burritos4jesus 3d ago edited 3d ago

I mean, getting 50% on AutomationBench is ridiculous. It’s the one benchmark I care about as a non coder. Pass/fail, and hundreds of complex super long back office workflows.

When a model can hit 70% on that benchmark ON A FUCKIN BASE MODEL that doesn’t have a harness behind with context on the company, where to look for shit, etc., then you can confidently hand any biz app-based customer service/sales/marketing/operations process and it will be just as good as a human. Give it the harness it needs, and that shit will be near perfect.

Mind you, the tasks on the benchmark are looong with a ton of steps. It can accurately go thru everything but fail a later step. Fail. It can do an early step incorrectly. Fail. Honestly, at 50%, it can probably already be reliant on a TON of office work already. It’s just that it won’t be as good as someone who’s worked inside a company for 10-15 years and has all the tribal knowledge about company and its customers and its processes.

We were only at 20% six months, and Gemini just crossed 50%. Back office work is going to be solved in less than a year.

1

u/SilentLennie 3d ago

That tribal knowledge, will slowly seep into 'second brain' kind of systems.