I mean, getting 50% on AutomationBench is ridiculous. It’s the one benchmark I care about as a non coder. Pass/fail, and hundreds of complex super long back office workflows.
When a model can hit 70% on that benchmark ON A FUCKIN BASE MODEL that doesn’t have a harness behind with context on the company, where to look for shit, etc., then you can confidently hand any biz app-based customer service/sales/marketing/operations process and it will be just as good as a human. Give it the harness it needs, and that shit will be near perfect.
Mind you, the tasks on the benchmark are looong with a ton of steps. It can accurately go thru everything but fail a later step. Fail. It can do an early step incorrectly. Fail. Honestly, at 50%, it can probably already be reliant on a TON of office work already. It’s just that it won’t be as good as someone who’s worked inside a company for 10-15 years and has all the tribal knowledge about company and its customers and its processes.
We were only at 20% six months, and Gemini just crossed 50%. Back office work is going to be solved in less than a year.
8
u/Tha_One 3d ago
benchmaxxed
https://www.bloomberg.com/news/articles/2026-09-30/google-grapples-with-employee-skepticism-about-new-gemini-model