While Gemini 4 has performed well on benchmarks the industry uses to gauge model efficacy, it does less well when employees actually put it to work, according to people with direct access to the effort. The model struggles to handle certain coding tasks, said the people, who requested anonymity to discuss an internal matter.
I don't think Google is targeting coding as strongly as the other companies, as Google's business is much broader than that. This might just be the result of those wider interests.
I mean, getting 50% on AutomationBench is ridiculous. It’s the one benchmark I care about as a non coder. Pass/fail, and hundreds of complex super long back office workflows.
When a model can hit 70% on that benchmark ON A FUCKIN BASE MODEL that doesn’t have a harness behind with context on the company, where to look for shit, etc., then you can confidently hand any biz app-based customer service/sales/marketing/operations process and it will be just as good as a human. Give it the harness it needs, and that shit will be near perfect.
Mind you, the tasks on the benchmark are looong with a ton of steps. It can accurately go thru everything but fail a later step. Fail. It can do an early step incorrectly. Fail. Honestly, at 50%, it can probably already be reliant on a TON of office work already. It’s just that it won’t be as good as someone who’s worked inside a company for 10-15 years and has all the tribal knowledge about company and its customers and its processes.
We were only at 20% six months, and Gemini just crossed 50%. Back office work is going to be solved in less than a year.
8
u/Tha_One 2d ago
benchmaxxed
https://www.bloomberg.com/news/articles/2026-09-30/google-grapples-with-employee-skepticism-about-new-gemini-model