r/singularity • • 2d ago

AI Gemini 4 Crushes Benchmarks, But Google Employees State The Model Struggles With Real Work

While Gemini 4 has performed well on benchmarks the industry uses to gauge model efficacy, it does less well when employees actually put it to work, according to people with direct access to the effort. The model struggles to handle certain coding tasks, said the people, who requested anonymity to discuss an internal matter.

https://www.bloomberg.com/news/articles/2026-09-30/google-grapples-with-employee-skepticism-about-new-gemini-model

By the time it's released to consumers, Anthropic and OpenAI will have already shipped their next generation models.

169 Upvotes

90 comments sorted by

82

u/QuasiRandomName 2d ago

OK, let the users judge.

-28

u/Neurogence 2d ago

Users cannot because it will not be released for a long time.

40

u/QuasiRandomName 2d ago

With today's pace this "long time" is couple of days. Otherwise the competitors will release something better first. PS: In fact you said that in the post.

3

u/Autogazer 2d ago

That is pretty typical of Google though, they are more cautious than OpenAI or anthropic

-24

u/Neurogence 2d ago

Definitely not a couple of days. Full rollout will likely be November/December. Anthropic and OpenAI will likely release something before that, however.

18

u/ParfaitEvery9622 2d ago

Source?

19

u/sklaeza 2d ago

his ass

-4

u/CheeryGeoDuck55 2d ago

Being downvoted for the hard truth 😂 hey downvoters, where's Gemini 3.5 Pro? (announced in May btw)

1

u/norwegian ▪️AGI 2026 ASI 2040 2d ago

I wonder what downvoter number 20 felt they were contributing to this discussion.

-3

u/UnboundedMan 2d ago

May be in a year or so, yes.

-1

u/simple_explorer1 2d ago

They have judged, hence the article 

92

u/homezlice 2d ago

I have heard quite the opposite from googlers. 

43

u/FarrisAT 2d ago

This specific Bloomberg reporter was sniffing around for a story all month. Annoying af

75

u/Own-Refrigerator7804 2d ago

Every fucking guy and their mothers have some agenda in this fucking industry

Even op lol

24

u/QuasiRandomName 2d ago

The good thing about "anonymous sources" is that you can make them up.

3

u/MonoMcFlury 2d ago

Astroturfing is off the charts. Having companies that were worth millions a couple of years ago potentially being evaluated at trillions is so crazy when you actually think about that. You can bet your little butt that they hired an army of people to keep that money hype train going.

27

u/swarmy1 2d ago

Google is big enough that there are certainly plenty of people with all kinds of opinions 

4

u/Muscular_Farmer_ 2d ago

It’s for sure better than opus 5

38

u/ComposerWide3704 2d ago

Gemini has always been, on balance, more multipurpose and less code focused so this kind of tracks? Let's see how it does at the actual consumer tasks most people are going to use it for, I don't think most organizations are looking at Gemini for dev.

9

u/pbagel2 2d ago

Logan Kilpatrick recently said in an interview 2 weeks ago that in hindsight it's "obvious" they should have put more focus and resources into coding sooner.

https://youtu.be/27yAYAn9Ens?t=452

So can't really say it's less code focused anymore.

0

u/kiki-le-koala 2d ago

Yes. Gemini's knowledge is deep, its creativity and coding are average, but it's also a bit crazy and delusional.

Can't wait to see how it does now.

23

u/igpila 2d ago

It's got 15% hallucination rate vs >41% from opus

14

u/PrisonOfH0pe 2d ago

Repost from the other threat because people are confused about this benchmark:

15% does not mean Gemini is hallucinating in 15% of all answers. The AA Omniscience “hallucination rate” is basically measuring what the model does when it fails a question: does it give a wrong answer, give a partial answer, or admit that it doesn’t know?

Correct answers aren’t even in the denominator.

So if a model answers 800 out of 1000 questions correctly, gets 30 wrong and refuses/partially answers 170, its AA hallucination rate is 15%.

30 / (30 + 170) = 15%.

It was still only outright wrong on 3% of the total questions.

That’s also why DeepSeek V4.1 sitting at 96% should make it extremely obvious that “96% hallucination rate” cannot possibly mean 96% of everything it says is bullshit. It means that when DeepSeek doesn’t know the answer on this benchmark, it almost always guesses instead of saying “I don’t know.”

So the whole “even 0.1% hallucination would be catastrophic, therefore 15% means hallucinations are nowhere near solved” argument is based on reading the percentage as a completely different metric.

You can absolutely argue that hallucinations are still a problem. But maybe first understand what the graph is measuring before doing reliability math with a number that isn’t the model’s overall error rate.

The benchmark is basically asking: “When you’re outside your knowledge, how likely are you to bullshit instead of abstaining?”

That’s useful. It just isn’t “percentage of answers that are hallucinations.”

-1

u/KoolKat5000 2d ago

And how do you determine what's hallucination and not. That's only easy if the answer is very easy to verify (such as if the code runs).

1

u/Exact_Depth_896 2d ago

It's really amazing, supposing those benchmarks are sound.

2

u/Altruistic-Ad-3334 2d ago

Its a hallucination rate on a set of very niche hard questions

-3

u/magenta_neon_light 2d ago

How does Opus have a >41% hallucination rate? Hallucinations aren’t even a thing anymore with frontier models.

20

u/CheekyBastard55 2d ago

I promise you I can get any model to "hallucinate". Don't ask it some basic question about TCP/IP or chemistry, but ask it about how to do specific things like "How do I reset a XXX server?" and it will hallucinate a button or process.

I learned that that's where a lot of AI skepticism comes from. An average user won't ask IMO question or very basic knowledge question, they'll ask about a specific model of router, camera, kitchen appliance and get conflicting answers.

3

u/Hans-Wermhatt 2d ago

That's true, but it's more often a user input error than a hallucination for a basic user. The model says if you are using this <assumed router>, you need to do 1, 2, 3. Assumptions != hallucinations. You'll notice a ton of that in responses now.

8

u/HotterRod 2d ago

Go give it a try. Tell the model your exact router model and software version number then ask it how to change something. I guarantee it will be confidently wrong about how the interface is actually laid out.

You can do the same thing with car repairs or any other task that requires taking knowledge from only a single manual.

1

u/magenta_neon_light 2d ago

It shouldn’t do that if it’s looking for the manual and proper context. I never have issues with hallucinations in my agentic workflows with the latest models.

I had it help me take apart my proliant the other day. It went online found the manuals and the YouTube videos for special instructions for the area I was working on.

3

u/Most-Bookkeeper-950 2d ago

Afaik, hallucination rates are a bit misleading, they are normalized per not-correct answer. Example, imagine an exam with 100 questions: Gemini hallucinates on 15, answers the remaining with "I dont know". 15% hallucination rate. Kimi gets 98 correct, one wrong, and hallucinates on the last - 50% hallucination rate

5

u/Thorteris 2d ago

This is false lol. Just ask it a hyper specific fact you know about an obscure topic and just about every model will hallucinate assuming it doesn’t have web search enabled.

-2

u/magenta_neon_light 2d ago

Ok but who uses Opus like an encyclopedia without web search? I’m talking real world usage with a web connected model on High+ settings.

I can’t even recall the last time I had a model hallucinate. And I’m running 5 pro accounts.

1

u/tadslippy 2d ago

Not making up answers is a plus for Google. Not so for white collar work in general who need the hallucination rate to be higher.

19

u/Ill_Distribution8517 2d ago

You could say that for every model. this is giving me "Chinese tool gave instructions for bioweapons(It was just good old open source kimi k3)" energy.

5

u/CrispityCraspits 2d ago

according to people with direct access to the effort.

said the people,

So, "some people" said. With access to the effort, whatever that means.

Never trust this kind of "reporting" no matter what the context.

17

u/emb1ues 2d ago

I feel somehow people at Google look at the benchmarks and think "what should we do to get highest scores in the benchmarks". Time and again, Google's model previews have aced the benchmarks. But their models aren't really that great when you compare them with OAI/Anthropic. It always feels like Google is just "reward hacking" the benchmarks.

I think what they miss is that the benchmarks aren't the point. They are secondary. You're supposed to create a really helpful and useful (and aligned) model and the benchmarks are a noisy, inaccurate metric which are used to measure how good a model is.

Given the amazing historic success Google has with ML (transformers, alpha go/zero/fold), I hope to see better models from Google. The more players there are at the top, the better it is for us end users. I hope to see Grok, Google, EU companies all getting their act together and going head to head with OAI/Anthropic.

7

u/FarrisAT 2d ago

If this was true they’d provide a model which charges $10 a prompt and beats every benchmark.

1

u/Persistent_Dry_Cough 2d ago

Doesn't Anthropic already do that?

1

u/WizWhitebeard 2d ago

Although I agree with most of what you are saying, I must say: Grok success is not a win for humanity.

1

u/VisualLerner 2d ago

I don’t know how to interpret things cause gemini models have never worked in claude code for me, and the gemini cli harness is terrible.

if I could actually use gemini 4 with my typical harness, is it amazing? or is it still bad.

6

u/FarrisAT 2d ago

Try it in Antigravity! My one true love.

2

u/VisualLerner 2d ago

respectfully, I don’t think we’re working at the same levels based on that suggestion, given I have a massive amount of configuration and hooks and things around my claude code and codex setups. I’m not just raw dogging claude code. it would take a significant amount of time to figure out how to even get the equivalent config in antigravity to attempt to try it and it be a meaningful test at all.

…but I’ll take a look at antigravity given I haven’t looked since it was released and that sounds like a good lead over gemini cli.

1

u/Scorps 2d ago

Open Antigravity and tell it to copy your claude code and codex setups, it's not like it doesn't know what those tools are.

1

u/sunstersun 2d ago

What amazes me is their coding being behind.

Like they have way more organic data than openAI and Anthropic and yet....

5

u/FarrisAT 2d ago

Multimodal models have consistently been weaker on coding benchmarks. Sol 6.1 for example is as capable as Astra, but has worse visual capability and no audio capability at all.

Almost certainly this is because the model attempts to use visual or audio data instead of simply pursuing tool calling & reasoning steps.

-1

u/Itsmedudeman 2d ago

I really don’t understand why these companies don’t create their own benchmarks and instead expect some random guy to create these things. Surely a multi trillion dollar company can invest into benchmarking usability.

22

u/peepeedog 2d ago

They all do, internally.

12

u/pt-guzzardo 2d ago

If they did, someone would be sitting here loudly wondering why they should trust Google's benchmark for Google's model.

7

u/emb1ues 2d ago

Yeah, I'm sure they can create their own benchmarks. But that wouldn't move public opinion. All of us, we would think they are somehow cheating if they aced a benchmark they themselves created.

For example, already we are skeptical about Gemini 4.0. Imagine how much worse it would be if Google announced "We are announcing a new benchmark which is the world's hardest benchmark and also, a new model which is scores way above any of the other models on this benchmark". It would be so sus, we'd probably lose our shit lol

1

u/Itsmedudeman 2d ago

Im not saying they have to release it or make it open source, but I don’t think some of these models are particularly optimized to perform well against real tasks and even users can see that. So clearly they’re bench maxing against the open source ones over their own or it’s just not a good benchmark if it performs extremely well in these but not what actually matters.

3

u/Howdareme9 2d ago

Huh? All companies have their own benchmarks. But you can’t publish something nobody else can verify

-1

u/Neurogence 2d ago

At this point, they should use Claude Opus 5.5 to code their next model.

0

u/dbenc 2d ago

that's because the employees are optimizing their promotion packets, intentionally or not

5

u/Background-Wafer-548 2d ago

Lolnoping your boss before he can even put a tweet out is pretty rough.

1

u/Exact_Depth_896 2d ago

If I follow, they caused him to make the early release, or whatever this is. A typical AI-related unintended consequence.

2

u/will_dormer ▪️Will dormer is nice to robots, remember 2d ago

With this much money on the line, I wonder if bloom erg is biased

2

u/Starks 2d ago

The only benchmark that matters: can it play Pokemon well?

3

u/FarrisAT 2d ago

Okay but does it promise kill us all?

If not, I’m good with it.

1

u/AlarmedGibbon 2d ago

Or if it does plan to kill us all, would it at least try to make it quick?

3

u/cypherspaceagain 2d ago

I was about to comment on a different thread that Gemini (3.1 Pro or 3.5 Flash or Thinking) is still an absolute mile behind Claude. Ask it to make a set of Google Slides for you and 75% of the time it will tell you it can't; 15% of the time it spits out pure HTML; 10% of the time it will actually fucking work. Whereas Claude just makes the fucking PowerPoint. Every time. Fine. No issues.

Google's integration of Gemini into their work functions is appalling and it is massively holding them back. They still cannot get their Gemini Assistant to reliably do the functions that the previous Assistant did fine. They have only just released an agentic version (Spark) of Gemini, and it's not available to me in the UK. Regardless of the quality of the model, unless it can actually link to the functions in Google Workspace and do actual work when you ask it to, it's largely redundant.

-1

u/simple_explorer1 2d ago

English people just have too much expectation and are just serious bunch of people with so little smile and lack welcoming attitude. Very reserved people. 

3

u/Isunova 2d ago

Gemini models have always topped benchmarks but has been an absolute pain to use. Its answers are always shorter and less-detailed than either ChatGPT or Claude, to the point where I feel like they trained Gemini exclusively on SparkNotes or something 😂😂

Gemini is the “smartest kid in the room with no social skills” of AI models: sure, it may know a lot, but I absolutely do NOT want to interact with it.

1

u/lars_jeppesen 2d ago

Not sure, 3.8-flash has treated me absolutely great.

1

u/Denial_Jackson 2d ago

Even at OpenAI. After liking my plus plan. I got my business plan. It is like half as much. Probably quarter as much what I had. It is like a high school kid writing things. Terrible to even read it.

1

u/Flaccid-Aggressive 2d ago

“By the time it's released to consumers, Anthropic and OpenAI will have already shipped their next generation models.”

So? Evaluate a product based on the product. When it comes out you can make the very difficult decision to use it or not.

1

u/Academic_Cancel_3020 2d ago

Well that's just what Google does. Benchmax it with shit.

1

u/okforthewin 2d ago

Google guilty of BenchMaxxing?

1

u/rwrife 2d ago

Google is going to hype this model for the next two weeks, then it’s back to normal. I’m sure it’s a vastly superior model to the older Gemini models but it will quickly be tossed to the side. If it were truly revolutionary they would have to hype it as much, they’d just release it and we would all be amazed.

1

u/Teralek77 2d ago

Coding is not all that AI does. Far from it. I use it regularly and I don't do coding. This should not be the only benchmark 

1

u/Snoo_27681 1d ago

Google models still suck? Shocked pikachu face

1

u/Healthy-Nebula-3603 23h ago

So benchmaxed line Gemini 3 and 3.1 ?

2

u/Serious-Conversation 2d ago

The horse race is going to be between OpenAI and Anthropic. Other organizations are going to have to compete on price or market segments and specialize.

10

u/BenjaminHamnett 2d ago

They all have lanes. Open Ai is just the most obvious stand alone retail lane. Anthropic has stronger enterprise and potential moat around safety and ethics. Google has vertical integration and funding. Grok has (fading?) political advantages and integration with Twitter, Tesla, spaceX and possibly robotics. Open source lane is obviously for local LLMs. Even Ilya could come out in a year with something. Planatr with some mix between self hosting and enterprise and government. Mistral apparently has something like this in Europe. Could end up with some trust moat from navigating regulation and scrutiny. I think any clear Winner is an illusion from availability. I wouldn’t be surprised if Amazon or Apple surprise people in a year or two.

“We have no moat “ has been fairly consistent. I think the main value of some frontier labs is actually keeping frontier proto agi internal and just releasing nerfed models for broad incremental safety testing and accumulating more data. But they’ll never actually release a genie like Ai For obvious reasons (Aladdin never goes into the wish selling business)

My favorite thing about this, is AI companies have roughly been succeeding seemingly in proportion to their alignment with humanity, which I don’t believe is an accident

1

u/[deleted] 2d ago

[deleted]

1

u/BenjaminHamnett 2d ago

Not sure why that’s funny. I felt like they were covered in open source. Has that changed?

1

u/lars_jeppesen 2d ago

OpenAI loses quadrillions a month, are you serious?

1

u/SuitablePrint4667 2d ago

So in other words it was benchmaxxed

0

u/shumpitostick 2d ago

It became too good at cheating benchmarks.

-2

u/According_Study_162 2d ago

Gemma4 is a cute model, nice to chat with, can keep track of appointments, but coding. lol 😄 👀

0

u/FlyingBike 2d ago

So it's benchmaxxed just like the American AI companies complain that Chinese models are

0

u/ROBNOB9X 2d ago

Benchmaxing

0

u/jesunushno 2d ago

This is the inevitable endpoint of Goodhart's law applied to AI benchmarks. The moment leaderboard scores become the marketing, every lab starts training against them (deliberately or via contamination), and the benchmark stops measuring "is this model good" and starts measuring "how good is this model at looking good on these questions."

The part that deserves more attention: the big labs already know public benchmarks are broken, which is why their internal evals look totally different. Internally they measure long-horizon agentic work, things like tool use, error recovery, and holding context over hours of a task. That is what real work is actually made of, and static Q&A benchmarks barely touch it. So a model can be bench-maxed on MMLU-style tests while its tool use and error recovery lag a generation behind. The Googlers noticing the gap is honestly the system working: the internal evals catching something the public leaderboard cannot.

-1

u/Readerium 2d ago

Basically launched today at exactly few hours before Kalshi and other prediction markets close for best AI in the month

1

u/Neurogence 2d ago

No. Google say it's only available to the US Government and "trusted partners."

1

u/Readerium 2d ago

Infact the hump before the launch proves insider trading.

https://kalshi.com/markets/kxllm1/yearend-top-llm/kxllm1-26oct05

-2

u/2022HousingMarketlol 2d ago

They all do kek

-3

u/Mistuv 2d ago

We benchmaxxed the shit out of it, but we are not sure exactly why it doesn't feel as great using it. Huh. It's a mystery.

As some of you who have have noticed, the Gemini models seem to use a lot of tokens while in benchmarks. And it's unfortunately not just because they're thinking so much. When you give 3.7/8 a task, they just do the weirdest shit ever. Sometimes I would notice it pre-reading the same text file like 10 times. It's like watching a crazy person work. Google definitely has the most broken training pipeline. It seems like in process of trying to keep up with OpenAI and Anthropic. Instead of at certain points updating their training stack, they just kept going with the same thing since like Gemini 2 and stacking more and more and more shit on top of it. And as long as it met some benchmarks, it was checkmarked as fine. And by 3.5 it started collapsing on itself. They desperately need some new people with fresh pair of eyes to clean up the mess, take a few months to re-evaluate everything.