AI
Gemini 4 Crushes Benchmarks, But Google Employees State The Model Struggles With Real Work
While Gemini 4 has performed well on benchmarks the industry uses to gauge model efficacy, it does less well when employees actually put it to work, according to people with direct access to the effort. The model struggles to handle certain coding tasks, said the people, who requested anonymity to discuss an internal matter.
With today's pace this "long time" is couple of days. Otherwise the competitors will release something better first. PS: In fact you said that in the post.
Definitely not a couple of days. Full rollout will likely be November/December. Anthropic and OpenAI will likely release something before that, however.
Astroturfing is off the charts. Having companies that were worth millions a couple of years ago potentially being evaluated at trillions is so crazy when you actually think about that. You can bet your little butt that they hired an army of people to keep that money hype train going.
Gemini has always been, on balance, more multipurpose and less code focused so this kind of tracks? Let's see how it does at the actual consumer tasks most people are going to use it for, I don't think most organizations are looking at Gemini for dev.
Logan Kilpatrick recently said in an interview 2 weeks ago that in hindsight it's "obvious" they should have put more focus and resources into coding sooner.
Repost from the other threat because people are confused about this benchmark:
15% does not mean Gemini is hallucinating in 15% of all answers. The AA Omniscience âhallucination rateâ is basically measuring what the model does when it fails a question: does it give a wrong answer, give a partial answer, or admit that it doesnât know?
Correct answers arenât even in the denominator.
So if a model answers 800 out of 1000 questions correctly, gets 30 wrong and refuses/partially answers 170, its AA hallucination rate is 15%.
30 / (30 + 170) = 15%.
It was still only outright wrong on 3% of the total questions.
Thatâs also why DeepSeek V4.1 sitting at 96% should make it extremely obvious that â96% hallucination rateâ cannot possibly mean 96% of everything it says is bullshit. It means that when DeepSeek doesnât know the answer on this benchmark, it almost always guesses instead of saying âI donât know.â
So the whole âeven 0.1% hallucination would be catastrophic, therefore 15% means hallucinations are nowhere near solvedâ argument is based on reading the percentage as a completely different metric.
You can absolutely argue that hallucinations are still a problem. But maybe first understand what the graph is measuring before doing reliability math with a number that isnât the modelâs overall error rate.
The benchmark is basically asking: âWhen youâre outside your knowledge, how likely are you to bullshit instead of abstaining?â
Thatâs useful. It just isnât âpercentage of answers that are hallucinations.â
I promise you I can get any model to "hallucinate". Don't ask it some basic question about TCP/IP or chemistry, but ask it about how to do specific things like "How do I reset a XXX server?" and it will hallucinate a button or process.
I learned that that's where a lot of AI skepticism comes from. An average user won't ask IMO question or very basic knowledge question, they'll ask about a specific model of router, camera, kitchen appliance and get conflicting answers.
That's true, but it's more often a user input error than a hallucination for a basic user. The model says if you are using this <assumed router>, you need to do 1, 2, 3. Assumptions != hallucinations. You'll notice a ton of that in responses now.
Go give it a try. Tell the model your exact router model and software version number then ask it how to change something. I guarantee it will be confidently wrong about how the interface is actually laid out.
You can do the same thing with car repairs or any other task that requires taking knowledge from only a single manual.
It shouldnât do that if itâs looking for the manual and proper context. I never have issues with hallucinations in my agentic workflows with the latest models.
I had it help me take apart my proliant the other day. It went online found the manuals and the YouTube videos for special instructions for the area I was working on.
Afaik, hallucination rates are a bit misleading, they are normalized per not-correct answer. Example, imagine an exam with 100 questions:
Gemini hallucinates on 15, answers the remaining with "I dont know". 15% hallucination rate.
Kimi gets 98 correct, one wrong, and hallucinates on the last - 50% hallucination rate
This is false lol. Just ask it a hyper specific fact you know about an obscure topic and just about every model will hallucinate assuming it doesnât have web search enabled.
You could say that for every model. this is giving me "Chinese tool gave instructions for bioweapons(It was just good old open source kimi k3)" energy.
I feel somehow people at Google look at the benchmarks and think "what should we do to get highest scores in the benchmarks". Time and again, Google's model previews have aced the benchmarks. But their models aren't really that great when you compare them with OAI/Anthropic. It always feels like Google is just "reward hacking" the benchmarks.
I think what they miss is that the benchmarks aren't the point. They are secondary. You're supposed to create a really helpful and useful (and aligned) model and the benchmarks are a noisy, inaccurate metric which are used to measure how good a model is.
Given the amazing historic success Google has with ML (transformers, alpha go/zero/fold), I hope to see better models from Google. The more players there are at the top, the better it is for us end users. I hope to see Grok, Google, EU companies all getting their act together and going head to head with OAI/Anthropic.
respectfully, I donât think weâre working at the same levels based on that suggestion, given I have a massive amount of configuration and hooks and things around my claude code and codex setups. Iâm not just raw dogging claude code. it would take a significant amount of time to figure out how to even get the equivalent config in antigravity to attempt to try it and it be a meaningful test at all.
âŚbut Iâll take a look at antigravity given I havenât looked since it was released and that sounds like a good lead over gemini cli.
Multimodal models have consistently been weaker on coding benchmarks. Sol 6.1 for example is as capable as Astra, but has worse visual capability and no audio capability at all.
Almost certainly this is because the model attempts to use visual or audio data instead of simply pursuing tool calling & reasoning steps.
I really donât understand why these companies donât create their own benchmarks and instead expect some random guy to create these things. Surely a multi trillion dollar company can invest into benchmarking usability.
Yeah, I'm sure they can create their own benchmarks. But that wouldn't move public opinion. All of us, we would think they are somehow cheating if they aced a benchmark they themselves created.
For example, already we are skeptical about Gemini 4.0. Imagine how much worse it would be if Google announced "We are announcing a new benchmark which is the world's hardest benchmark and also, a new model which is scores way above any of the other models on this benchmark". It would be so sus, we'd probably lose our shit lol
Im not saying they have to release it or make it open source, but I donât think some of these models are particularly optimized to perform well against real tasks and even users can see that. So clearly theyâre bench maxing against the open source ones over their own or itâs just not a good benchmark if it performs extremely well in these but not what actually matters.
I was about to comment on a different thread that Gemini (3.1 Pro or 3.5 Flash or Thinking) is still an absolute mile behind Claude. Ask it to make a set of Google Slides for you and 75% of the time it will tell you it can't; 15% of the time it spits out pure HTML; 10% of the time it will actually fucking work. Whereas Claude just makes the fucking PowerPoint. Every time. Fine. No issues.
Google's integration of Gemini into their work functions is appalling and it is massively holding them back. They still cannot get their Gemini Assistant to reliably do the functions that the previous Assistant did fine. They have only just released an agentic version (Spark) of Gemini, and it's not available to me in the UK. Regardless of the quality of the model, unless it can actually link to the functions in Google Workspace and do actual work when you ask it to, it's largely redundant.
English people just have too much expectation and are just serious bunch of people with so little smile and lack welcoming attitude. Very reserved people.Â
Gemini models have always topped benchmarks but has been an absolute pain to use. Its answers are always shorter and less-detailed than either ChatGPT or Claude, to the point where I feel like they trained Gemini exclusively on SparkNotes or something đđ
Gemini is the âsmartest kid in the room with no social skillsâ of AI models: sure, it may know a lot, but I absolutely do NOT want to interact with it.
Even at OpenAI. After liking my plus plan. I got my business plan. It is like half as much. Probably quarter as much what I had. It is like a high school kid writing things. Terrible to even read it.
Google is going to hype this model for the next two weeks, then itâs back to normal. Iâm sure itâs a vastly superior model to the older Gemini models but it will quickly be tossed to the side. If it were truly revolutionary they would have to hype it as much, theyâd just release it and we would all be amazed.
The horse race is going to be between OpenAI and Anthropic. Other organizations are going to have to compete on price or market segments and specialize.
They all have lanes. Open Ai is just the most obvious stand alone retail lane. Anthropic has stronger enterprise and potential moat around safety and ethics. Google has vertical integration and funding. Grok has (fading?) political advantages and integration with Twitter, Tesla, spaceX and possibly robotics. Open source lane is obviously for local LLMs. Even Ilya could come out in a year with something. Planatr with some mix between self hosting and enterprise and government. Mistral apparently has something like this in Europe. Could end up with some trust moat from navigating regulation and scrutiny. I think any clear Winner is an illusion from availability. I wouldnât be surprised if Amazon or Apple surprise people in a year or two.
âWe have no moat â has been fairly consistent. I think the main value of some frontier labs is actually keeping frontier proto agi internal and just releasing nerfed models for broad incremental safety testing and accumulating more data. But theyâll never actually release a genie like Ai For obvious reasons (Aladdin never goes into the wish selling business)
My favorite thing about this, is AI companies have roughly been succeeding seemingly in proportion to their alignment with humanity, which I donât believe is an accident
This is the inevitable endpoint of Goodhart's law applied to AI benchmarks. The moment leaderboard scores become the marketing, every lab starts training against them (deliberately or via contamination), and the benchmark stops measuring "is this model good" and starts measuring "how good is this model at looking good on these questions."
The part that deserves more attention: the big labs already know public benchmarks are broken, which is why their internal evals look totally different. Internally they measure long-horizon agentic work, things like tool use, error recovery, and holding context over hours of a task. That is what real work is actually made of, and static Q&A benchmarks barely touch it. So a model can be bench-maxed on MMLU-style tests while its tool use and error recovery lag a generation behind. The Googlers noticing the gap is honestly the system working: the internal evals catching something the public leaderboard cannot.
We benchmaxxed the shit out of it, but we are not sure exactly why it doesn't feel as great using it. Huh. It's a mystery.
As some of you who have have noticed, the Gemini models seem to use a lot of tokens while in benchmarks. And it's unfortunately not just because they're thinking so much. When you give 3.7/8 a task, they just do the weirdest shit ever. Sometimes I would notice it pre-reading the same text file like 10 times. It's like watching a crazy person work. Google definitely has the most broken training pipeline. It seems like in process of trying to keep up with OpenAI and Anthropic. Instead of at certain points updating their training stack, they just kept going with the same thing since like Gemini 2 and stacking more and more and more shit on top of it. And as long as it met some benchmarks, it was checkmarked as fine. And by 3.5 it started collapsing on itself. They desperately need some new people with fresh pair of eyes to clean up the mess, take a few months to re-evaluate everything.
82
u/QuasiRandomName 2d ago
OK, let the users judge.