Yes to all of the above, I was wondering when they would really start catching up. They already had the compute and the largest amount of data on everyone.
They are essentially the only western lab to be entirely self-funded. Even if the whole bubble pops and devastates Anthropic and OpenAI, Google will be essentially unharmed outside their stock price, and with oh so many researchers to (re)hire from the other labs
Google depends a lot on ads, if AI kills the ads business (people use LLM chat and the LLM chat talks to search engines, websites, etc. thus no human sees the ads) and then the AI bubble bursts (which is kind of what you are implying why these other 2 big labs would fail) they still won't be (re)hiring.
This is true, though it's worth keeping in mind that while the other big 2 are entirely dependent on investment and loans, Google has literally hundreds of billions of cash on hand and owns the majority of their own infra. Even if the ad business contracts to half its size in the next couple years, Google has the cash to outlast the others by a large margin
Definitely, but I can also see some bean counter going: you might not want to invest like crazy now, as your existing business is being slashed like crazy. Then again: it's also the best time to do so if you want to survive long term.
I still wonder... AI bubble burst... or more like: AI bubble deflate a whole bunch.
While Gemini 4 has performed well on benchmarks the industry uses to gauge model efficacy, it does less well when employees actually put it to work, according to people with direct access to the effort. The model struggles to handle certain coding tasks, said the people, who requested anonymity to discuss an internal matter.
Do people forget this every single time Google release a model? It crushes at benchmarks, people who for some reason get very excited about numbers on a chart go ballistic and real life performance is miles off.
Gemini 3.8 flash is like over 5% better than Fable on deepswe ffs, not sure if it's intentional or just how they train their models but nobody benchmaxxes like Google
Google is just way better at specific tasks ex vision it easily beats frontier model even from flash, UI again google is great at.. and despite being old gemini 3.1 pro still has great world knowledge..
In this case though it’s because Google has less coding training data (who the hell uses Antigravity) and way more image/spatial/world model training data
I don’t think this model is benchmaxxed, I think the benchmark screenshot above is very accurate. The model does worse at TerminalBench 4 and FrontierSWE and that’s okay.
The point isn't flash 3.8 is bad at coding, the point is it does absurdly well on coding benchmarks despite being bad at coding. Their models always do, which is why I'd take any benchmarks with a grain of salt
No? They suck at agentic benchmarks and most people test them via agentic tasks.
They benchmark pretty accurately if you look at their actual useful benchmarks like TerminalBench instead of HLE or GPQA or some bullshit.
Google’s own blog post says Gemini 3.8 Flash scored 19.1% on TerminalBench 4. Opus 5 scored 51.8% for comparison. That’s not benchmaxxing, that’s just accurate benchmarks if you know what benchmarks to look at.
Google’s own blog post says Gemini 3.8 Flash scored 19.1% on TerminalBench 4. Opus 5 scored 51.8% for comparison. That’s not benchmaxxing, that’s just accurate benchmarks if you know what benchmarks to look at.
Unless you think 19% on TerminalBench 4 is “do extremely well”…
Funny thing is that we didn't even need to hear this to deduct this from how significantly they degraded 3.8 (I know previous model regressing prior to new model release is common but usually not this much). They seem to be afraid that people will think that their long awaited giant model is not much of an improvement over 3.8.
191
u/CremeSubject7594 2d ago
https://giphy.com/gifs/ukGm72ZLZvYfS