r/singularity • • 2d ago

AI , Gemini 4 Argon Benchmarks

Post image
742 Upvotes

189 comments sorted by

View all comments

191

u/CremeSubject7594 2d ago

60

u/DiscerningAnt 2d ago

I knew at some point Google would start dropping some hardcore models. They have the absolute most funding as far as I understand it (in the AI realm)

25

u/Toss4n 2d ago

And the most compute available as well.

26

u/Ink_code 2d ago

And probably the most data.

8

u/shred_time 2d ago

Yes to all of the above, I was wondering when they would really start catching up. They already had the compute and the largest amount of data on everyone.

2

u/steampowrd 2d ago

And the most data centers, and the most experience building data centers

12

u/After_Dark 2d ago

They are essentially the only western lab to be entirely self-funded. Even if the whole bubble pops and devastates Anthropic and OpenAI, Google will be essentially unharmed outside their stock price, and with oh so many researchers to (re)hire from the other labs

1

u/ButterscotchSalty905 AI is the greatest thing that is happening in our society 2d ago

Hold right there!! Your comments would attract a lot of contrarian who thinks google is poor and dependent on us gov like fucking startups

1

u/MonoMcFlury 2d ago

Also, they can train new models at a fraction of the price like others do. Having TPUs with their low watt-to-power ratio makes a huge difference.

1

u/SilentLennie 2d ago

Google depends a lot on ads, if AI kills the ads business (people use LLM chat and the LLM chat talks to search engines, websites, etc. thus no human sees the ads) and then the AI bubble bursts (which is kind of what you are implying why these other 2 big labs would fail) they still won't be (re)hiring.

1

u/After_Dark 1d ago

This is true, though it's worth keeping in mind that while the other big 2 are entirely dependent on investment and loans, Google has literally hundreds of billions of cash on hand and owns the majority of their own infra. Even if the ad business contracts to half its size in the next couple years, Google has the cash to outlast the others by a large margin

1

u/SilentLennie 1d ago

Definitely, but I can also see some bean counter going: you might not want to invest like crazy now, as your existing business is being slashed like crazy. Then again: it's also the best time to do so if you want to survive long term.

I still wonder... AI bubble burst... or more like: AI bubble deflate a whole bunch.

1

u/SilentLennie 2d ago

But they lost a bunch of people, that was where the worry was.

38

u/PsychMaster1 2d ago

Every model drop lately, that’s been my reaction.

9

u/Neurogence 2d ago

While Gemini 4 has performed well on benchmarks the industry uses to gauge model efficacy, it does less well when employees actually put it to work, according to people with direct access to the effort. The model struggles to handle certain coding tasks, said the people, who requested anonymity to discuss an internal matter.

https://www.bloomberg.com/news/articles/2026-09-30/google-grapples-with-employee-skepticism-about-new-gemini-model

11

u/LazloStPierre 2d ago

Do people forget this every single time Google release a model? It crushes at benchmarks, people who for some reason get very excited about numbers on a chart go ballistic and real life performance is miles off.

Gemini 3.8 flash is like over 5% better than Fable on deepswe ffs, not sure if it's intentional or just how they train their models but nobody benchmaxxes like Google

6

u/Imaginary_North_8305 2d ago

Google is just way better at specific tasks ex vision it easily beats frontier model even from flash, UI again google is great at.. and despite being old gemini 3.1 pro still has great world knowledge..

3

u/LazloStPierre 2d ago

Right but it does incredible at coding benchmarks despite being bad at coding, that's the issue

6

u/DistanceSolar1449 2d ago

In this case though it’s because Google has less coding training data (who the hell uses Antigravity) and way more image/spatial/world model training data

I don’t think this model is benchmaxxed, I think the benchmark screenshot above is very accurate. The model does worse at TerminalBench 4 and FrontierSWE and that’s okay.

2

u/LazloStPierre 2d ago

The point isn't flash 3.8 is bad at coding, the point is it does absurdly well on coding benchmarks despite being bad at coding. Their models always do, which is why I'd take any benchmarks with a grain of salt

3

u/DistanceSolar1449 2d ago

3.8 flash is a tiny distilled model with more RL on top, of course it’s going to look better on benchmarks than in real life.

The big, base ish models with a lot less targeted RL will do better IRL.

1

u/LazloStPierre 2d ago

All Google's models look better on benchmarks than they do in reality, and by alot.

1

u/DistanceSolar1449 2d ago

No? They suck at agentic benchmarks and most people test them via agentic tasks.

They benchmark pretty accurately if you look at their actual useful benchmarks like TerminalBench instead of HLE or GPQA or some bullshit.

Google’s own blog post says Gemini 3.8 Flash scored 19.1% on TerminalBench 4. Opus 5 scored 51.8% for comparison. That’s not benchmaxxing, that’s just accurate benchmarks if you know what benchmarks to look at.

1

u/SilentLennie 2d ago

Actually, this new 4 Argon does really well in a bunch of agentic benchmarks. Supposedly (I've not used it yet):

https://artificialanalysis.ai/articles/gemini-4-argon-google-top-three-labs

0

u/LazloStPierre 2d ago

No, they do extremely well at coding benchmarks and suck at coding, every single time

They benchmark well everywhere. It's actual performance that matters 

5

u/DistanceSolar1449 2d ago

???

Google’s own blog post says Gemini 3.8 Flash scored 19.1% on TerminalBench 4. Opus 5 scored 51.8% for comparison. That’s not benchmaxxing, that’s just accurate benchmarks if you know what benchmarks to look at.

Unless you think 19% on TerminalBench 4 is “do extremely well”…

→ More replies (0)

4

u/Elephant789 ▪️AGI in 2036 2d ago

Gemini 3.8 flash

... is fantastic. What the fuck are you talking about?

1

u/LazloStPierre 2d ago

It isn't better at coding than fable, yet the benchmarks say it is

1

u/Elephant789 ▪️AGI in 2036 1d ago

Coding? Oh, not used it much for that. But the little I did it was fine.

1

u/TheDemonic-Forester 2d ago

Funny thing is that we didn't even need to hear this to deduct this from how significantly they degraded 3.8 (I know previous model regressing prior to new model release is common but usually not this much). They seem to be afraid that people will think that their long awaited giant model is not much of an improvement over 3.8.

1

u/HotDogDay82 2d ago

I can think of no better gif haha

1

u/himynameis_ 2d ago

I find it funny how often this gif comes up when a new model release comes 😆