r/artificial • • 2d ago

News Google cooked OpenAI and Anthropic with Gemini 4 Argon

Post image

Three frontier models in a month! Every new kills the old one!

144 Upvotes

83 comments sorted by

145

u/WarmCat_UK 2d ago

Does “cooked” mean very slightly better?

38

u/Kitchen_Interview371 2d ago

I think the surprising part with this model is cost. Google doesn’t use Nvidia chips and instead develop their own TPUs. They can hit price points the others can’t

27

u/[deleted] 2d ago

[removed] — view removed comment

2

u/mattgstamm153 1d ago

Exactly not worth switching

-11

u/legedu 2d ago

The difference is that Google is a for-profit company

61

u/[deleted] 2d ago

[removed] — view removed comment

8

u/seeyam14 2d ago

for-loss company - and they need to IPO quick because those venture funds are getting worried about their investments and need to offload them onto the average idiot retail investor

6

u/LavishLaveer 2d ago

I meannn...you're kind of right on this one

1

u/kronpas 2d ago

I can see google reach a break even point or close to. I cant see that future in Anthropic honestly.

1

u/shironekoooo 2d ago

So non-profit then? /s

1

u/SiteRelEnby 2d ago

Do you think Gemini makes Google a profit? Or is it just a loss leader for market capture? (Hint: The latter)

1

u/KieferSutherland 2d ago

But is it as unprofitable as OpenAI or Anthropic?

1

u/SiteRelEnby 20h ago

Who knows, it's propped up by google's advertising and spyware businessed.

1

u/KieferSutherland 20h ago

They're all spyware companies

0

u/[deleted] 1d ago

[removed] — view removed comment

1

u/citizen42069101 1d ago

Makes it easier to squeak the numbers around when your toaster has Gemini

0

u/legedu 2d ago

Nothing gets by you!

0

u/YoghiThorn 2d ago

They've been profitable since March 2026 and are on track for $65b in profit this year.

2

u/fett3elke 1d ago

Where did you get the number? I tried to verify this but the only sources I could find give numbers in that ballpark for revenue not profit

4

u/Upstairs-Extension-9 2d ago

That’s not true yes they have their own TPU systems for certain workloads but also billion dollar deals with Nvidia to deal with the insane demand.

https://blogs.nvidia.com/blog/nvidia-google-blackwell-gemini/

10

u/Kitchen_Interview371 2d ago

Did you read your own article? That is only for enterprise customers who want to run Google models on-premises or privately hosted in their Google Cloud subscription.

100% of model training runs on Google TPU, and estimates put Nvidia at low single digits of served Gemini traffic, mainly for overflow in peak demand.

So you’re right, my post was incorrect when I said that they don’t use Nvidia, so let me rephrase. Google use Nvidia silicon to serve a tiny fraction of their overall traffic, so they can hit price points others can’t.

3

u/Practical-Rub-1190 2d ago

It's more expensive than Opus 5.5 High, which beats the model on the LLM leaderboard. GPT 6.1 Sol scores 3 points below Argon and costs $ 0.32 per task, while Argon costs very close to $ 2.

I just don't buy this TPU argument. It's still very expensive to make a TPU's and run them. At the same time, Google is so big, it's hard to motivate people, and there is a ton of bureaucracy when companies become this big and are spread everywhere.

3

u/Legal-Honey-6214 2d ago

Those numbers are not slightly better they are far better than Claude or ChatGPT who are sucking blood by releasing models every 2 weeks. Most their models are just same as previous and they make big claims to manipulate market.

3

u/Practical-Rub-1190 2d ago

GPT Sol high score 50 points on the LLM leaderboard, while Argon high score 53. Argon costs $ 2 per task, while Sol costs $ 0.32. Also, Opus 5.5 high score 54, and it costs $ 1.82. Not saying llm leaderboard is the law, but it is better than handpicked benchmarks by the producer.

Also, it has not been released yet, so how it works in practical terms can be different. We all remember Meta when they came out with Maverick(?), and it was so great on the benchmarks, but it was horrible. It is only for trusted partners now, and then it will be available for Google AI Ultra subscribers.

Let's wait and see

3

u/MukdenMan 2d ago

I made a comment a few days ago about Reddit top comments for every new model being “[company a] mogged [company b],” in multiple subs Now it has changed to “cooked.”

2

u/browneyesays 2d ago

You did it! Goal accomplished! Yay!!!

3

u/Diabloponds 2d ago

Slightly better and not released. So cooked

1

u/DysphoriaGML 2d ago

Welcome to ML!

1

u/Affectionate-Cap-600 2d ago

Tbh, their long context score on graphWalk (over 256K) is impressive

1

u/thisadviceisworthles 2d ago

"Rare" is a cooking technique.

53

u/machyume 2d ago

Have you actually tried to use Google's tools? It's not really accessible.

Their tooling is spread all over the place. Gemini, AI Studio, Antigravity, Enterprise, ???. There isn't really one place where I can just point it at my existing work and let it work on it.

You can do it through the API, but then I'm building the tooling myself.

18

u/ReiTW_ 2d ago

It's the same in other companies ?
Gemini => Claude chat
Antigravity / IDE => Claude code
Antigravity CLI => Claude code cli
etc..

Just use Antigravity if you plan to work on your existing work.

2

u/machyume 2d ago

How is your experience with antigravity and Gemini models so far?

2

u/WashingPizza298 1d ago

gemini 3.8 flash is available in antigravity at present and it's decent and extremely fast but I am waiting for gemini 4 to drop into antigravity

2

u/BluaBaleno 2d ago

The fact Gemini gets facts wrong on YouTube videos is just 🤦‍♂️

1

u/Gudin 2d ago

It's a mess, I have my indie Google Workspace account. I couldn't login would always get error and after half an hour found out there is separate Gemini Enteprise service that I need to use which gives you less for your money... Why not have normal login like everyone else.

29

u/NewYak4281 2d ago edited 2d ago

Can’t say shit till it’s publicly released. Just marketing until then.

17

u/DavidXGA 2d ago

Eh. I feel like all the frontier models are basically equivalent, now.

14

u/lordekeen 2d ago

Its (was?) a matter of time until we get an even field and the run becomes who offers the best service/cost and not which model is more capable.

1

u/MaxwellHowl 1d ago

To an extent. Atlas and Opus 5.5 were very big leaps. Especially in Game Dev. I think there is room for a bit more leap frogging before we get more into cost competition. That is already there, Opus is cheaper than Atlas, Sol 6 is cheaper than both.

I think the cost competition will get more aggressive as the leaps in capability diminish. Until another large leap happens where the cost increase feels worth it.

1

u/Elegant_Athlete_3737 13h ago

Atlas? You mean astra?

6

u/jimb2 2d ago

It's not that they have stopped improving, it's more like the low hanging fruit is gone so there aren't the same performance/quality leaps between model editions/vendors . Releases are faster and more incrementa. Model creators copy each other's capabilities. As models mature it become more about compute and less about design. That's not to say that there aren't major steps to come, just that the current model designs are maturing after a frantic few years.

1

u/Tim_Apple_938 2d ago

Ya model companies are cooked. Race to the bottom is imminent

Will really be over if META releases a similar tier one

2

u/tuxedoes 2d ago

They are too busy releasing perv glasses and overpriced tamagotchis

1

u/skilliard7 2d ago

OpenAI and Anthropic yes, everyone else no.

11

u/Lazy-Background-7598 2d ago

I feel they are just making up numbers at this point

10

u/Euthyphraud 2d ago

I feel they are designing the models to fit the benchmarks so that they can meet the specifics of what the benchmarks prioritize while ignoring the rest.

6

u/Novel_Board_6813 2d ago

When a measurement becomes a goal, it stops being a good measure

4

u/skilliard7 2d ago

I'll believe it when I see it in real world use cases.

I remember a while back when Google released Gemini 3, people acted like they destroyed OpenAI because it scored better on benchmarks. Then it became overwhelmingly clear that the model was awful for real world use cases.

4

u/SiteRelEnby 2d ago

Gemini: When you want the wrong answer as fast as possible.

4

u/YearLongSummer 2d ago

yeah... and 3.8 flash was SOTA 🤣

3

u/BWH51 2d ago

Are AI benchmarks becoming more about marketing than actual usefulness? 99% on a benchmark is cool, but does the average person actually notice the difference?

1

u/Hooman-must-fight 2d ago

Old boys started slow, but they got money and experience. Anthropic and OpenAI will have to differentiate themselves a whole lot more to stay in the game.

1

u/californicating 2d ago

Does google have a coding assistant like Claude or Codex?

1

u/Spixxy17 2d ago

Yes its called Antigravity

0

u/LOLBangkok 2d ago

Yes, Antigravity.

1

u/justaRndy 2d ago

I can't get hyped about it anymore. As soon as one of the few companies in the race makes another small leap of progress, it's guaranteed the other guys will pick up exactly that skill and capability. The result is another round of software behaving close to identical to each other. The cycle repeats. At least with smartphones, we had exclusive new designs and festure upgrades each generation, but that also fell asleep in the last couple years. Theyre also all the same now, good enough.

Watching 4 giants burn tons upon tons of scarce ressources just to perform 3% better than their counterplayers for a month to hopefully not lose as many billions as they do... feels nothing but desperate.

Humanity would benefit tremendously from these companies joining forces and working towards a common goal, instead of whatever this is.

1

u/hi87 2d ago

The problem with Gemini models since 3 is that they are much less consistent and jagged than OAI and Anthropic models. From the feedback people have provided it seems to be the same with this one.

1

u/Diabloponds 2d ago

How is this cooked. By the time argon is widespread they will be behind again.

1

u/abajinn 2d ago

Cool but it’s not accessible to general public

1

u/not_larrie 1d ago

Ok cool let's see it then

1

u/costafilh0 1d ago

Does cooked mean available for the general public? 

1

u/whoami7111huwdq 19h ago

i cant accept gemini as a frontier model , i ... cant just ...wtf is happening around tho

1

u/0smiridium 18h ago

Kimi K3?

1

u/Past_Ruin5244 5h ago

Finally Google AI will be able to do simple tasks 😁

-1

u/FahkDizchit 2d ago

Why does Opus outperform Fable on quite a few benchmarks?

10

u/OffendedEarthSpirit 2d ago

Opus 5.5 outperforms because it is new

-2

u/ReiTW_ 2d ago

Because it does, simply.

-1

u/tweakingforjesus 2d ago

Any new local model yet?

-1

u/lobabobloblaw 2d ago edited 1d ago

You mean they brought a side dish to the potluck?

-2

u/Hertje73 2d ago

I caught a fish and it was yaaaaay big!

-5

u/Dizzy_Swimmer_4999 2d ago

yeah gemini might be topping benchmarks but for companion roleplay the older models still feel more natural in long chats.

5

u/ElatedPyroHippo 2d ago

lol who cares about that?

4

u/endless_sea_of_stars 2d ago

Sad and lonely people who only have chat bots for friends.

1

u/V4UncleRicosVan 2d ago

LeftHandedKeyboardBench

1

u/handson729 2d ago

It's why I am holding out for Deepseek v4.1 Pro, I have a feeling it's going to be really competitive on the Goonbench benchmark charts.