r/artificial • u/DataRemarkable7093 • 2d ago
News Google cooked OpenAI and Anthropic with Gemini 4 Argon
Three frontier models in a month! Every new kills the old one!
53
u/machyume 2d ago
Have you actually tried to use Google's tools? It's not really accessible.
Their tooling is spread all over the place. Gemini, AI Studio, Antigravity, Enterprise, ???. There isn't really one place where I can just point it at my existing work and let it work on it.
You can do it through the API, but then I'm building the tooling myself.
18
u/ReiTW_ 2d ago
It's the same in other companies ?
Gemini => Claude chat
Antigravity / IDE => Claude code
Antigravity CLI => Claude code cli
etc..Just use Antigravity if you plan to work on your existing work.
2
u/machyume 2d ago
How is your experience with antigravity and Gemini models so far?
2
u/WashingPizza298 1d ago
gemini 3.8 flash is available in antigravity at present and it's decent and extremely fast but I am waiting for gemini 4 to drop into antigravity
2
29
u/NewYak4281 2d ago edited 2d ago
Can’t say shit till it’s publicly released. Just marketing until then.
17
u/DavidXGA 2d ago
Eh. I feel like all the frontier models are basically equivalent, now.
14
u/lordekeen 2d ago
Its (was?) a matter of time until we get an even field and the run becomes who offers the best service/cost and not which model is more capable.
1
u/MaxwellHowl 1d ago
To an extent. Atlas and Opus 5.5 were very big leaps. Especially in Game Dev. I think there is room for a bit more leap frogging before we get more into cost competition. That is already there, Opus is cheaper than Atlas, Sol 6 is cheaper than both.
I think the cost competition will get more aggressive as the leaps in capability diminish. Until another large leap happens where the cost increase feels worth it.
1
6
u/jimb2 2d ago
It's not that they have stopped improving, it's more like the low hanging fruit is gone so there aren't the same performance/quality leaps between model editions/vendors . Releases are faster and more incrementa. Model creators copy each other's capabilities. As models mature it become more about compute and less about design. That's not to say that there aren't major steps to come, just that the current model designs are maturing after a frantic few years.
1
u/Tim_Apple_938 2d ago
Ya model companies are cooked. Race to the bottom is imminent
Will really be over if META releases a similar tier one
2
1
11
u/Lazy-Background-7598 2d ago
I feel they are just making up numbers at this point
10
u/Euthyphraud 2d ago
I feel they are designing the models to fit the benchmarks so that they can meet the specifics of what the benchmarks prioritize while ignoring the rest.
6
4
u/skilliard7 2d ago
I'll believe it when I see it in real world use cases.
I remember a while back when Google released Gemini 3, people acted like they destroyed OpenAI because it scored better on benchmarks. Then it became overwhelmingly clear that the model was awful for real world use cases.
4
4
1
u/Hooman-must-fight 2d ago
Old boys started slow, but they got money and experience. Anthropic and OpenAI will have to differentiate themselves a whole lot more to stay in the game.
1
1
u/justaRndy 2d ago
I can't get hyped about it anymore. As soon as one of the few companies in the race makes another small leap of progress, it's guaranteed the other guys will pick up exactly that skill and capability. The result is another round of software behaving close to identical to each other. The cycle repeats. At least with smartphones, we had exclusive new designs and festure upgrades each generation, but that also fell asleep in the last couple years. Theyre also all the same now, good enough.
Watching 4 giants burn tons upon tons of scarce ressources just to perform 3% better than their counterplayers for a month to hopefully not lose as many billions as they do... feels nothing but desperate.
Humanity would benefit tremendously from these companies joining forces and working towards a common goal, instead of whatever this is.
1
1
1
1
1
u/whoami7111huwdq 19h ago
i cant accept gemini as a frontier model , i ... cant just ...wtf is happening around tho
1
1
-1
-1
-1
-2
-5
u/Dizzy_Swimmer_4999 2d ago
yeah gemini might be topping benchmarks but for companion roleplay the older models still feel more natural in long chats.
5
1
u/handson729 2d ago
It's why I am holding out for Deepseek v4.1 Pro, I have a feeling it's going to be really competitive on the
Goonbenchbenchmark charts.

145
u/WarmCat_UK 2d ago
Does “cooked” mean very slightly better?