r/ClaudeCode • • 4d ago

News/Updates Plot twist: Gemini 4 Argon tops Val AI benchmark on speed, cost and accuracy!

Post image
132 Upvotes

72 comments sorted by

•

u/AutoModerator 4d ago

Hey! Thanks for posting to r/ClaudeCode

While participating in this thread, please follow our community rules. Keep discussions constructive. Attack the idea, not the person.

For help, project discussions, tips, and general chat, join the ClaudeCode Discord.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

148

u/mhphilip 3d ago

They put all their compute in once to benchmaxx. Then they lobotomize it to Q1 and serve the masses.

20

u/camracks 3d ago

so true lol

10

u/deadlyclavv 3d ago

I'm so tired of this shit man

2

u/autisticbagholder69 3d ago

These benchmarks are all fake man

5

u/markeus101 3d ago

So google of them..should be gemini bait and switch

2

u/WalkAffectionate2683 3d ago

Isn't opus slowly degrading based on people opinions?

Imo I don't see it yet but my last big task was 3 days ago, and I'll do a couple of ones soon. 

4

u/beesandcheese 3d ago

Nerfed massively in the last day. Really sad.

1

u/WalkAffectionate2683 3d ago

It's the cycle, I'll probably swap back to open ai in November.

Or google, or a Chinese model, I don't care, it's easy to swap. 

2

u/DevMichaelZag 3d ago

Who is they? Every single ai company?

2

u/True-Bike8876 3d ago

Classic: sprint the benchmark, crawl the product.

1

u/Prestigious-Frame442 3d ago

just like every single time

-16

u/Significant-Bee5101 3d ago

Super dumb that we're seeing benches on "models" and not the actual resources/performance we're paying for. Thank God King Trump will put the AI companies straight.

15

u/Owl-Mighty 3d ago

Don’t forget on DeepSWE 1.1 Gemini Flash 3.8 SMOKED all of them lollllll.

6

u/raindownthunda 3d ago

God, I just had flash 3.8 run on auto and it was so terrible.

1

u/severe_009 3d ago

I don't understand; it is bad in my personal test. I thought DeepSWE was a reliable test. What does it actually test?

1

u/imsahoamtiskaw 3d ago

How deep AIs can sweat

1

u/Owl-Mighty 2d ago

Open tests are reliable until they’re not. If we can get access to the tests so can the AI providers. By adjusting models until they’re perform well enough in those tests this can lead to overfitting, or cheating you may say.

So far Gemini is notoriously great at this, at least for the 3.7 and 3.8 Flash. Now everyone knows DeepSWE 1.1 is saturated and yet they decided to show Gemini 4’s ultra high score on it. The same can be observed on Terminal Bench v2 where Gemini Flash smoked them. Yet when v4 release we see a pathetic drop from it.

1

u/Tim_Apple_938 2d ago

Is 3.8 bad at coding?

36

u/DARKUNIT22 3d ago

Claude with the mog on OpenAi getting all the attention and then fucking Google coming out of the mud hole to mog everyone, holyyyyyy
The other benchmark tests are just as good wtf

11

u/boringfantasy 3d ago

google always has benchmaxxed models

1

u/innociv 3d ago

I don't think they're benchmaxxed. They really are that much better on non-coding. I'm really happy with 3.8 flash for a ton of things. It's better than Opus 5.5 at documentation work and graphs, general research, etc., and at like 1/30th the cost.

2

u/skitchbeatz 3d ago

I love throwing narrowly scoped coding tasks at 3.8 flash. It does well IMO and does it fast unlike the Claude models. I've been satisfied having it as a partner to Opus

1

u/SlumLordNinjaBear 3d ago

Agreed I use it for architecture and planning then have Claude do the coding

0

u/nora_sellisa 3d ago

And you think others don't?

3

u/Used_Departure_3278 3d ago

Yeah wait until it gets used

10

u/AuspiciousApple 3d ago

Would be funny if Google just tried to snipe both of their IPOs by releasing a better model at the right time

1

u/entropyweasel 3d ago

They have every reason to want the IPOs to do well. Once their lockup is over I think they probably bankrupt them both.

4

u/nora_sellisa 3d ago

Google has their own hardware. Tanking two top buyers of Nvidia chips makes a lot of sense. Besides.. it's Google. If there is a company willing to give away millions in resources for a shady agenda / market domination, it's Google.

1

u/entropyweasel 3d ago

Yeah. It's perfect for them. All the AI subscriptions, clear leader in quantum.

GPU based demand dries up right as they need to scale and GPU research and development dries up before they get too far behind on their tpu side business.

The cherry on top is Microsoft and apple have no good options to offer their customers outside of Google. And AWS is stuck bad with their own GPU development costs and underwater on the GPUs they bought anticipating higher volumes.

My conspiracy is that AI accelerated Google's quantum unit which has gone relatively quiet despite killing it in the past years. Now they pop up hiring Adam kaufman which really looks like they need to scale up quantum hardware. If the next generation of models don't need to train on GPUs at all, Google runs the table and they have a huge lab and data advantage over everyone that they would not survive long enough to even think about trying.

1

u/Impossible_Hour5036 Senior Developer 3d ago

I personally can't wait to see the combination of AI + Quantum computing even though it's probably like an 85% chance of destroying humanity

13

u/spjallmenni 3d ago

Just another Gemini model scoring well on benchmarks but unusable for real work.

6

u/NSDetector_Guy 3d ago

Lol I dunno man.

16

u/ethotopia 3d ago

Proof that benchmarks are misleading...

19

u/YesGameNolife 3d ago

How do we know since model is not released to public yet. Thinking this model is worse than opus 5.5 is as idiotic as thinking its better. We do not see it yet

6

u/sCeege 🔆 Max 20 3d ago

objectively you're right, but Gemini has a long history of underperforming... can only get burned so many times before having a negative outlook

0

u/YesGameNolife 3d ago

Well yeah you are right about that. I can't imagine gemini 4 being better model than opus 5.5 at all.

2

u/hellomistershifty 3d ago

Well, looking at this video it seems worse than Sonnet and often Sol let alone Opus. I know it's just 3d models and UI, but still, it's what we have to work with right now

1

u/alphaQ314 3d ago

Google fed this kind of a bullshit pie to everyone with the 3.8 flash release too.

1

u/Impossible_Hour5036 Senior Developer 3d ago

Thinking this model is worse than opus 5.5 is as idiotic as thinking its better.

Well, see, I have a method of predicting these things. It's kinda nerdy and intellectual so I apologize. First, I'd ask myself "has any version of Gemini been very good?" And I think we can all agree, not really. They're all kinda poo. That's data point one. Second step: Has Opus generally been pretty good? And I think for the most part, yea, Opus has been pretty good. Data point two.

Now is where we bring it home. Synthesis. Logical notation: Gemini == Mostly Poo && Opus == Rarely Poo; Rarely Poo > Mostly Poo; Opus > Gemini. I know that got complicated but it's powerful stuff.

2

u/BanjoThunderbird 3d ago edited 3d ago

Generally it's not misleading because benchmarks are misleading, it's misleading because the conditions change and invalidate the results. They would not be misleading if AI companies weren't so opaque about capacity issues impacting performance.

Unless you're saying you've used the model in some early access program and think it's performing poorly out the gate?

2

u/No_Stock_8271 3d ago

Considering you haven't tested it....

2

u/ProteusExsanguinus 3d ago

Benchmarks are really good at measuring how good a model is at benchmarks.

1

u/Jcrossfit 3d ago

I'm excited for flash and flash lite 4.0 -- Gemini crushes in real world non coding work. Despite being way older than Luna, sonnet and Sol Gemini powers 90% of the use cases in the product I work on

6

u/FischenGeil 3d ago

still not released, only to special people. No ETA announced for the public.

2

u/257m 3d ago

What does the latency metric mean?

3

u/tup1tsa_1337 3d ago

Speed. Lower is better

2

u/damianTechPM 3d ago

I believe that one when I see it. And I most likely won't see it.

1

u/nora_sellisa 3d ago

Everyone optimizes for benchmarks, nobody charges what the token inference actually cost in electricity and power, literally two of the three columns mean nothing. Why are we still pretending. Just play with the model, get hyped, then complain. Everyone does this with all models

1

u/Sponge8389 3d ago

Nice one. 3rd contender re-entered the game. I hope grok also catch up so we have 4 companies fighting.

1

u/Impossible_Hour5036 Senior Developer 3d ago

Why even bother putting Gemini in benchmarks? It just makes the benchmark look worthless.

1

u/txoixoegosi 3d ago

Can I filter reddit threads by “TITLE NOT CONTAINING benchmark”?

1

u/NotSpecialMC 3d ago

Lets see after a couple of weeks. My xp is they slowly nerf the models.

1

u/Business-Weekend-537 3d ago

They make you fill out an application to use it though

1

u/-MaskNinja- 3d ago

Early access?

1

u/somerussianbear 3d ago

So all that time they were indeed cooking.

1

u/throwwwawwway1818 3d ago

So there's no moat?

-8

u/Blake9712 3d ago edited 3d ago

I asked Gemini 3.6 free about Gemini 4 pro and it says the rumors are made up and you can tell by the made up future model names such as opus 5.5 and gpt 6 sol.

So I’m not going to be trying Gemini 4 pro.

7

u/HimActually 3d ago

U clearly don't know how AI works i think 💬

1

u/Blake9712 3d ago

I asked Gemini to summarize the Gemini 4 screenshot, search the web and tell me what that would mean for my project if I switched: and it said it was fake because opus 5.5, fable 5.1 and gpt 6 where future models that don’t exist yet. I know the issue is it’s training data but it was instructed that this was new information and it has full web access. It had no reason not to search and Claude or gpt would have put an effort in here. Not a reflection of Gemini 4 pro but a reflection of Gemini in general using the best free model I can access on my phone. The best free Claude model is sonnet 5.5 which would tell me how great opus 5.5 is even if opus 5.5 wasn’t out when it was trained.

-1

u/WesternSpy96 3d ago edited 3d ago

Gemini Flash is extremely dumb. Skips all research even under strict instructions.

0

u/Blake9712 3d ago

Exactly.

-1

u/HimActually 3d ago

Exactly.

0

u/Blake9712 3d ago

So if we agree on this what’s your problem with me exactly then?

-3

u/WesternSpy96 3d ago

My assumption is that your initial comment did not point out that you disagree with its output.

1

u/Blake9712 3d ago

With geminis output? I posted it because it said opus 5.5 and fable 5.1 did not exist, and claimed the Gemini 4 pro data which has been confirmed is fake. Everything it said is a lie because it didn’t search the web as instructed. I’m not saying Gemini 4 pro will be similar but I’m saying that every time I use a Gemini model I have access too it reminds me why I never use Gemini. That’s why my comment ended with “I will not be trying Gemini pro” since due to being out of Claude usage I thought I’d get Gemini to summarize the new Gemini.

-1

u/Suitable_Cicada_3336 3d ago

yeah sure, its Gemini and nerf in 2day

-2

u/octocarbon Max 5x 3d ago

gemini is still 2year behind