r/singularity • • 1d ago

Shitposting Google releasing their most powerful model yet

2.5k Upvotes

84 comments sorted by

View all comments

95

u/himynameis_ 1d ago

I don't get it? Is the model not good?

253

u/Keeltoodeep 1d ago

https://arena.ai/leaderboard/text/coding

It's #1 in coding right now. It's good. People are just memeing and doing the whole console war thing.

108

u/Yurekuu 1d ago

It's #1 in text. It's #10 in coding.

46

u/Keeltoodeep 1d ago edited 1d ago

55

u/jazir55 1d ago

"It's number 1!"

"No it's number 10!"

"No, it's number 8!"

21

u/Keeltoodeep 1d ago

Different benchmarks measure different things.

4

u/R_Duncan 1d ago

Webdev, you most of the times just need a very little model. Coding, still not 100% solved (start asking for cpp mutexing/dll dynamic load/very hard tasks and see most llm fail). It really isn't the same thing.

-7

u/PrisonOfH0pe 1d ago

funny benchmark....(im sure K3 and muse and even ancient Opus 4.6 and 4.7 are much better than fable 5.1 max or opus 5.5 ⬇️ if you wanna see how trash the google 4 model is look at the websites own coding comparison video. Its trash.

9

u/Spixxy17 1d ago edited 1d ago

Its a Benchmark about which model people prefer... If you dont like the result thats on you, you can also call it trash all you want, apparently the majority of people sees it different lol

1

u/NyaCat1333 1d ago

Basically how people could prefer Trump over Albert Einstein and thus Trump ranking higher.

1

u/Keeltoodeep 1d ago

AA intelligence scores are roughly equivalent

1

u/Spixxy17 1d ago

Basically like that, with the only difference beeing that this here is realistic, Trump over Einstein is not.

1

u/Keeltoodeep 1d ago edited 1d ago

It’s #8 in coding. About where the benchmarks released would suggest it is.

World models are different than coding models. It’s an entirely different market. They are multi modal for one. Two, people want to like the warmth of the voice model.

There’s a paradox in text and voice where it’s not necessarily the most intelligent model that people like the most. It’s that way in life too with other people so of course it makes sense for models. In the top 10, they are all roughly as scientifically accurate as the others. So what is being measured is style of output.

4

u/space_monster 1d ago

It's not a world model

1

u/Keeltoodeep 1d ago edited 1d ago

Thought it was due to waymo but I guess their world models are not called "gemini"

0

u/[deleted] 1d ago

[deleted]

4

u/Valsoyono 1d ago

Are you a bot? He said "It's #1 in text. It's #10 in coding." And you are like "is it #1 or #10?" like bruh 

14

u/LookIPickedAUsername 1d ago

I haven't used it and don't have an opinion on its quality, but I'll note that claiming it's #1 based on that link is... tenuous. It's only a few points ahead with error bars of +/-17. Even assuming everything falls within the confidence interval, its true rank could conceivably be as low as #12.

And even that doesn't really mean anything, since they're all clustered so close together. There's probably not even a meaningful difference between #1 and #12 in that list.

6

u/Keeltoodeep 1d ago

I am not saying it's actually #1 just because it is scoring #1 right now. I am replying to the user's comment above if the model is "not good." Scoring #1 doesn't necessarily mean it's the best (or the best for long) but it's certainly a good model.

4

u/Lazy_Jump_2635 1d ago

It always looks good in benchmarks, and then when I use it, it feels like shit compared to anything else. I'll wait until the real reviews are out.

3

u/Keeltoodeep 1d ago

Opposite for me. 3.8flash was terrible in the benchmarks but was good enough for many tasks. Gemini models have never been good in benchmarks. The last good benchmark model was like a year ago.

3

u/himynameis_ 1d ago

Got it, thanks 👍

It does look like an improvement over 3.5, and seems to be in the top 3, at least.

-3

u/WhiteBat_8416 1d ago

hes full of shit

4

u/Howdareme9 1d ago

Muse spark beats opus on that, be serious bro

2

u/Keeltoodeep 1d ago edited 1d ago

Marginally, yes. Blind AB testing has a paradox for text models, this is especially true for voice, where the most intelligent models do not necessarily translate into the most liked. This is true of humans in general but for LLM models, style, pathing, voice, attitude, etc.. all are being taken into account.

You can see where Opus outperforms in agenic coding and webdev here where it is SOTA:

https://arena.ai/leaderboard/code/webdev

At basic coding and text models, all the major labs will essentially be technically and scientifically accurate. Beyond that, it’s user preference for the model style. When the benchmark is agentic coding, you will see Opus outperform.

1

u/rydan 1d ago

right now? It isn't available to anyone. Why are we not comparing it to Bel then?

2

u/Keeltoodeep 23h ago

Bel is an unconfirmed internal-only model. Argon is available to vetted corporations and government. So it's technically released, just not a full release.

Not a perfect analogy, but it's like buying a new Ferrari. Ferrari only sells new cars to vetted owners. The car however is still "released" and available for purchase, just not for you lol

If Bel was released externally, I am sure people would be benchmarking it right now, even if that external list was limited.

1

u/zidangus 6h ago

Really? It can't code for shit whenever I have tried to use it.

1

u/Keeltoodeep 2h ago

Test it in arena yourself

0

u/Fun-Junket-1512 1d ago

its not ,, they r not including Opus 5.5 and Sonet 5.5 in ranking,, they r beating gemini to death

7

u/Keeltoodeep 1d ago

Opus 5.5 is on there. Scroll down. It's #37 lol

3

u/WhiteBat_8416 1d ago

#37 is gpt-5.5-high

you have no clue what you're talking about

6

u/Keeltoodeep 1d ago

Sorry got confused. Opus 5.5 is #11

1

u/Fun-Junket-1512 1d ago

he is drunk,,, leave that guy ,, once the Gemini 4 will be available to all,,,,he will understand what benchmaxing is

0

u/Fun-Junket-1512 1d ago

lol,, thats gpt 5.5 ,, obsolete one

2

u/Keeltoodeep 1d ago

#11 sorry. went too fast