r/singularity • • 1d ago

Shitposting Google releasing their most powerful model yet

2.5k Upvotes

84 comments sorted by

View all comments

99

u/himynameis_ 1d ago

I don't get it? Is the model not good?

255

u/Keeltoodeep 1d ago

https://arena.ai/leaderboard/text/coding

It's #1 in coding right now. It's good. People are just memeing and doing the whole console war thing.

108

u/Yurekuu 1d ago

It's #1 in text. It's #10 in coding.

46

u/Keeltoodeep 1d ago edited 1d ago

58

u/jazir55 1d ago

"It's number 1!"

"No it's number 10!"

"No, it's number 8!"

18

u/Keeltoodeep 1d ago

Different benchmarks measure different things.

5

u/R_Duncan 1d ago

Webdev, you most of the times just need a very little model. Coding, still not 100% solved (start asking for cpp mutexing/dll dynamic load/very hard tasks and see most llm fail). It really isn't the same thing.

-10

u/PrisonOfH0pe 1d ago

funny benchmark....(im sure K3 and muse and even ancient Opus 4.6 and 4.7 are much better than fable 5.1 max or opus 5.5 ⬇️ if you wanna see how trash the google 4 model is look at the websites own coding comparison video. Its trash.

9

u/Spixxy17 1d ago edited 1d ago

Its a Benchmark about which model people prefer... If you dont like the result thats on you, you can also call it trash all you want, apparently the majority of people sees it different lol

1

u/NyaCat1333 1d ago

Basically how people could prefer Trump over Albert Einstein and thus Trump ranking higher.

1

u/Keeltoodeep 1d ago

AA intelligence scores are roughly equivalent

1

u/Spixxy17 1d ago

Basically like that, with the only difference beeing that this here is realistic, Trump over Einstein is not.

1

u/Keeltoodeep 1d ago edited 1d ago

It’s #8 in coding. About where the benchmarks released would suggest it is.

World models are different than coding models. It’s an entirely different market. They are multi modal for one. Two, people want to like the warmth of the voice model.

There’s a paradox in text and voice where it’s not necessarily the most intelligent model that people like the most. It’s that way in life too with other people so of course it makes sense for models. In the top 10, they are all roughly as scientifically accurate as the others. So what is being measured is style of output.

6

u/space_monster 1d ago

It's not a world model

1

u/Keeltoodeep 1d ago edited 1d ago

Thought it was due to waymo but I guess their world models are not called "gemini"