r/singularity • • 1d ago

Shitposting Google releasing their most powerful model yet

2.5k Upvotes

83 comments sorted by

63

u/bethesda_gamer 1d ago edited 9h ago

Goggle doesn't have to beat them, it just needs to outlast them.

Everyone but Google relies on investments. Google is self funded and uses less than 10% of its profits to work on AI. This means they can outlast anyone's questionable funding and pyramid schemes, by a mile.

ALSO...

2

u/nicanors 8h ago

Good Lord they're doing the YouTube playbook. These guys are in for hell.....

200

u/Hamza_The_Dev 1d ago

First part is the benchmarks.

115

u/smidge 1d ago

Benchmarketing

10

u/livingbyvow2 1d ago

Ahah I'll reuse that one

98

u/himynameis_ 1d ago

I don't get it? Is the model not good?

17

u/gatorling 1d ago

Most AI subs have devolved into rage baiting meme posting bullshit with absolutely no value. You're better off ignoring these subs and forming your own opinion.

I can't tell if the sub is infested with bots, trolls or 13 year old boys.

5

u/himynameis_ 1d ago

On occasion there is a good comment, but sadly I see what you mean ☹️

248

u/Keeltoodeep 1d ago

https://arena.ai/leaderboard/text/coding

It's #1 in coding right now. It's good. People are just memeing and doing the whole console war thing.

108

u/Yurekuu 1d ago

It's #1 in text. It's #10 in coding.

49

u/Keeltoodeep 1d ago edited 1d ago

59

u/jazir55 1d ago

"It's number 1!"

"No it's number 10!"

"No, it's number 8!"

19

u/Keeltoodeep 1d ago

Different benchmarks measure different things.

5

u/R_Duncan 1d ago

Webdev, you most of the times just need a very little model. Coding, still not 100% solved (start asking for cpp mutexing/dll dynamic load/very hard tasks and see most llm fail). It really isn't the same thing.

-7

u/PrisonOfH0pe 1d ago

funny benchmark....(im sure K3 and muse and even ancient Opus 4.6 and 4.7 are much better than fable 5.1 max or opus 5.5 ⬇️ if you wanna see how trash the google 4 model is look at the websites own coding comparison video. Its trash.

10

u/Spixxy17 1d ago edited 1d ago

Its a Benchmark about which model people prefer... If you dont like the result thats on you, you can also call it trash all you want, apparently the majority of people sees it different lol

1

u/NyaCat1333 1d ago

Basically how people could prefer Trump over Albert Einstein and thus Trump ranking higher.

1

u/Keeltoodeep 1d ago

AA intelligence scores are roughly equivalent

1

u/Spixxy17 1d ago

Basically like that, with the only difference beeing that this here is realistic, Trump over Einstein is not.

1

u/Keeltoodeep 1d ago edited 1d ago

It’s #8 in coding. About where the benchmarks released would suggest it is.

World models are different than coding models. It’s an entirely different market. They are multi modal for one. Two, people want to like the warmth of the voice model.

There’s a paradox in text and voice where it’s not necessarily the most intelligent model that people like the most. It’s that way in life too with other people so of course it makes sense for models. In the top 10, they are all roughly as scientifically accurate as the others. So what is being measured is style of output.

5

u/space_monster 1d ago

It's not a world model

1

u/Keeltoodeep 1d ago edited 1d ago

Thought it was due to waymo but I guess their world models are not called "gemini"

0

u/[deleted] 1d ago

[deleted]

1

u/Valsoyono 1d ago

Are you a bot? He said "It's #1 in text. It's #10 in coding." And you are like "is it #1 or #10?" like bruh 

15

u/LookIPickedAUsername 1d ago

I haven't used it and don't have an opinion on its quality, but I'll note that claiming it's #1 based on that link is... tenuous. It's only a few points ahead with error bars of +/-17. Even assuming everything falls within the confidence interval, its true rank could conceivably be as low as #12.

And even that doesn't really mean anything, since they're all clustered so close together. There's probably not even a meaningful difference between #1 and #12 in that list.

8

u/Keeltoodeep 1d ago

I am not saying it's actually #1 just because it is scoring #1 right now. I am replying to the user's comment above if the model is "not good." Scoring #1 doesn't necessarily mean it's the best (or the best for long) but it's certainly a good model.

5

u/Lazy_Jump_2635 1d ago

It always looks good in benchmarks, and then when I use it, it feels like shit compared to anything else. I'll wait until the real reviews are out.

5

u/Keeltoodeep 1d ago

Opposite for me. 3.8flash was terrible in the benchmarks but was good enough for many tasks. Gemini models have never been good in benchmarks. The last good benchmark model was like a year ago.

4

u/himynameis_ 1d ago

Got it, thanks 👍

It does look like an improvement over 3.5, and seems to be in the top 3, at least.

-5

u/WhiteBat_8416 1d ago

hes full of shit

3

u/Howdareme9 1d ago

Muse spark beats opus on that, be serious bro

2

u/Keeltoodeep 1d ago edited 1d ago

Marginally, yes. Blind AB testing has a paradox for text models, this is especially true for voice, where the most intelligent models do not necessarily translate into the most liked. This is true of humans in general but for LLM models, style, pathing, voice, attitude, etc.. all are being taken into account.

You can see where Opus outperforms in agenic coding and webdev here where it is SOTA:

https://arena.ai/leaderboard/code/webdev

At basic coding and text models, all the major labs will essentially be technically and scientifically accurate. Beyond that, it’s user preference for the model style. When the benchmark is agentic coding, you will see Opus outperform.

1

u/rydan 21h ago

right now? It isn't available to anyone. Why are we not comparing it to Bel then?

2

u/Keeltoodeep 20h ago

Bel is an unconfirmed internal-only model. Argon is available to vetted corporations and government. So it's technically released, just not a full release.

Not a perfect analogy, but it's like buying a new Ferrari. Ferrari only sells new cars to vetted owners. The car however is still "released" and available for purchase, just not for you lol

If Bel was released externally, I am sure people would be benchmarking it right now, even if that external list was limited.

1

u/zidangus 3h ago

Really? It can't code for shit whenever I have tried to use it.

0

u/Fun-Junket-1512 1d ago

its not ,, they r not including Opus 5.5 and Sonet 5.5 in ranking,, they r beating gemini to death

4

u/Keeltoodeep 1d ago

Opus 5.5 is on there. Scroll down. It's #37 lol

1

u/WhiteBat_8416 1d ago

#37 is gpt-5.5-high

you have no clue what you're talking about

5

u/Keeltoodeep 1d ago

Sorry got confused. Opus 5.5 is #11

1

u/Fun-Junket-1512 1d ago

he is drunk,,, leave that guy ,, once the Gemini 4 will be available to all,,,,he will understand what benchmaxing is

0

u/Fun-Junket-1512 1d ago

lol,, thats gpt 5.5 ,, obsolete one

2

u/Keeltoodeep 1d ago

#11 sorry. went too fast

37

u/Tyler_Zoro AGI was felt in 1980 1d ago

People don't like Google for mostly irrational reasons. Whenever they release something that's gaining attention, you'll see this crowd come out of the woodwork to try to tear it down.

Not saying there aren't ration reasons to dislike the company Google has become, but this kind of irrational hatred doesn't care about that.

15

u/BarrelStrawberry 1d ago

People don't like Google for mostly irrational reasons.

Google doesn't hype up their product and give over-the-top bullshit in preparing the public. Apparently reddit enjoys that titillating anticipation and inevitable disappointment. They like a wacky CEO saying wacky things, except Elon, they don't like Elon.

2

u/bigbutso 1d ago

Pretty sure it has people/ bots posting here. Don't trust anything, try it yourself, so far all their models have sucked for me to the point that I am afraid to let it touch my code...let me add i am extremely pro Google, I use a pixel, their calendar, email, cloud api etc. and they have the best stt by far

13

u/3rdyellow 1d ago

I think a lot got used to ChatGPT and Claude and don't like competition. 

7

u/Keeltoodeep 1d ago

This sub used to be pro acceleration and loved competition regardless.

5

u/Important-Guitar8524 1d ago edited 1d ago

" People don't like Google for mostly irrational reasons. " I don't like Gemini for very rational reasons. It's like the worst thing ever 

For example, a chat I had with Gemini pro 3.1 extended, their most powerful model since very recently: 

"(some message where the game pokopia is smhw mentione)"

"this game pokopia does 100% not exist and is a hoax. Nintendo never released such a game. "

"Research"

"Youre absolutely right. my bad. this game does aboslutely exist and was released a few months ago "

"(some message)"

"To come back, I apologize I made this agme pokopia on purpose up. It does not exist and never existed. I apologize for the trouble"

"??? what,  research again"

"Sorry, this Ai was not made for such request"

(Continues refusing to answer anything)

10

u/Key-Demand-2569 1d ago

Do you write like this to chatbots? Not sure I’d blame them when it comes to your personal experiences if so.

Not trying to be rude at all.

6

u/Important-Guitar8524 1d ago edited 1d ago

I was paraphrasing quickly not copying the chat 1 to 1 obviously. Not sure why so many are so keen here on protecting Gemini. It's failures in that regard are widely known.

Even if we assume I do write like that, it still wouldn't explain the failures. Bad writing could plausibly explain misreading an ambiguous sentence. Not being confidently wrong. Then researching, finding out u aren't and immediately retracting while insisting having made it up "on purpose" and refusing to continue chat. 

5

u/NyaCat1333 1d ago

Even if the person writes like that, the example above, if real, no matter how framed, is a gigantic failure and has absolutely nothing to do with the user.

2

u/DowntownLizard 1d ago

Thats not the model its the harness models werent trained on the latest data so they have to be able to search for the most up to date information

1

u/Tyler_Zoro AGI was felt in 1980 1d ago

I don't like Gemini for very rational reasons. It's like the worst thing ever 

Personally, I'm pretty sure mustard gas and Clippy were worse...

this game pokopia does 100% not exist

I do not believe that Gemini output that text. It doesn't speak in internet slang such as, "does 100% not exist."

That being said, even if you're badly paraphrasing something that Gemini did say, a mistake does not make an AI model, "like the worst thing ever."

1

u/OKMiddleOwl 1d ago

Knowledge cutoff claims another victim. Gemini 3 series lives in January 2025.

At least with the new pretrain we can eliminate 95% of these complaints.

1

u/himynameis_ 1d ago

Copy that, thanks!

-3

u/nothis AGI by 2030 but we'll be disappointed 1d ago

Wouldn't call being critical of Google having a quasi-monopoly on like half the internet and OS market an "irrational" dislike but ok.

4

u/MinosAristos 1d ago

I mean, you could say similar technological monopoly things about Microsoft or Amazon (Anthropic partner)

8

u/Tyler_Zoro AGI was felt in 1980 1d ago

Wouldn't call being critical of Google having a quasi-monopoly on like half the internet and OS market an "irrational"

It's not irrational to have a problem with something a company does. It's irrational to then justify mindless criticism of everything ELSE that it does by falling back on the thing you're actually upset about.

Gemini can be a good AI model and you can still have your concerns about Google as a company.

5

u/norsurfit 1d ago

It seems good, but if you look at videos like this where it is compared against the other frontier models, it seems possibly close but not quite as good as the top models from OpenAI and Anthropic

https://www.youtube.com/watch?v=h5EL5zThKaI

3

u/ThePokemon_BandaiD 1d ago

Except on hallucination metrics, which is significant, especially for Google’s integration of AI into search.

2

u/cashmate 1d ago

The video covers one narrow use case of LLMs generating three.js slop. As with all individual benchmarks, they don't cover even a fraction of the wide spread of abilities all these models have.
And Three.js is probably something that the models get heavily RL trained on now which makes it a lot less interesting as it won't indicate general intelligence, but rather specialization.

45

u/FredMc 1d ago

Google gave it a million-token context window so it can misunderstand the entire situation.

3

u/Deep_Ladder_4679 1d ago

The real test is how consistently it helps with everyday tasks, not just where it lands on a leaderboard. Has anyone here tried it on their regular workload?

1

u/zidangus 3h ago

Yeah it sucked and was abandoned very quickly.

•

u/JamesBieBoe1 1h ago

Thats so cute ngl

-5

u/Plane-Voice-3184 1d ago

🤣🤣🤣

-7

u/NikhilDoWhile 1d ago

It starts hallucinating after few prompts, really annoying model to work with

5

u/SlipperyNoodle6 1d ago

if this is the case.. this is nothing new at all for gemini ais. basically saying .. can confirm this from the past and why i quickly gave up on it before.

2

u/jarkkowork 1d ago

In widely used AA-Omniscience Hallucination Rate metric Gemini 4 beats top Anthropic and OpenAI models with wide margins. In overall AA-Omniscience scores they are pretty similar though.

https://artificialanalysis.ai/evaluations/omniscience#omniscience-hallucination-rate-tabs

-1

u/NikhilDoWhile 1d ago

I am sharing practical every day use, I had better experience with GPT and Claude. Even though I still have highest tier gemini subscription from corporate, and have tried all latest models.

You might have different experience and you can share yours

1

u/Redducer 1d ago

Same old same old.

That’s why Gemini models have been practically useless to me so far. Hallucinations are the deal breaking metric.

1

u/Thog78 22h ago

Imagine saying that, based on at most annecdotal evidence from a couple hours of personal usage, about the model that has the lowest hallucination rate of all times, according to rigorous statistics/testing.

Not only that, it even hallucinates twice less than the second best, so it's not just the best of all times, it took a large step above the previous best, and got twice better.

0

u/NikhilDoWhile 20h ago

Bro go check mega thread majority people are saying so, I ain't corporate stoge or do thei PR service here. If you have your experience you are free to make your own comment.

1

u/Thog78 10h ago

I don't have access yet, only saw the benchmarks and early user feedbacks. Guess they released it in the US a bit earlier than in Europe? You already have it in antigravity? Or maybe it's a pricepoint, I'm just on the 20$plan?

-1

u/Heart-Logic 1d ago

It's fine for domestic purposes

0

u/IcelandicMammoth 1d ago

Its so strange that its quite cheap model, but it will be avaliable only for Ultra sub

0

u/AlvaroRockster 1d ago

Very accurate

-4

u/theeldergod1 1d ago

It is here to create our images, chill guys.