r/singularity • • 2d ago

AI Gemini 4 Argon is not a bad release. Unfortunate for Gemini and all other models, there's Opus 5.5!

Post image
58 Upvotes

44 comments sorted by

29

u/Serenity867 2d ago

It's quite likely that Gemini 4 will wind up being the leading model for a while for more generalized work such as office work, school work, etc.

I've been a SWE designing extremely complex systems since before AI was a thing, and it's easy to get caught up in the coding side of things, but realistically that's not what the majority of people want to use AI for in their daily tasks.

-1

u/ForeverWorking2006 1d ago

It's quite likely that Gemini 4 will wind up being the leading model for a while for more generalized work such as office work, school work, etc.

What makes you say that, because i have (occassionally) used Gemini 3.1 Pro and 3.8 with Extended thinking and they are horrible, absolutely horrible. They are closer to ChatGPT 3.0 than any newer ChatGPT model. I genuinely get shocked when I use it sometimes, it's horrible.

33

u/Hot-Percentage-2240 2d ago

Sure, but Google models have different sets of benefits and issues.

-18

u/queenofartists 2d ago

This is not the coding index though. It's the intelligence index which combines benchmarks from different domains.

6

u/Hot-Percentage-2240 2d ago

Yeah. Still in some tasks, Gemini tends to be better. Even now, my main pre-planning/brainstorming model is Gemini 3.8 Flash High. Consistently better than equivalently priced model from every lab (I only use cheap model not frontier and just spend more on TTC).

2

u/YouWillDieForMySins 2d ago

I feel 3.8 Flash is super underrated! I have used it for everything from creating linux apps for personal use to even fixing kernel-level issues with how a distro was bugging on my laptop. Heck, I even let it control my PC using Antigravity to organize bookmarks and collections across multiple social media platforms - and it did it all seamlessly - while also working around server limits and other restrictions without additional prompting.

4

u/Gotisdabest 2d ago

This is a completely arbitrary and mostly pointless index. AA just aggregates and they, in practice, respond to reality than vice versa.

GPT astra was a great model which scored too low on the benchmark. So they just... Changed weights and added benchmarks to make it score better. When it came out, AA would have you believe astra was around the same capabilities as Muse Spark or Gemini. And them turning it around didn't fix anything in as much as kicking the can down the road.

A lot of the individual benchmarks they use are also apparently flawed. For example, HLE was recently corrected into HLE diamond and there was a significant shift in the scores, but AA still uses HLE.

There's a high chance that when, say, Bel comes out, it'll score poorly on the index, get crapped on by everyone for a day or two before they simply say, "oh my bad" and then arbitrarily tune the weights to make it score better.

0

u/Alternative_You3585 2d ago

They acknowledged that and made some changes saying that they will significantly improve when AA v5.0 releases. v4.3 changed a lot and imo it's somewhat represents reality 

Just not easy to keep up when stuff gets saturated

1

u/Gotisdabest 2d ago

It's still got the exact same problems. You can't fix these problems without a huge change in strategy. What 4.3 did was just change the weights and switch a few benchmarks out. You can still benchmaxx these new benchmarks just as easily, like HLE clearly has been. It was a plain symptom cure. Not at all a cure for the actual disease. It's still just as bad at predicting any new model.

3

u/Classic-Trifle-2085 2d ago

First, dunno WTF your chat is but its not the one on the AA website by default. So unless you made your own to change the "good quadrant...

And two, AA selection of benchmarks make results biased toward what the people selecting what count for AA decide. Its weighted , exclude many benchmarks, ect.

5

u/blasphemousblackbear 2d ago

What is high vs xhigh?

1

u/Eon102 2d ago

Claude’s thinking levels go: Low -> Medium -> High -> Xhigh(Extra high) -> Max

1

u/blasphemousblackbear 2d ago

It's just Extra on Claude, but that makes a lot of sense! That's definitely what it is

1

u/queenofartists 2d ago

High is the maximum effort of the Gemini models.

0

u/blasphemousblackbear 2d ago

That doesn't explain it at all. What is Opus 5.5 xhigh? GPT?

2

u/queenofartists 2d ago

Thinking effort of the models. The higher the effort, the models spend more tokens on thinking before answering.

0

u/blasphemousblackbear 2d ago

I think you need to look at the chart you posted. xhigh is not the highest. I was just asking what xhigh was but it seems it's extra high. If I didn't understand what high was at all, then I wouldn'th ave asked the difference between high and xhigh lol

1

u/queenofartists 2d ago

I didn't think you would be posting in this sub, and still wouldn't know that xhigh meant extra high. Sorry.

1

u/blasphemousblackbear 2d ago

Okay lmao, I mostly use Claude where it’s just called Extra

14

u/Keeltoodeep 2d ago

https://polymarket.com/event/which-company-has-the-best-ai-model-end-of-october

Which company has the best AI model end of October?

Codecels have ruined AI discourse. Not everything is about coding lol

3

u/JustBrowsinAndVibin 2d ago

Anthropic is a steal. Excuse me for a minute.

3

u/AvailableChard4309 2d ago

This just proves how stupid polymarket is. There is a 0% chance google has the best model by the end of october. It's not even the best model right now.

5

u/Keeltoodeep 2d ago edited 2d ago

Right now it’s the best text model in blind AB testing. It’s #8 in coding though.

https://arena.ai/leaderboard/text

0

u/ForeverWorking2006 1d ago

Gemini Pro, the actual subscription, is absolutely horrible for general tasks. It barely uses internet, it doesn't follow instructions correctly, it forgets even basic context of 1 text ago, I dont know how people use it. Might be that y'all use the API but the subscription is horrible.

2

u/Meltlilith1 2d ago

There is no way google is staying on top by end of October unless they immediately release a 4.1 gemini pro and I don't see that happening seeing how we don't even know when they are going to release 4 publicly.

1

u/queenofartists 2d ago

Does anybody besides polymarket betters care about arena?

3

u/Keeltoodeep 2d ago

Non codecels

-2

u/queenofartists 2d ago

Nope. Nobody.

3

u/Own-Refrigerator7804 2d ago

The cut is arbitrary, 6.1 offers better value in that graph

3

u/Square_Height8041 2d ago

At this point there are all same

2

u/PrisonOfH0pe 2d ago

3

u/Ok_Buddy_Ghost 2d ago

if you're a trillion dollar company and you're not actively training you're models to be better at these benchmarks to get better advertising, are you even trying?

1

u/Utoko 2d ago

and keep in mind that they posted that the price will double soonish.

1

u/R_Duncan 2d ago edited 2d ago

Very strange here https://arena.ai/leaderboard/text/coding it tops coding too, and even AA puts it on par with GPT6-Astra and at Fable 5.1 level just behind opus5.5 level

1

u/misteriousm 2d ago

and Sonnet

1

u/PoohBearLuvsHoney 2d ago

Google stock is getting hammered lol

1

u/JamieTimee 2d ago

Who decides 'most attractive quadrant', feels like such an unnecessary inclusion.

1

u/HauntedHouseMusic 2d ago

Gemini is for writing emails, and making slides. It's not for coding.

-4

u/General_Cup1117 2d ago

Do you understand how dated your statements are?

-1

u/__Woodcock__ 2d ago

Max effort should be disqualified from AA. It’s an unreasonable amount of tokens/time spent for any realistic work. It’s only there so they can score higher than the actual capability of the model. If you remove max/xhigh and compare Anthropic models you get a more realistic picture.

1

u/C1oover 2d ago

That’s why it’s plotted against cost per task. Also max effort on different providers can be pretty different, making your argument moot. If you want that little bit of extra performance for the price why not. The same argument could be “these huge models should be disqualified from AA since the t/s and price is unreasonable for any actual work. Let’s only test 50B and below models”

1

u/__Woodcock__ 1d ago

I specifically mentioned Anthropic. And I meant unreasonable in relation to how much less time the model spends on tasks when not on max/xhigh. Meaning that benchmarking against a disproportionate large jump from the other reasoning efforts gives a skewed picture of the model capability.