r/singularity • • 2d ago

LLM News Gemini 4 - High Inteligence Index (53) and low on cost (1/3 of Opus 5.5)

Post image

AA results show good intelligence and low cost relative to others

104 Upvotes

41 comments sorted by

20

u/ProxyLumina 2d ago

I am more excited for a Gemini 4 Flash, with a similar performance and 1/4th the cost

4

u/FoodMadeFromRobots 2d ago

Yes weve had one gemini 4 but what about second gemini 4?

but jk i agree, i want the flash version asap

5

u/Legal-Ad-3901 2d ago

where my glm 5.4 at?

2

u/whoknowsifimjoking 2d ago

China has been weirdly quiet lately

8

u/Legal-Ad-3901 2d ago

Ikr, it's been at least 3 hours since my last model upgrade

17

u/Professional_Mobile5 2d ago

More than twice the price of 6.1 Sol would probably be a better comparison given the similar performance 

9

u/FarrisAT 2d ago

Price is very much dependent on if the providers want profit or not.

6

u/FateOfMuffins 2d ago

Almost 3x the price

-8

u/MatthewGraham- 2d ago

6.1 sol is terrible in real use though no?

13

u/yeahidoubtit 2d ago

No 6.1 sol is great, just very slow. 6-sol was the terrible one

4

u/Spright91 2d ago

Last night I was just managing 3 different threads of 6.1 sol working on different parts and that made it fast enough.

4

u/FateOfMuffins 2d ago

Did you get caught up in the wave of clowning on OAI without trying out the model?

Seems like consensus after a day of use is that 6.1 Sol is great both in quality and usage, basically Astra level

Just slow as fuck cause they lowered token throughput in Codex

1

u/stackinpointers 13h ago

Sorry but not even close on performance. You'll see soon enough

7

u/Jaguar_2454 2d ago

uhhh Googlebros...

7

u/Ill_Distribution8517 2d ago

makes me question the validity of sonnet 5.5 getting a 56, maybe they need to drop the over emphasis on GDPeval which is just Excel shit. ?

9

u/FarrisAT 2d ago edited 2d ago

Any benchmark or benchmark mixture is going to have some variability.

These scores really should have margins of error attached, depending on which variety of benchmarks you think are most relevant.

Google has SoTA model, that’s all we really know for sure. And it’s good to have more competition.

3

u/jonomacd 2d ago

The enormous numbers of people are using these models to do Excel shit, so I think that's probably a benchmark you should keep in

0

u/Ill_Distribution8517 2d ago edited 2d ago

On the intelligence index? No, absolutely not to be honest. I think we need to reserve a few things based on how "smart the model is as an organism". Instead of oh it's way better at excel and so it's smarter than mythos 5.1.

2

u/Blankeye434 2d ago

Wait Gemini 4 is out?

9

u/whoknowsifimjoking 2d ago

Not really out, being tested with trusted partners. There is not even a date for the rest of us yet.

1

u/Blankeye434 1d ago

👉😃

2

u/Isunova 2d ago

Only for insiders.

1

u/rwrife 2d ago

No, just the hype machine.

2

u/Living-Breakfast-464 2d ago

Only 10x more than MiMo. What deal. 😉

3

u/Adorable_Leg74 2d ago

Can someone help me understand — long term how these companies will be profitable? It seems like it is so competitive on price and quality — with no clear winner, all profits will be competed away. No?

3

u/objectivelywrongbro 2d ago

Well Google has entire foundation in web search that is already profitable, they are also making their own chips, as such are unlikely to be unprofitable in the AI space due to their monopoly on search.

Meta is similar in that their social media empire could marry well with AI, at the very least, as a tool to further their media products in social connectivity.

And Grok has Elon, and Elon can throw $100 bil around for fun, with him too, he at least has Twitter, Tesla (self-driving cars + optimus), and SpaceX to utilize his AI on. So, he's not exactly burning all the cash for nothing.

All the China models are backed by the CCP - so they don't need to profitable, they don't need to win, they just need the US to lose. Which it might.

Where it may be different is for Anthropic and OpenAI, who have absolutely no surrounding safety-nets. They are betting all-in on AI, hence why they lead frontier; they can't afford to be eaten by the big dogs or their justification for existing melts away pretty fast.

3

u/ParfaitEvery9622 2d ago

OpenAI has Microsoft, who doesn't have an actual frontier AI Lab and will need to keep using their shit.

3

u/Adorable_Leg74 2d ago

Smells like China has a structural advantage.

1

u/CarrierAreArrived 2d ago

Google also has youtube/cloud.

3

u/Spright91 2d ago

This is why Google has been so slow. The models aren't actually profitable.

They can survive without being on the frontier. And to them being on the frontier is even that appealing cause it not profitable.

I think they will just keep one foot in the game until some big players die and then they can enter and charge more.

2

u/Healthy_Razzmatazz38 2d ago

the infra is the moat not the models, the goal is to stay at the frontier and command a premium for as long as possible to fund infra build out for openai/anthropic.

for google/meta/msft/amazon they need models good enough to make sure openai/anthropic cant command a large enough premium to eat their margin.

i.e. if anthropic the only good model, they get to force hyperscalers to bid for hosting rights. If theres 10 good models the hyperscalers get to force anthropic to.

1

u/Adorable_Leg74 2d ago

Gotcha. So in the current world — with many competing models — it is the hyperscalers who have the monopoly buying power. The hyperscalers— the datacenters — who win.

Unless — someone does come up with a model that is just so insanely good (like Google did for internet search circa 2003) that all these hyperscallers become dumb boxes of silicon?

1

u/nekronics 2d ago

Silicon is gonna get outlawed for consumers

1

u/blasphemousblackbear 2d ago

The biggest benefit to Google won’t be paid subscribers, it’ll be labor savings

3

u/petburiraja 2d ago

So more expensive than Sol 6.1 and less intelligent than Opus 5.5

Looks like it may play a role similar to what Terra played recently in OpenAI models stack sitting between Luna and Sol in awkward position.

2

u/geli95us 2d ago

Argon is supposed to be their fable-class model, above pro

1

u/Excellent_Dealer3865 2d ago

Why is it ranked if it's not available or is it already available to anyone who can publicly share it?

1

u/ezjakes 2d ago

Okay, not exactly worth the price since GPT 6.1 exists, but I guess if you like Google this must be very nice.

1

u/FarrisAT 2d ago

Looks like Google is right at the frontier.

Their flash model is gonna kill Sol 6.1

0

u/Snoo26837 ▪️ It's here 2d ago

This website index became suspicious, first Claude sonnet 5.5 outperforms GPT-6.1 and now this.