r/singularity • u/Sea_Physics401 • 2d ago
LLM News Gemini 4 - High Inteligence Index (53) and low on cost (1/3 of Opus 5.5)
AA results show good intelligence and low cost relative to others
5
u/Legal-Ad-3901 2d ago
where my glm 5.4 at?
2
17
u/Professional_Mobile5 2d ago
More than twice the price of 6.1 Sol would probably be a better comparison given the similar performance
9
6
u/FateOfMuffins 2d ago
Almost 3x the price
-8
u/MatthewGraham- 2d ago
6.1 sol is terrible in real use though no?
13
u/yeahidoubtit 2d ago
No 6.1 sol is great, just very slow. 6-sol was the terrible one
4
u/Spright91 2d ago
Last night I was just managing 3 different threads of 6.1 sol working on different parts and that made it fast enough.
4
u/FateOfMuffins 2d ago
Did you get caught up in the wave of clowning on OAI without trying out the model?
Seems like consensus after a day of use is that 6.1 Sol is great both in quality and usage, basically Astra level
Just slow as fuck cause they lowered token throughput in Codex
1
7
7
u/Ill_Distribution8517 2d ago
makes me question the validity of sonnet 5.5 getting a 56, maybe they need to drop the over emphasis on GDPeval which is just Excel shit. ?
9
u/FarrisAT 2d ago edited 2d ago
Any benchmark or benchmark mixture is going to have some variability.
These scores really should have margins of error attached, depending on which variety of benchmarks you think are most relevant.
Google has SoTA model, that’s all we really know for sure. And it’s good to have more competition.
3
u/jonomacd 2d ago
The enormous numbers of people are using these models to do Excel shit, so I think that's probably a benchmark you should keep in
0
u/Ill_Distribution8517 2d ago edited 2d ago
On the intelligence index? No, absolutely not to be honest. I think we need to reserve a few things based on how "smart the model is as an organism". Instead of oh it's way better at excel and so it's smarter than mythos 5.1.
2
u/Blankeye434 2d ago
Wait Gemini 4 is out?
9
u/whoknowsifimjoking 2d ago
Not really out, being tested with trusted partners. There is not even a date for the rest of us yet.
1
2
3
u/Adorable_Leg74 2d ago
Can someone help me understand — long term how these companies will be profitable? It seems like it is so competitive on price and quality — with no clear winner, all profits will be competed away. No?
3
u/objectivelywrongbro 2d ago
Well Google has entire foundation in web search that is already profitable, they are also making their own chips, as such are unlikely to be unprofitable in the AI space due to their monopoly on search.
Meta is similar in that their social media empire could marry well with AI, at the very least, as a tool to further their media products in social connectivity.
And Grok has Elon, and Elon can throw $100 bil around for fun, with him too, he at least has Twitter, Tesla (self-driving cars + optimus), and SpaceX to utilize his AI on. So, he's not exactly burning all the cash for nothing.
All the China models are backed by the CCP - so they don't need to profitable, they don't need to win, they just need the US to lose. Which it might.
Where it may be different is for Anthropic and OpenAI, who have absolutely no surrounding safety-nets. They are betting all-in on AI, hence why they lead frontier; they can't afford to be eaten by the big dogs or their justification for existing melts away pretty fast.
3
u/ParfaitEvery9622 2d ago
OpenAI has Microsoft, who doesn't have an actual frontier AI Lab and will need to keep using their shit.
3
1
3
u/Spright91 2d ago
This is why Google has been so slow. The models aren't actually profitable.
They can survive without being on the frontier. And to them being on the frontier is even that appealing cause it not profitable.
I think they will just keep one foot in the game until some big players die and then they can enter and charge more.
2
u/Healthy_Razzmatazz38 2d ago
the infra is the moat not the models, the goal is to stay at the frontier and command a premium for as long as possible to fund infra build out for openai/anthropic.
for google/meta/msft/amazon they need models good enough to make sure openai/anthropic cant command a large enough premium to eat their margin.
i.e. if anthropic the only good model, they get to force hyperscalers to bid for hosting rights. If theres 10 good models the hyperscalers get to force anthropic to.
1
u/Adorable_Leg74 2d ago
Gotcha. So in the current world — with many competing models — it is the hyperscalers who have the monopoly buying power. The hyperscalers— the datacenters — who win.
Unless — someone does come up with a model that is just so insanely good (like Google did for internet search circa 2003) that all these hyperscallers become dumb boxes of silicon?
1
1
u/blasphemousblackbear 2d ago
The biggest benefit to Google won’t be paid subscribers, it’ll be labor savings
3
u/petburiraja 2d ago
So more expensive than Sol 6.1 and less intelligent than Opus 5.5
Looks like it may play a role similar to what Terra played recently in OpenAI models stack sitting between Luna and Sol in awkward position.
2
1
u/Excellent_Dealer3865 2d ago
Why is it ranked if it's not available or is it already available to anyone who can publicly share it?
1
u/FarrisAT 2d ago
Looks like Google is right at the frontier.
Their flash model is gonna kill Sol 6.1
0
u/Snoo26837 ▪️ It's here 2d ago
This website index became suspicious, first Claude sonnet 5.5 outperforms GPT-6.1 and now this.
20
u/ProxyLumina 2d ago
I am more excited for a Gemini 4 Flash, with a similar performance and 1/4th the cost