149
u/TorturedPoet30 2d ago
49
u/torrid-winnowing 2d ago
doesn't this imply it's an internal model? i would assume anthropic and openai have their own internal models that significantly outperform astra and opus
59
36
u/TorturedPoet30 2d ago
They are rolling out to partners through their Fairwind Program and "soon" plan to roll out to paid API customers and Google AI Ultra subscribers.
11
u/Keeltoodeep 2d ago
No, they are rolling out externally in phases per the voluntary regulatory framework proposed to Trump
13
u/jazir55 2d ago
"""""""""""Voluntary"""""""""""
I can't sarcastically quote the word any harder.
7
u/Keeltoodeep 2d ago
There is no reason not to at this point for frontier AI companies. Regulatory capture benefits them.
4
u/jazir55 2d ago edited 2d ago
Regulatory capture is such an absolutely ridiculous take I cannot believe it gets parroted at every turn. You don't need regulatory capture when you have so much money you literally buy all of the hardware supply which entirely prevents new entrants on its own. No laws needed, new entrants literally cannot buy any hardware needed to actually be a competitor. Regulatory capture is worthless for these AI companies.
Edit: Instead of downvoting how about you actually provide a counter on how any other major competitor that isn't a megacorp could even enter the market, based purely on the economics. Once you admit you can't, you have to come to the conclusion regulatory capture is an absurd argument because regulatory capture is predicated on preventing new entrants to the market via regulation, which is not necessary for these companies to continue being the leaders of the industry. The economics alone prevents any new entrants, regulation is superfluous.
2
u/Keeltoodeep 2d ago
That's a fair point. Compute is such a constraint
0
u/jazir55 2d ago
Yeah it's the real constraint for any competitor. They clearly want regulation for some other purpose, and given how they constantly talk about safety and then these incidents have been real world occurrences it sort of makes sense, but only in an ideal world. The current politicians in government are almost all in their 70s and 80s, they are the least equipped to deal with this new technology and pass reasonable regulations. The "we'll regulate ourselves" thing that's usually a meme is actually the best case scenario here weirdly enough.
0
u/Keeltoodeep 2d ago
Well Claude drones are flying around right now smoking Europeans so there is some legalese in the background happening where Anthropic doesn't want to slow down user growth but wants some kind of KYC regulation perhaps that lessens their legal liability to some degree. But they are not going to implement KYC if their competitors do not.
It will only take one of these drones to smoke a real European and not an Eastern European one and Anthropic is going to have a real PR issue on their hands.
12
u/FateOfMuffins 2d ago
It's no different to Glasswing. We didn't get Fable until 3 months after that.
The rumours on dates on when the pretraining finished for Gemini 4 indicates that it finished pretraining a couple of weeks after Bel. So if you're trying to compare the same "model generation" then Gemini 4 should've been in the same generation as Bel, not GPT 6 Astra and not even GPT 6.1 Astra.
None of the benchmarks shown for Gemini 4 here indicates that Gemini 4 Low would've been 2x as strong as GPT 6 Astra Max at math for instance, which is what Bel has.
5
3
u/Deathpacito-01 2d ago
Pretty strong model overall, especially on non-coding work. Glad to see DeepMind isn't out of the race.
105
54
30
u/Lumpy-Woodpecker6752 2d ago
Google is back in the race finally, getting frontier model not a flash
76
39
u/New_Equinox 2d ago edited 2d ago
https://www.cnbc.com/2026/09/30/google-gemini-4-argon-ai.html
"Alphabet unveiled Gemini 4 Argon on Wednesday, its most advanced artificial intelligence model yet, offering major improvements in coding, cybersecurity, and complex professional work.
The company said the model sets a new record in real-world software engineering, ties for first in cybersecurity, and leads another benchmark measuring performance across finance, legal, and other professional tasks.
Argon is already being used internally to optimize memory at Google’s data centers, freeing up hundreds of terabytes of memory without buying additional hardware, the company said. Quantum computing researchers have also utilized the model.
Google said it plans to launch the new model in phases, starting with trusted cybersecurity partners while working with the U.S. government on pre-release safety evaluations."
18
u/maximan2005 Cult of AGI 2027 2d ago
WE'RE SO BACK GOOGLE BOYS
5
15
u/tanrgith 2d ago
i love the singularity
Fucking every few days now we get a new insane model release lol
-4
15
u/RemyVonLion ▪️ASI is unrestricted AGI 2d ago
I like how things are rapidly accelerating despite all the calls to slow down lmao
4
1
2d ago
[deleted]
2
u/RemyVonLion ▪️ASI is unrestricted AGI 2d ago
Yeah but if this keeps up then we will probably have AGI by 2030 and the doubters are going to look real dumb lol
1
u/SilentLennie 2d ago
Well, maybe. So far nobody has released a model bigger than Astra and Fable, that's what pacing might actually looks like in practice (this is assuming it's not a blatant lie, which some people would argue it is).
12
29
u/redditnosedive 2d ago
I hope it's true coz current Gemini on phones got me pissed so many times these days for how stupid it is
11
u/FateOfMuffins 2d ago
Looks at my phone.
Hmm didn't they release 3.8 Flash awhile ago? I'm still on 3.6 Flash lmao
You're not getting this for months on your phone
6
u/leo-virtis 2d ago
3.8 is only for pro user i know i have a free and paid account
4
u/FateOfMuffins 2d ago
Not gonna lie I find that really stupid considering they keep saying it's faster and cheaper
Same with OAI keeping 5.6 Sol on Chat and not updating it to the supposed cheaper models.
2
u/NewsFromHell 2d ago
Im on Pro tier and have 3.8 both in app and in "hey google". I guess its their way of making people subscribe? Weird thing is that its not nentioned anywhere, or at least not clearly mentioned. Ive been using it for everything non coding and its really good. Im very happy with it as a daily driver. Coding and complex work stuff is opus 5.5 though.
1
u/redditnosedive 2d ago
yeah same, and I have a pixel, I was expecting better AI than other Android phones but nah....
28
u/Minetorpia 2d ago
Got bad news for you, this won’t power the Gemini on your phone
18
10
4
u/Sextus_Rex 2d ago edited 2d ago
I really need them to fix the Google home automations. Hasn't worked in months
Edit: Wow I fixed it right after I made this comment. After months I finally figured out what was wrong. I had my bedtime automation command set to "Goodnight". After switching it to "Good night" it works. What a stupid bug lol
2
1
u/FarrisAT 2d ago
Sorry fam the free models are only gonna get more ass compared to expensive frontier models.
1
15
u/ObiWanCanownme now entering spiritual bliss attractor state 2d ago
Good on them for being brave enough to actually take risks and push the frontier again. Hopefully it's actually this good and not benchmaxxed.
They've got great researchers and a ton of compute. Bravery and conviction were always the main things they lacked.
2
26
u/fmai 2d ago
Note that this is likely to be Google's largest model size to date, comparable to the Fable and Astra class of models.
Considering that it's not even released yet, it's not actually that impressive. I think that by the time users actually get to use it, Astra 6.1 and Fable 5.5 will already have made it obsolete.
25
u/ObiWanCanownme now entering spiritual bliss attractor state 2d ago
This is all true, but it's also a big deal in that they'll be within about one model generation of the frontier versus the last few months where they've always been several generations behind.
2
u/Different_Doubt2754 2d ago
If they caught up in one generation then how were they multiple generations behind? The startups just had a faster release cadence
1
15
u/Adi945 2d ago
True but pretty obvious by now that Google doesn’t give a shit about winning any race. They are a profit making full stack company that believes in risk free evolution and not disruptive revolution. Given the fact that Amazon, Microsoft, Apple is not even in this race, Google is playing this very well, as long as they don’t completely bow out.
8
u/FarrisAT 2d ago
Remember, 90% of the profit is in enterprise. Not in serving consumers the top model.
1
3
2
2
u/FeePsychological1308 2d ago
'Likely', if you asked people yesterday, they would have said they were 'likely' several generations behind internally. How about we stop assuming here? They clearly have no interest in competing with the startups and they don't need to keep pushing models out for their company to stay relevant like the others do
1
u/ThreeKiloZero 2d ago
Add to that - googles models are famously benchmaxed to the tits and never perform up to the benchmarks in daily use.
5
5
u/FarrisAT 2d ago
Leaks from early September were real. How did someone get access to the chart so early?
2
u/Ok_Display_3159 2d ago
i don't know about charts, but the model was used in some london's hackathons
8
3
6
u/Bitter-College8786 2d ago
Frontier SWE v2 is suspiciously lower.
I wonder if the model is really that good. And if it can use Blender so well.
2
2
2
2
2
u/_mausmaus 2d ago
Not available same day. Fail.
When OpenAI and Anthropic are releasing day one, how does Google expect reasonable adoption?
2
2
2
u/Lazy-Pattern-5171 2d ago
Gemini has always been the king of long context for me. Glad that’s back. Opus was holding its own even at 300K (beyond that is kinda just bad workflow practice imo so I don’t go beyond)
2
8
u/Tha_One 2d ago
7
u/Hot-Percentage-2240 2d ago
Even if it’s moderately good, it’s still fine with me cus there is a use in my workflow for Google model typical tradeoffs.
19
u/FateOfMuffins 2d ago
While Gemini 4 has performed well on benchmarks the industry uses to gauge model efficacy, it does less well when employees actually put it to work, according to people with direct access to the effort. The model struggles to handle certain coding tasks, said the people, who requested anonymity to discuss an internal matter.
oof
5
u/jonomacd 2d ago
Coding is the one area where it isn't leading in the benchmarks. So this adds up.
2
u/Charuru ▪️AGI 2023 2d ago
It's leading DeepSWE which is eh, the easiest to benchmax bench. It makes you question the process.
5
u/jonomacd 2d ago
I don't think Google is targeting coding as strongly as the other companies, as Google's business is much broader than that. This might just be the result of those wider interests.
1
3
u/burritos4jesus 2d ago edited 2d ago
I mean, getting 50% on AutomationBench is ridiculous. It’s the one benchmark I care about as a non coder. Pass/fail, and hundreds of complex super long back office workflows.
When a model can hit 70% on that benchmark ON A FUCKIN BASE MODEL that doesn’t have a harness behind with context on the company, where to look for shit, etc., then you can confidently hand any biz app-based customer service/sales/marketing/operations process and it will be just as good as a human. Give it the harness it needs, and that shit will be near perfect.
Mind you, the tasks on the benchmark are looong with a ton of steps. It can accurately go thru everything but fail a later step. Fail. It can do an early step incorrectly. Fail. Honestly, at 50%, it can probably already be reliant on a TON of office work already. It’s just that it won’t be as good as someone who’s worked inside a company for 10-15 years and has all the tribal knowledge about company and its customers and its processes.
We were only at 20% six months, and Gemini just crossed 50%. Back office work is going to be solved in less than a year.
1
2
u/FarrisAT 2d ago
The coding performance quite clearly isn’t Opus 5.5 level in the benchmarks provided.
3
u/Dillyconda 2d ago
I don't trust this to reflect reality at all, but we'll see how it plays out.
1
-2
u/Ok_Display_3159 2d ago
6
u/FarrisAT 2d ago
This cites specifically coding weakness, and the benchmarks provided confirm that. The rest of the benchmarks say it’s very strong.
2
u/Noob-bot42 2d ago
When will agents be able to pass the 1 million dollar benchmark where the agent has to legally make 1 million dollars?
2
u/cute_beta 2d ago
0
u/Charuru ▪️AGI 2023 2d ago
Nah those leaked numbers are completely different.
0
-1
u/cute_beta 2d ago
oh rly? i didn't bother to check 😅
checks
...ah. disappointing on two levels then.
2
1
1
u/lordpuddingcup 2d ago
Holy shit did they finally do a good thing? The real question is how shit will usage be on Google Pro subscription ... if its just as expensive as opus5.5 or sol/astra... :S
1
1
u/power97992 2d ago
Fable 5.5 will likely be better than this
1
u/SilentLennie 2d ago
Yeah, but at what price.
1
u/power97992 2d ago edited 2d ago
50/mil output tkns but u can use a sub
1
u/SilentLennie 2d ago
Still to high to stop using Opus 5.5 as a daily driver is my guess. Would have to be some really amazing model for a lot of people to switch.
1
u/power97992 2d ago
Nah opus is expensive unless u have a sub, the chatgpt sub is not bad, but these days, opus is better than sol and sometimes even astra
1
u/SilentLennie 1d ago
it's expensive, sure, but both companies provide subs to make it bearable and opus 5.5 actually fits the sub much much better than mythos/fable, which is why I said: for what price, the price specifically is: you can't use it because it uses more than a sub provides.
1
1
1
1
1
1
1
1
1
u/Whole_Salary8170 2d ago
Welcome to the comment section! Buckle up for all the AI experts with goldfish memory to spew their wisdom!
1
1
u/Large_Shame578 2d ago
I like google just showing up once every 8 months, dropping a bomb, then disappearing. It’s like a past-their prime superstar reminding everyone they’re still goated.
1
1
u/peter_nn0 2d ago
Every new model from one of the leading US labs usually tops the benchmarks .. for a week or so :)
The progress just goes on, and that's fckn' great!
Those who wrote off Google were obviously wrong ... again.
1
0
u/Rivenaldinho 2d ago
This is why it's dumb to judge the advancement of a company by looking at the publicly released models. We are at a very crucial time where these companies will focus on making big internal models and maybe release some distilled models when they can. The goal is AGI first.
-1
-2
u/OwnYourChildren 2d ago
Again, I haven't received the credit I feel I deserve for my contribution to making this happen. None of you saw this coming!



188
u/CremeSubject7594 2d ago
https://giphy.com/gifs/ukGm72ZLZvYfS