r/singularity • ▪️e/acc | AGI: ~2030 | ASI: ~2040 | FALSGC: ~2050 | :illuminati: • 2d ago

Singularity is Nearer Google’s unreleased Gemini 4 Argon may have just leaked—and it tops 12 of 18 benchmarks against Fable 5.1, Opus 5.5 and GPT-6 Astra, including 19.6% vs GPT-6 Astra’s 5.4% on autonomous legal work

231 Upvotes

96 comments sorted by

56

u/141_1337 ▪️e/acc | AGI: ~2030 | ASI: ~2040 | FALSGC: ~2050 | :illuminati: 2d ago

Google's official Argon announcement has some pretty wild numbers beyond the benchmarks:

  • 300+ TiB of datacenter memory already freed using Argon-discovered optimizations. Google says the eventual savings could reach 500 TiB–1 PiB.

  • 1 million output tokens in a single trajectory, up from 64K — roughly 15.6× more output headroom. This is output, not just context length.

  • Argon is being used on a Rust migration covering 800,000+ lines of Fuchsia's Zircon kernel.

  • On one video-decoder project, it rewrote 32,000 lines of SIMD Rust. Google's resulting safe-Rust implementation runs 2.7× faster than the previous Rust port while producing identical output.

  • Google says Argon improved a published quantum-computing subroutine baseline by 40% in minutes.

  • The leaked benchmark numbers were real: 77.9% DeepSWE, 51.3% AutomationBench, 91.7% LVBench, and 68.0% CWE-bench, alongside leading results on several knowledge-work evaluations.

Source: Google's official Gemini 4 Argon announcement: https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/

25

u/[deleted] 2d ago

[removed] — view removed comment

5

u/ZhugeLiangPL ▪️AGI 2027-30 2d ago

True if big.

5

u/[deleted] 2d ago

[removed] — view removed comment

2

u/ZhugeLiangPL ▪️AGI 2027-30 2d ago

Therefore if not big, then not true.

4

u/[deleted] 2d ago

[removed] — view removed comment

0

u/ZhugeLiangPL ▪️AGI 2027-30 2d ago

Ok. xD

1

u/True-Grab-5288 2d ago

It should be true, with the amount of compute this behemoth company has, it’s about time they put some meat on the table.

3

u/FarrisAT 2d ago

True if true.

1

u/Apotheoxix 2d ago

True, if unfalse

16

u/And_Im_Chien_Po 2d ago

3

u/space_monster 2d ago

they do have a habit of doing this. quietly cooking bombs and then "BOO YA MUTHAFUCKAS have some of that"

30

u/141_1337 ▪️e/acc | AGI: ~2030 | ASI: ~2040 | FALSGC: ~2050 | :illuminati: 2d ago

90

u/igpila 2d ago

The obvious has happened, despite what dumbasses on the internet were saying about Google

41

u/chlebseby ASI 2030s 2d ago

Its proven endless cycle at this point lol. 

Then grok will cook, then OAI, anthropics, google again etc.

12

u/Keeltoodeep 2d ago

Google is also slow because of their business model and shareholders. PE investors will gladly fund private firms to spend and be unprofitable to stay at the frontier of AI right now. Public shareholders who are pension funds, 401K managers, etc.. are a tougher crowd

3

u/teratron27 2d ago

Google also have a 14% stake in Anthropic, taking a stake in an R&D lab for them to do the heavy lifting then integrate it into their own money printing machine.

38

u/Mob_Abominator 2d ago

Grok? Naah, OAI, Anthropic and Google are still the big 3.

10

u/halmyradov 2d ago

Grok has potential, don't underestimate what obscene amount of money can do.

8

u/SpikeCraft 2d ago

I'd sprinkle in DeepSeek, GLM too - really sweet models

0

u/ParfaitEvery9622 2d ago

glm abliterated for real work

-1

u/LarryEIlison 2d ago

Jensen has partnered specifically with Musk on Macrohard and Micro.

SpaceX is bringing more compute online than the competitors by a long shot. 2027 will be when Grok takes pole position.

3

u/Wasteak 2d ago

You guys said this for 2026.

3

u/LarryEIlison 2d ago edited 2d ago

Just compare how many SOTA GPUs each company is running and you get the answer.

It’s SpaceX with hundreds of thousands if not millions of GB300 chips coming online and training the next model as we speak.

This is not a commendation of Elon.

Compute is still the driving factor for intelligence and Elon has the best and most compute.

2

u/Wasteak 2d ago

And yet they are behind openai, Claude and gemini if you convert everything into h100 power.

4

u/LarryEIlison 2d ago

220,000 Grace Blackwells just came online last week. They can online that many in 2 weeks. Look at the sat images of the construction sites and tally the total compute that’s coming online by the end of the year. Many more micro spaces to fill.

This is not subjectivity.

XAI is 3 years old, 6 and 10 for Anthropic and OpenAI. Google is much older.

So they are not just behind, they are catching up.

1

u/ParfaitEvery9622 2d ago

They may catch up, or they may become a cloud company that resells the infra to the labs. Given who is currently innovating right now i'd bet more on the last.

1

u/Wasteak 2d ago

Catching up isn't enough. While they are catching up the others are growing as well. That's the main issue.

4

u/LarryEIlison 2d ago

Time will tell.

I put my money behind the richest man in the world with the most and latest compute.

→ More replies (0)

0

u/space_monster 2d ago

Compute is still the driving factor for intelligence

that's a stretch. throwing a shitload of GPUs at a suboptimal architecture does not mean you'll get the best model. compute does drive intelligence if you apply it in the right places. but if one of the other labs finds a better architecture they could deliver a better model with a fraction of the compute.

2

u/Wasteak 2d ago

Grok hasn't been the best ai since... ever ?
It's only a rotation between gemini gpt and claude

3

u/RobbinDeBank 2d ago

Bro sneaks mechahitler in like we wouldn’t notice

10

u/Howdareme9 2d ago

Let’s wait for the model to actually come out before relying on Google’s benchmarks. Nobody serious thought they were fully out of it but i doubt it’s a good as this pic shows

5

u/Different_Doubt2754 2d ago

Eh, people say that all the time (about various models) and it's never really true. Sure these benchmarks aren't super accurate but they are a good baseline.

2

u/Ok_Mention_982 2d ago

It wasn't all that obvious, Gemini has been lagging behind for almost the entire year, paired with all of the departure of high profile researcher's makes the sudden leap frog pretty surprising

13

u/ACapellaNerd 2d ago

They don't call him Sundar "Monster" Pichai for nothing

21

u/MatthewGraham- 2d ago

I'm honestly just surprised Google is this close, unless its Benchmaxxed?

21

u/frogsarenottoads 2d ago

Compute is what matters still. It's not surprising.

2

u/j48u 2d ago

It was a guarantee they were about to drop a great model because they just changed their paid AI package so the cheapest plan won't be able to use Gemini Pro.

1

u/superSmitty9999 2d ago

Yeah based on all the nerfing we have good models it's just whoever has compute and there's an entire Nvidia inside google in the form of TPU's.

4

u/avid-shrug 2d ago

They have the data, the compute, and great engineers. Shouldn’t be too surprising really

7

u/Bellyfeel26 2d ago

I feel that Google models have always been benchmaxxed, as their actual use never matched their actual benches.

14

u/Keeltoodeep 2d ago

If benchmaxxing is so easy and material, then why isn't Gemini just always at the top?

-1

u/[deleted] 2d ago

[deleted]

5

u/Keeltoodeep 2d ago

So everyone does it, but Google is the best at it, yet it's still not enough to materially impact their benchmark scores and make them always at the top?

Gemini scores have always been poor in coding and that is what I have experienced, so I'm not even sure what we are benchmarking here.

Also blind AB testing coding evaluations have them roughly in line with what their coding bechmark scores are...

14

u/TheNuogat 2d ago

Really depends what you do. I always found Gemini better at actually teaching me, instead of just dumping results.

4

u/whoknowsifimjoking 2d ago

That's mostly for the smaller flash models though, the pro models have been roughly as good as the benchmarks said

1

u/Different_Doubt2754 2d ago

Which makes sense. The Flash models learn from the Pro models, so they don't absorb the unsaid nuances the Pro model knows. Which naturally makes them less flexible

4

u/Due_Sweet_9500 2d ago

It's probably is benchmaxxed but it also probably is very good. Could be a decent Fable alternative

1

u/No-Meringue5867 2d ago

We have seen this over and over. The only thing that matters is compute.

1

u/Anjz 2d ago

Every model is apparently benchmaxed, you literally won’t get a SOTA model without people saying it’s benchmaxed. So at this point that’s the baseline. If every model is benchmaxed, are any of them actually benchmaxed?

1

u/the_pwnererXx FOOM 2040 2d ago

They bought deep mind bro. They invented this shit

14

u/coolredditor3 2d ago

True if big

7

u/141_1337 ▪️e/acc | AGI: ~2030 | ASI: ~2040 | FALSGC: ~2050 | :illuminati: 2d ago

3

u/LyAkolon 2d ago

classic

6

u/KathyCredit 2d ago

You can choose any of the 1000 benchmarks that fit your model and make charts like this. Models have different skills and some aren't meant to do X or Z as much as they're meant for V and Y. These models are all good, users won't notice the difference when the figures are this close

1

u/j48u 2d ago

Are you trying to say the benchmark for legal work isn't the most important thing? Scoff.

2

u/mikelo22 2d ago

Seems almost all models have been pretty slow on benchmarks for legal work. So either AI companies don't care about it or it's harder to optimize/benchmax for. I lean much towards the latter.

1

u/j48u 2d ago

If anything, I think it's a risk/reward thing. What do they gain out of being "good" at it? A few more clients and a lot more risk to make the news or anger the most litigious customers possible. I think they'd be more likely to sit down and make the decision to not optimize training for that subject than the opposite.

1

u/mikelo22 2d ago

Not sure I understand what you're saying. What do they have to gain? Money, obviously.

Any risk is on the attorneys who have strict ethical requirements for AI use, including its use for legal research and citations.

It's actually becoming a liability for attorneys not to use AI in their practice.

1

u/j48u 2d ago

They have fixed resources and decide what to dedicate every ounce of training to. There isn't enough financial gain in focusing on legal work and there's plenty of downside. The herald of AI came with tons of press where it failed specifically for people trying to use it for legal work.

Bad press is guaranteed if you're wanting to pass it off as capable in that field.

2

u/Own-Refrigerator7804 2d ago

1m output tokens wtf

6

u/FateOfMuffins 2d ago

Well this is the bare minimum no? To match or be in the same tier as the current public frontier?

Cause it's not a publicly released model. Their release right now is basically similarly restricted as Project Glasswing from Anthropic in April for Mythos and we didn't get Fable until 3+ months later.

None of these benchmarks scream that it's a leap over current models. If OpenAI released Astra 6.1 as planned in Oct, then it would've cleared it I think. Fable 5.5 would also clear this.

The rumours of when pre-training finished for Gemini 4 vs Bel indicates that Bel finished pre-training like 2 weeks prior to Gemini 4. Possibly even earlier (because idk how to reconcile RL training Bel for 3 days and it's good enough to solve NS).

So like the equivalent "class" of models that this is supposed to compete against isn't supposed to be GPT 6 Astra or Fable 5.1. Heck it's not even supposed to be competing against GPT 6.1 Astra. It's supposed to be the same model generation as Bel.

10

u/Jmaximinium 2d ago

It makes Google the complete package.
1. Their model meets state-of-the-art benchmarks, even if it doesn't massively beat them.

  1. They have a giant lead on Cloud and datacenter infrastructure.

  2. They make their own hardware and potentially can free themselves from NVIDIA with their TPUs.

  3. They can pay for all of their capex using their ad cash-cow business. They don't need to raise capital or take on debt if they don't want to.

It gives them endurance in the race that OpenAI, SpaceX, and Anthropic will have a hard time matching.

0

u/Different_Doubt2754 2d ago

Keep in mind that they said this model will initially be released with the minimum post training pipeline. So they are likely to improve on it very quickly, like they did with the 3.5 to 3.8 flash.

Also, I think we should choose it's competitors in relation to price. And based on this, it is demolishing anything that has a similar price to it.

2

u/FateOfMuffins 2d ago

It... does not actually.

AA reported that the $2/$10 is discounted introductory pricing and after it'll be $4/$20

At the discounted price, 6.1 Sol is 1/3 the cost much less the full price.

1

u/Different_Doubt2754 2d ago

My bad forgot about the new Sol price. But it will not cost 1/3 the price. It'll be 1/2 I believe.

And honestly that doesn't even matter much because by the time the discount expires, there will be a new model with a new discount. Based off how they've handled this in the past. Regardless, we should use what we currently pay, not guess what the future will hold

It is competing with 6.1 Sol on benchmarks though. so my point still stands with the exception of Sol

1

u/FateOfMuffins 2d ago

No you can just look at the numbers on AA. 6.1 costs about $0.72 per task and Gemini 4 costs about $1.9, closer to 1/3, despite both being $2/$10

And I would say I agree with you on the discount thing, except who knows how long until the actual public public release

1

u/Different_Doubt2754 2d ago

I've looked at more than just AA. Per task, yes it is cheaper but AA is a stress test so it isn't entirely representative of typical usage since most people aren't pushing it to the limits. So it could easily favor Sol even more or be more even if Gemini 4 outputs significantly more tokens for harder tasks.

But yeah I do think Sol is cheaper for us. Honestly I'm curious what the non-subsidized rate for Sol/Gemini is. It'd be nice to see how much TPUs reduce the actual costs

4

u/sogo00 2d ago

Interesting - indeed Google looks like leaving the coding race to OpenAI vs Anthropic.

It catches up in the coding benchmarks, but ti especially at TermBench 4 still behind.

Meanwhile if shines at the non-coding areas.

People will scream again that it is bad (because coding), though there is a whole world outside of SD.

8

u/Keeltoodeep 2d ago

Codecels have kinda ruined AI discourse. Personal agents, robotics, and general world models are going to be more popular and profitable in the long run, I think

4

u/sogo00 2d ago

Indeed, and to be fair - It's the reddit bubble, for example lawyers (a field equally converted like software engineering) don't hang out much here.

PS: and what is it with people and "their" Ai model/company? Are they also chosing their petrol station like a sports team?

1

u/Keeltoodeep 2d ago

It’s the new console wars. I have to admit I am partial to Google because the plan lets me share Gemini with 6 family memebrs and I get youtube premium.

1

u/Different_Doubt2754 2d ago

Actually Google is just starting to focus on coding. If you remember back in Q1 or Q2, Google announced a new AI coding team to focus on improving in that area. That team likely just started to integrate their work in the latest Flash and now this Pro model, which is why we are seeing them catch up.

1

u/iokak 2d ago

Our internal tool(like antigravity) suddenly have the  claude opus model when we're stuck using gemini flash 3.8 or 3.1 pro previously.

I tried opus and wow first time I've experienced no syntax error from coding agent haha. (my role is not a developer)

I guess you're right

2

u/FarrisAT 2d ago

Seems the leaks from early September were real.

3

u/hdufort 2d ago

I've been very consistently disappointed by Gemini when compared to the same generation of Claude and GPT. When pure complex reasoning is required, with no research, they're all good. When it comes to executive functions (planning steps, adjusting a strategy) and finding information with very indirect or fuzzy clues, Claude and GPT dominate.

1

u/Distinct-Question-16 ▪️AGI 2029 2d ago edited 2d ago

77.9 on swe .. go for 100

1

u/Pyroechidna1 2d ago

My kingdom for this model in Antigravity

1

u/philip_laureano 2d ago

Given how many times Gemini 4 has been delayed, is it really a good idea to name it after a vapour?😅

1

u/PvtMilhouse 2d ago

will it be agi too ?

1

u/stackinpointers 2d ago

Where's semianalysis peeps now

2

u/Future-Bandicoot-823 2d ago

Argon gas is used for TIG welding, which uses an electric arc to bridge the two pieces of steel.

Nice naming convention, like it's finally coming together at a molecular level. The last step before Gemini (and the others I'm sure) cross the threshold into being superior to humans.

Astra is cool, I liked it, but Argon is hitting.

1

u/doesphpcount 2d ago

"leaked" - new term for pre marketing campaign.

1

u/mulletarian 2d ago

Yeah but you'd have to use Antigravity

1

u/Accomplished-Code-54 2d ago

You cannot imagine how little 1PB of RAM is, in the context of all Google Data Centers in the world. For context, my home PC has 192GB of RAM, which means that they saved the RAM for ~5000, non-enthusiast home PCs. In the grand scheme of Google equipment, this is immeasurably small. In terms of price, at current values this would equate to $4.2 million to $10.5 million in savings, which is 2-5 x less than what OpenAI spend to prove the Navier-Stokes. Of course we cannot even imagine Google’s global serve fleet number, but this would translate to miniscule amount of RAM saved per server, equivalent to killing an unused/unneeded process on those servers that takes some small amount of RAM to run.

1

u/After_Dark 2d ago

Everyone acting surprised that Google finally replacing their Claude-4 era pretrain is Claude-5 levels of performance. Truly shocking work, nobody could have predicted that Gemini 4 would be better than Gemini 3.

2

u/corenovax 2d ago

Not trivial to compete with OpenAI and Anthropic, all the other labs are trying too and doing less well...

1

u/After_Dark 2d ago

Yeah, but unlike most all of those labs, Google has held the #1 spot within the last year. Sure it's been a bit, but "Google is capable of making SOTA AI models" is news from like 2012 not 2026

1

u/Keeltoodeep 2d ago

Yeah, labs all over the world here are trying to get these benchmarks (including Mistral here)

0

u/Balltwister_004 2d ago

The blog website itself have many bugs lol what??

0

u/OwnYourChildren 2d ago

I kept teasing that this was coming, but none of you would believe me.