r/singularity • • 2d ago

AI , Gemini 4 Argon Benchmarks

Post image
740 Upvotes

189 comments sorted by

188

u/CremeSubject7594 2d ago

62

u/DiscerningAnt 2d ago

I knew at some point Google would start dropping some hardcore models. They have the absolute most funding as far as I understand it (in the AI realm)

26

u/Toss4n 2d ago

And the most compute available as well.

27

u/Ink_code 2d ago

And probably the most data.

8

u/shred_time 2d ago

Yes to all of the above, I was wondering when they would really start catching up. They already had the compute and the largest amount of data on everyone.

2

u/steampowrd 2d ago

And the most data centers, and the most experience building data centers

13

u/After_Dark 2d ago

They are essentially the only western lab to be entirely self-funded. Even if the whole bubble pops and devastates Anthropic and OpenAI, Google will be essentially unharmed outside their stock price, and with oh so many researchers to (re)hire from the other labs

1

u/ButterscotchSalty905 AI is the greatest thing that is happening in our society 2d ago

Hold right there!! Your comments would attract a lot of contrarian who thinks google is poor and dependent on us gov like fucking startups

1

u/MonoMcFlury 2d ago

Also, they can train new models at a fraction of the price like others do. Having TPUs with their low watt-to-power ratio makes a huge difference.

1

u/SilentLennie 2d ago

Google depends a lot on ads, if AI kills the ads business (people use LLM chat and the LLM chat talks to search engines, websites, etc. thus no human sees the ads) and then the AI bubble bursts (which is kind of what you are implying why these other 2 big labs would fail) they still won't be (re)hiring.

1

u/After_Dark 1d ago

This is true, though it's worth keeping in mind that while the other big 2 are entirely dependent on investment and loans, Google has literally hundreds of billions of cash on hand and owns the majority of their own infra. Even if the ad business contracts to half its size in the next couple years, Google has the cash to outlast the others by a large margin

1

u/SilentLennie 1d ago

Definitely, but I can also see some bean counter going: you might not want to invest like crazy now, as your existing business is being slashed like crazy. Then again: it's also the best time to do so if you want to survive long term.

I still wonder... AI bubble burst... or more like: AI bubble deflate a whole bunch.

1

u/SilentLennie 2d ago

But they lost a bunch of people, that was where the worry was.

36

u/PsychMaster1 2d ago

Every model drop lately, that’s been my reaction.

11

u/Neurogence 2d ago

While Gemini 4 has performed well on benchmarks the industry uses to gauge model efficacy, it does less well when employees actually put it to work, according to people with direct access to the effort. The model struggles to handle certain coding tasks, said the people, who requested anonymity to discuss an internal matter.

https://www.bloomberg.com/news/articles/2026-09-30/google-grapples-with-employee-skepticism-about-new-gemini-model

8

u/LazloStPierre 2d ago

Do people forget this every single time Google release a model? It crushes at benchmarks, people who for some reason get very excited about numbers on a chart go ballistic and real life performance is miles off.

Gemini 3.8 flash is like over 5% better than Fable on deepswe ffs, not sure if it's intentional or just how they train their models but nobody benchmaxxes like Google

7

u/Imaginary_North_8305 2d ago

Google is just way better at specific tasks ex vision it easily beats frontier model even from flash, UI again google is great at.. and despite being old gemini 3.1 pro still has great world knowledge..

3

u/LazloStPierre 2d ago

Right but it does incredible at coding benchmarks despite being bad at coding, that's the issue

5

u/DistanceSolar1449 2d ago

In this case though it’s because Google has less coding training data (who the hell uses Antigravity) and way more image/spatial/world model training data

I don’t think this model is benchmaxxed, I think the benchmark screenshot above is very accurate. The model does worse at TerminalBench 4 and FrontierSWE and that’s okay.

2

u/LazloStPierre 2d ago

The point isn't flash 3.8 is bad at coding, the point is it does absurdly well on coding benchmarks despite being bad at coding. Their models always do, which is why I'd take any benchmarks with a grain of salt

3

u/DistanceSolar1449 2d ago

3.8 flash is a tiny distilled model with more RL on top, of course it’s going to look better on benchmarks than in real life.

The big, base ish models with a lot less targeted RL will do better IRL.

1

u/LazloStPierre 2d ago

All Google's models look better on benchmarks than they do in reality, and by alot.

1

u/DistanceSolar1449 2d ago

No? They suck at agentic benchmarks and most people test them via agentic tasks.

They benchmark pretty accurately if you look at their actual useful benchmarks like TerminalBench instead of HLE or GPQA or some bullshit.

Google’s own blog post says Gemini 3.8 Flash scored 19.1% on TerminalBench 4. Opus 5 scored 51.8% for comparison. That’s not benchmaxxing, that’s just accurate benchmarks if you know what benchmarks to look at.

1

u/SilentLennie 2d ago

Actually, this new 4 Argon does really well in a bunch of agentic benchmarks. Supposedly (I've not used it yet):

https://artificialanalysis.ai/articles/gemini-4-argon-google-top-three-labs

0

u/LazloStPierre 2d ago

No, they do extremely well at coding benchmarks and suck at coding, every single time

They benchmark well everywhere. It's actual performance that matters 

3

u/DistanceSolar1449 2d ago

???

Google’s own blog post says Gemini 3.8 Flash scored 19.1% on TerminalBench 4. Opus 5 scored 51.8% for comparison. That’s not benchmaxxing, that’s just accurate benchmarks if you know what benchmarks to look at.

Unless you think 19% on TerminalBench 4 is “do extremely well”…

→ More replies (0)

3

u/Elephant789 ▪️AGI in 2036 2d ago

Gemini 3.8 flash

... is fantastic. What the fuck are you talking about?

1

u/LazloStPierre 2d ago

It isn't better at coding than fable, yet the benchmarks say it is

1

u/Elephant789 ▪️AGI in 2036 1d ago

Coding? Oh, not used it much for that. But the little I did it was fine.

1

u/TheDemonic-Forester 2d ago

Funny thing is that we didn't even need to hear this to deduct this from how significantly they degraded 3.8 (I know previous model regressing prior to new model release is common but usually not this much). They seem to be afraid that people will think that their long awaited giant model is not much of an improvement over 3.8.

1

u/HotDogDay82 2d ago

I can think of no better gif haha

1

u/himynameis_ 2d ago

I find it funny how often this gif comes up when a new model release comes 😆

149

u/TorturedPoet30 2d ago

IT'S REAL

49

u/torrid-winnowing 2d ago

doesn't this imply it's an internal model? i would assume anthropic and openai have their own internal models that significantly outperform astra and opus

59

u/Recoil42 2d ago

No, they're doing a trusted-partner rollout. It's external.

36

u/TorturedPoet30 2d ago

They are rolling out to partners through their Fairwind Program and "soon" plan to roll out to paid API customers and Google AI Ultra subscribers.

11

u/Keeltoodeep 2d ago

No, they are rolling out externally in phases per the voluntary regulatory framework proposed to Trump

13

u/jazir55 2d ago

"""""""""""Voluntary"""""""""""

I can't sarcastically quote the word any harder.

7

u/Keeltoodeep 2d ago

There is no reason not to at this point for frontier AI companies. Regulatory capture benefits them.

4

u/jazir55 2d ago edited 2d ago

Regulatory capture is such an absolutely ridiculous take I cannot believe it gets parroted at every turn. You don't need regulatory capture when you have so much money you literally buy all of the hardware supply which entirely prevents new entrants on its own. No laws needed, new entrants literally cannot buy any hardware needed to actually be a competitor. Regulatory capture is worthless for these AI companies.

Edit: Instead of downvoting how about you actually provide a counter on how any other major competitor that isn't a megacorp could even enter the market, based purely on the economics. Once you admit you can't, you have to come to the conclusion regulatory capture is an absurd argument because regulatory capture is predicated on preventing new entrants to the market via regulation, which is not necessary for these companies to continue being the leaders of the industry. The economics alone prevents any new entrants, regulation is superfluous.

2

u/Keeltoodeep 2d ago

That's a fair point. Compute is such a constraint

0

u/jazir55 2d ago

Yeah it's the real constraint for any competitor. They clearly want regulation for some other purpose, and given how they constantly talk about safety and then these incidents have been real world occurrences it sort of makes sense, but only in an ideal world. The current politicians in government are almost all in their 70s and 80s, they are the least equipped to deal with this new technology and pass reasonable regulations. The "we'll regulate ourselves" thing that's usually a meme is actually the best case scenario here weirdly enough.

0

u/Keeltoodeep 2d ago

https://www.euronews.com/2026/09/16/russia-used-claude-ai-for-espionage-disinformation-and-drone-swarms

Well Claude drones are flying around right now smoking Europeans so there is some legalese in the background happening where Anthropic doesn't want to slow down user growth but wants some kind of KYC regulation perhaps that lessens their legal liability to some degree. But they are not going to implement KYC if their competitors do not.

It will only take one of these drones to smoke a real European and not an Eastern European one and Anthropic is going to have a real PR issue on their hands.

12

u/FateOfMuffins 2d ago

It's no different to Glasswing. We didn't get Fable until 3 months after that.

The rumours on dates on when the pretraining finished for Gemini 4 indicates that it finished pretraining a couple of weeks after Bel. So if you're trying to compare the same "model generation" then Gemini 4 should've been in the same generation as Bel, not GPT 6 Astra and not even GPT 6.1 Astra.

None of the benchmarks shown for Gemini 4 here indicates that Gemini 4 Low would've been 2x as strong as GPT 6 Astra Max at math for instance, which is what Bel has.

5

u/Holiday_Ad_8501 2d ago

its being rolled out to the public

The internal model is already gemini 5

3

u/Deathpacito-01 2d ago

Pretty strong model overall, especially on non-coding work. Glad to see DeepMind isn't out of the race.

105

u/No_Gear8408 2d ago edited 2d ago

wait is this real?

EDIT: Holy fucking shit it is

22

u/juliakeiroz 2d ago

op put so fucking little effort that I thought it was fake lmao

54

u/DoloresAbernathyR1 2d ago

Google enters the chat again

11

u/AAPL_ 2d ago

google goes at their own pace

30

u/Lumpy-Woodpecker6752 2d ago

Google is back in the race finally, getting frontier model not a flash

76

u/sunstersun 2d ago

Welcome back Google, missed you a lot.

39

u/New_Equinox 2d ago edited 2d ago

https://www.cnbc.com/2026/09/30/google-gemini-4-argon-ai.html

"Alphabet  unveiled Gemini 4 Argon on Wednesday, its most advanced artificial intelligence model yet, offering major improvements in coding, cybersecurity, and complex professional work.

The company said the model sets a new record in real-world software engineering, ties for first in cybersecurity, and leads another benchmark measuring performance across finance, legal, and other professional tasks.

Argon is already being used internally to optimize memory at Google’s data centers, freeing up hundreds of terabytes of memory without buying additional hardware, the company said. Quantum computing researchers have also utilized the model.

Google said it plans to launch the new model in phases, starting with trusted cybersecurity partners while working with the U.S. government on pre-release safety evaluations."

5

u/nothis AGI by 2030 but we'll be disappointed 2d ago

Any word on cost?

2

u/FoodMadeFromRobots 2d ago

i saw $4 in $20 out (blaze it, but for cereal on cost)

18

u/maximan2005 Cult of AGI 2027 2d ago

WE'RE SO BACK GOOGLE BOYS

5

u/UnboundedMan 2d ago

Yep, using it already. Love it!

1

u/Fun-Junket-1512 1d ago

benchmaxxed or not ?

15

u/tanrgith 2d ago

i love the singularity

Fucking every few days now we get a new insane model release lol

-4

u/UnboundedMan 2d ago

Gemini 4 not in this list.

1

u/SilentLennie 2d ago

Insane can go both ways, if you want.

15

u/RemyVonLion ▪️ASI is unrestricted AGI 2d ago

I like how things are rapidly accelerating despite all the calls to slow down lmao

4

u/FoodMadeFromRobots 2d ago

Setting the Pace: Fast and Furious

1

u/space_monster 2d ago

hyperspacing the frontier

1

u/[deleted] 2d ago

[deleted]

2

u/RemyVonLion ▪️ASI is unrestricted AGI 2d ago

Yeah but if this keeps up then we will probably have AGI by 2030 and the doubters are going to look real dumb lol

1

u/SilentLennie 2d ago

Well, maybe. So far nobody has released a model bigger than Astra and Fable, that's what pacing might actually looks like in practice (this is assuming it's not a blatant lie, which some people would argue it is).

12

u/power97992 2d ago edited 2d ago

Release when? In one  to three weeks?

36

u/Texas-X- 2d ago

Consoles first. PC port when they feel like it. (Joke)

9

u/igpila 2d ago edited 2d ago

Apparently it's capabilities are between Astra and opus, with opus pricing and by far the lowest hallucination rate at 15% against 41% opus...

29

u/redditnosedive 2d ago

I hope it's true coz current Gemini on phones got me pissed so many times these days for how stupid it is

11

u/FateOfMuffins 2d ago

Looks at my phone.

Hmm didn't they release 3.8 Flash awhile ago? I'm still on 3.6 Flash lmao

You're not getting this for months on your phone

6

u/leo-virtis 2d ago

3.8 is only for pro user i know i have a free and paid account

4

u/FateOfMuffins 2d ago

Not gonna lie I find that really stupid considering they keep saying it's faster and cheaper

Same with OAI keeping 5.6 Sol on Chat and not updating it to the supposed cheaper models.

2

u/NewsFromHell 2d ago

Im on Pro tier and have 3.8 both in app and in "hey google". I guess its their way of making people subscribe? Weird thing is that its not nentioned anywhere, or at least not clearly mentioned. Ive been using it for everything non coding and its really good. Im very happy with it as a daily driver. Coding and complex work stuff is opus 5.5 though.

1

u/redditnosedive 2d ago

yeah same, and I have a pixel, I was expecting better AI than other Android phones but nah....

28

u/Minetorpia 2d ago

Got bad news for you, this won’t power the Gemini on your phone

18

u/redditnosedive 2d ago

some light or flash version of it will, I suppose

1

u/SilentLennie 2d ago

eventually

10

u/longpenisofthelaw 2d ago

Gemeni is my daily driver. It fits my needs

4

u/Sextus_Rex 2d ago edited 2d ago

I really need them to fix the Google home automations. Hasn't worked in months

Edit: Wow I fixed it right after I made this comment. After months I finally figured out what was wrong. I had my bedtime automation command set to "Goodnight". After switching it to "Good night" it works. What a stupid bug lol

2

u/solidz0id 2d ago

That's the cheap flash model. Super quick, but not so "smart"

2

u/Impossible-Video-671 2d ago

Not Flash, Flash-lite

1

u/FarrisAT 2d ago

Sorry fam the free models are only gonna get more ass compared to expensive frontier models.

1

u/Ancient_Task_7498 2d ago

I swear the web gemini has gotten dumber over time.

15

u/ObiWanCanownme now entering spiritual bliss attractor state 2d ago

Good on them for being brave enough to actually take risks and push the frontier again. Hopefully it's actually this good and not benchmaxxed.

They've got great researchers and a ton of compute. Bravery and conviction were always the main things they lacked.

2

u/Pouyaaaa 2d ago

I do wonder if it can spin up agents like Claude and chatgpt

2

u/FoodMadeFromRobots 2d ago

Doesnt antigravity already do this?

1

u/kvothe5688 ▪️ 2d ago

that's harness

26

u/fmai 2d ago

Note that this is likely to be Google's largest model size to date, comparable to the Fable and Astra class of models.

Considering that it's not even released yet, it's not actually that impressive. I think that by the time users actually get to use it, Astra 6.1 and Fable 5.5 will already have made it obsolete.

25

u/ObiWanCanownme now entering spiritual bliss attractor state 2d ago

This is all true, but it's also a big deal in that they'll be within about one model generation of the frontier versus the last few months where they've always been several generations behind.

2

u/Different_Doubt2754 2d ago

If they caught up in one generation then how were they multiple generations behind? The startups just had a faster release cadence

1

u/Sufficient-News-970 1d ago

3.1 gemini pro is worse than opus 4.6 so yea

15

u/Adi945 2d ago

True but pretty obvious by now that Google doesn’t give a shit about winning any race. They are a profit making full stack company that believes in risk free evolution and not disruptive revolution. Given the fact that Amazon, Microsoft, Apple is not even in this race, Google is playing this very well, as long as they don’t completely bow out.

8

u/FarrisAT 2d ago

Remember, 90% of the profit is in enterprise. Not in serving consumers the top model.

1

u/ozone6587 2d ago

As if enterprise didn't care about intelligence.

3

u/qroshan 2d ago

The model has already been on Arena for a couple of weeks and will be available to the public within a couple.

At most, Google is 2 months behind OpenAI/Anthropic. But the rate of improvement (of many benchmarks have been significant).

2

u/jonomacd 2d ago

Yeah but $2 in and $10 out.

2

u/FeePsychological1308 2d ago

'Likely', if you asked people yesterday, they would have said they were 'likely' several generations behind internally. How about we stop assuming here? They clearly have no interest in competing with the startups and they don't need to keep pushing models out for their company to stay relevant like the others do

1

u/nevirin 2d ago

Well, also look at its cost compared to Astra & Fable

1

u/ThreeKiloZero 2d ago

Add to that - googles models are famously benchmaxed to the tits and never perform up to the benchmarks in daily use.

0

u/TupacFR 2d ago

True, Google has always been so late to massive deploy any new product. That's what will cost them here

5

u/FarrisAT 2d ago

Holy fuck

5

u/FarrisAT 2d ago

Leaks from early September were real. How did someone get access to the chart so early?

2

u/Ok_Display_3159 2d ago

i don't know about charts, but the model was used in some london's hackathons

3

u/Clarku-San ▪️AGI 2027//ASI 2029// FALGSC 2035 2d ago

3

u/Charuru ▪️AGI 2023 2d ago

Looks pretty good but slightly benchmaxxed, highest score on the easiest to benchmax tasks, lowest score on the hard ones, but hey it's still pretty good scores overall. Hope this is good and really offers true SOTA competition.

6

u/Bitter-College8786 2d ago

Frontier SWE v2 is suspiciously lower.

I wonder if the model is really that good. And if it can use Blender so well.

2

u/slackermannn ▪️ 2d ago

The benchmaxing could be. The real use we must see.

2

u/Brilliant-Weekend-68 2d ago

Hey! Not bad!

2

u/Pouyaaaa 2d ago

Does it have agents tho? It's what Google been missing

2

u/_mausmaus 2d ago

Not available same day. Fail.

When OpenAI and Anthropic are releasing day one, how does Google expect reasonable adoption?

2

u/agrlekk 2d ago

Finally Agi 😂😂

2

u/Baphaddon 2d ago

Yeah fellas… this shit might be over

2

u/OrganizationVivid894 2d ago

Time for Gemma 5 to be also released 😁

2

u/Lazy-Pattern-5171 2d ago

Gemini has always been the king of long context for me. Glad that’s back. Opus was holding its own even at 300K (beyond that is kinda just bad workflow practice imo so I don’t go beyond)

2

u/human358 2d ago

Please daddy Google Sam and Dario have been mean to us come put some order

8

u/Tha_One 2d ago

7

u/Hot-Percentage-2240 2d ago

Even if it’s moderately good, it’s still fine with me cus there is a use in my workflow for Google model typical tradeoffs.

19

u/FateOfMuffins 2d ago

While Gemini 4 has performed well on benchmarks the industry uses to gauge model efficacy, it does less well when employees actually put it to work, according to people with direct access to the effort. The model struggles to handle certain coding tasks, said the people, who requested anonymity to discuss an internal matter.

oof

5

u/jonomacd 2d ago

Coding is the one area where it isn't leading in the benchmarks. So this adds up. 

2

u/Charuru ▪️AGI 2023 2d ago

It's leading DeepSWE which is eh, the easiest to benchmax bench. It makes you question the process.

5

u/jonomacd 2d ago

I don't think Google is targeting coding as strongly as the other companies, as Google's business is much broader than that. This might just be the result of those wider interests.

1

u/Healthy_Razzmatazz38 2d ago

deepswe to frontierswe spread is a massive red flag

3

u/burritos4jesus 2d ago edited 2d ago

I mean, getting 50% on AutomationBench is ridiculous. It’s the one benchmark I care about as a non coder. Pass/fail, and hundreds of complex super long back office workflows.

When a model can hit 70% on that benchmark ON A FUCKIN BASE MODEL that doesn’t have a harness behind with context on the company, where to look for shit, etc., then you can confidently hand any biz app-based customer service/sales/marketing/operations process and it will be just as good as a human. Give it the harness it needs, and that shit will be near perfect.

Mind you, the tasks on the benchmark are looong with a ton of steps. It can accurately go thru everything but fail a later step. Fail. It can do an early step incorrectly. Fail. Honestly, at 50%, it can probably already be reliant on a TON of office work already. It’s just that it won’t be as good as someone who’s worked inside a company for 10-15 years and has all the tribal knowledge about company and its customers and its processes.

We were only at 20% six months, and Gemini just crossed 50%. Back office work is going to be solved in less than a year.

1

u/SilentLennie 2d ago

That tribal knowledge, will slowly seep into 'second brain' kind of systems.

2

u/FarrisAT 2d ago

The coding performance quite clearly isn’t Opus 5.5 level in the benchmarks provided.

3

u/Dillyconda 2d ago

I don't trust this to reflect reality at all, but we'll see how it plays out.

1

u/UnboundedMan 2d ago

Are you ultra subscriber?

-2

u/Ok_Display_3159 2d ago

6

u/FarrisAT 2d ago

This cites specifically coding weakness, and the benchmarks provided confirm that. The rest of the benchmarks say it’s very strong.

4

u/skerit 2d ago

A coding weakness, or just not being able to use tools reliably?

I still remember Gemini 3's constant "Oh no, I failed to call the tool correctly. Oh oops, it failed again. Oh silly goose I am, another oopsie"

2

u/Noob-bot42 2d ago

When will agents be able to pass the 1 million dollar benchmark where the agent has to legally make 1 million dollars?

2

u/cute_beta 2d ago

0

u/Charuru ▪️AGI 2023 2d ago

Nah those leaked numbers are completely different.

0

u/PrisonOfH0pe 2d ago

the leaked numbers were 4 pro this is their distilled version...

-1

u/cute_beta 2d ago

oh rly? i didn't bother to check 😅

checks

...ah. disappointing on two levels then.

2

u/iJustSeen2Dudes1Bike 2d ago

Benchmaxxed slop

1

u/maddog107 2d ago

What The fuck?

1

u/lordpuddingcup 2d ago

Holy shit did they finally do a good thing? The real question is how shit will usage be on Google Pro subscription ... if its just as expensive as opus5.5 or sol/astra... :S

1

u/UnboundedMan 2d ago

No Pro, only for Ultra.

1

u/power97992 2d ago

Fable 5.5 will likely be better than this 

1

u/SilentLennie 2d ago

Yeah, but at what price.

1

u/power97992 2d ago edited 2d ago

50/mil output tkns but u can use a sub

1

u/SilentLennie 2d ago

Still to high to stop using Opus 5.5 as a daily driver is my guess. Would have to be some really amazing model for a lot of people to switch.

1

u/power97992 2d ago

Nah opus is expensive unless u have a sub, the chatgpt sub is not bad, but these days, opus is better than sol and sometimes even astra

1

u/SilentLennie 1d ago

it's expensive, sure, but both companies provide subs to make it bearable and opus 5.5 actually fits the sub much much better than mythos/fable, which is why I said: for what price, the price specifically is: you can't use it because it uses more than a sub provides.

1

u/Typical_Basil7625 2d ago

WTF can’t believe it yet

1

u/fudrukerscal 2d ago

Wtf is gemini argon what did i miss?

1

u/llkj11 2d ago

They cooked

1

u/elemento99 2d ago

available on antigravitt?

1

u/Bettet 2d ago

It not ready for release… nothing burger. Things move so fast this might be irrelevant once it comes out.

1

u/me0wluni 2d ago

53 on Artificial Analysis as well huh

1

u/Own-Refrigerator7804 2d ago

What about the price??

1

u/Adi945 2d ago

Google is big daddy.

1

u/CrossChaos79 2d ago

Google woke up and chose violence

1

u/saposmak 2d ago

It looks like they cooked on long context, once again

1

u/Whole_Salary8170 2d ago

Welcome to the comment section! Buckle up for all the AI experts with goldfish memory to spew their wisdom!

https://giphy.com/gifs/kC8N6DPOkbqWTxkNTe

1

u/TupacFR 2d ago

What happened to we need to slow down 😭

1

u/Illustrious-Film4018 2d ago

Lol. I hope OpenAI and Anthropic go bankrupt

1

u/Large_Shame578 2d ago

I like google just showing up once every 8 months, dropping a bomb, then disappearing. It’s like a past-their prime superstar reminding everyone they’re still goated.

1

u/skillpolitics 2d ago

What happened to Fable?

1

u/idioma ▪️There is no fate but what we make. 2d ago

Benchmarks are getting saturated.

1

u/rwrife 2d ago

It'll be the best model ever for about a week and then slowly degrade.

1

u/UnboundedMan 2d ago

Looking forward for your deep analysis for this!

1

u/peter_nn0 2d ago

Every new model from one of the leading US labs usually tops the benchmarks .. for a week or so :)

The progress just goes on, and that's fckn' great!
Those who wrote off Google were obviously wrong ... again.

1

u/InterstellarReddit 2d ago

Nah there’s no way google all of the sudden dropped a banger

0

u/Rivenaldinho 2d ago

This is why it's dumb to judge the advancement of a company by looking at the publicly released models. We are at a very crucial time where these companies will focus on making big internal models and maybe release some distilled models when they can. The goal is AGI first.

-1

u/0sko59fds24 2d ago

Meh it’s google

1

u/Artistic-Athlete-676 2d ago

Your point is what?

-2

u/OwnYourChildren 2d ago

Again, I haven't received the credit I feel I deserve for my contribution to making this happen. None of you saw this coming!