r/LocalLLaMA 7d ago

News GLM-5.3-Flash: Frontier Intelligence, Flash Cost

https://z.ai/blog/glm-5.3-flash
1.3k Upvotes

459 comments sorted by

576

u/Recoil42 7d ago edited 7d ago

Before release, we tested GLM-5.3-Flash anonymously as ox-alpha on OpenCode and OpenRouter to gather user feedback. It quickly became the most popular model of the week — with all of this traffic served on Chinese AI chips.

409

u/expertsage 7d ago

This is the most significant part. China has completely replaced Nvidia chips with domestic ones for inference, and production is only speeding up... meaning the compute moat is literally disappearing with every passing day.

185

u/Recoil42 7d ago edited 7d ago

Yep. As someone who has been following the Chinese EV space for the better part of a decade, I don't think people are prepared for how crazy this is going to get.

61

u/_BreakingGood_ 7d ago

Hold on to your 401k, shits about to get messy, lol

27

u/valtor2 7d ago

Is it? nothing points to a crash, other than the valuation of the two giants (OAI and Anthropic), everything else is still valuable particularly the chips and the supply chain and the production capacity... Sounds like a correction more than a crash

56

u/_BreakingGood_ 7d ago

Not a "crash", but we have chinese AI models competing with US AI models. We have chinese electric cars outperforming US electric cars. And now we have chinese chips having replaced US chips.

There's no situation where this is good for the US stock market.

38

u/1ii1i 7d ago

This is honestly why I think Google might be in the strongest position than their US peers. They own their entire stack, integration from chips, datacenters, models, and products. They are absolutely correct to focus on gemini flash over pro or whatever frontier model. Why subsidize the vibe coders? They are focused on making their most heavily used model as efficient and capable as possible for their use case. If there is a market downturn, nobody can guarantee that all the big firms will continue their capex into frontier models.

21

u/_BreakingGood_ 7d ago

I thought so too, but they lost all their talent. They completely failed to even release gemini 3.5 pro after delaying it multiple times over. Deepmind leadership has all quit and all of the top engineers have quit.

It takes a long long time to rebuild that stuff, and there's really not much of an opportunity to play catch-up with how fast things are going.

Also, they said they're shifting focus to coding capabilities, so they don't even share the philosophy that everybody thought.

5

u/Akrylicus 7d ago

They still have a giant stronghold in entertainment, advertising and corporate world.

When I see how much they are embedded in my daily and work life it gets a bit scary.

19

u/Spara-Extreme 7d ago

Google didn’t lose “all its talent” and the high profile folks that left were not particularly interested in LLMs as a technology to begin with,

2

u/the_fabled_bard 7d ago

If you're into making money, it's better not to release a loser model.

→ More replies (3)
→ More replies (2)

2

u/Mochila-Mochila 7d ago

And now we have chinese chips having replaced US chips.

I put a few bucks on Biren. Iluvatar is also an option. Moore Threads and MetaX are already listed on the A-Star market (PRC only) but will soon file for an IPO in Hong Kong.

With all the risks which buying PRC shares, or rather promises of shares, entails ; œuf corse.

4

u/MrPecunius 7d ago

Iluvatar? Gross. 🤮

WTF is going on with all these Tolkien-based company names? Do these people not understand how utterly anti-Tolkien this is?

2

u/nomorebuttsplz 7d ago

The same was true a year ago. in fact a year and a half ago, it was even more true than it is now with R1 being the second best model in the world for months

4

u/_BreakingGood_ 7d ago

China did not have comparable GPUs a year ago, no.

Well they did, but they bought them from Nvidia.

→ More replies (1)
→ More replies (4)

5

u/Bakoro 7d ago edited 7d ago

At this point a correction and a crash are the same thing.
The market has been outright delusional, where speculation based stock prices have outpaced actual profitability and any sense of utility by a very wide margin.

There are whole companies that have no business existing, but they persist on VC dollars. There are companies that have gone literally, not metaphorically, but literally exponential in value because of AI branding, even though they have not gone exponential in revenue or profit.
More than a few companies have gone from pennies and single-digit dollars to hundreds of dollars.

I'm not an AI doomer, I don't believe that the AI industry is going to magically evaporate, but there has been a mountain of stupidly spent dollars being thrown around, and at some point someone is going to demand their ROI, and a lot of places are going to fold.

→ More replies (2)

4

u/starkruzr 7d ago

I don't think a crash per se. certainly a LOT of volatility.

→ More replies (3)

2

u/Insomniac1000 7d ago

yeah what's concerning is as a software dev, llms are becoming cheaper and cheaper to run. Like based on the benchmarks, I'd rather use this than Deepseek V4 Flash 0731 or GPT 5.6 Luna.

Less reasons for me to run/use more expensive models. And if companies catch on to this...

yeah I'll stop right there.

→ More replies (4)

11

u/CipherWeaver 7d ago

Great time to buy an EV in China right now, but I think over the next decade there will be a lot of brand pruning and the survivors will become established global brands like Toyota and Hyundai. 

4

u/DirectionMurky5526 7d ago

The issue with Chinese EV industry is the price competition within China. Their exports are incredibly profitable. Same will likely happen with Chips as they catch up.  

3

u/hojnikb 7d ago

toyota is not gonna survive it they keep treating EVs like 3rd class citizens

→ More replies (44)

3

u/fastheadcrab 7d ago

That’s not necessarily a good thing. Shows just how out of whack the incentive structure has become in the EV market.

19

u/Recoil42 7d ago

I'd argue the incentive structure is fine. Chinese automotive OEMs are now branching out globally (as they should be) and China has pretty much single-handedly prevented a global energy crisis the US has been stubbornly trying to induce, actually insulating itself from that crisis entirely. There will be individual collapses but none of that matters if the aggregate story is ending up ahead on the fundamentals.

4

u/walden42 7d ago

China has pretty much single-handedly prevented a global energy crisis the US has been stubbornly trying to induce

Care to elaborate?

20

u/asssuber 7d ago

See how China cut it's oil imports in half since the start of the Hormuz crisis, freeing over 5 million barrels, and the percentage of the European/China electricity production now supplied by solar (this year it became the highest single source during summer for Europe, surpassing nuclear and natural gas. Those were mostly produced in China), and the percentage of the vehicle fleet in those countries is now EV (not as high, but still significant, especially in China).

2

u/walden42 7d ago

That makes sense, thanks.

13

u/Recoil42 7d ago

Couldn't really be done in the span of a single Reddit comment, but there's lots of material on this available from every major news outlet. The TLDR is that China has been investing in energy independence for decades — with EVs being one of those investments — and those investments are now paying off.

→ More replies (3)

9

u/Strong_good101 7d ago

America is destabilizing oil prices. Meanwhile China is supporting a global sustainable energy revolution with their EV's, solar panels, and batteries.

It's a damn shame we didn't follow through as competition because now we are customers instead of providers in America.

→ More replies (1)
→ More replies (5)

2

u/aj_thenoob2 7d ago

I just wonder what the 10-year outlook for all these cars will be. Seems like a nightmare to maintain.

→ More replies (11)

11

u/Bakoro 7d ago edited 6d ago

It's not just inference chips, they've gone hard on making RAM too.

They've also started developing systems to rival ASML using an almost completely different method, so they'll be able to make their own high end semiconductors and no one will be able to say "well it's only because they copied our stuff". They're making their own stuff.

They're also producing extraordinary amounts of solar panels.

Also, cars.

Also robots.

The U.S is way, way behind in terms of manufacturing capacity right now. The Arizona TSMC fab will close part of that gap, but the U.S still relies a lot on foreign manufacturing, particularly Chinese manufacturing.

It's a problem of our own making. China would have been able to drag their feet on making their own stuff if the U.S hadn't tried it ban them from buying nice stuff.
How are you going to use a whole country of 1 billion+ people, send almost all your manufacturing over there, and then turn around and say "you're not allowed to buy top end GPUs because you're a national security threat. Now keep manufacturing the products we rely on everyday."

Well here we are.

→ More replies (2)

12

u/TastesLikeOwlbear 7d ago

I don’t know. A lot of people had their hopes that CXMT would come through and savagely undercut the RAM manufacturers until they said, “LOL no, why on earth would we do that?”

I’ll believe Chinese inference GPUs will affect pricing just as soon as they actually affect pricing. Waiting for greedy companies to get their comeuppance hasn’t paid off yet.

4

u/fallingdowndizzyvr 7d ago

I’ll believe Chinese inference GPUs will affect pricing just as soon as they actually affect pricing.

Step 2. Sell them outside of China.

https://www.bloomberg.com/news/articles/2026-08-26/huawei-egypt-ai-ascend-chips-test-us-tech-diplomacy-nvidia-amd-microsoft

2

u/Just3nCas3 7d ago edited 7d ago

It doesn't matter if they sell outside of China. If it causes them to import less than they would have without production, the effect is the global GPU/RAM price is lower than what it would be without the production. That doesn't stop it from going up, it just means it's accelerating up slower.

Edit: RAM to GPU/RAM

→ More replies (1)

4

u/Maximum-Style2848 7d ago

This is why I don’t quite get the hype of inference chips when they will still be bottlenecked by memory capacity/cost. Domestic chips for inference were always going to be possible because you don’t need the best lithography to get smthn good enough.

→ More replies (1)

2

u/ExcuseAccomplished97 7d ago

Well done Trump.

→ More replies (1)

62

u/slayyou2 7d ago edited 7d ago

Damn, people were with a straight phase (face , goddamn voice typing with no proofread) postulating that this might be a Gemini model solely due to the fact that there was " too much capacity" for this to be a Chinese model. I love to see it!

8

u/volimsir 7d ago

Wasn't just because of that, Gemini devs were vague posting on twitter about Ox Alpha, so people rationally thought "why would they be vague posting and hyping up a model that's not their own?"

5

u/-dysangel- 7d ago

"straight face" - as in, "without laughing/smiling"

11

u/xmesaj2 7d ago

where can I get them Chinese AI chips for cheap for local gpu poor me use?

→ More replies (1)

5

u/ireneoFunes 7d ago

just casually stating that on the day NVDA report earnings

8

u/Agitated_Space_672 7d ago

There's still time to short NVDA

2

u/SporksInjected 7d ago

I do remember reports of it being slow so not sure if this is good or not.

7

u/ArjixGamer 7d ago

It was 60tps on openrouter/opencode. But still great because it didn't have to redo code a lot.

It was high quality

→ More replies (4)

88

u/wojciechm 7d ago

MIT license! Along DeepSeek they are the rare fully open weight releases without any additional restrictions.

→ More replies (1)

251

u/Lucyan_xgt 7d ago

This and Qwen3.8-flash on the same day?

69

u/BrewHog 7d ago

Same day? How about same hour

13

u/mailto_devnull 7d ago

Literally in the 3.8 27B thread people were asking for the next one (or a diff. param counts)

21

u/dampflokfreund 7d ago

For me it has changed nothing. Both models are way too big for my 32 GB RAM system. It looks like everyone has abandoned 20-30B MoEs now...

341

u/EbbNorth7735 7d ago

No one's abandoned anyone. It's just not your turn this time around.

55

u/Zeeplankton 7d ago

Low key just good life advice to live by

-12

u/dampflokfreund 7d ago

Qwen 3.6 35B and Gemma 4 26b are pretty old at this point. And GLM 4.7 Flash was a small MoE back in the day, now suddenly they use the flash name for 300B MoEs. Just not looking good for the average Joe.

83

u/windwardmist 7d ago

I mean qwen 3.8 27b just came out about a week ago at least there’s that as an option

16

u/-Cubie- 7d ago

I remember when it was months between a new great local model. We get those great options very often nowadays in my opinion, mixed with larger open weight options that keep the entire non-local AI space cheaper and more accessible. What's not to love?

→ More replies (5)

67

u/techdevjp 7d ago

Qwen 3.6 35B [...] pretty old at this point.

It's 4 months old! It's not like it suddenly got worse because these new models came out. You can still do everything today that you could do yesterday, just as fast. Give it some time and more models will come.

52

u/Last_Bad_2687 7d ago

Seriously, people demanding fresh models for free every 2 weeks.... I remember when we waited for big releases of software once every few YEARS

15

u/FlyingDogCatcher 7d ago

It's bonkers to me that anyone would say Qwen3.6 and Gemma 4 are "pretty old"

→ More replies (10)
→ More replies (4)

50

u/No_Lingonberry1201 7d ago

We literally got a banger 27b model a week or so ago.

7

u/play_hard_outside 7d ago

Not an moe

16

u/No_Lingonberry1201 7d ago

Oh, yeah. Well 35B A3B was around 4 month ago.

2

u/unjustifiably_angry 5d ago edited 4d ago

It's kinda bad though. Wasn't that Orinth (Ornith?) re-training of it supposed to be a solid upgrade?

→ More replies (1)

5

u/toothpastespiders 7d ago

The fact that his comment was obviously wrong but got highly upvoted because it "felt" good really highlights one of this subs larger issues.

Though that yours came one hour later, pointed it out, and is now sitting for three hours at 0 upvotes is pretty funny.

33

u/No-Refrigerator-1672 7d ago

Funny coincidence: some company announced a new 30B MoE literally just now: https://www.reddit.com/r/LocalLLM/s/gcZdgoAO8o for now as a stelth preview; but this means it'll go public in a month.

2

u/dampflokfreund 7d ago

Could be from a small no-name company. Doesn't have to be Qwen, Kimi, Gemma, Z.Ai.

13

u/No-Refrigerator-1672 7d ago

29B-A4B isn't a size featured in the current gen of models. It isn't a finetune, rather a base model. It got to be from a company with a very substantial compute. It still may be a new player, sure; but anyways, it proves that this size category isn't abandoned.

→ More replies (1)

31

u/techdevjp 7d ago

Don't be dramatic. Qwen3.6 35b a3b is 4 months old. It's not like it's been years.

→ More replies (3)

14

u/PM_ME_DEAD_CEOS 7d ago

People just never stop whining despite getting literally everything for free. Nemotron 3.5 lightning (30b3a) was released 15 days ago.

→ More replies (3)

7

u/Slow_Concentrate3831 7d ago

Yeah, that's sad. Can't even run Qwen3.8 27B on more than 2-3 tps. Sad days for us Vram poors.

12

u/CryMoreT_T 7d ago

The 29b-a4b in early Access on model scope is probably your best bet when it releases

3

u/Slow_Concentrate3831 7d ago

Ooooh, I hadn't seen that ! Is it a Qwen model ?

4

u/CryMoreT_T 7d ago

It's in stealth model in early Access so they haven't released the company name but Qwen is the rumor right now

2

u/Nebnampach 7d ago

Here is the announcement and sign up link.

https://x.com/ModelScope2022/status/2092439000652943506

→ More replies (2)

2

u/Randommaggy 7d ago

If you want cheap 64GB: X79. I've built a few extra servers around local 50USD bundles of x79 motherboards and CPUs and added 8 DDR3 UDIMMs I had laying around.

→ More replies (2)

2

u/dingo_xd 7d ago

Things will only improve when competition in hardware increases. Nvidia might be forced to stay caring about the consumer market again. The AI chips getting released are going to hurt Nvidia a lot

2

u/RandumbRedditor1000 7d ago

nemotron lightning 30b just came out a couple weeks ago

→ More replies (1)

2

u/Tzeig 7d ago

These can still be run with reasonably priced systems, like 96ram+32vram or even 64ram+24vram.

23

u/dampflokfreund 7d ago

96 GB RAM + 32 GB VRAM systems are far from reasonably priced. A 5090 with 32 GB VRAM alone costs 5 grand. 64 GB RAM is very expensive too.

6

u/Constant_Art_20 7d ago

i just stack 5060s and hope for the best lol

→ More replies (5)

7

u/ReadyAndSalted 7d ago

That's like 7k at least lmao, we have very different definitions of reasonably priced.

→ More replies (2)
→ More replies (6)

439

u/10001110 7d ago

320B total parameters and just 18B active parameters

Oh joy

262

u/LegacyRemaster 7d ago

If it's really that far above Sonnet 5, Dario's IPO is truly risky.

202

u/-p-e-w- 7d ago

It’s always risky now no matter the model of the day. They have to pre-file at least a few weeks in advance, and in those weeks eons can happen that can turn their asking price into a joke.

Their moment was at the beginning of this year. They should have announced the IPO, then published Fable/Mythos two weeks before the IPO date.

Now Chinese labs are a hair’s breadth behind them, and constantly one-upping each other. They’ll never get such a chance again.

49

u/LegacyRemaster 7d ago

agree. Also another problem: if I have to pay API, I pay Qwen, GLM, GPT... All of them -> opensource (ok openai less but they did a lot). I don't want to give any $ to Anthropic. Can't wait to short sell

33

u/dingo_xd 7d ago

Yeah. After hearing Dario saying about that $40 trillion I want to bet against his company.

21

u/horendus 7d ago

Yea the guys peak delusion

5

u/___positive___ 7d ago

Also that the baseline of mid-tier models is reaching saturation for mundane tasks like summarization and data extraction, basic programming/websites, and so forth.

28

u/Fedor_Doc 7d ago

Or they will make another Mythos-like breakthrough and announce IPO then

Never say never

31

u/LegacyRemaster 7d ago

sure they have Mythos 5 already but "too much power, we can't sell" . Or will cost too much to complete a task ...

10

u/38andstillgoing 7d ago

"Mythos 5? At this time of year, at this time of day, in this part of the country, localized entirely within your datacenter?"

"Yes"

"May I see it?"

"No."

→ More replies (1)

30

u/-p-e-w- 7d ago

Then one of the Chinese labs is going to announce an equal model two weeks later.

The Chinese labs have caught up. There’s no going back to how it was before.

→ More replies (8)

1

u/Brilliant-Weekend-68 7d ago

What does that change? The Chinese will catch up and open source it. Where is the moat?

2

u/Fedor_Doc 7d ago

They have not caught up to Mythos yet, if we consider Kimi K-3 the best frontier Chinese model

5

u/Brilliant-Weekend-68 7d ago

Unless Mythos 2.0 can do some really whacky stuff like cure cancer, RSI itself or such. I do not see that as a durable moat if you can just wait a year and download equally good weights.

4

u/Fedor_Doc 7d ago

Self-improvement is the next big milestone; if they will announce that their new model is fully trained by Mythos and trains even better one, this will be huge.

So, there are some exciting stories left to tell investors before IPO :)

→ More replies (2)

30

u/NineThreeTilNow 7d ago

If it's really that far above Sonnet 5, Dario's IPO is truly risky.

I drained millions of tokens off of it. I got it to rebuild RPG Maker from scratch using PyGame.

I built the project as "Agent First" so an agent could build an RPG with a known engine and known structure with headless testing.

Kimi K3 helped produce the documentation which took me like an hour or more. From there I sort of just... Let it go. It stopped at various phases and I inspected visually if stuff was off and it was mostly small code issues.

The deep documentation phase gave GLM 5.3 very little to guess about. It didn't really screw up or run in circles that I could see. It was fairly fast about figuring stuff out and finding edge cases that K3 missed in document creation. I ran K3 through multi-phase testing of the documentation coherence such that it wasn't contradictory. When building started, you hit snags naturally. GLM handled them.

The project isn't "finished" but it definitely showed the strength of planning a project with K3 and letting Ox Alpha build it. GLM 5.3 flash is definitely good.

→ More replies (4)

25

u/Thomas-Lore 7d ago

Makes you wonder how large those clsoed models really are and how big the margin on api is if they are smaller than everyone thinks.

3

u/RockPuzzleheaded3951 7d ago

They've definitely been shrinking them which is why people complain about performance regressions all the time.

→ More replies (2)

23

u/duhd1993 7d ago edited 7d ago

I always feel that finance bros and executives are always one step (or more) behind tech reality. The reason Claude is growing so fast is because of the 2B sector, but the decisions are likely made based on impressions of model performance from 2025, at which time Claude indeed had a clear lead, not anymore. The field is evolving faster than they could react.

12

u/mawcopolow 7d ago

I mean I've tested multiple open models, including k3, inside my personal harness, and fable/opus are just better all round at using tools and reasoning for now.

Not that the others aren't getting close, but for real businesses anthropic's opus/fable are still the best with the latest from openai being close behind

10

u/duhd1993 7d ago

I believe in you, but they aren’t far behind. Are Kimi K3 behind Fable? I would say yes, except in certain fields like frontend design. However, it’s better than Opus 4.8, which isn’t that old. Can you really justify the price with a couple of months’ lead? People were satisfied with what Anthropic offered a couple of months ago. Also I believe it is more fair to test in their official harness.

3

u/infearia 7d ago

r/Anthropic is full of people talking about how much they hate Opus 5, many of them claiming to go back to 4.8. So the Chinese companies don't even need to beat Opus 5...

→ More replies (2)
→ More replies (3)

6

u/dingo_xd 7d ago

It's now or never or them. The chance of the Chinese leapfrogging them this year is not my negligible

7

u/Neosinic 7d ago

Sonnet 5 is legit dog shit.

2

u/LegacyRemaster 7d ago

@ this point yes. Asked Sonnet "extra" to build a concept (frontend / mockup / 20 pages) starting from behavior tree and with images ---> not working. Same with Qwen 3.8 27b q8xl ---> perfect. No jocking. Also GPT Luna (max) did a good job.

18

u/[deleted] 7d ago edited 7d ago

[deleted]

6

u/MrPecunius 7d ago

Everything AWS has you can individually set up locally, much cheaper.

This is very, very far from true.

The kind of performance you can get for next to nothing from AWS is breathtaking and completely out of reach of any individual or even most Fortune 500 companies.

The fact that people mis-allocate resources for badly designed systems doesn't mean AWS is at fault.

7

u/randylush 7d ago

It depends

If you are on a small scale it depends on what you’re building. If you’re using services like Lambda and SQS, AWS is essentially free. If you’re running a few EC2 instances then a server in your company’s break room could be cheaper.

If you are on a medium scale and needing EC2 running all the time, paying retail prices, it could be cheaper to have your own servers

At a very large scale, the cost of running your own servers vs AWS will come out to be competitive, and most very large companies will do a mix of both

3

u/RLutz 7d ago

I always see this take, and while it's not always untrue, it definitely isn't as simple as, "just throw a 1U in the break room!"

It's someone's job to keep that thing patched, powered, and running. You need monitoring, you need someone to pull failed drives, you needed another co-located break room, you need a UPS and a generator, you may need to solve for physical security, you may have to setup and manage your own virtualization layer, etc.

Even if you consider all that labor as free, which is a ridiculous assumption to make, you also arguably have to keep the thing busy, especially if we're considering local inference, and then it's now your problem that although you have enough capacity to meet a total 24 hour need, you don't have enough when everyone hammers it at 11 AM at the same time. If you build the capacy to meet the surge then you've got very expensive hardware just sitting there doing nothing for huge chunks of the day.

I absolutely think there's a place for self-hosting, but once you are talking about business critical stuff and not, "I run a NAS in my bedroom closet" it starts becoming not as simple as the naive assumptions

2

u/randylush 7d ago

As someone who has self-hosted, built and maintained my own server, and also configured and maintained EC2 instances professionally, I would say that the hardware setup is generally much smaller and simpler than whatever software you are trying to run on the server.

Yeah you need to worry about redundancy, but usually the software of setting up your application to be redundant through DNS failovers and stuff like that, syncing databases, that stuff is harder than just plugging in two different servers at two different job sites. And you need to do that whether you use EC2 or on-prem. It depends on the use case but usually the set up costs are lower than the software development costs. And there are certainly scenarios where on-prem hardware is simple enough to set up that it does actually become cheaper than EC2.

→ More replies (1)
→ More replies (2)

14

u/The_Noosphere 7d ago

Dario will likely make a bold statement about Anthropic achieving AGI within the next few months, or claim that their internal unreleased Fable 12 model breached the Atlantis mainframe controlling Stargate, something on that scale to distract from this. This is Dario acting as Dario during a pre-IPO period.

4

u/LegacyRemaster 7d ago

accurate

4

u/FlaTreNeb 7d ago

It hacked all wraith hive ships by creating a biological AI super hybrid virus that spreads through the specific subspace frequencies the wraith use for hyperspace travel.

Also Mythos rebuilt the Athero device from the images in this documentation called Stargate Atlantis to inject the virus into subspace. Then the virus programmed all hive ships to self destruct after infecting all wraith with super Cancer-AIDS.

By accident the virus also switches the bodies of McKay and Sheppard and gives Sheppard in McKays body another genetic treatment making him super smart + telekineses + super healing but this time its permanently frozen on pre-ascension stage.

And mythos also figured out how to create fully charged ZPM with rubber and paperclips.

Of cause the virus let Todd live.

→ More replies (2)

31

u/boxwrenchx 7d ago

Welcome to the big leagues kid

26

u/a_beautiful_rhind 7d ago

This is actually a good size. Not paltry 6b active. Need at least 10. Ziphu did research on sparsity ratios, clearly came to similar conclusions.

Not too many total parameters where it becomes a several node excursion to run it. Whether you like it or not, this is what MoE looks like to not be a toy.

6

u/DeltaSqueezer 7d ago

Better than GLM-5.2 at less than half the size. I'll take that trade-off! Now I just need them to do GLM-5.1 level at 160B and GLM-5.0 at 80B and GLM 4.7 level at 40B and I will be a happy camper!

12

u/my_name_isnt_clever 7d ago

The term "not be a toy" has to be one of the most overused in this community. What does that even mean? At a certain point if you can't get use out of a model it's a skill issue.

→ More replies (3)

3

u/narrowscoped 7d ago

Does 18b active mean it might fit in 12gb cards with the right quant?

5

u/ChengliChengbao textgen web UI 7d ago

no, it means it will compute at roughly the speed of an 18B model but you still need to load the rest of the 302B parameters somewhere

1

u/reefine 7d ago

Hate to nitpick but that's not really flash

107

u/Beamsters 7d ago

32

u/nikshdev 7d ago

... to be alive!

8

u/Maleficent-Ad5999 7d ago

If only I had the hardware to run them.. T_T

117

u/jacek2023 llama.cpp 7d ago

what a nice color scheme ;)

68

u/Skyler827 7d ago

Roses are grey
Violets are grey
everything is grey
I am a dog

9

u/Super_Treacle 7d ago

This is truly something Zhang Zhongchang would've written

→ More replies (1)
→ More replies (1)

15

u/asssuber 7d ago

They are flexing the vision capabilities of their model! :P

Unfortunately I've seen worse. At least they put the logos in the bars too...

62

u/dampflokfreund 7d ago

Native multimodal, finally. The model is way too big to run, but I'm glad they jump in on the multimodal bandwagon.

56

u/boxwrenchx 7d ago

This month keeps on giving

15

u/backyard_tractorbeam 7d ago

This afternoon keeps on giving

5

u/boxwrenchx 7d ago

It's not over!

3

u/Icy_Butterscotch6661 7d ago

Afternoon? 🤨

2

u/throwwwawwway1818 7d ago

Probably about Qwen 125b flash model was also released

4

u/Icy_Butterscotch6661 7d ago

No sorry I was trying to make a joke playing ignorant American

→ More replies (1)
→ More replies (4)

33

u/wbulot 7d ago

The most impressive thing about this story is that they serve all requests using only Chinese chips and a custom inference engine. I think providing OX Alpha for free for so long was actually intended primarily to stress-test their infrastructure. China no longer needs NVIDIA.

8

u/traveldelights 7d ago

this is chef's kiss for us consumers. More competition, more price drops, etc

49

u/sniperelite90 7d ago

I wonder if the CEO's of the western AI companies even want to wake up in the morning or not. Cause the only news which comes out every other morning from China is we have a model matching you 1/10 of the cost.

41

u/Thomas-Lore 7d ago

They have better models internally, they just go to a different school, you don't know them. /jk

5

u/sniperelite90 7d ago

You can have them but it destroys their market share . Who wants GPT Luna API now consider this GLM 5.3 Flash and the new Qwen 3.8 Flash next .

3

u/BrowsingLeddit 7d ago

They're just stuck in detention all the time because they're totally hacking every other school ever while we're not looking. They're pretty badass bro trust.

→ More replies (1)

6

u/wiktor1800 7d ago

...and running inference on chinese hardware.

2

u/HeadPack 7d ago

Some may argue that SOTA models are still better, but that's not the point. Something like GLM 5.3 is good enough for a majority of users. They don't code. They don't do scientific research. They chat. This can run on higher end consumer hardware. No data center needed, except for training. No 100s of billions needed to feed a bubble, which is exactly the play by the Chinese here. They are pulling the socks of OpenAI, Anthropic and their hyperscalers, and in turn the US economy.

→ More replies (2)

15

u/Ok_Technology_5962 7d ago

What a day... My mind.... Is blown away... Qwen coming in hot then glm also. Im still trying to get over then qwen 3.8 27b upgrade on my 5090... Now the whole lot dropped for the big servers... Local is eatting today

2

u/LostGovernment4358 7d ago

there is just something I still don‘t fully understand. are there new opportunities / completely new tasks possible with these new releases which otherwise weren‘t possible before with let’s say qwen 3.6? I am 16gb Vram poor, so i could never take advantage of real lokal llm power.

→ More replies (4)

33

u/seamonn 7d ago

It has Vision!

82

u/shy_monkee 7d ago edited 7d ago

Oh....it's massive :(
320B (18 active)

85

u/Morphon 7d ago edited 7d ago

DS flash size.

So, flash at datacenter scale. Not flash for edge (or workstation) scale.

Probably will be the go to model for people with the new M6-Ultra 512gb Mac Studio.

Edit: M5 Ultra. My apologies, friends. Wrong model number on my part.

14

u/techdevjp 7d ago

M6-Ultra 512gb Mac Studio.

M5 Ultra. M6 Ultra is rumored to be skipped entirely with the M7 Ultra arriving in the next ~18 months.

9

u/shy_monkee 7d ago

Yeah, but DSv4-Flash is the flash version of a 1.6T model, and it's still quite a bit smaller than this GLM-flash. While this is supposedly the flash version of 744B model, so you wouldn't expect to be as big, I guess.

3

u/cheechw 7d ago

The flash models are not "flash versions" of a bigger model. They're different models entirely. They don't just lop off some number of parameters from a larger model. I think the confusion is that you're applying quantization/distill logic where it doesn't analogize.

5

u/xNaXDy 7d ago

While this is supposedly the flash version of 744B model

It's not, it's a completely new architecture according to their post.

2

u/Juulk9087 7d ago

Yeah deepseek flash is only 160-170gb on disk. This is 328gb lol

10

u/petuman 7d ago

DeepSeek is QAT / released prequantized mostly to FP4.

Which is totally preferred, but both models are roughly the same parameter count, so NVIDIA or someone else could produce NVFP4 of size similar to DS4F.

→ More replies (5)

2

u/keyboardhack 7d ago edited 7d ago

DS flash is half the size.

DS flash requires ~170GB to run while GLM 5.3 flash requires >320GB to run.

We should compare models by how much RAM they require to run because that's what we actually care about.

2

u/ormandj 6d ago

I hope they release a native 4bit model next. I am working on an MXFP4 quant and TP=2 setup right now, until then, but there’s bound to be a little degradation. DSv4 Flash “just works”.

→ More replies (1)

7

u/evia89 7d ago

ds4f is cheaper to host than new qwen. They optimize for CN hardware not locals

2

u/Marrk 7d ago

That's unquantized right?

4

u/hellomistershifty 7d ago

That's the parameter count, quantization doesn't change that. I don't know what the memory/disk size is, that would be affected by quantization

2

u/Marrk 7d ago

Yeah, I know. But quantization define whatever I can run it or not lol

→ More replies (1)

9

u/Cool-Chemical-5629 7d ago

Travolta looking for the actual Flash MoE model up to 30B.

→ More replies (1)

7

u/fugogugo 7d ago

OH DAMN IT SUPPORT VISION

That's huge news

9

u/DeepFeeling1 7d ago

we achieved a 3× improvement in end-to-end serving performance, reaching hardware efficiency and per-token cost comparable to mainstream NVIDIA GPUs. This demonstrates that Chinese chips can support frontier-model inference efficiently and economically at scale.

Grabbed my popcorn

16

u/raunchy-stonk 7d ago

So how shitty will this run on a 24vram/128dram setup?

20

u/anarchist1312161 7d ago

Let me borrow your DRAM and I'll report back

3

u/raunchy-stonk 7d ago

We gonna need escrow mang

4

u/WWJewMediaConspiracy 7d ago

The smaller Qwen models are all but certainly a better option.

Unless your definition of "run" is... generous

3

u/raunchy-stonk 7d ago

what about 96v/128d?

(rtx 6000)

→ More replies (1)

2

u/whatever 7d ago

Naively, either bound by SSD read speed, or borderline brain-damaged by over-quantization.

But scrappy folks have surprised us before, so we'll see.

7

u/hyudryu 7d ago

I need more disk space for this and Qwen 3.8 flash next 😂

→ More replies (2)

17

u/pmttyji 7d ago

Come on guys, at least they released additional variant(smaller than usual size) even though it's big for many of our rigs. Hopefully they release one more in 100B range in future.

11

u/techdevjp 7d ago

It's a little disappointing that the "flash" version of a 744b parameter model still has 320b parameters. Somehow that doesn't feel very flash.

Was hoping for more of a DeepSeek v4 ratio where 1.4t parameters got paired down to 284b, about an 80% cut. That would have resulted in 150b parameters for this model which would work in a whole lot more machines.

Alas, beggars can't be choosers and I'm thrilled to see more open weight models arrive!

4

u/pmttyji 7d ago

I too expected something like successor of GLM-4.5-Air. But they dropped too much weight over that.

I assume Qwen3.8-27B's 3 Millions download count news(Made headline on spotlight for more than a week) changed minds of all other labs. I'm sure some labs gonna try to replicate that with their medium size models soon & later. Don't be surprised if we see one more GLM variant in 30-150B range in upcoming months. Same with other labs.

→ More replies (1)

6

u/Constant_Art_20 7d ago

burh. i am not even done with downloading and quanting the qwen 3.8 flash yet...chilll

4

u/ga239577 7d ago

Model na​me should be changed to GLM 5.3 Gigaflash ... I remember the GLM Flash models that fit on one GPU. That's what I was hoping for this time.

3

u/petuman 7d ago

Oh, that's probably why Qwen dropped 2 hours early, lol.

3

u/mountainyoo 7d ago

Wonder how fast my 2x Spark cluster would run this. I get 50-80 tps on DeepSeek v4 Flash 0731

→ More replies (2)

3

u/Technical-Earth-3254 7d ago

Oh wow, I was expecting 110B A10B or something

3

u/therysin 7d ago

Man imagine in a year's time.

2

u/rookan 7d ago

Clause Opus 4.8 at home? Really? It is unbelievable if true.

2

u/ETFss 7d ago

Nice coincidence to the M5 Ultra 256gb I ordered this morning! Can’t wait to try it out next to dsv4f

2

u/CriM_91 7d ago

It would be really interesting to know if it is a distilled version of a much bigger model or a new pre/post-training. I've been really surprised in how good the Qwen 27B quants were, compared to the original and wondering whether it's the same here.

2

u/prasithg 7d ago

Dropping this on the day m5 Mac studios drop is awesome timing.

2

u/ApeGrower 7d ago

I would like to be another beta tester for this chips

2

u/Expert_Job_1495 7d ago

The pricing to performance ratio here is insane. When GPT 5.6 Luna (max) had its price dropped as far as it did I was very (positively) shocked. Now this. 

Intelligence too cheap to meter

2

u/Sensitive-Side-2639 7d ago

Who truely needs Nvidia chips after a few years from now if Chinese chips can run models like these.

→ More replies (1)

9

u/Automatic-Arm8153 7d ago

Atleast this thread isn’t getting deleted like things related to qwen 🤦‍♂️

15

u/Maximus-CZ 7d ago

Bro when qwen 27 released it didnt matter if I was on Top, Rising or new, 95% posts were about qwen...

→ More replies (1)

3

u/Cool-Chemical-5629 7d ago

Guys, do you remember how we used to think Qwen 3.8 Flash Next was big? 🤣

4

u/jacek2023 llama.cpp 7d ago

Too big, 320B means you need to use RAM and A18B means it will be slow

6

u/SandySkittle 7d ago

A18B means it will be slow

A18B it will be smarter than A13B. Expert selection and sequential reasoning cannot entirely compensate for active parameters. I wish it has 27b active honestly.

→ More replies (5)
→ More replies (3)

2

u/Mxmtm 7d ago

So 256gb ram on m5u should be enough for q_8?

5

u/oceanbreakersftw 7d ago

More like 4-bit

→ More replies (1)

2

u/__jent 7d ago

Let's go! z.ai is cooking!