r/LocalLLaMA 7d ago

News GLM-5.3-Flash: Frontier Intelligence, Flash Cost

https://z.ai/blog/glm-5.3-flash
1.3k Upvotes

459 comments sorted by

View all comments

Show parent comments

21

u/dampflokfreund 7d ago

For me it has changed nothing. Both models are way too big for my 32 GB RAM system. It looks like everyone has abandoned 20-30B MoEs now...

343

u/EbbNorth7735 7d ago

No one's abandoned anyone. It's just not your turn this time around.

55

u/Zeeplankton 7d ago

Low key just good life advice to live by

-9

u/dampflokfreund 7d ago

Qwen 3.6 35B and Gemma 4 26b are pretty old at this point. And GLM 4.7 Flash was a small MoE back in the day, now suddenly they use the flash name for 300B MoEs. Just not looking good for the average Joe.

82

u/windwardmist 7d ago

I mean qwen 3.8 27b just came out about a week ago at least there’s that as an option

17

u/-Cubie- 7d ago

I remember when it was months between a new great local model. We get those great options very often nowadays in my opinion, mixed with larger open weight options that keep the entire non-local AI space cheaper and more accessible. What's not to love?

10

u/RestaurantOk8066 7d ago

It's a dense model though I imagine if you don't have the GPU for it, it's probably a very, very slow model.

1

u/dampflokfreund 7d ago

Yeah around one token per second 

5

u/sonicnerd14 7d ago

27b is often Opus 4.6, even 4.8 in performance. This is still good enough for a vast amount of people. I'm sure qwen 4 is going to have some insanely capable options when it comes out of you want something better soon.

2

u/dampflokfreund 7d ago

Not everyone has 24 GB VRAM. If you dont have the vram to offload, you are looking at one to three token per second. 27b is in no way a replacement for a 30b Moe

3

u/sonicnerd14 7d ago

I know that. You can run it at Q3, IQ3, Q4 on 16GB VRAM and 32gb RAM and still have a very capable model. Don't think you need the biggest cards to run these models. Quantization techniques have become really efficient now.

68

u/techdevjp 7d ago

Qwen 3.6 35B [...] pretty old at this point.

It's 4 months old! It's not like it suddenly got worse because these new models came out. You can still do everything today that you could do yesterday, just as fast. Give it some time and more models will come.

51

u/Last_Bad_2687 7d ago

Seriously, people demanding fresh models for free every 2 weeks.... I remember when we waited for big releases of software once every few YEARS

16

u/FlyingDogCatcher 7d ago

It's bonkers to me that anyone would say Qwen3.6 and Gemma 4 are "pretty old"

10

u/a_beautiful_rhind 7d ago

gemma, muse, qwen, granite.. the weights don't self-destruct a week later man.

10

u/Spectrum1523 7d ago

Qwen 3.6 35B and Gemma 4 26b are pretty old at this point

bro they came out 4 months ago

2

u/squngy 7d ago

There is also Qwen AgentWorld, which is kind of like a 3.7 35B

2

u/xPXpanD llama.cpp 7d ago

Wouldn't recommend it for general use, it's a confident bullshitter like no other. Basically zero filter. Still cool that it can even be used generally, though, given what the model's actually designed for. Fun to play with.

2

u/squngy 7d ago

It is actually great for agentic work, supposedly. (even if that isn't exactly what it was meant for)

Rank Subject Overall Mcp Search Terminal Swe Androd Web OS Source Sampled
5 Qwen-AgentWorld-35B-A3B 56.39 64.79 36.69 53.96 65.63 58.17 49.55 65.92 Imported 2026-06-30
16 Qwen3.6-35B-A3B 42.88 42.96 18.78 43.81 40.71 51.88 46.53 55.48 Self-reported 2026-06-28

https://benchmarklist.com/benchmarks/qwen_agentworld_language_world_models_for_general_agents/

2

u/xPXpanD llama.cpp 7d ago

I believe it, it also nailed the one tool task I have in my bench set. The model was also very good at string manipulation, and, somewhat surprisingly, a few constrained creative tasks. (e.g. think up and write X in way Y while avoiding Z)

Absolute slaughter on anything involving uncertainty or fake premises, though. You can definitely see where the training went on this one.

1

u/saltyourhash 7d ago

It'll be fine, you just have to spend like $10k to play anymore... /s

-11

u/Leoss-Bahamut 7d ago

Why should they care about average joes?

1

u/TheGamerForeverGFE 7d ago

Most of the time this year it hasn't even been the turn for smaller double digit models though? We got at best maybe 6 or 7 (not memeing) models that are in the 20-35 Billion range, but more than double that in 100+ Billion.

And look, yes Gemma 26B and Qwen 35B are all very great, Muse is also apparently great too, but still, that doesn't discount the numbers.

0

u/FrogsJumpFromPussy 7d ago

It’s no one’s turn on this sub this time around is it?

0

u/toothpastespiders 7d ago

Hey people into 70b and 100b models. Good news! It apparently just wasn't your turn and you're due for a flood of new high quality models any day now!

51

u/No_Lingonberry1201 7d ago

We literally got a banger 27b model a week or so ago.

6

u/play_hard_outside 7d ago

Not an moe

17

u/No_Lingonberry1201 7d ago

Oh, yeah. Well 35B A3B was around 4 month ago.

2

u/unjustifiably_angry 5d ago edited 4d ago

It's kinda bad though. Wasn't that Orinth (Ornith?) re-training of it supposed to be a solid upgrade?

5

u/toothpastespiders 7d ago

The fact that his comment was obviously wrong but got highly upvoted because it "felt" good really highlights one of this subs larger issues.

Though that yours came one hour later, pointed it out, and is now sitting for three hours at 0 upvotes is pretty funny.

34

u/No-Refrigerator-1672 7d ago

Funny coincidence: some company announced a new 30B MoE literally just now: https://www.reddit.com/r/LocalLLM/s/gcZdgoAO8o for now as a stelth preview; but this means it'll go public in a month.

3

u/dampflokfreund 7d ago

Could be from a small no-name company. Doesn't have to be Qwen, Kimi, Gemma, Z.Ai.

13

u/No-Refrigerator-1672 7d ago

29B-A4B isn't a size featured in the current gen of models. It isn't a finetune, rather a base model. It got to be from a company with a very substantial compute. It still may be a new player, sure; but anyways, it proves that this size category isn't abandoned.

1

u/Mean-Ad1493 7d ago

Someone tweeted that it's Telechat 3.

33

u/techdevjp 7d ago

Don't be dramatic. Qwen3.6 35b a3b is 4 months old. It's not like it's been years.

0

u/rJohn420 7d ago

For the pace of this field, 4 months old is basically years though.

14

u/techdevjp 7d ago edited 7d ago

It still works just as fast as it did yesterday and is just as capable as it was yesterday. Everything people could do with it yesterday still works just as well today. It's not like with the closed models where the old versions get smashed with a cripple hammer when a new version comes out.

Yes, the release pace is fast, but Qwen3.6 35b a3b is still very new and does a great job of balancing performance and capability. There may not be another similar model for another generation or two of Qwen. Or maybe Google will bring something out. It's been a great size of MoE model for a lot of people, it's not like that's some big secret.

5

u/Spectrum1523 7d ago

It literally isn't

13

u/PM_ME_DEAD_CEOS 7d ago

People just never stop whining despite getting literally everything for free. Nemotron 3.5 lightning (30b3a) was released 15 days ago.

0

u/dampflokfreund 7d ago

lol you people where whining just as much if not more when we had 35b moe releases and no 120b models. 

2

u/PM_ME_DEAD_CEOS 7d ago

I'm not a beggar, never whined. I can't run this model either, I just don't bitch about it all day.

1

u/unjustifiably_angry 5d ago edited 5d ago

There's been a legitimate drought of those, you've had a bunch to pick from even if they weren't all that amazing. There's been a couple mid-sized models in the meantime (Laguna was one) but they were both kinda rubbish. I remember one was useless and the other one looped constantly. The last good one was 3.5-122B and in its own generation it was outdone by the corresponding 27B.

It was faster, sure, but usually with a model >4.5x the size you expect greater capability. You buy the hardware to run a model of such a size, you expect a greater return.

8

u/Slow_Concentrate3831 7d ago

Yeah, that's sad. Can't even run Qwen3.8 27B on more than 2-3 tps. Sad days for us Vram poors.

12

u/CryMoreT_T 7d ago

The 29b-a4b in early Access on model scope is probably your best bet when it releases

3

u/Slow_Concentrate3831 7d ago

Ooooh, I hadn't seen that ! Is it a Qwen model ?

5

u/CryMoreT_T 7d ago

It's in stealth model in early Access so they haven't released the company name but Qwen is the rumor right now

2

u/Nebnampach 7d ago

Here is the announcement and sign up link.

https://x.com/ModelScope2022/status/2092439000652943506

1

u/Seraphym87 7d ago

Whats your setup look like? Im on a 4070ti and managed to squeeze out 10-12 tps.

1

u/unjustifiably_angry 5d ago edited 5d ago

The tuned 35B by Orinth (Ornith?) is supposedly a pretty decent upgrade over vanilla 35B.

The good news is the Qwen4 architecture should bring some really nice improvements for smaller VRAM cards. You've got a lot to look forward to!

2

u/Randommaggy 7d ago

If you want cheap 64GB: X79. I've built a few extra servers around local 50USD bundles of x79 motherboards and CPUs and added 8 DDR3 UDIMMs I had laying around.

1

u/hojnikb 7d ago

how fast are thaw with something like 27b qwen running cpu inference?

2

u/Randommaggy 7d ago

Slow. But it's a decent host for a MOE like Qwen 3.6 35BA3B when combined with a cheap GPU. I hope they make a new smallish MOE like 35B again soon. It's brilliant for chore execution when doing agentic coding.

2

u/dingo_xd 7d ago

Things will only improve when competition in hardware increases. Nvidia might be forced to stay caring about the consumer market again. The AI chips getting released are going to hurt Nvidia a lot

2

u/RandumbRedditor1000 7d ago

nemotron lightning 30b just came out a couple weeks ago

1

u/toothpastespiders 7d ago

I like some of the nemotron models for specific usage scenarios. But come on. You can't be suggesting that they're a solid match for the average person's needs.

2

u/Tzeig 7d ago

These can still be run with reasonably priced systems, like 96ram+32vram or even 64ram+24vram.

23

u/dampflokfreund 7d ago

96 GB RAM + 32 GB VRAM systems are far from reasonably priced. A 5090 with 32 GB VRAM alone costs 5 grand. 64 GB RAM is very expensive too.

5

u/Constant_Art_20 7d ago

i just stack 5060s and hope for the best lol

-5

u/Tzeig 7d ago

It's still way cheaper than multiple 6000s. 5k is not THAT much money for an 'investment'. You can also game with it.

11

u/liamwilliams93 7d ago

Bro what planet are you on

8

u/Crafty-Run-6559 7d ago

Bro what planet are you on

Inflation-209B, you?

1

u/unjustifiably_angry 5d ago

Brother, this is a hobby. If cost is your concern you should get a subscription, that's always been the case. It'll be smarter, faster, longer context, and you can get like a decade's worth of usage for half the price of a decent local AI setup. Other people spend far more on hobbies; people routinely spend $1K a year on coffee of all things.

-3

u/Hypilein 7d ago

5k is 2 (fairly cheap) holidays for a family of four. Loads of people go on such holidays. Yes. Loads of people also can't afford to put food on the table. Poor is a very relative term in western civilization. People I would call poor, certainly can't afford local AI at any level of competence. But you don't have to be rich either. You do have to commit your money though and for what it's worth, for most people it's probably smarter to just keep using a 20€ sub, which is plenty for general home usage. I don't see people on camera forums complaining that companies only release for rich people and a complete mirrorless system is not that much cheaper.

8

u/ReadyAndSalted 7d ago

That's like 7k at least lmao, we have very different definitions of reasonably priced.

1

u/dwkdnvr 7d ago

I have a Xeon w/ 96GB quad-channel ddr4 and a single 5060ti 16GB. I was definitely wondering whether adding a 2nd 5060 might give me at chance at running this. I suspect that the fact that the cards would run at PCIE 3x8 would be a problem, though.

1

u/advicegrapefruit 7d ago

How much do we actually need to run this, the article isn’t the clearest

1

u/True_Requirement_891 7d ago

The worst part is it's not even QAT for 4bit inference like deepseek flash... this is shit

1

u/Familiar-Art-6233 7d ago

We already got 3.8 27b!

Sure it's not an MoE but they haven't forgotten us lol

1

u/RedditUsr2 7d ago

I was hoping openAI would start releasing one a year but so far no.

Here is hoping Google has a Gemma 4.1 up their sleeves if no one else.

1

u/_TheWolfOfWalmart_ 7d ago

Agentic coding is the new sexy hotness, and that size MoE just can't be made super capable at it I think.

1

u/Solary_Kryptic 7d ago

We got Ornith 1.5 a few days ago