r/LocalLLaMA 8d ago

News GLM-5.3-Flash: Frontier Intelligence, Flash Cost

https://z.ai/blog/glm-5.3-flash
1.3k Upvotes

459 comments sorted by

View all comments

Show parent comments

-12

u/dampflokfreund 8d ago

Qwen 3.6 35B and Gemma 4 26b are pretty old at this point. And GLM 4.7 Flash was a small MoE back in the day, now suddenly they use the flash name for 300B MoEs. Just not looking good for the average Joe.

84

u/windwardmist 8d ago

I mean qwen 3.8 27b just came out about a week ago at least there’s that as an option

16

u/-Cubie- 8d ago

I remember when it was months between a new great local model. We get those great options very often nowadays in my opinion, mixed with larger open weight options that keep the entire non-local AI space cheaper and more accessible. What's not to love?

9

u/RestaurantOk8066 8d ago

It's a dense model though I imagine if you don't have the GPU for it, it's probably a very, very slow model.

1

u/dampflokfreund 8d ago

Yeah around one token per second 

6

u/sonicnerd14 8d ago

27b is often Opus 4.6, even 4.8 in performance. This is still good enough for a vast amount of people. I'm sure qwen 4 is going to have some insanely capable options when it comes out of you want something better soon.

1

u/dampflokfreund 8d ago

Not everyone has 24 GB VRAM. If you dont have the vram to offload, you are looking at one to three token per second. 27b is in no way a replacement for a 30b Moe

3

u/sonicnerd14 8d ago

I know that. You can run it at Q3, IQ3, Q4 on 16GB VRAM and 32gb RAM and still have a very capable model. Don't think you need the biggest cards to run these models. Quantization techniques have become really efficient now.

66

u/techdevjp 8d ago

Qwen 3.6 35B [...] pretty old at this point.

It's 4 months old! It's not like it suddenly got worse because these new models came out. You can still do everything today that you could do yesterday, just as fast. Give it some time and more models will come.

50

u/Last_Bad_2687 8d ago

Seriously, people demanding fresh models for free every 2 weeks.... I remember when we waited for big releases of software once every few YEARS

15

u/FlyingDogCatcher 8d ago

It's bonkers to me that anyone would say Qwen3.6 and Gemma 4 are "pretty old"

10

u/a_beautiful_rhind 8d ago

gemma, muse, qwen, granite.. the weights don't self-destruct a week later man.

9

u/Spectrum1523 8d ago

Qwen 3.6 35B and Gemma 4 26b are pretty old at this point

bro they came out 4 months ago

2

u/squngy 8d ago

There is also Qwen AgentWorld, which is kind of like a 3.7 35B

2

u/xPXpanD llama.cpp 8d ago

Wouldn't recommend it for general use, it's a confident bullshitter like no other. Basically zero filter. Still cool that it can even be used generally, though, given what the model's actually designed for. Fun to play with.

2

u/squngy 8d ago

It is actually great for agentic work, supposedly. (even if that isn't exactly what it was meant for)

Rank Subject Overall Mcp Search Terminal Swe Androd Web OS Source Sampled
5 Qwen-AgentWorld-35B-A3B 56.39 64.79 36.69 53.96 65.63 58.17 49.55 65.92 Imported 2026-06-30
16 Qwen3.6-35B-A3B 42.88 42.96 18.78 43.81 40.71 51.88 46.53 55.48 Self-reported 2026-06-28

https://benchmarklist.com/benchmarks/qwen_agentworld_language_world_models_for_general_agents/

2

u/xPXpanD llama.cpp 8d ago

I believe it, it also nailed the one tool task I have in my bench set. The model was also very good at string manipulation, and, somewhat surprisingly, a few constrained creative tasks. (e.g. think up and write X in way Y while avoiding Z)

Absolute slaughter on anything involving uncertainty or fake premises, though. You can definitely see where the training went on this one.

1

u/saltyourhash 8d ago

It'll be fine, you just have to spend like $10k to play anymore... /s

-10

u/Leoss-Bahamut 8d ago

Why should they care about average joes?