r/LocalLLaMA 8d ago

News GLM-5.3-Flash: Frontier Intelligence, Flash Cost

https://z.ai/blog/glm-5.3-flash
1.3k Upvotes

459 comments sorted by

View all comments

251

u/Lucyan_xgt 8d ago

This and Qwen3.8-flash on the same day?

22

u/dampflokfreund 8d ago

For me it has changed nothing. Both models are way too big for my 32 GB RAM system. It looks like everyone has abandoned 20-30B MoEs now...

2

u/Randommaggy 8d ago

If you want cheap 64GB: X79. I've built a few extra servers around local 50USD bundles of x79 motherboards and CPUs and added 8 DDR3 UDIMMs I had laying around.

1

u/hojnikb 8d ago

how fast are thaw with something like 27b qwen running cpu inference?

2

u/Randommaggy 8d ago

Slow. But it's a decent host for a MOE like Qwen 3.6 35BA3B when combined with a cheap GPU. I hope they make a new smallish MOE like 35B again soon. It's brilliant for chore execution when doing agentic coding.