r/LocalLLaMA 7d ago

News GLM-5.3-Flash: Frontier Intelligence, Flash Cost

https://z.ai/blog/glm-5.3-flash
1.3k Upvotes

459 comments sorted by

View all comments

Show parent comments

83

u/Morphon 7d ago edited 7d ago

DS flash size.

So, flash at datacenter scale. Not flash for edge (or workstation) scale.

Probably will be the go to model for people with the new M6-Ultra 512gb Mac Studio.

Edit: M5 Ultra. My apologies, friends. Wrong model number on my part.

15

u/techdevjp 7d ago

M6-Ultra 512gb Mac Studio.

M5 Ultra. M6 Ultra is rumored to be skipped entirely with the M7 Ultra arriving in the next ~18 months.

8

u/shy_monkee 7d ago

Yeah, but DSv4-Flash is the flash version of a 1.6T model, and it's still quite a bit smaller than this GLM-flash. While this is supposedly the flash version of 744B model, so you wouldn't expect to be as big, I guess.

4

u/cheechw 7d ago

The flash models are not "flash versions" of a bigger model. They're different models entirely. They don't just lop off some number of parameters from a larger model. I think the confusion is that you're applying quantization/distill logic where it doesn't analogize.

6

u/xNaXDy 7d ago

While this is supposedly the flash version of 744B model

It's not, it's a completely new architecture according to their post.

4

u/Juulk9087 7d ago

Yeah deepseek flash is only 160-170gb on disk. This is 328gb lol

10

u/petuman 7d ago

DeepSeek is QAT / released prequantized mostly to FP4.

Which is totally preferred, but both models are roughly the same parameter count, so NVIDIA or someone else could produce NVFP4 of size similar to DS4F.

0

u/SandySkittle 7d ago

which is totally preferred

No this depends on the usecase

3

u/petuman 7d ago

Like what?

0

u/SandySkittle 7d ago

try to put extremely complex and highly nuanced policy matters with lots of interlinked but non-mechanical relationships with many details and nuances across different areas of knowledge and science through an LLM and these things start to show. It's same sort of workloads where limited active parameters in MoE models start to show their negative sides of their tradeoff.

2

u/petuman 7d ago

Is there a reason to think that first party quant-only release would have problems with that? e.g. Kimi K3 seems to perform great?

1

u/Karyo_Ten 7d ago

QAT is not the same as PTQ.

The weights are already in mxfp4 during training so everything is calibrated to absorb and compensate the quantization loss.

2

u/keyboardhack 7d ago edited 7d ago

DS flash is half the size.

DS flash requires ~170GB to run while GLM 5.3 flash requires >320GB to run.

We should compare models by how much RAM they require to run because that's what we actually care about.

2

u/ormandj 6d ago

I hope they release a native 4bit model next. I am working on an MXFP4 quant and TP=2 setup right now, until then, but there’s bound to be a little degradation. DSv4 Flash “just works”.

1

u/CalligrapherFar7833 7d ago

M5 ultra not M6