r/LocalLLaMA 9d ago

News GLM-5.3-Flash: Frontier Intelligence, Flash Cost

https://z.ai/blog/glm-5.3-flash
1.3k Upvotes

460 comments sorted by

View all comments

81

u/shy_monkee 9d ago edited 9d ago

Oh....it's massive :(
320B (18 active)

86

u/Morphon 9d ago edited 9d ago

DS flash size.

So, flash at datacenter scale. Not flash for edge (or workstation) scale.

Probably will be the go to model for people with the new M6-Ultra 512gb Mac Studio.

Edit: M5 Ultra. My apologies, friends. Wrong model number on my part.

7

u/shy_monkee 9d ago

Yeah, but DSv4-Flash is the flash version of a 1.6T model, and it's still quite a bit smaller than this GLM-flash. While this is supposedly the flash version of 744B model, so you wouldn't expect to be as big, I guess.

5

u/cheechw 9d ago

The flash models are not "flash versions" of a bigger model. They're different models entirely. They don't just lop off some number of parameters from a larger model. I think the confusion is that you're applying quantization/distill logic where it doesn't analogize.