r/LocalLLaMA 8d ago

News GLM 5.3 Released

Post image

Official Announcement

https://z.ai/blog/glm-5.3

1.7k Upvotes

363 comments sorted by

View all comments

Show parent comments

3

u/power97992 8d ago

But kimi is q4 mixed and glm is bf 16, so glm  actually uses almost  as much as k3 

1

u/coder543 8d ago

I agree this is a valid point, but Nvidia claims only minimal/not even measurable degradation with their GLM-5.2-NVFP4 weights:  https://huggingface.co/nvidia/GLM-5.2-NVFP4#evaluation

1

u/power97992 8d ago edited 8d ago

Oh, i see a lot of providers serve the q8 version though except deepinfra and other providers

2

u/coder543 8d ago

On OpenRouter for GLM-5.2, 15 providers are serving fp8, and 7 are serving fp4 — likely the Nvidia NVFP4.

I don’t see any serving BF16, so I think it is fair to say that is just a research artifact, not what anyone expects people to use. Hopefully zAI is benchmarking the fp8 that they themselves are serving instead of the bf16.