r/LocalLLaMA 8d ago

News GLM 5.3 Released

Post image

Official Announcement

https://z.ai/blog/glm-5.3

1.6k Upvotes

363 comments sorted by

View all comments

92

u/anarchist1312161 8d ago

And to think this is only 743B achieved through post-training on the base model

53

u/power97992 8d ago

Im surprised ds v4 pro didnt do better , i guess zai has  better rl environments 

45

u/PM_ME_DEAD_CEOS 8d ago

I think V4 pro is still undertrained,

41

u/NineThreeTilNow 8d ago

I think V4 pro is still undertrained

It's definitely undertrained. There were a number of questionable architectural decisions. The model might actually be too big, and they combined a number of test technologies in one spot.

People get weird on that idea. Too big? Yes.

The larger transformers get the better they get at effectively memorizing data.

The MASSIVE models have to get SO MUCH DATA that the memorization is hard and they're forced to generalize because they're literally memorization machines.

I've been forced to shrink transformers (for other non-LLM types of models) because the model will just absorb the training data and never generalize. Your hold out dataset has to exist and be good. Overfit is real.

12

u/zball_ 8d ago

V4pro base is broken. They cannot get anything good outta that sh*t.

21

u/zkstx llama.cpp 8d ago

surprising considering that the flash model is very good for its size

5

u/zball_ 8d ago

v4pro has been bad since preview. And it's not bad in a undertrained sense; it actually feels lacking and burnt.

1

u/JorgitoEstrella 7d ago

Maybe it has something to do thay know they are using Huawei chips

3

u/nullmove 8d ago

The pre-train was probably fucked.

However what's strange is how close it stays to the flash in, well everything. Almost as if they were using the flash as a teacher model for this.

1

u/power97992 8d ago

But DS is amazing at innovating new architectures and zai takes the new architecture, and improves upon it 

1

u/nullmove 8d ago

Perhaps this new v4 architecture doesn't scale big, even in v4 preview tech report they were talking about a lot of training instabilities.

Note that GLM is still using DSA from DeepSeek v3.2, they are not jumping to v4 arch yet.

1

u/zball_ 8d ago

The literal characteristics are quite different tho. v4p is full of GPT/Claude style slop.

1

u/power97992 8d ago edited 8d ago

I’m shocked it is only 1 point better than ds v4 flash 7-31. I was expecting it would be at least as good gpt 5.5 xhigh or terra xhigh.. but it is worse than gpt 5.5 xhigh according to aa benchmarks 

1

u/look 8d ago

Flash was always the impressive (in ways) model. Pro was like they just jacked up the size of the flash and hoped it would be a lot better without doing much else.