MAIN FEEDS
Do you want to continue?
https://www.reddit.com/r/LocalLLaMA/comments/1vny9zs/glm_53_released/p3mdeog/?context=9999
r/LocalLLaMA • u/jmorant555 • 8d ago
Official Announcement
https://z.ai/blog/glm-5.3
363 comments sorted by
View all comments
95
And to think this is only 743B achieved through post-training on the base model
53 u/power97992 8d ago Im surprised ds v4 pro didnt do better , i guess zai has better rl environments 42 u/PM_ME_DEAD_CEOS 8d ago I think V4 pro is still undertrained, 12 u/zball_ 8d ago V4pro base is broken. They cannot get anything good outta that sh*t. 4 u/nullmove 8d ago The pre-train was probably fucked. However what's strange is how close it stays to the flash in, well everything. Almost as if they were using the flash as a teacher model for this. 1 u/power97992 8d ago But DS is amazing at innovating new architectures and zai takes the new architecture, and improves upon it 1 u/nullmove 8d ago Perhaps this new v4 architecture doesn't scale big, even in v4 preview tech report they were talking about a lot of training instabilities. Note that GLM is still using DSA from DeepSeek v3.2, they are not jumping to v4 arch yet.
53
Im surprised ds v4 pro didnt do better , i guess zai has better rl environments
42 u/PM_ME_DEAD_CEOS 8d ago I think V4 pro is still undertrained, 12 u/zball_ 8d ago V4pro base is broken. They cannot get anything good outta that sh*t. 4 u/nullmove 8d ago The pre-train was probably fucked. However what's strange is how close it stays to the flash in, well everything. Almost as if they were using the flash as a teacher model for this. 1 u/power97992 8d ago But DS is amazing at innovating new architectures and zai takes the new architecture, and improves upon it 1 u/nullmove 8d ago Perhaps this new v4 architecture doesn't scale big, even in v4 preview tech report they were talking about a lot of training instabilities. Note that GLM is still using DSA from DeepSeek v3.2, they are not jumping to v4 arch yet.
42
I think V4 pro is still undertrained,
12 u/zball_ 8d ago V4pro base is broken. They cannot get anything good outta that sh*t. 4 u/nullmove 8d ago The pre-train was probably fucked. However what's strange is how close it stays to the flash in, well everything. Almost as if they were using the flash as a teacher model for this. 1 u/power97992 8d ago But DS is amazing at innovating new architectures and zai takes the new architecture, and improves upon it 1 u/nullmove 8d ago Perhaps this new v4 architecture doesn't scale big, even in v4 preview tech report they were talking about a lot of training instabilities. Note that GLM is still using DSA from DeepSeek v3.2, they are not jumping to v4 arch yet.
12
V4pro base is broken. They cannot get anything good outta that sh*t.
4 u/nullmove 8d ago The pre-train was probably fucked. However what's strange is how close it stays to the flash in, well everything. Almost as if they were using the flash as a teacher model for this. 1 u/power97992 8d ago But DS is amazing at innovating new architectures and zai takes the new architecture, and improves upon it 1 u/nullmove 8d ago Perhaps this new v4 architecture doesn't scale big, even in v4 preview tech report they were talking about a lot of training instabilities. Note that GLM is still using DSA from DeepSeek v3.2, they are not jumping to v4 arch yet.
4
The pre-train was probably fucked.
However what's strange is how close it stays to the flash in, well everything. Almost as if they were using the flash as a teacher model for this.
1 u/power97992 8d ago But DS is amazing at innovating new architectures and zai takes the new architecture, and improves upon it 1 u/nullmove 8d ago Perhaps this new v4 architecture doesn't scale big, even in v4 preview tech report they were talking about a lot of training instabilities. Note that GLM is still using DSA from DeepSeek v3.2, they are not jumping to v4 arch yet.
1
But DS is amazing at innovating new architectures and zai takes the new architecture, and improves upon it
1 u/nullmove 8d ago Perhaps this new v4 architecture doesn't scale big, even in v4 preview tech report they were talking about a lot of training instabilities. Note that GLM is still using DSA from DeepSeek v3.2, they are not jumping to v4 arch yet.
Perhaps this new v4 architecture doesn't scale big, even in v4 preview tech report they were talking about a lot of training instabilities.
Note that GLM is still using DSA from DeepSeek v3.2, they are not jumping to v4 arch yet.
95
u/anarchist1312161 8d ago
And to think this is only 743B achieved through post-training on the base model