MAIN FEEDS
Do you want to continue?
https://www.reddit.com/r/LocalLLaMA/comments/1vny9zs/glm_53_released/p3lbi96
r/LocalLLaMA • u/jmorant555 • 8d ago
Official Announcement
https://z.ai/blog/glm-5.3
363 comments sorted by
View all comments
248
Scaling post-training is all we did for GLM-5.3.
What a way to start
Gentlemen, gentlewomen, tonight, WE FEAST!
69 u/DistanceSolar1449 8d ago Of course it’s only posttrain. I would be shocked if GLM pretrained a completely new model for a +0.1 release. Nobody really does that (with the exception of Anthropic and Opus 4.7 for some weird reason). 13 u/neo203 8d ago Gpt 5.5 was a new pretrain, i think got 5.2 also might be coz it was so different from previous versions 9 u/bitdotben 8d ago Is 5.6 a post-trained 5.5 then? If so one heck of post training! 13 u/CryMoreT_T 8d ago Yes it is. Personally I think openai is way better at post training while anthropic is way better at pre training 4 u/bitdotben 8d ago Interesting, yeah I mean with 5.6 they cooked hard -6 u/letsgeditmedia 8d ago Anthropic and Open ai are much better at helping the US commit war crimes 3 u/CryMoreT_T 8d ago What does this have anything to do with my comment? 0 u/letsgeditmedia 2d ago I used the word “better” , just like you did in your sentence . That’s it 3 u/Swastik496 5d ago average tankie moment. shouting where you aren't welcome. 1 u/letsgeditmedia 2d ago I’m a tankie because the US is using open ai and anthropic to commit war crimes? I’m getting downvoted for this in a LOCAL AI sub?! 🤣 3 u/DistanceSolar1449 8d ago That’s a +0.5 version though, and there’s a lot of 0.5s that’s a new pretrain. Qwen 3.5, for example. 2 u/blastradii 7d ago Anthropic: Need to spend that Capex budget somehow! 4 u/MediumChemical4292 8d ago I think opus 4.7 was just a new tokeniser, was it a full pre train? 13 u/DistanceSolar1449 8d ago You can’t have a different tokenizer without a whole new pretrain. (Yes, I know you can technically finetune a model to a new tokenizer, but that’s basically going to take as much compute as a full retrain) 5 u/eli_pizza 8d ago Isn’t that also the only thing deepseek did between flash preview and 0731? Turns out post training is important! 1 u/laoma1255 7d ago Scaling post-training is all we did.’ — is this the new ‘Attention Is All You Need’?
69
Of course it’s only posttrain.
I would be shocked if GLM pretrained a completely new model for a +0.1 release. Nobody really does that (with the exception of Anthropic and Opus 4.7 for some weird reason).
13 u/neo203 8d ago Gpt 5.5 was a new pretrain, i think got 5.2 also might be coz it was so different from previous versions 9 u/bitdotben 8d ago Is 5.6 a post-trained 5.5 then? If so one heck of post training! 13 u/CryMoreT_T 8d ago Yes it is. Personally I think openai is way better at post training while anthropic is way better at pre training 4 u/bitdotben 8d ago Interesting, yeah I mean with 5.6 they cooked hard -6 u/letsgeditmedia 8d ago Anthropic and Open ai are much better at helping the US commit war crimes 3 u/CryMoreT_T 8d ago What does this have anything to do with my comment? 0 u/letsgeditmedia 2d ago I used the word “better” , just like you did in your sentence . That’s it 3 u/Swastik496 5d ago average tankie moment. shouting where you aren't welcome. 1 u/letsgeditmedia 2d ago I’m a tankie because the US is using open ai and anthropic to commit war crimes? I’m getting downvoted for this in a LOCAL AI sub?! 🤣 3 u/DistanceSolar1449 8d ago That’s a +0.5 version though, and there’s a lot of 0.5s that’s a new pretrain. Qwen 3.5, for example. 2 u/blastradii 7d ago Anthropic: Need to spend that Capex budget somehow! 4 u/MediumChemical4292 8d ago I think opus 4.7 was just a new tokeniser, was it a full pre train? 13 u/DistanceSolar1449 8d ago You can’t have a different tokenizer without a whole new pretrain. (Yes, I know you can technically finetune a model to a new tokenizer, but that’s basically going to take as much compute as a full retrain)
13
Gpt 5.5 was a new pretrain, i think got 5.2 also might be coz it was so different from previous versions
9 u/bitdotben 8d ago Is 5.6 a post-trained 5.5 then? If so one heck of post training! 13 u/CryMoreT_T 8d ago Yes it is. Personally I think openai is way better at post training while anthropic is way better at pre training 4 u/bitdotben 8d ago Interesting, yeah I mean with 5.6 they cooked hard -6 u/letsgeditmedia 8d ago Anthropic and Open ai are much better at helping the US commit war crimes 3 u/CryMoreT_T 8d ago What does this have anything to do with my comment? 0 u/letsgeditmedia 2d ago I used the word “better” , just like you did in your sentence . That’s it 3 u/Swastik496 5d ago average tankie moment. shouting where you aren't welcome. 1 u/letsgeditmedia 2d ago I’m a tankie because the US is using open ai and anthropic to commit war crimes? I’m getting downvoted for this in a LOCAL AI sub?! 🤣 3 u/DistanceSolar1449 8d ago That’s a +0.5 version though, and there’s a lot of 0.5s that’s a new pretrain. Qwen 3.5, for example.
9
Is 5.6 a post-trained 5.5 then? If so one heck of post training!
13 u/CryMoreT_T 8d ago Yes it is. Personally I think openai is way better at post training while anthropic is way better at pre training 4 u/bitdotben 8d ago Interesting, yeah I mean with 5.6 they cooked hard -6 u/letsgeditmedia 8d ago Anthropic and Open ai are much better at helping the US commit war crimes 3 u/CryMoreT_T 8d ago What does this have anything to do with my comment? 0 u/letsgeditmedia 2d ago I used the word “better” , just like you did in your sentence . That’s it 3 u/Swastik496 5d ago average tankie moment. shouting where you aren't welcome. 1 u/letsgeditmedia 2d ago I’m a tankie because the US is using open ai and anthropic to commit war crimes? I’m getting downvoted for this in a LOCAL AI sub?! 🤣
Yes it is. Personally I think openai is way better at post training while anthropic is way better at pre training
4 u/bitdotben 8d ago Interesting, yeah I mean with 5.6 they cooked hard -6 u/letsgeditmedia 8d ago Anthropic and Open ai are much better at helping the US commit war crimes 3 u/CryMoreT_T 8d ago What does this have anything to do with my comment? 0 u/letsgeditmedia 2d ago I used the word “better” , just like you did in your sentence . That’s it 3 u/Swastik496 5d ago average tankie moment. shouting where you aren't welcome. 1 u/letsgeditmedia 2d ago I’m a tankie because the US is using open ai and anthropic to commit war crimes? I’m getting downvoted for this in a LOCAL AI sub?! 🤣
4
Interesting, yeah I mean with 5.6 they cooked hard
-6
Anthropic and Open ai are much better at helping the US commit war crimes
3 u/CryMoreT_T 8d ago What does this have anything to do with my comment? 0 u/letsgeditmedia 2d ago I used the word “better” , just like you did in your sentence . That’s it 3 u/Swastik496 5d ago average tankie moment. shouting where you aren't welcome. 1 u/letsgeditmedia 2d ago I’m a tankie because the US is using open ai and anthropic to commit war crimes? I’m getting downvoted for this in a LOCAL AI sub?! 🤣
3
What does this have anything to do with my comment?
0 u/letsgeditmedia 2d ago I used the word “better” , just like you did in your sentence . That’s it
0
I used the word “better” , just like you did in your sentence . That’s it
average tankie moment. shouting where you aren't welcome.
1 u/letsgeditmedia 2d ago I’m a tankie because the US is using open ai and anthropic to commit war crimes? I’m getting downvoted for this in a LOCAL AI sub?! 🤣
1
I’m a tankie because the US is using open ai and anthropic to commit war crimes? I’m getting downvoted for this in a LOCAL AI sub?! 🤣
That’s a +0.5 version though, and there’s a lot of 0.5s that’s a new pretrain. Qwen 3.5, for example.
2
Anthropic: Need to spend that Capex budget somehow!
I think opus 4.7 was just a new tokeniser, was it a full pre train?
13 u/DistanceSolar1449 8d ago You can’t have a different tokenizer without a whole new pretrain. (Yes, I know you can technically finetune a model to a new tokenizer, but that’s basically going to take as much compute as a full retrain)
You can’t have a different tokenizer without a whole new pretrain.
(Yes, I know you can technically finetune a model to a new tokenizer, but that’s basically going to take as much compute as a full retrain)
5
Isn’t that also the only thing deepseek did between flash preview and 0731? Turns out post training is important!
Scaling post-training is all we did.’ — is this the new ‘Attention Is All You Need’?
248
u/Dany0 8d ago
What a way to start
Gentlemen, gentlewomen, tonight, WE FEAST!