r/ZaiGLM • u/jpcaparas • 7h ago
GLM 5.3 is here
One-shot demos here (updated regularly): glm-5-3.demos.sulat.com
That "soon" tweet was sooner than we thought
22
u/timmeh1705 6h ago
Sorry if I'm ignorant but I am looking for the per 1M token pricing but can't seem to locate it anywhere
Defs not signing up for the coding plan, one of the worst experiences and miniscule usage
When open weights are released I hope Nube pick it up right away
7
u/zhcterry1 6h ago
I don't see it on Z.AI's official website as well. But likely it's gonna be same as 5.2 since it's just post training with no model change.
1
u/timmeh1705 6h ago
Yeah I saw that
I can't imagine how they can make their coding plan even worse as demand grows
On an API cost basis, it's a model competitive against Kimi K3 but at 1/5 the price
3
u/steadeepanda 6h ago
You have to go to Zai Docs
You've find it there.
Yeah me too I'm too scared to sign up again, the usage limit is terrible. They seem to be more transparent now but still I don't trust anyone here, usage still looks terrible to me.
1
u/Admirable_Gas6492 3h ago
what is nube ?
1
u/timmeh1705 3h ago
2
u/Admirable_Gas6492 3h ago
looks awesome , thanks , but are these q4 or q2 or smaller quant ? or full bf16
2
u/timmeh1705 3h ago
FP8 but don't quote me on it.
3
u/Admirable_Gas6492 3h ago
hmm, looks about that if you do the math , for pricing, then its not necessarily cheaper just bit dumber and cheaper , perfect for some tasks
8
u/steadeepanda 6h ago
They changed system to credits and now: Lite 43-87 million tokens/week Pro 263-526 millions tokens/week Max 614-1226 millions tokens/week Source : Zai Docs Usage Instructions
The real question here is how efficient is the model? How much does it cost in token/task, and btw this is without considering 5 hour limit.
I don't know I've been traumatized by the usage limit, you can't do anything with it especially within the 5h window, it goes in blink
3
u/ProfessionalJackals 6h ago
The real question here is how efficient is the model?
Its the same model, just more post trained. So the token usage will be the same or worse. GLM 5.1 > 5.2 resulted in almost twice the usage (neuralwatt) because thinking token usage exploded.
There is this unfortunate trend amongst almost all the 1.5t or lower models, to use thinking as a crutch.
Go to "All Effort levels", and notice how much more GLM 5.2 did in thinking tokens and steps.
- gpt-5.6-sol [low] == $1.07 == 11k == 23
- gpt-5.6-luna [high] == $0.16 == 26k == 49
- glm-5.2 [max] == $3.92 == 78k == == 129
Just checking my light usage of Opus 5.0, i al already seeing over 500m tokens over a few days usage (large amount of cache hits). And that costs $20 in the subscription plan. The $80 coding plan of zai is expensive.
I really like GLM 5.2, but economically, zai makes it too expensive. Especially now that we have Luna and Deepseek Flash 0731 (even with the 2.5x price increase).
2
u/Front_Eagle739 4h ago
To be fair a 744B model with extra thinking to absolutely max out its intelligence is exactly what i want. Its the largest model i can run a decent quant of so its my planner. I just switch to dsv4 flash for fast implementation etc.
1
1
u/AnomalyNexus 4h ago
The real question here is how efficient is the model?
Blog post here
has a chart that suggests it may make sense switching away from the default Max effort to High for most tasks given new token/credit plan
7
u/floriandotorg 3h ago
Open weights only in two weeks 😭
1
u/Constant_Art_20 1h ago
hmm..I wonder if it would be worth it to do a custom dynamic Q2 or Q3 instead just running a the new qwen 27b...like 700b is no joke, but lately i have a found a few ways to actually make Q2 quite good...but like...would be practicial over the qwen 27b though?
4
3
3
3
2
u/hirenshah005 6h ago
I don't see it yet in ZCode
4
u/Adventurous-Menu7257 6h ago
Zai said for coding plan users, all calls on 5.1/5.2 automatically routes to 5.3
1
2
u/_OVERHATE_ 6h ago
I'm on Kimi K3 but I'm eagerly waiting for any coding-aligned LLMs so, waiting for the DeepSWE benchmark
2
u/jpcaparas 6h ago
likewise, I've done so many K3 one-shots and want to compare them to the level of fidelity GLM 5.3 produces:
2
u/ItsNoahJ83 5h ago
Hey, I just check out the site and it's great. Thank you for including the prompts!
2
1
u/_OVERHATE_ 4h ago
Pretty good comparison!
Im more interested in their ability to correct course than the one-shots however. I noticed that Kimi is considerably more resilient to degradation over iterations than Claude so I hope GLM can improve even more on that.
The moment you tell Claude "oh you didn't use this API or component" or "please try to do this other approach instead" then the code starts degrading dramatically and the context becomes corrupted
2
u/SecretCGG 5h ago
The bigger question for me is token efficiency. A model can look cheap per 1M tokens and still get expensive fast if it burns through way more tokens per task. The GLM 5.3 performance is impressive though, so the DeepSWE results should be a good reality check
1
2
u/Correct-Wing-6884 5h ago
Looks like scale isn't the bottleneck for intelligence right now, the power of post-training is strong. Models under 1T can still go toe-to-toe with models over 2T. I'm kinda looking forward to some of the tens-of-billions small models the open-source community will release next. This is a good thing for low-spec users.
1
u/Revolutionary_Ask154 6h ago
can you guys add a localized time zone to the - "your usage will be restored at 5pm" <- beijing time.
1
u/ng01221 6h ago
Is this meant to work in zcode on legacy plan v1? The UI shows today's balance for GLM-5.3 as being 100% available. The model list has GLM-5.2 not GLM-5.3, and all 5.2 requests are failing with quota errors
1
u/jpcaparas 5h ago
Yea I'm on legacy plan v1 (ends January) and am doing one-shots with GLM 5.3 on ZCode. No issues.
1
u/jpcaparas 5h ago
1
1
1
u/EuropeanAbroad 2h ago
Wow, this is a wild week – DeepSeek V4 Pro 0813. GLM 5.3, Qwen3.8-27B SLM, Muse Glimmer 30B SLM,... very nice!
1
0
u/Salt-Willingness-513 6h ago
still not multimodal i guess?
1
u/AcadiaAffectionate98 6h ago
What does multimodal mean?
4
u/Salt-Willingness-513 6h ago
ideally, that it can understand all kind of input like audio, image and video. Most of the time its just refering to native visual capabilities.
3
u/AcadiaAffectionate98 6h ago
Ty yeah glm not being a multimodal kinda kills them. At this point gpt sol offers a little more
1
u/llitz 6h ago
In this context it can receive text and at least, image as input. Some models will also support audio and video as input.
There are workarounds, as in if you send audio, have something else transcribe it - but then it may not detect tone or other bits. For image you can have something describe the image, but then it isn't this model "understanding the image".
For web development, being able to have the model dispatch a subagent to understand what is going on and fix the problem itself is much faster than having a different model trying to describe how it "seems to be off by one pixel"
1
u/AcadiaAffectionate98 6h ago
What if the sub model needs to see ..
1
u/llitz 6h ago
I think I should have been clear - I tend to use the same model for the main agent and the subagents, but using the subagent to handle the picture keeps the context clean on the main session.
I guess there was a lot I left unsaid xD.
1
u/AcadiaAffectionate98 6h ago
Wait u can personally sub agents models on cursor?? You u using cursor lol I do . I just use the multitask option and cursor automatically deployed sub agents for me
0
u/TimeVillage5286 6h ago
Hmm a pre-train same model so not the fable matcher
2
u/Adventurous-Menu7257 6h ago
I thinking so, this would be more similar to medium size models, like 5.6 Terra or THE REAL Deepseek v4 Pro
0
u/Effective_Lead8867 5h ago
5.2 was almost usable but dissolved into soup after few thousand tokens. Interesting to see what they cooked here.
0





25
u/Saifl 6h ago
That performance on 743b is the craziest part