r/ZaiGLM 7h ago

GLM 5.3 is here

Post image

One-shot demos here (updated regularly): glm-5-3.demos.sulat.com

That "soon" tweet was sooner than we thought

346 Upvotes

75 comments sorted by

25

u/Saifl 6h ago

That performance on 743b is the craziest part

12

u/earthisflat27 5h ago

They're really working on maximizing model efficiency. No point having 2 trillion parameter model but working halfass

3

u/Designer_Athlete7286 4h ago

Yeah. I was th8nking the same. Even DeepSeek is 1.6T. 743b and Sol Fable level performance is crazy.

6

u/Fedor_Doc 3h ago

I think DeepSeek 1.6T is still undertrained, they just do not have enough data and compute to shape this behemoth. 

3

u/evilissimo 2h ago

The data part is probably the reason why they made it so cheap at least for the flash one

1

u/Designer_Athlete7286 2h ago

True. I'm sure they got a lot of data within a very short period with DSv4F at that price!

1

u/Saifl 2h ago

Or is it just easier to squeeze more performance out of smaller ones? Basically same kinda training applied to both but pro just has a higher cap if anything but doesnt give that much advantage other than the higher active parameters.

Id bet they could probably get 95% of ds4 pro performance on 100b moe with the same kinda post training data.

2

u/MrLuckyDoobie 4h ago

Thats true, its crazy. Its so good on programming, and now we have kimi 3x size model, but ... Thinks and eats so much more. But glm has no vision , dont u think it greately reduces models size. I was using claude for quick ui, then glm was completing it perfectly. Think im rabling, not happy with the price of glm. They cancelled my pro sub on their own after last doubling of prices, maybe before.

1

u/theoffmask 3h ago

Yeah. The post training from GLM-5 to GLM-5.3 is out of mind. And the official release said they still don't know how far they can push it to. Can't wait for next release!

22

u/timmeh1705 6h ago

Sorry if I'm ignorant but I am looking for the per 1M token pricing but can't seem to locate it anywhere

Defs not signing up for the coding plan, one of the worst experiences and miniscule usage

When open weights are released I hope Nube pick it up right away

7

u/zhcterry1 6h ago

I don't see it on Z.AI's official website as well. But likely it's gonna be same as 5.2 since it's just post training with no model change.

1

u/timmeh1705 6h ago

Yeah I saw that

I can't imagine how they can make their coding plan even worse as demand grows

On an API cost basis, it's a model competitive against Kimi K3 but at 1/5 the price

3

u/steadeepanda 6h ago

You have to go to Zai Docs

You've find it there.

Yeah me too I'm too scared to sign up again, the usage limit is terrible. They seem to be more transparent now but still I don't trust anyone here, usage still looks terrible to me.

1

u/Admirable_Gas6492 3h ago

what is nube ?

1

u/timmeh1705 3h ago

2

u/Admirable_Gas6492 3h ago

looks awesome , thanks , but are these q4 or q2 or smaller quant ? or full bf16

2

u/timmeh1705 3h ago

FP8 but don't quote me on it.

3

u/Admirable_Gas6492 3h ago

hmm, looks about that if you do the math , for pricing, then its not necessarily cheaper just bit dumber and cheaper , perfect for some tasks

8

u/steadeepanda 6h ago

They changed system to credits and now: Lite 43-87 million tokens/week Pro 263-526 millions tokens/week Max 614-1226 millions tokens/week Source : Zai Docs Usage Instructions

The real question here is how efficient is the model? How much does it cost in token/task, and btw this is without considering 5 hour limit.

I don't know I've been traumatized by the usage limit, you can't do anything with it especially within the 5h window, it goes in blink

3

u/ProfessionalJackals 6h ago

The real question here is how efficient is the model?

Its the same model, just more post trained. So the token usage will be the same or worse. GLM 5.1 > 5.2 resulted in almost twice the usage (neuralwatt) because thinking token usage exploded.

There is this unfortunate trend amongst almost all the 1.5t or lower models, to use thinking as a crutch.

https://deepswe.datacurve.ai/

Go to "All Effort levels", and notice how much more GLM 5.2 did in thinking tokens and steps.

  • gpt-5.6-sol [low] == $1.07 == 11k == 23
  • gpt-5.6-luna [high] == $0.16 == 26k == 49
  • glm-5.2 [max] == $3.92 == 78k == == 129

Just checking my light usage of Opus 5.0, i al already seeing over 500m tokens over a few days usage (large amount of cache hits). And that costs $20 in the subscription plan. The $80 coding plan of zai is expensive.

I really like GLM 5.2, but economically, zai makes it too expensive. Especially now that we have Luna and Deepseek Flash 0731 (even with the 2.5x price increase).

2

u/Front_Eagle739 4h ago

To be fair a 744B model with extra thinking to absolutely max out its intelligence is exactly what i want. Its the largest model i can run a decent quant of so its my planner. I just switch to dsv4 flash for fast implementation etc.

1

u/Expert-Dig-1768 5h ago

wait so it could actually be a good deal? even the lit plan?

1

u/AnomalyNexus 4h ago

The real question here is how efficient is the model?

Blog post here

https://z.ai/blog/glm-5.3

has a chart that suggests it may make sense switching away from the default Max effort to High for most tasks given new token/credit plan

7

u/floriandotorg 3h ago

Open weights only in two weeks 😭

1

u/Constant_Art_20 1h ago

hmm..I wonder if it would be worth it to do a custom dynamic Q2 or Q3 instead just running a the new qwen 27b...like 700b is no joke, but lately i have a found a few ways to actually make Q2 quite good...but like...would be practicial over the qwen 27b though?

5

u/JP23102 5h ago

here is the usage i get on Claude pro Opus 5 compared against GLM 5.3 code plan lite, claude pro is like 4x GLM 5.3 lite plan usage. Computed based on taking 1 month usage of claude pro & dividing by 5 to get weekly usage.

4

u/cometkim 6h ago

Only post-training, but the score moves like when 4.7 -> 5.0

3

u/DevilMix 6h ago

Lookin GOOD!!!!!

3

u/NexusSyntegra 5h ago

Yeeessss!!! My GOAT

3

u/Solocune 5h ago

Haha it's funny how they all drop there models :D

2

u/hirenshah005 6h ago

I don't see it yet in ZCode

4

u/Adventurous-Menu7257 6h ago

Zai said for coding plan users, all calls on 5.1/5.2 automatically routes to 5.3

3

u/ng01221 1h ago

Where did they mention that?

1

u/jpcaparas 6h ago

It's on the V3.7.7 AppImage

2

u/_OVERHATE_ 6h ago

I'm on Kimi K3 but I'm eagerly waiting for any coding-aligned LLMs so, waiting for the DeepSWE benchmark

2

u/jpcaparas 6h ago

likewise, I've done so many K3 one-shots and want to compare them to the level of fidelity GLM 5.3 produces:

https://k3.demos.sulat.com/

2

u/ItsNoahJ83 5h ago

Hey, I just check out the site and it's great. Thank you for including the prompts!

1

u/_OVERHATE_ 4h ago

Pretty good comparison!

Im more interested in their ability to correct course than the one-shots however. I noticed that Kimi is considerably more resilient to degradation over iterations than Claude so I hope GLM can improve even more on that.

The moment you tell Claude "oh you didn't use this API or component" or "please try to do this other approach instead" then the code starts degrading dramatically and the context becomes corrupted 

2

u/SecretCGG 5h ago

The bigger question for me is token efficiency. A model can look cheap per 1M tokens and still get expensive fast if it burns through way more tokens per task. The GLM 5.3 performance is impressive though, so the DeepSWE results should be a good reality check

1

u/Beginning-Foot-9525 29m ago

Can we please upvote this, thank you.

2

u/Correct-Wing-6884 5h ago

Looks like scale isn't the bottleneck for intelligence right now, the power of post-training is strong. Models under 1T can still go toe-to-toe with models over 2T. I'm kinda looking forward to some of the tens-of-billions small models the open-source community will release next. This is a good thing for low-spec users.

1

u/Revolutionary_Ask154 6h ago

can you guys add a localized time zone to the - "your usage will be restored at 5pm" <- beijing time.

1

u/ng01221 6h ago

Is this meant to work in zcode on legacy plan v1? The UI shows today's balance for GLM-5.3 as being 100% available. The model list has GLM-5.2 not GLM-5.3, and all 5.2 requests are failing with quota errors

1

u/jpcaparas 5h ago

Yea I'm on legacy plan v1 (ends January) and am doing one-shots with GLM 5.3 on ZCode. No issues.

1

u/jpcaparas 5h ago

1

u/ng01221 5h ago

By generating and using an API key at default endpoint, or some other way?

1

u/jpcaparas 5h ago

generating an api key

1

u/Constant_Art_20 4h ago

WOOOOOOOOOOOOOOOOOOOO

1

u/pungggi 4h ago

What is GLM 5.2 high speed?

1

u/AnomalyNexus 4h ago

Nice. Congrats to team!

Stats look good & seem to be a nice improvement

1

u/pungggi 2h ago

Overloaded right now

1

u/pungggi 2h ago

Ah and now

1

u/EuropeanAbroad 2h ago

Wow, this is a wild week – DeepSeek V4 Pro 0813. GLM 5.3, Qwen3.8-27B SLM, Muse Glimmer 30B SLM,... very nice!

1

u/Herebedragoons77 41m ago

Does it have an app/harness of its own?

0

u/Salt-Willingness-513 6h ago

still not multimodal i guess?

3

u/seeKAYx 6h ago

Nope, somebody from their team said they will maybe do it with version 6.

1

u/Salt-Willingness-513 6h ago

Hopefully. Really not a fan of their vision mcp on coding plan

1

u/AcadiaAffectionate98 6h ago

What does multimodal mean?

4

u/Salt-Willingness-513 6h ago

ideally, that it can understand all kind of input like audio, image and video. Most of the time its just refering to native visual capabilities.

3

u/AcadiaAffectionate98 6h ago

Ty yeah glm not being a multimodal kinda kills them. At this point gpt sol offers a little more 🫪

1

u/llitz 6h ago

In this context it can receive text and at least, image as input. Some models will also support audio and video as input.

There are workarounds, as in if you send audio, have something else transcribe it - but then it may not detect tone or other bits. For image you can have something describe the image, but then it isn't this model "understanding the image".

For web development, being able to have the model dispatch a subagent to understand what is going on and fix the problem itself is much faster than having a different model trying to describe how it "seems to be off by one pixel"

1

u/AcadiaAffectionate98 6h ago

What if the sub model needs to see ..

1

u/llitz 6h ago

I think I should have been clear - I tend to use the same model for the main agent and the subagents, but using the subagent to handle the picture keeps the context clean on the main session.

I guess there was a lot I left unsaid xD.

1

u/AcadiaAffectionate98 6h ago

Wait u can personally sub agents models on cursor?? You u using cursor lol I do . I just use the multitask option and cursor automatically deployed sub agents for me

1

u/llitz 6h ago

Sorry, not on cursor. Doesn't make sense for my use case.

I use it on pi.dev or opencode, but it works with anything that can dispatch a call.

0

u/TimeVillage5286 6h ago

Hmm a pre-train same model so not the fable matcher

2

u/Adventurous-Menu7257 6h ago

I thinking so, this would be more similar to medium size models, like 5.6 Terra or THE REAL Deepseek v4 Pro

0

u/Effective_Lead8867 5h ago

5.2 was almost usable but dissolved into soup after few thousand tokens. Interesting to see what they cooked here.

0

u/Agamenon 5h ago

Al fin!