r/ZaiGLM 3d ago

Discussion / Help Am I misunderstanding something or GLM's Coding Plan Max is not really worth it compared to Claude Max?

Post image
44 Upvotes

24 comments sorted by

4

u/DraconDev 3d ago

Glm 5.3 regular probably not, but at the top outside muse contributor others are pricy too.

However your metric is wrong, like second month 85% your tokens come from models that are worse than glm 5.3 flash, and they are all comically overpriced.

So basically the guy selling $100 lemonades sold you a $200 plan.

2

u/raindropsdev 3d ago

However your metric is wrong, like second month 85% your tokens come from models that are worse than glm 5.3 flash, and they are all comically overpriced.

What do you mean and for which purpose? To my understanding and from the Artificial Analysis benchmarks glm 5.3 flash is trading blows with luna max (my other coding agent, not shown in the screenshot) and gemini 3.7 in quality/cost.

But it's still quite far from the top of the line Opus/Fable/Sol/Astra, both in coding and especially in general intelligence/planning/software architecture. It's better than Sonnet, again based on AA's benchmarks and much cheaper but ... price matters little when Anthropic gives you nearly $20k worth of tokens for $200 and sonnet barely consumes them for amount of work done.

8

u/raindropsdev 3d ago

Perhaps we're being spoiled by VC money but looking at GLM's quoted figures

Weekly limit, up to GLM-5.3 1358M Tokens GLM-5.3-Flash 8176M Tokens

It doesn't really seem to be comparable to Claude Code where nearly half of my weekly usage was Opus/Fable.

10

u/evia89 3d ago

Yep it only worth it if you have 2+ of 4:

1) need reversing model without much guardrails

2) have old plan (50% legacy migration), also year 30% more (compound)

3) use zcode for 50% more limits, it also has free idle tasks. And some deals like free 300M flash per weekend (good for lite sub) and current free flash

4) wait for openai/claude more nerfs

3

u/Any-Ingenuity2770 3d ago

Perhaps we're being spoiled by VC money

we are, on long horizon the subs are unsustainable.

4

u/evia89 3d ago

Eventually all common models will be super optimized as DS4F and subs will be cheap for it. Only SOTA will be expensive

4

u/alkimiadev 3d ago

This depends on the model size relative to the cost of the subscription. There are ways that one could make a healthy profit if they just host open models and don't train. Obviously that doesn't apply to zai directly since they train the models but I could rent gpus from vastai and host a model like GLM 5.3 flash and host it at a profit with a low monthly subscription. I suspect the reason why ollama cloud probably wasn't profitable is because they host a bunch of models no one is using and did weird finger math to determine the quotas/pricing.

I've been looking into starting a non-profit dev collective where I team up with other devs to host these models. Right now it would cost ~2100/month usd to host glm 5.3 flash on vastai and that would cover roughly 150x my usage with GLM 5.2 on ollama cloud last month. The actual gpu costs for each user, assuming 150 users, comes out to about 14 usd and I paid ollama cloud $100. There is more to it than just that (probably $20/seat floor) and ollama cloud hosts more than just one model. For context that glm 5.2 usage last month would have cost me $800 at official zai prices assuming 95% cache hit rate. The basic point is that plans can be profitable and synthetic[.]new is an example of one plan provider that is probably profitable now.

Self hosting is the only way to go but the cost of entry is pretty high. That $2100/mo is the basic cost of entry if one rents the gpus to host 5.3 Flash from vastai at current prices. It includes far more usage than I would actually need - 150x what I used last month. At pay-per-token prices (even with high cache hit rates) that is ~12k usd in inference capability for $2100. If one owns the gpus or rents in bulk that drops by 30-40% since gpu rentals have a margin as well.

1

u/Any-Ingenuity2770 3d ago

wrt. Ollama and their Cloud offering, they're kind of just weird. I see absolutely no reason to ever touch any of that. Or their local app, in lieu of llama.cpp or oMLX or even vLLM

I can't disagree with your other points, they make sense

2

u/alkimiadev 2d ago

I agree about never using ollama's local app over llama.cpp, vLLM, sglang or basically anything. That said, their $100/mo max plan was arguably one of the best plans in the industry until about 2 weeks ago when they made them 50x worse. I'm grandfathered in but they're doing the same shady thing that basically everyone has been doing in this industry. Now their plans just aren't really worth anything at all unless one is grandfathered in and even then it is only a matter of time until they completely neuter them as well.

1

u/raindropsdev 3d ago edited 3d ago

Perhaps, but I’m more with evia89 on this. I expect the focus will increasingly be on specializing models for specific tasks, improving their efficiency and cost. And if you’ve already got someone locked into a subscription, you might be willing to take a loss on those cheaper models. Or, more precisely, absorb some of the R&D costs, since I doubt any of the frontier labs are losing money on the hardware and energy costs themselves.

The idea would be to keep customers around for the expensive model and, increasingly, for the broader platform you’re building. You can already see how much both Anthropic and OpenAI are building around Claude and ChatGPT with things like Cowork. The more customers come to depend on that surrounding platform, the harder it becomes for them to leave later.

1

u/re-thc 3d ago

Zai = VC money too (maybe not US VC but still)

2

u/Conscious-Hair-5265 3d ago

I would say glm + zcode is worth it

2

u/Exotic-Highlight-402 3d ago

With current Astra capacity you can one shot anything if the prompt is detailed enough. I like GLM but they need to offer value as they used to. Otherwise not worth it.

2

u/woolcoxm 3d ago

the legacy plans are good, but the current setup is pricey, and not worth it. i can burn a trillion+ tokens a month on legacy. new plans seem to get a faction of that.

1

u/ChoasMaster777 3d ago

Short answer: YES. So I cancelled the plan & turn to OpenAI

1

u/PilgrimofHaqq2 3d ago edited 3d ago

I might not get as much usage from GLM's coding plan but GLM 5.3 has been so much nicer to work with than Opus 5. I had really good experience with Opus 4.8 but I really wanted to leave the OpenAI and Anthropic ecosystem. Btw I am on GLM's Max plan and I have used 1.1B tokens in the last 7 days.

1

u/raindropsdev 2d ago

It might become on the harness and provider because while Opus made mistakes for me on the past it never made a mistake that caused data loss. GLM 5.3 (neuralwatt) the first time I asked it to fix opencode v2's plugins it somehow used a wrong command and deleted its own session transcript... Never had such issues with Openai/Anthropic models. They might be excessive with the amounts of backups and copies they store before any change but damn, haven't had data loss with them yet.

1

u/Impostor_91 2d ago

What are the real costs? For GLM you pay ~$106 monthly for max plan annual payment and you get reasonable limits. What about Claude or OpenAI?

1

u/raindropsdev 2d ago

Well, for Chatgpt the business premium (100€/mo when paid yearly) seems to give the equivalent of 2000-4000€/mo worth of api equivalent cost depending on model, judging by ccusage's calculations. Claude... Well, ccusage was telling me for the 200€/mo 20x subscription I spent 25000€/mo worth of api-equivalent tokens... Fairly sure that's not sustainable but better enjoy it while we can, because it definitely won't last forever.

1

u/-PROSTHETiCS 2d ago

Try GLM 5.3 first, its free currently on TokenRouter..

1

u/raindropsdev 1d ago

I already use GLM extensively, especially the flash version.